Digital signal processing load balancing method, scheduling digital signal processor, system, chip and readable medium

By transferring the task scheduling function from the host CPU to the scheduling DSP in the SOC system, and allocating tasks according to the current load (processing time) of the DSP, the problems of excessive host CPU load and unbalanced load are solved, achieving more efficient load balancing and meeting real-time requirements.

CN120849070APending Publication Date: 2025-10-28SANECHIPS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410440034.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing technologies, load balancing for multiple instance CV DSPs is mainly handled by the host CPU. This increases the burden on the host CPU when there are many CV DSPs, and the load calculation is not scientific enough to meet the needs of scenarios with high real-time requirements.

Method used

The task scheduling function is transferred from the host CPU to the scheduling DSP. The load balancer determines the target DSP based on the current load of each DSP (representing the total processing time of the tasks to be processed) and sends tasks to it, thereby reducing the burden on the host CPU and achieving a more scientific load balance.

Benefits of technology

It reduces the burden on the host CPU, improves the overall performance of the SOC, and achieves load balancing for multi-instance computer vision DSPs in scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849070A_ABST
    Figure CN120849070A_ABST
Patent Text Reader

Abstract

The invention provides a digital signal processing load balancing method, which is applied to scheduling digital signal processors (DSPs), and comprises the following steps: receiving a target task sent by a host central processing unit (Host CPU), determining the current load of each DSP, the load being used for representing the total processing duration of a task to be processed, and the DSPs at least comprising a calculation DSP; determining a target DSP according to the current load of each DSP, and issuing a target task to the target DSP; according to the embodiment of the invention, the task scheduling function is transferred from the Host CPU to the scheduling DSP, so that the burden of the Host CPU can be reduced; besides, the execution duration of the task is considered during load balancing, so that the load balancing is more reasonable, the overall performance of the SOC is improved, and the load balancing requirement of the multi-instance computer vision DSP in a scene with a high real-time requirement can be met. The invention also provides a scheduling digital signal processor, a system, a chip and a readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to a digital signal processing load balancing method, a scheduling digital signal processor, a system, a chip, and a readable medium. Background Technology

[0002] As chips specifically designed to accelerate computer vision algorithms, CV (Computer Vision) DSPs (Digital Signal Processing) have been widely used in various computer vision applications. Compared to traditional CPU computing platforms, CV DSPs employ a Harvard architecture with data and program separation, and also feature SIMD (Single Instruction Multiple Data) and VLIW (Very Long Instruction Word) characteristics, resulting in faster instruction execution speeds. Since applications in smart cockpits and assisted driving systems typically require multiple CV algorithms to run simultaneously, the performance of a single CV DSP is insufficient. Multi-instance CV DSPs, however, can perform parallel computation of multiple algorithms, significantly improving computational efficiency and enabling more flexible task allocation.

[0003] As the number of CV DSPs in a System-on-Chip (SoC) continues to increase, load imbalance may occur among multiple CV DSPs. Therefore, an efficient scheduling mechanism is needed to ensure load balancing among multiple CV DSP instances during parallel computing. In related technologies, the host CPU counts the number of tasks in each CV DSP task queue and performs load balancing across multiple CV DSP instances based on the number of tasks. However, as the number of CV DSPs increases, the burden on the host CPU significantly increases, impacting the overall SoC performance. Furthermore, in practical applications, the load on a CV DSP does not solely depend on the number of tasks. Therefore, existing DSP scheduling schemes are insufficient to meet the demands of scenarios with high real-time requirements. Summary of the Invention

[0004] This disclosure provides a digital signal processing load balancing method, a scheduling digital signal processor, a system, a chip, and a readable medium.

[0005] In a first aspect, embodiments of this disclosure provide a digital signal processing load balancing method, the method being applied to scheduling a digital signal processor (DSP), the method comprising:

[0006] Receive the target task sent by the host CPU;

[0007] Determine the current load of each DSP, the load being used to characterize the total processing time of the task to be processed, wherein the DSP includes at least a computation DSP;

[0008] The target DSP is determined based on the current load of each DSP.

[0009] The target task is sent to the target DSP.

[0010] In another aspect, embodiments of this disclosure also provide a scheduling digital signal processor, comprising: at least one processor; a memory storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the digital signal processing load balancing method as described above; and at least one I / O interface connected between the processor and the memory, configured to enable information interaction between the processor and the memory.

[0011] In another aspect, embodiments of this disclosure also provide a digital signal processing system, including at least two computational digital signal processors and a scheduling digital signal processor as described above, wherein each computational digital signal processor and the scheduling digital signal processor are connected via an internal bus.

[0012] In another aspect, embodiments of this disclosure also provide a chip including a host CPU, a storage device, and a digital signal processing system as described above, wherein the host CPU, the storage device, and the digital signal processing system are connected via a chip bus.

[0013] In another aspect, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed, implements the digital signal processing load balancing method as described above.

[0014] The digital signal processing load balancing method provided in this disclosure is applied to scheduling a digital signal processor (DSP). The method includes: receiving a target task sent by a host CPU; determining the current load of each DSP, where the load characterizes the total processing time of the task to be processed; and specifying that each DSP includes at least a computational DSP. The method then determines a target DSP based on the current load of each DSP and sends the target task to that target DSP. This disclosure transfers the task scheduling function from the host CPU to the scheduling DSP, reducing the burden on the host CPU. Furthermore, the method considers the task execution time during load balancing, making the load balancing more reasonable, improving the overall performance of the system-on-a-chip (SoC), and meeting the load balancing requirements of multi-instance computer vision DSPs in scenarios with high real-time requirements. Attached Figure Description

[0015] Figure 1This is a schematic diagram of the SOC system architecture according to an embodiment of the present disclosure;

[0016] Figure 2 A schematic diagram of the digital signal processing load balancing process provided in the embodiments of this disclosure. Figure 1 ;

[0017] Figure 3 This is a schematic diagram illustrating the process of determining the current load of each DSP, provided in an embodiment of this disclosure.

[0018] Figure 4 A schematic diagram of the digital signal processing load balancing process provided in the embodiments of this disclosure. Figure 2 ;

[0019] Figure 5 A schematic diagram illustrating communication between the DSP system and the Host CPU provided in an embodiment of this disclosure;

[0020] Figure 6 This is a schematic diagram of the process of issuing a target task to a target DSP according to an embodiment of this disclosure;

[0021] Figure 7 A schematic diagram illustrating the task queue maintenance of the scheduling DSP provided in an embodiment of this disclosure;

[0022] Figure 8 A schematic diagram of the structure of the scheduling DSP provided in the embodiments of this disclosure;

[0023] Figure 9 This is a schematic diagram of the structure of a DSP system provided in an embodiment of this disclosure. Detailed Implementation

[0024] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, these exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this disclosure.

[0025] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the said feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded.

[0027] The embodiments described herein can be described with reference to plan views and / or cross-sectional views using the ideal schematic diagrams of this disclosure. Therefore, the example illustrations can be modified according to manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to those shown in the drawings, but include modifications to configurations formed based on manufacturing processes. Therefore, the areas illustrated in the drawings are schematic in nature, and the shapes of the areas shown in the figures illustrate specific shapes of areas of an element, but are not intended to be limiting.

[0028] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0029] In related technologies, load balancing for multi-instance CV DSPs is mostly achieved in the following way: the host CPU is responsible for counting the number of tasks in each CV DSP's task queue. When distributing tasks, the host CPU prioritizes distributing tasks to CV DSPs with fewer tasks in their queues. This approach has two major drawbacks: 1. Load balancing for CV DSPs is entirely handled by the host CPU. When there are many CV DSPs, this significantly increases the load on the host CPU, affecting the overall performance of the SoC. 2. The load of a CV DSP is entirely determined by the number of tasks in its task queue. However, in actual use, the processing time of various CV algorithms varies greatly. Sometimes, the processing time of a complex task exceeds the processing time of dozens or even hundreds of simple tasks. Therefore, the number of tasks does not accurately reflect the load.

[0030] To address the aforementioned problems, this disclosure provides a digital signal processing load balancing method. Figure 1 This is a schematic diagram of the SOC system architecture according to an embodiment of this disclosure. Figure 1 As shown, the SOC system includes: a Host CPU, a NOC (network-on-chip), DDR (Double Data Rate), and a DSP system. The DSP system includes a scheduling DSP, N computing DSPs (DSP 1, DSP 2, ..., DSP N), an Inter NOC (internal network-on-chip), and shared memory.

[0031] The host CPU initiates tasks, distributing upper-layer tasks to the scheduling DSP and receiving computation results from the scheduling DSP. The NOC (System-on-Chip) is the SOC system bus, used to connect various devices for transmitting data, addresses, and control information. DDR (External Memory Storage) stores task descriptors for the CV DSPs, describing task information such as data storage address and estimated processing time. The scheduling DSP is the hardware entity for load balancing and scheduling multiple CV DSP instances. It receives tasks from the host CPU and distributes them to less heavily loaded compute DSPs through load balancing. The scheduling DSP itself can also handle some computational tasks. Each compute DSP is the primary CV DSP executing computational tasks, responsible for executing CV algorithms. Internally, it includes cache units, data transfer modules, and vector acceleration computation units. Each compute DSP has a unique and fixed ID. The Inter NOC is the internal bus of the DSP system, connected to both the scheduling DSP and each compute DSP. Shared memory is the physical memory shared by the scheduling DSP and compute DSPs, used to store intermediate computation results and temporarily store final computation results.

[0032] The digital signal processing load balancing method provided in this disclosure is applied to scheduling a DSP. Figure 2 A schematic diagram of the digital signal processing load balancing process provided in the embodiments of this disclosure. Figure 1 ,like Figure 2 As shown, the digital signal processing load balancing method includes the following steps:

[0033] Step S11: Receive the target task sent by the Host CPU.

[0034] The host CPU issues a new CV algorithm task to the scheduling DSP, which then further distributes the new CV algorithm task through load balancing. This new CV algorithm task is the target task. In this embodiment, transferring the scheduling task of the multi-instance CV DSP from the host CPU to the CV DSP can reduce the load on the host CPU.

[0035] Step S12: Determine the current load of each DSP. The load is used to characterize the total processing time of the task to be processed. The DSPs include at least the computation DSP.

[0036] The computational DSP is the primary executor of the CV algorithm tasks; therefore, scheduling the DSP requires calculating the current load of each computational DSP. In this embodiment, the load represents the total duration of the tasks to be processed, not the total number of tasks. In this step, the load of each computational DSP at the current moment is calculated at least for each DSP. In this embodiment, the calculation of DSP load takes into account the processing time of the tasks executed by the DSP. This load calculation method is more scientific and more conducive to achieving load balancing of multiple CV DSP instances.

[0037] Step S13: Determine the target DSP based on the current load of each DSP.

[0038] The scheduling of DSPs includes a load balancer, which is used to calculate the load on each DSP and distribute the target task to the DSP with the least load.

[0039] Step S14: Send the target task to the target DSP.

[0040] Select the DSP with the lowest current load as the target DSP, and distribute the target task to the target DSP with the lowest load.

[0041] The digital signal processing load balancing method provided in this disclosure is applied to scheduling a digital signal processor (DSP). The method includes: receiving a target task sent by a host CPU; determining the current load of each DSP, where the load characterizes the total processing time of the task to be processed; and specifying that each DSP includes at least a computational DSP. The method then determines a target DSP based on the current load of each DSP and sends the target task to that target DSP. This disclosure transfers the task scheduling function from the host CPU to the scheduling DSP, reducing the burden on the host CPU. Furthermore, the method considers the task execution time during load balancing, making the load balancing more reasonable, improving the overall performance of the system-on-a-chip (SoC), and meeting the load balancing requirements of multi-instance computer vision DSPs in scenarios with high real-time requirements.

[0042] Figure 3 This is a schematic diagram illustrating the process of determining the current load of each DSP provided in embodiments of this disclosure. In some embodiments, such as... Figure 3 As shown, determining the current load of each DSP (i.e., step S12) includes the following steps:

[0043] Step S121: Calculate the time interval between the target task and the previously received task.

[0044] After the scheduling DSP receives the target task, it records the time when the target task is received and calculates the time interval between the time when the target task is received and the time when the previous task is received. Specifically, the time interval between the target task and the received previous task can be calculated according to the following formula (1):

[0045] Δt = time now -time pre (1)

[0046] Where Δt is the time interval between the target task and the previously received task, time now The time when the target task is received. pre This refers to the moment when the previous task was received.

[0047] Step S122: For each DSP, determine the current load of the DSP based on the load of the DSP when receiving the previous task and the time interval.

[0048] In this step, the load of each DSP is calculated separately. Specifically, the current load of each DSP can be calculated according to the following formula (2):

[0049]

[0050] Among them, load old To receive the payload from the previous task, load now This represents the current load, i.e., the load when the target task is received.

[0051] The load when the DSP receives the previous task old If the load is less than the time interval Δt, the current load of the DSP is determined to be 0; the load when the DSP received the previous task is also considered. old If the time interval Δt is greater than or equal to the current DSP load, the current DSP load is determined as the difference between the load when the DSP received the previous task and the time interval. old -Δt).

[0052] Load the payload received from the previous task. old Subtracting the time interval Δt between the target task and the previous task gives the current load. If the time interval Δt exceeds the load when receiving the previous task... old This indicates that all tasks of the DSP have been completed, and at this point, the load is updated to 0. If the time interval Δt is less than the load when the previous task was received... old This indicates that the DSP still has unfinished tasks, so the load is updated to (load old -Δt).

[0053] Since the current load status of each DSP must be considered each time load balancing is performed, in order to ensure the accuracy of load balancing when new tasks are sent out, the current load of the target DSP can be further updated after the target task is sent to the target DSP.

[0054] Figure 4 A schematic diagram of the digital signal processing load balancing process provided in the embodiments of this disclosure. Figure 2 In some embodiments, such as Figure 4 As shown, after issuing the target task to the target DSP (i.e., step S14), the digital signal processing load balancing method may further include the following steps:

[0055] Step S15: Update the current load of the target DSP according to the estimated processing time of the target task. The estimated processing time of the target task is sent by the Host CPU to the scheduling DSP when sending the target task.

[0056] After the scheduling DSP sends the target task to the target DSP, it is necessary to further update the load size of the target DSP. Specifically, the updated load of the target DSP can be calculated according to the following formula (3):

[0057] load new =load now +time (3)

[0058] Among them, load new The load is the updated payload for the target DSP, i.e., the payload after receiving the target task; load now The load of the target DSP before receiving the target task is specified, i.e., the load after receiving the previous task; time is the estimated processing time of the target task, which can be determined by simulation software and sent to the scheduling DSP by the host CPU when sending the target task.

[0059] In some embodiments, the DSP used to perform CV DSP tasks can be not only a computation DSP but also a scheduling DSP. That is, the scheduling DSP can also undertake the computation tasks of the CV algorithm. This can make full use of the processing resources of the scheduling DSP and improve task processing efficiency when there are many CV DSP tasks on each computation DSP.

[0060] In some embodiments, receiving the target task sent by the Host CPU (i.e., step S11) includes the following steps: receiving a first task notification sent by the Host CPU, and obtaining the task identifier of the target task from a first storage area in the storage device corresponding to the scheduling DSP; the task identifier is used to instruct the target DSP to obtain the input data of the target task according to the task identifier. In this embodiment of the disclosure, the storage device is DDR, and the input data refers to the original data used to execute the target task.

[0061] In some embodiments, the task identifier of the target task includes at least: the storage address of the input data and calculation results of the target task, and the storage address of the intermediate calculation results of the target task, and may also include the estimated processing time of the target task. By storing the task identifier of the target task in DDR, and using the task identifier to represent the relevant information of the target task, the target DSP can obtain the original data and calculation results of the target task based on the task identifier.

[0062] Figure 5 This is a schematic diagram illustrating communication between the DSP system and the host CPU provided in an embodiment of this disclosure. Figure 5 As shown, communication between the Host CPU and the scheduling DSP in the DSP system, as well as between the scheduling DSP and each computing DSP in the DSP system, is achieved through an interrupt dispatcher. The interrupt dispatcher is used for interrupt routing and distribution among heterogeneous computing cores. DDR is divided into multiple memory areas (comm 0, comm 1, ..., comm N), with one memory area corresponding to one DSP. The memory area is shared memory in DDR used for communication. Communication data between the Host CPU and the scheduling DSP, and between the scheduling DSP and each computing DSP, is stored in the corresponding memory area. comm 0 is the first memory area between the Host CPU and the scheduling DSP, and comm 1-comm N are the second memory areas between the scheduling DSP and computing DSPs 1-N.

[0063] In some embodiments, receiving the first task notification sent by the Host CPU includes: receiving the first task notification sent by the Host CPU through an interrupt dispatcher. That is, the Host CPU notifies the scheduling DSP of a new CV algorithm task, i.e., the target task, through the interrupt dispatcher, and the Host CPU writes the task descriptor of the target task to comm 0 so that the scheduling DSP can obtain the task descriptor of the target task from comm 0.

[0064] In some embodiments, after the target task is sent to the target DSP (i.e., step S14), the digital signal processing load balancing method may further include the following steps: receiving a first task completion notification sent by the target DSP; the first task completion notification is sent by the target DSP after completing the target task; sending a second task completion notification to the host CPU; the second task completion notification is used to instruct the host CPU to obtain the task identifier of the target task from the first storage area so as to obtain the calculation result of the target task according to the task identifier; the task identifier of the target task is copied by the target DSP and stored in the first storage area after completing the target task.

[0065] After the target DSP completes the target task, it stores the calculation result in DDR and sends a first task completion notification to the scheduling DSP via the interrupt dispatcher to inform the scheduling DSP that the target task has been completed. It also writes the storage address of the calculation result into the target task's task identifier, copies this task identifier, and stores it in the first memory area of ​​DDR, i.e., comm 0. Upon receiving the first task completion notification, the scheduling DSP sends a second task completion notification to the host CPU to inform the host CPU that the target task has been completed and instructs the host CPU to retrieve the target task's task identifier from the first memory area. Based on this task identifier, the host CPU then retrieves the calculation result of the target task from the corresponding storage location in DDR.

[0066] Figure 6 This is a schematic diagram of the process of issuing a target task to a target DSP according to an embodiment of this disclosure. In some embodiments, such as... Figure 6 As shown, receiving the target task sent by the Host CPU (i.e., step S14) includes the following steps:

[0067] Step S141: Obtain the task identifier from the first storage area and add the task identifier to the task queue of the target DSP.

[0068] Figure 7 This is a schematic diagram of the scheduling DSP maintaining the task queue provided in an embodiment of this disclosure, as shown below. Figure 7 As shown, the scheduling DSP maintains a task queue for itself and for each computation DSP. Each task queue is actually a linked list, and the elements in the linked list are task descriptors, i.e., task 0, task 1, ..., task n. Each DSP executes the corresponding CV computation task sequentially according to the order of the task descriptors in the task queue. It should be noted that the scheduling DSP also maintains information such as the number of tasks in the queue and the total processing time (i.e., load) for each task queue.

[0069] In this step, the scheduling DSP obtains the task identifier of the target task from comm 0 of the DDR and adds the obtained task identifier to the tail of the task queue of the target DSP.

[0070] Step S142: If the target DSP is a computing DSP, copy the task identifier and store the copied task identifier in the second storage area of ​​the storage device corresponding to the target DSP.

[0071] If the target DSP is a computation DSP, the scheduling DSP copies the task identifier obtained from comm 0 (i.e., the first memory area) and stores the copied task identifier in the second memory area of ​​DDR corresponding to the target DSP. For example, assuming the target DSP is computation DSP 3, the scheduling DSP stores the copied task identifier of the target task in comm 3 of DDR corresponding to computation DSP 3.

[0072] Step S143: Send a second task notification to the target DSP. The second task notification is used to instruct the target DSP to obtain the task identifier from the second memory area.

[0073] In some embodiments, combined with Figure 5 As shown, sending the second task notification to the target DSP includes: sending the second task notification to the target DSP via an interrupt dispatcher. That is, the scheduling DSP sends the second task notification to the target DSP via the interrupt dispatcher to instruct the target DSP to retrieve the task identifier of the target task from the second memory area corresponding to the target DSP.

[0074] Step S144: If the target DSP is a scheduling DSP, obtain the task identifier from the first storage area.

[0075] If the target DSP is the scheduling DSP itself, the scheduling DSP directly obtains the task identifier of the target task from comm 0 in DDR, without having to copy and store the task identifier of the target task, or send a second task notification.

[0076] After steps S143 and S144, once the target DSP obtains the task identifier of the target task, it can parse the task identifier to obtain the storage address of the input data of the target task. Based on this storage address, it retrieves the input data of the target task and executes the corresponding CV calculation task. The storage address of the final calculation result and the storage address of the intermediate calculation results are then written into the task identifier of the target task. In some embodiments, techniques such as DMA (Direct Memory Access) can be used to accelerate the acquisition of external data (i.e., input data, such as image data) from the CV DSP. It should be noted that the storage address of the intermediate calculation results can be a shared memory address, meaning that the intermediate calculation results can be stored in shared memory.

[0077] To clearly illustrate the DSP load balancing scheduling scheme of this disclosure, the overall process of DSP load balancing scheduling is described in detail below. The DSP load balancing scheduling method includes the following steps:

[0078] Step 1, Task Issuance. The Host CPU notifies the scheduling DSP of a new CV algorithm task via the interrupt dispatcher. The Host CPU then writes the task descriptor of the new CV algorithm task to the comm 0 area of ​​the DDR.

[0079] Step 2, Load Calculation. After receiving the notification from the Host CPU, the scheduling DSP obtains the task descriptor from comm 0, records the time when the new CV algorithm task was received, calculates the time interval since the last task was received, and calculates the current load of all DSPs.

[0080] Step 3, Task Allocation. The scheduling DSP selects the DSP with the lowest current load as the target DSP to execute the new CV algorithm task, and adds the new CV algorithm task to the target DSP's task queue. If the target DSP is a computation DSP, the scheduling DSP notifies the target DSP of the new CV algorithm task via the interrupt dispatcher, and copies the task descriptor of the new CV algorithm task to the corresponding comm area in DDR for the target DSP.

[0081] Step 4, Task Execution. The target DSP executes each task sequentially according to its own task queue. Specifically, it retrieves the task identifier of the task to be executed from the corresponding comm area in the DDR, obtains the input data based on the task identifier, and performs calculations.

[0082] Step 5: Feedback of Calculation Results. When the CV algorithm task is completed, the target DSP notifies the scheduling DSP via the interrupt dispatcher to add the storage address of the calculation result to the task identifier and copy the task identifier to the comm 0 area. Then, the scheduling DSP notifies the host CPU via the interrupt dispatcher that the calculation is complete, and the host CPU retrieves the calculation result from the comm 0 area.

[0083] The following simulation experiments clearly illustrate the load balancing effect of the embodiments of this disclosure. In the simulation experiments, the SOC system integrates two CV DSPs, and the following six test tasks are to be executed: reszie (image resizing), normalize (image pixel normalization), hwc2hcw (hwc format to hcw format conversion), bilateral filter (bilateral filtering), canny (edge ​​detection), and pyrup (image magnification). These six test tasks are all high-frequency CV tasks in real-world scenarios, and it is assumed that the above six tasks are issued simultaneously. The test image size is 1920*1080. First, software testing was conducted to obtain the time taken by each of the above test tasks on a single CV DSP: 0.52ms, 0.42ms, 0.1ms, 8.63ms, 17.2ms, and 2.8ms, respectively. Table 1 lists the load of each CV DSP in the traditional load balancing scheme and the load balancing scheme of the embodiments of this disclosure.

[0084] Table 1

[0085] Using traditional load balancing solutions This publicly disclosed load balancing solution is adopted. DSP1 load (ms) 1.04 9.15 DSP2 load (ms) 28.63 20.52 DSP2 load / DSP1 load 27.5 2.24

[0086] As shown in Table 1, using the traditional load balancing scheme, the total processing time of DSP1 is 1.04ms, and the total processing time of DSP2 is 28.63ms. The ratio of the total processing time of DSP2 to DSP1 is 27.5, indicating a large ratio and unbalanced DSP load. Using the load balancing scheme disclosed in this invention, the total processing time of DSP1 is 9.15ms, and the total processing time of DSP2 is 20.52ms. The ratio of the total processing time of DSP2 to DSP1 is 2.24, indicating a small ratio. Compared to the traditional load balancing scheme, the ratio of the total processing time of DSP2 to DSP1 is reduced by more than 10 times, resulting in a more balanced DSP load.

[0087] This disclosure relates to a scheduling scheme for load balancing of multiple CV DSP instances. When multiple CV DSP instances are processing data in parallel or processing multiple tasks simultaneously, this disclosure can be used to schedule multiple tasks among the DSPs, achieving load balancing among the multiple CV DSP instances. This scheme has two main features: First, since a CV DSP is an "enhanced" CPU with its own scheduling capabilities, one of the multiple CV DSP instances can be used as a scheduling DSP, responsible for scheduling CV algorithms across the multiple CV DSP instances. Furthermore, the scheduling resources and computing resources of the DSP are independent of each other; executing scheduling tasks does not affect its own computing resources. The remaining CV DSPs can act as computing DSPs, fully responsible for computing tasks. Second, since there is dedicated simulation software for testing CV DSP performance, the estimated processing time of each CV task can be obtained in advance through simulation. The sum of the estimated processing times of all tasks in each CV DSP task queue can be used as the load of that CV DSP. The embodiments disclosed herein transfer the scheduling task of multi-instance CV DSP from the host CPU to the CV DSP, which can reduce the load on the host CPU. Furthermore, the calculation of the load takes into account the processing time required for the task. This load calculation method is more scientific and more conducive to achieving load balancing of multi-instance CV DSP.

[0088] The embodiments disclosed herein can be applied to the following scenarios: intelligent vehicle and intelligent cockpit and assisted driving scenarios, smartphone scenarios, smart home scenarios, and other scenarios involving the deployment of a large number of visual algorithm applications and requiring multi-instance CVDSP load balancing support.

[0089] This disclosure also provides a scheduling digital signal processor, such as... Figure 8 As shown, it includes:

[0090] At least one processor 801;

[0091] The memory 802 stores at least one program, which, when executed by the at least one processor, enables the at least one processor to implement the digital signal processing load balancing method provided in the foregoing embodiments.

[0092] At least one I / O interface 803 is configured to enable information exchange between the processor and memory.

[0093] Among them, processor 801 is a device with data processing capabilities, including but not limited to central processing unit (CPU); memory 802 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); I / O interface (read-write interface) 803 enables information interaction between processor 801 and memory 802.

[0094] In some embodiments, the processor 801, memory 802, and I / O interface 803 are interconnected via a bus, and thus connected to other components of the computing device.

[0095] This disclosure also provides a digital signal processing system, such as... Figure 9 As shown, the digital signal processing system includes at least two computational DSPs and a scheduling DSP as described above, with each computational DSP and the scheduling DSP connected via an internal bus.

[0096] This disclosure also provides a chip, such as... Figure 1 As shown, the chip includes a Host CPU, a storage device, and a digital signal processing system as described above. The Host CPU, storage device, and digital signal processing system are connected via a chip bus. The storage device can be DDR, and the functions of the digital signal processing system are as described previously and will not be repeated here.

[0097] This disclosure also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed, implements the digital signal processing load balancing method provided in the foregoing embodiments.

[0098] It will be understood by those skilled in the art that all or some of the steps in the methods disclosed above, and the functional modules / units in the apparatus, can be implemented as software, firmware, hardware, and suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0099] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A digital signal processing load balancing method, the method being applied to scheduling a digital signal processor (DSP), the method comprising: Receive the target task sent by the host CPU; Determine the current load of each DSP, the load being used to characterize the total processing time of the task to be processed, wherein the DSP includes at least a computation DSP; The target DSP is determined based on the current load of each DSP. The target task is sent to the target DSP.

2. The method as described in claim 1, characterized in that, Determining the current load of each DSP includes: Calculate the time interval between the target task and the previously received task; For each DSP, the current load of the DSP is determined based on the load of the DSP when it received the previous task and the time interval.

3. The method as described in claim 2, characterized in that, The step of determining the current load of the DSP based on the load when the DSP received the previous task and the time interval includes: If the load of the DSP when receiving the previous task is less than or equal to the time interval, the current load of the DSP is determined to be 0. If the load of the DSP when receiving the previous task is greater than the time interval, the current load of the DSP is determined to be the difference between the load of the DSP when receiving the previous task and the time interval.

4. The method as described in claim 1, characterized in that, After issuing the target task to the target DSP, the method further includes: The current load of the target DSP is updated based on the estimated processing time of the target task, which is sent by the Host CPU to the scheduling DSP when the target task is sent.

5. The method according to any one of claims 1-4, characterized in that, The DSP also includes the scheduling DSP.

6. The method as described in claim 5, characterized in that, The target task sent by the host CPU includes: Receive the first task notification sent by the Host CPU; The task identifier of the target task is obtained from the first storage area corresponding to the scheduling DSP in the storage device; the task identifier is used to instruct the target DSP to obtain the input data of the target task according to the task identifier.

7. The method as described in claim 6, characterized in that, After issuing the target task to the target DSP, the method further includes: Receive a first task completion notification sent by the target DSP; the first task completion notification is sent by the target DSP after completing the target task. A second task completion notification is sent to the Host CPU; the second task completion notification is used to instruct the Host CPU to obtain the task identifier of the target task from the first storage area, so as to obtain the calculation result of the target task according to the task identifier; the task identifier of the target task is copied by the target DSP and stored in the first storage area after the target task is completed.

8. The method as described in claim 6, characterized in that, Sending the target task to the target DSP includes: Obtain the task identifier from the first storage area and add the task identifier to the task queue of the target DSP; If the target DSP is a computing DSP, the task identifier is copied and the copied task identifier is stored in the second storage area of ​​the storage device corresponding to the target DSP. A second task notification is sent to the target DSP, which instructs the target DSP to retrieve the task identifier from the second storage area.

9. The method as described in claim 6, characterized in that, Sending the target task to the target DSP includes: Obtain the task identifier from the first storage area and add the task identifier to the task queue of the target DSP; If the target DSP is the scheduling DSP, the task identifier is obtained from the first storage area.

10. A scheduling digital signal processor, characterized in that, include: At least one processor; A memory having stored at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the digital signal processing load balancing method according to any one of claims 1-9; At least one I / O interface is configured to enable information interaction between the processor and the memory.

11. A digital signal processing system, characterized in that, It includes at least two computational digital signal processors and a scheduling digital signal processor as described in claim 10, wherein each of the computational digital signal processors is connected to the scheduling digital signal processor via an internal bus.

12. A chip, characterized in that, It includes a host CPU, a storage device, and a digital signal processing system as described in claim 11, wherein the host CPU, the storage device, and the digital signal processing system are connected via a chip bus.

13. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed, it implements the digital signal processing load balancing method as described in any one of claims 1-9.