Digital signal processing load balancing method, scheduling digital signal processor, system, chip, and readable medium
By introducing a scheduling DSP into the SOC system, load balancing scheduling is performed based on the total processing time of the DSP tasks, which solves the problems of excessive host CPU load and unscientific load calculation, achieves efficient load balancing of multiple instance CV DSPs, and improves system performance.
Patent Information
- Application Number
- PCT/CN2025/085004
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-12
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-16
AI Technical Summary
In existing technologies, load balancing for multiple instance CV DSPs is mainly handled by the host CPU. This increases the burden on the host CPU when there are many CV DSPs, and the load calculation is not scientific enough to meet the needs of scenarios with high real-time requirements.
By introducing a scheduling DSP into the SOC system, tasks are received from the Host CPU and load-balanced scheduling is performed based on the current load of each DSP (i.e., the total processing time of the task). The tasks are distributed to the DSP with the least load, reducing the burden on the Host CPU, and the task execution time is taken into account to achieve a more scientific load balance.
It effectively reduces the host CPU load, improves the overall performance of the SOC, achieves load balancing of multiple CV DSP instances, and meets the needs of scenarios with high real-time requirements.
Smart Images

Figure CN2025085004_16102025_PF_FP_ABST
Abstract
Description
Digital signal processing load balancing method, scheduling digital signal processor, system, chip and readable medium
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202410440034.6, filed on April 12, 2024, with the Chinese Patent Office, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to, but is not limited to, the technical field of computer technology. BACKGROUND
[0004] As a chip specially designed for accelerating computer vision algorithms, CV (Computer Vision) DSP (Digital Signal Processing) has been widely used in various computer vision applications. Compared with traditional CPU operation platforms, CV DSP adopts a Harvard architecture that separates data and programs, and also has two major features of SIMD (Single Instruction Multiple Data) and VLIW (Very Long Instruction Word), and the execution speed of instructions is faster. Due to the application programs of intelligent cockpit and auxiliary driving, multiple CV algorithms usually need to run simultaneously, and the performance of a single CV DSP is difficult to meet the requirements, while multiple instance CV DSP can perform parallel computing of multiple algorithms, which can significantly improve the computing efficiency and realize more flexible task allocation.
[0005] With the increasing number of CV DSPs in SOC (System on Chip), load imbalance may occur between multiple CV DSPs, and therefore an efficient scheduling mechanism is needed to ensure load balancing of multiple CV DSP instances during parallel computing. In related technologies, the Host CPU (Central Processing Unit) counts the number of tasks in each CV DSP task queue, and performs load balancing of multiple instance CV DSPs according to the number of tasks. However, with the increasing number of CV DSPs, the burden on the Host CPU will be greatly increased, affecting the overall performance of the SOC, and in actual use, the load of the CV DSP is not only determined by the number of tasks. Therefore, the existing DSP scheduling scheme is difficult to meet the requirements of scenarios with high real-time requirements. SUMMARY
[0006] The present disclosure provides a digital signal processing load balancing method, a scheduling digital signal processor, a system, a chip and a readable medium.
[0007] In a first aspect, the embodiments of the present disclosure provide a digital signal processing load balancing method, which is applied to a scheduling digital signal processor (DSP), and the method comprises the following steps: receiving a target task sent by a host central processing unit (Host CPU); determining a current load of each DSP, wherein the load is used to represent a total processing time length of a task to be processed, and the DSPs at least include a computing DSP; determining a target DSP according to the current load of each DSP; and issuing the target task to the target DSP.
[0008] In another aspect, the embodiments of the present disclosure further provide a scheduling digital signal processor, which comprises: at least one processor; a memory, which stores at least one program, and when the at least one program is executed by the at least one processor, the at least one processor implements any digital signal processing load balancing method as described herein; and at least one I / O interface, which is connected between the processor and the memory and is configured to realize information interaction between the processor and the memory.
[0009] In another aspect, the embodiments of the present disclosure further provide a digital signal processing system, which comprises at least two computing digital signal processors and any scheduling digital signal processor as described herein, and each computing digital signal processor is connected to the scheduling digital signal processor through an internal bus.
[0010] In another aspect, the embodiments of the present disclosure further provide a chip, which comprises a host central processing unit (Host CPU), a storage device and a digital signal processing system as described herein, and the Host CPU, the storage device and the digital signal processing system are connected through a chip bus.
[0011] In another aspect, the embodiments of the present disclosure further provide a computer readable medium, which stores a computer program, wherein the program is executed by a processor to implement any digital signal processing load balancing method as described herein. BRIEF DESCRIPTION OF DRAWINGS
[0012] FIG. 1 is a schematic diagram of a SOC system architecture according to an embodiment of the present disclosure;
[0013] FIG. 2 is a schematic diagram of a digital signal processing load balancing process according to an embodiment of the present disclosure;
[0014] FIG. 3 is a schematic diagram of a process for determining a current load of each DSP according to an embodiment of the present disclosure;
[0015] FIG. 4 is a schematic diagram of a digital signal processing load balancing process according to an embodiment of the present disclosure;
[0016] FIG. 5 is a schematic diagram of a communication between a DSP system and a Host CPU according to an embodiment of the present disclosure;
[0017] FIG. 6 is a flow diagram of a process for issuing a target task to a target DSP according to an embodiment of the present disclosure;
[0018] FIG. 7 is a diagram of a DSP maintaining a task queue according to an embodiment of the present disclosure;
[0019] FIG. 8 is a diagram of a structure of a DSP scheduler according to an embodiment of the present disclosure;
[0020] FIG. 9 is a diagram of a structure of a DSP system according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] In the following description, example embodiments will be described with reference to the accompanying drawings, but the example embodiments can be embodied in different forms and should not be construed as being limited to the embodiments set forth herein. Rather, the embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0022] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0024] The embodiments described herein can be described with reference to plan views and / or cross-sectional views by virtue of the ideal schematic representations of the disclosure. Accordingly, the example illustrations are not intended to be limiting in terms of the embodiments described herein. Thus, the embodiments are well suited to a wide variety of manufacturing processes and / or configurations of the embodiments. As the embodiments within the scope of the disclosure can take many different forms, the example embodiments that are illustrated and described are not intended to limit the spirit or the scope of the disclosure. Thus, the present disclosure is to be accorded the widest scope, and is to be limited only by the claims.
[0025] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.
[0026] In the related art, the load balancing of multi-instance CV DSPs is mostly realized in the following manner: the Host CPU is responsible for counting the number of tasks in each CV DSP task queue, and the Host CPU preferentially issues tasks to the CV DSPs with fewer tasks in the task queue when issuing tasks. This scheme has the following two major shortcomings: 1. The load balancing of CV DSPs is completely the responsibility of the Host CPU, and when the number of CV DSPs is large, the burden on the Host CPU is greatly increased, affecting the overall performance of the SOC. 2. The load of the CV DSP is completely determined by the number of tasks in its task queue, but in actual use, the processing time of various CV algorithms differs greatly, and sometimes the processing time of a complex task exceeds that of tens or even hundreds of simple tasks, so the number of tasks cannot truly reflect the load size.
[0027] To at least solve the above problems, the embodiments of the present disclosure provide a digital signal processing load balancing method. FIG. 1 is a schematic diagram of the SOC system architecture according to an embodiment of the present disclosure. As shown in FIG. 1, the SOC system includes a Host CPU, a NOC (network-on-chip), a DDR (Double Data Rate), and a DSP system, wherein the DSP system includes a scheduling DSP, N computing DSPs (DSP 1, DSP 2, …, DSP N), an Inter NOC (internal network-on-chip), and a Share memory (shared memory).
[0028] The Host CPU is the initiator of the task, used to distribute the upper-layer task to the scheduling DSP, and receive the returned calculation result from the scheduling DSP. The NOC is the bus of the SOC system, used to connect various devices and transmit data, address, and control information. The DDR is an external physical memory, used to store the task identifier of the CV DSP, and the task identifier is used to describe the related information of the task, such as the data storage address, the estimated processing time, etc. The scheduling DSP is a hardware entity for load balancing scheduling of multi-instance CV DSPs, used to receive the task issued by the Host CPU, and issue the task to the computing DSP with smaller load through load balancing, in addition, the scheduling DSP itself can also undertake some computing tasks. Each computing DSP is a CV DSP mainly executing computing tasks, responsible for undertaking the execution of the CV algorithm, and includes a cache unit, a data moving module, a vector acceleration computing unit, etc. in the inside, and each computing DSP has a unique and fixed ID. The Inter NOC is the internal bus of the DSP system, and the scheduling DSP and each computing DSP are connected to the Inter NOC. The Share memory is a physical memory shared by the scheduling DSP and the computing DSP, used to store the intermediate calculation result and temporarily store the final calculation result.
[0029] The digital signal processing load balancing method provided by the embodiments of the present disclosure is applied to a scheduling DSP. FIG. 2 is a schematic diagram of a digital signal processing load balancing process provided by the embodiments of the present disclosure. As shown in FIG. 2, the digital signal processing load balancing method can include the following steps S11-S14.
[0030] In step S11, a target task sent by a Host CPU is received.
[0031] The Host CPU issues a new CV algorithm task to the scheduling DSP, so that the scheduling DSP further issues the new CV algorithm task in a load balancing manner. The new CV algorithm task is the target task. In the embodiments of the present disclosure, the scheduling task of the multi-instance CV DSP is transferred from the Host CPU to the CV DSP, so as to reduce the load of the Host CPU.
[0032] In step S12, the current load of each DSP is determined. The load is used to represent the total processing time length of the tasks to be processed. The DSPs include at least a computing DSP.
[0033] The computing DSP is the main executor of the CV algorithm task. Therefore, the scheduling DSP at least calculates the current load of each computing DSP. In the embodiments of the present disclosure, the load is not used to represent the total number of tasks to be processed, but is used to represent the total time length of the tasks to be processed. In this step, the load of each computing DSP at the current time is calculated. In the embodiments of the present disclosure, the processing time length of the DSP in executing the task is considered when the load of the computing DSP is calculated. This load calculation manner is more scientific and is more conducive to achieving the load balancing of the multi-instance CV DSP.
[0034] In step S13, the target DSP is determined according to the current load of each DSP.
[0035] The scheduling DSP includes a load balancing scheduler. The load balancing scheduler is used to calculate the load of each DSP and distribute the target task to the DSP with the smallest load.
[0036] In step S14, the target task is issued to the target DSP.
[0037] The DSP with the smallest current load is selected as the target DSP, and the target task is distributed to the target DSP with the smallest load.
[0038] The method provided by the embodiment of the present disclosure is applied to scheduling a digital signal processor (DSP), and includes: receiving a target task sent by a host central processing unit (Host CPU), determining a current load of each DSP, the load being used to represent a total processing duration of a task to be processed, and the DSP including at least a computing DSP; determining a target DSP according to the current load of each DSP, and issuing the target task to the target DSP; the embodiment of the present disclosure transfers the task scheduling function from the Host CPU to the scheduling DSP, so that the burden of the Host CPU can be reduced; in addition, the execution duration of the task is considered when performing load balancing, so that the load balancing is more reasonable, the overall performance of the SOC is improved, and the load balancing requirement of a multi-instance computer vision DSP in a scene with high real-time requirement can be met.
[0039] FIG. 3 is a flowchart of determining the current load of each DSP provided by the embodiment of the present disclosure. In some embodiments, as shown in FIG. 3, the determination of the current load of each DSP (i.e., step S12) can include steps S121 and S122.
[0040] In step S121, the time interval between the target task and the previous task received is calculated.
[0041] After the scheduling DSP receives the target task, the time when the target task is received is recorded, and the time interval between the time when the target task is received and the time when the previous task is received is calculated. For example, the time interval between the target task and the previous task received can be calculated according to the following formula (1): now -time pre (1)
[0042] Wherein, Δt is the time interval between the target task and the previous task received, time now is the time when the target task is received, time pre is the time when the previous task is received.
[0043] In step S122, for each DSP, the current load of the DSP is determined according to the load of the DSP when the previous task is received and the time interval.
[0044] In this step, the load of each DSP is calculated respectively. For example, the current load of each DSP can be calculated according to the following formula (2):
[0045] Wherein, load old is the load when the previous task is received, and load nowis the current load, that is, the load when the target task is received.
[0046] The load when the DSP receives the previous task old When the time interval is less than Δt, the current load of the DSP is determined to be 0; when the DSP receives the previous task, the load old When the time interval Δt is greater than or equal to the time interval Δt, the current load of the DSP is determined to be the difference between the load when the DSP receives the previous task and the time interval (load old -Δt).
[0047] Will receive the load of the previous task old Subtract the time interval Δt between the target task and the previous task to get the current load. If the time interval Δt exceeds the load when receiving the previous task, old , indicating that all tasks of the DSP have been completed, at this time, the load is updated to 0. If the time interval Δt is less than the load when receiving the previous task old , indicating that the DSP has unfinished tasks, then the load is updated to (load old -Δt).
[0048] Since the current load of each DSP must be considered each time load balancing is performed, in order to ensure the accuracy of load balancing when new tasks are subsequently issued, the current load of the target DSP can be further updated after the target task is issued to the target DSP.
[0049] Figure 4 is a schematic diagram of the digital signal processing load balancing process provided by an embodiment of the present disclosure. In some embodiments, as shown in Figure 4, after the target task is sent to the target DSP (i.e., step S14), the digital signal processing load balancing method may further include the following steps: Step S15, updating the current load of the target DSP based on the estimated processing duration of the target task. The estimated processing duration of the target task is sent by the Host CPU to the scheduling DSP when sending the target task.
[0050] After the scheduling DSP sends the target task to the target DSP, it is necessary to further update the load size of the target DSP. For example, the current load of the target DSP after the update can be calculated according to the following formula (3): load new =load now +time (3)
[0051] Among them, load new The updated load of the target DSP, that is, the load after receiving the target task; loadnow Loadtarget is the load of the target DSP before receiving the target task, i.e., the load after receiving the previous task; time is the estimated processing time of the target task, which can be determined by simulation software and sent by the Host CPU to the scheduling DSP when sending the target task.
[0052] In some embodiments, the DSP for executing the CV DSP task can not only be a computing DSP, but also a scheduling DSP, i.e., the scheduling DSP can also undertake the computing task of the CV algorithm, so that the processing resources of the scheduling DSP can be fully utilized, and the task processing efficiency can be improved in the case that the CV DSP tasks of each computing DSP are relatively large.
[0053] In some embodiments, the receiving of the target task sent by the Host CPU (i.e., step S11) includes the following steps: receiving the first task notification sent by the Host CPU, and obtaining the task identifier of the target task from the first storage area corresponding to the scheduling DSP in the storage device; the task identifier is used to instruct the target DSP to obtain the input data of the target task according to the task identifier. In the embodiments of the present disclosure, the storage device is a DDR, and the input data refers to the original data used for executing the target task.
[0054] In some embodiments, the task identifier of the target task at least includes the storage addresses of the input data and the computing result of the target task, and the storage address of the intermediate computing result of the target task, and can further include the estimated processing time of the target task. By storing the task identifier of the target task in the DDR, the task identifier of the target task is used to represent the related information of the target task, so that the target DSP can obtain the original data of the target task and the computing result of the target task according to the task identifier.
[0055] FIG. 5 is a schematic diagram of the communication between the DSP system and the Host CPU according to an embodiment of the present disclosure. As shown in FIG. 5, the Host CPU and the scheduling DSP in the DSP system, and the scheduling DSP and each computing DSP in the DSP system communicate through an interrupt distributor, which is used for interrupt routing and distribution between the heterogeneous computing cores. The DDR is divided into a plurality of storage areas (Comm 0, Comm 1, …, Comm N), one storage area corresponding to one DSP, and the storage area is a shared memory in the DDR for communication. The communication data between the Host CPU and the scheduling DSP, and between the scheduling DSP and each computing DSP is stored in the corresponding storage area, wherein Comm 0 is the first storage area between the Host CPU and the scheduling DSP, and Comm 1-Comm N are the second storage areas between the scheduling DSP and the computing DSP 1-computing DSP N.
[0056] In some embodiments, the receiving the first task notification sent by the Host CPU comprises: receiving the first task notification sent by the Host CPU through the interrupt distributor. That is, the Host CPU notifies the scheduling DSP of a new CV algorithm task, i.e., a target task, through the interrupt distributor, and the Host CPU writes the task identifier of the target task into the Comm 0 so that the scheduling DSP obtains the task identifier of the target task from the Comm 0.
[0057] In some embodiments, after the target task is issued to the target DSP (i.e., step S14), the digital signal processing load balancing method can further comprise the following steps: receiving a first task completion notification sent by the target DSP; the first task completion notification is sent by the target DSP after completing the target task; sending a second task completion notification to the Host CPU; the second task completion notification is used to instruct the Host CPU to obtain the task identifier of the target task from the first storage area so as to obtain the calculation result of the target task according to the task identifier; the task identifier of the target task is copied and stored in the first storage area by the target DSP after completing the target task.
[0058] After the target DSP completes the target task, the calculation result is stored in the DDR, the first task completion notification is sent to the scheduling DSP through the interrupt distributor to notify the scheduling DSP that the target task has been completed, the storage address of the calculation result is written into the task identifier of the target task, and the task identifier is copied and stored in the first storage area of the DDR, i.e., the Comm 0. After the scheduling DSP receives the first task completion notification, the second task completion notification is sent to the Host CPU to notify the Host CPU that the target task has been completed, and the Host CPU is instructed to obtain the task identifier of the target task from the first storage area, and the calculation result of the target task is obtained from the corresponding storage position in the DDR according to the task identifier of the target task.
[0059] FIG. 6 is a flowchart of issuing a target task to a target DSP according to an embodiment of the present disclosure. In some embodiments, as shown in FIG. 6, the receiving the target task sent by the Host CPU (i.e., step S14) comprises the following steps S141 to S144.
[0060] In step S141, the task identifier is obtained from the first storage area and added to the task queue of the target DSP.
[0061] FIG. 7 is a schematic diagram of a task queue maintained by a scheduling DSP according to an embodiment of the present disclosure. As shown in FIG. 7, the scheduling DSP maintains a task queue for itself and each of the computing DSPs. Each task queue is actually a linked list, and the elements in the linked list are task identifiers, i.e., task 0, task 1, …, task n. Each DSP executes the corresponding CV computation task in the order of the task identifiers in the task queue. It should be noted that the scheduling DSP also maintains information such as the number of tasks in the queue, the total processing time (i.e., the load) of all tasks, etc. for each task queue.
[0062] In this step, the scheduling DSP obtains the task identifier of the target task from Comm 0 of the DDR, and adds the obtained task identifier to the tail of the task queue of the target DSP.
[0063] In step S142, if the target DSP is a computing DSP, the task identifier is copied, and the copied task identifier is stored in the second storage area corresponding to the target DSP in the storage device.
[0064] If the target DSP is a computing DSP, the scheduling DSP copies the task identifier obtained from Comm 0 (i.e., the first storage area), and stores the copied task identifier in the second storage area corresponding to the target DSP in the DDR. For example, if the target DSP is computing DSP 3, the scheduling DSP stores the copied task identifier of the target task in Comm 3 corresponding to computing DSP 3 in the DDR.
[0065] In step S143, a second task notification is sent to the target DSP, which instructs the target DSP to obtain the task identifier from the second storage area.
[0066] In some embodiments, in combination with FIG. 5, the sending of the second task notification to the target DSP includes sending the second task notification to the target DSP through the interrupt distributor. That is, the scheduling DSP sends the second task notification to the target DSP through the interrupt distributor, so as to instruct the target DSP to obtain the task identifier of the target task from the second storage area corresponding to the target DSP.
[0067] In step S144, if the target DSP is the scheduling DSP itself, the task identifier is obtained from the first storage area.
[0068] If the target DSP is the scheduling DSP itself, the scheduling DSP directly obtains the task identifier of the target task from Comm 0 in the DDR, without the need to copy and store the task identifier of the target task, or send a second task notification.
[0069] After step S143 and step S144, the target DSP can parse the task identifier of the target task after obtaining the task identifier of the target task, obtain the storage address of the input data of the target task, obtain the input data of the target task according to the storage address of the input data and perform the corresponding CV calculation task, and write the storage address of the final calculation result and the storage address of the intermediate calculation result into the task identifier of the target task. In some embodiments, DMA (Direct Memory Access) and other technologies can be used to speed up the acquisition of external data of the CV DSP (i.e., input data such as image data). It should be noted that the storage address of the intermediate calculation result can be a share memory address, i.e., the intermediate calculation result can be stored in the share memory.
[0070] To clearly illustrate the DSP load balancing scheduling scheme of the embodiments of the present disclosure, the overall flow of the DSP load balancing scheduling is described in detail below. The DSP load balancing scheduling method includes the following steps.
[0071] Step 1, task distribution. The Host CPU informs the scheduling DSP of a new CV algorithm task through the interrupt distributor, and the Host CPU writes the task identifier of the new CV algorithm task into the Comm 0 area of the DDR.
[0072] Step 2, load calculation. After receiving the notification sent by the Host CPU, the scheduling DSP obtains the task identifier from the Comm 0, records the time when the new CV algorithm task is received, calculates the time interval from the last time when the task is received, and calculates the current load of all DSPs.
[0073] Step 3, task allocation. The scheduling DSP selects the DSP with the smallest current load as the target DSP for executing the new CV algorithm task, and adds the new CV algorithm task to the task queue of the target DSP. In the case that the target DSP is a calculation DSP, the scheduling DSP informs the target DSP of the new CV algorithm task through the interrupt distributor, and copies the task identifier of the new CV algorithm task to the Comm area corresponding to the target DSP in the DDR.
[0074] Step 4, task execution. The target DSP executes each task in the task queue of the target DSP in turn, wherein the task identifier of the task to be executed is obtained from the corresponding Comm area in the DDR, the input data is obtained according to the task identifier, and the calculation is performed.
[0075] Step 5, feedback calculation result. When the CV algorithm task is completed, the target DSP notifies the scheduling DSP through the interrupt distributor, adds the storage address of the calculation result to the task identifier, and copies the task identifier to the Comm 0 area. The scheduling DSP notifies the Host CPU of the completion of calculation through the interrupt distributor, and the Host CPU obtains the calculation result from the Comm 0 area.
[0076] The following simulation experiment clearly shows the effect of load balancing of the embodiments of the present disclosure. In the simulation experiment, two CV DSPs are integrated in the SOC system, and the following six test tasks are to be executed: reszie (image size transformation), normalize (image pixel normalization), hwc2hcw (hwc format conversion to hcw format), bilateral filter (bilateral filtering), canny (edge detection), and pyrup (image enlargement). The above six test tasks are high-frequency CV tasks in actual scenarios, and it is assumed that the above six tasks are simultaneously issued. The test image size is 1920*1080. First, the time consumption of the above test tasks on a single CV DSP is obtained through software testing, which is 0.52 ms, 0.42 ms, 0.1 ms, 8.63 ms, 17.2 ms, and 2.8 ms, respectively. Table 1 lists the load of each CV DSP in the traditional load balancing scheme and the load balancing scheme of the embodiments of the present disclosure.
[0077] Table 1
[0078] As can be seen from Table 1, in the traditional load balancing scheme, the total task processing time of DSP1 is 1.04 ms, the total task processing time of DSP2 is 28.63 ms, and the ratio of the total task processing time of DSP2 to DSP1 is 27.5, that is, the ratio of the total task processing time of DSP is large, which indicates that the load of DSP is unbalanced. In the load balancing scheme of the present disclosure, the total task processing time of DSP1 is 9.15 ms, the total task processing time of DSP2 is 20.52 ms, and the ratio of the total task processing time of DSP2 to DSP1 is 2.24, that is, the ratio of the total task processing time of DSP is small. Compared with the traditional load balancing scheme, the ratio of the total task processing time of DSP2 to DSP1 is reduced by more than 10 times, and the load of DSP is more balanced.
[0079] The embodiment of the present disclosure relates to a multi-instance CV DSP load balancing scheduling scheme. When multiple CV DSP instances process data in parallel or multiple tasks at the same time, the embodiment of the present disclosure can be used to schedule multiple tasks between DSPs, so as to realize load balancing between multiple CV DSP instances. The scheme has two characteristics: first, since the CV DSP is a "enhanced version" of the CPU, it also has scheduling capability, so one of the multi-instance CV DSP can be used as a scheduling DSP, which is responsible for scheduling the CV algorithm on the multi-instance CV DSP. In addition, the scheduling resources and computing resources of the DSP are independent of each other, and performing scheduling tasks will not affect its own computing resources. The remaining CV DSP can be used as a computing DSP, which is fully responsible for computing tasks. Second, since there is simulation software for testing the performance of the CV DSP, the estimated processing time of each CV task can be obtained in advance through simulation, and the sum of the estimated processing time of all tasks in the task queue of each CV DSP is used as the load of the CV DSP. The embodiment of the present disclosure transfers the scheduling task of the multi-instance CV DSP from the Host CPU to the CV DSP, which can reduce the load of the Host CPU, and considers the processing time required by the task when calculating the load. This load calculation method is more scientific and more conducive to realizing the load balancing of the multi-instance CV DSP.
[0080] The embodiment of the present disclosure can be applied to the following scenarios: intelligent vehicle and intelligent cockpit and auxiliary driving scenarios, smart phone scenarios, smart home scenarios, and other scenarios involving a large number of visual algorithm application deployment and requiring multi-instance CV DSP load balancing support.
[0081] The embodiment of the present disclosure also provides a scheduling digital signal processor, as shown in FIG. 8, which includes: at least one processor 801; a memory 802 having at least one program stored thereon, when the at least one program is executed by the at least one processor, the at least one processor implements the digital signal processing load balancing method provided by the foregoing embodiments; and at least one I / O interface 803 configured to realize information interaction between the processor and the memory.
[0082] The processor 801 is a device with data processing capability, including but not limited to a central processing unit (CPU) and the like; the memory 802 is a device with data storage capability, including but not limited to a random access memory (RAM, more specifically SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), and a flash memory (FLASH); and the I / O interface (read-write interface) 803 can realize information interaction between the processor 801 and the memory 802 and the like.
[0083] In some embodiments, the processor 801, the memory 802 and the I / O interface 803 are connected with each other through a bus, and further connected with other components of the computing device.
[0084] The embodiments of the present disclosure further provide a digital signal processing system as shown in FIG. 9, which comprises at least two computing DSPs and a scheduling DSP as described above, each computing DSP being connected with the scheduling DSP through an internal bus.
[0085] The embodiments of the present disclosure further provide a chip as shown in FIG. 1, which comprises a Host CPU, a storage device and a digital signal processing system as described above, the Host CPU, the storage device and the digital signal processing system being connected through a chip bus. The storage device can be a DDR, and the function of the digital signal processing system is as described above, which will not be repeated here.
[0086] The embodiments of the present disclosure further provide a computer readable medium having a computer program stored thereon, wherein the computer program, when executed, implements the digital signal processing load balancing method provided by the above-mentioned embodiments.
[0087] Those of ordinary skill in the art will realize and understand, all or certain steps in the methods disclosed above and the functional modules / units in the devices can be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Certain physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Furthermore, it is well known to those of ordinary skill in the art that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and can include any information delivery media.
[0088] Example embodiments have been disclosed herein and, although the use of specific terms is expressly used herein, they are intended in a generic sense only and, unless expressly stated to the contrary, are not intended to limit the application of the present disclosure. In some instances, features, characteristics or aspects described in connection with a particular embodiment can be used, alone or in combination with other embodiments, and in conjunction with the description of other embodiments, unless expressly stated to the contrary. Accordingly, one of ordinary skill in the art will recognize that the various forms and details of the application can be varied without departing from the scope of the present application as set forth in the appended claims.
Claims
1. A digital signal processing load balancing method, the method being applied to scheduling a digital signal processor (DSP), the method comprising: Receive the target task sent by the host CPU; Determine the current load of each DSP, where the load is used to represent the total processing time of the task to be processed, and the DSP includes at least a computing DSP; Determine a target DSP based on the current load of each of the DSPs; Send the target task to the target DSP.
2. The method according to claim 1, wherein Determining the current load of each DSP includes: Calculating the time interval between the target task and the previous task received; For each DSP, the current load of the DSP is determined according to the load when the DSP receives the previous task and the time interval.
3. The method according to claim 2, wherein: The determining the current load of the DSP according to the load when the DSP receives the previous task and the time interval includes: If the load of the DSP when receiving the previous task is less than or equal to the time interval, determine that the current load of the DSP is 0; In a case where the load of the DSP when receiving the previous task is greater than the time interval, the current load of the DSP is determined to be the difference between the load when the DSP receives the previous task and the time interval.
4. The method according to claim 1, wherein After delivering the target task to the target DSP, the method further includes: The current load of the target DSP is updated according to the estimated processing duration of the target task, which is sent by the Host CPU to the scheduling DSP when sending the target task.
5. The method according to any one of claims 1 to 4, wherein: The DSP also includes the scheduling DSP.
6. The method according to claim 5, wherein: The receiving target task sent by the host central processing unit Host CPU includes: Receive a first task notification sent by the Host CPU; A task identifier of the target task is obtained from a first storage area corresponding to the scheduling DSP in a storage device; the task identifier is used to instruct the target DSP to obtain input data of the target task according to the task identifier.
7. The method of claim 6, wherein: After delivering the target task to the target DSP, the method further includes: receiving a first task completion notification sent by the target DSP; the first task completion notification is sent by the target DSP after completing the target task; A second task completion notification is sent to the Host CPU; the second task completion notification is used to instruct the Host CPU to obtain the task identifier of the target task from the first storage area, so as to obtain the calculation result of the target task according to the task identifier; the task identifier of the target task is copied by the target DSP and stored in the first storage area after completing the target task.
8. The method of claim 6, wherein: The sending of the target task to the target DSP includes: Retrieving the task identifier from the first storage area, and adding the task identifier to the task queue of the target DSP; In the case where the target DSP is a computing DSP, copying the task identifier and storing the copied task identifier in a second storage area corresponding to the target DSP in the storage device; A second task notification is sent to the target DSP, where the second task notification is used to instruct the target DSP to obtain the task identifier from the second storage area.
9. The method of claim 6, wherein: The sending of the target task to the target DSP includes: Retrieving the task identifier from the first storage area, and adding the task identifier to the task queue of the target DSP; In a case where the target DSP is the scheduling DSP, the task identifier is acquired from the first storage area.
10. A scheduling digital signal processor, comprising: at least one processor; a memory storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the digital signal processing load balancing method according to any one of claims 1 to 9; At least one I / O interface is configured to implement information interaction between the processor and the memory.
11. A digital signal processing system comprising at least two computing digital signal processors and the scheduling digital signal processor according to claim 10, wherein each of the computing digital signal processors is connected to the scheduling digital signal processor via an internal bus.
12. A chip comprising a host central processing unit (CPU), a storage device, and the digital signal processing system according to claim 11, wherein the CPU, the storage device, and the digital signal processing system are connected via a chip bus.
13. A computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the digital signal processing load balancing method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Task scheduling method and device, storage medium and electronic device
CN111506398A
Method and system for performing task scheduling among computing nodes by state monitoring chip
CN115080215A
Multi-task distributed scheduling load balancing method for heterogeneous computing platform
CN115292039A
Jenkins-based load balancing method and high-availability system
CN116166432A
Cross-cluster load balancing method and device, equipment and storage medium
CN117149445A