A computing system based on DSP chip array
By designing a computing system based on DSP chip array, using hardware links and specialized boot and data sharing solutions, the existing DSP computing array has solved the problems of high hardware cost and low computing efficiency, and has achieved efficient program loading, low-cost hardware, simple production and maintenance, and convenient version management, and has broken through the computing power limit of single-core modules.
Patent Information
- Application Number
- CN202111525994.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-14
AI Technical Summary
The existing DSP computing arrays have high hardware costs, low computing efficiency, slow program loading, and inconvenient program version management.
Design a computing system based on DSP chip array, connect multiple DSP chips and array support units through hardware links, and adopt special DSP array boot solution and data sharing solution to achieve efficient program loading, low hardware cost, simple production and maintenance, and convenient version management.
It improves the computing power of module-level computing units, loads programs faster, reduces hardware costs, simplifies production and maintenance and facilitates program version management, breaks through the upper limit of computing power of single-core module computing units, and has the ability to accelerate specific algorithms.
Smart Images

Figure CN114185599B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic communication technology, and in particular to a computing system based on a DSP chip array. Background Art
[0002] DSP operation arrays are used in the design of electronic equipment in various industries. DSP programs have the characteristics of strong portability, easy cutting, low development threshold, customizability, and considerable computing power. There is a wide range of DSP operation array application requirements in many electronic equipment such as supercomputing nodes, special equipment, computing services, and military products.
[0003] However, the existing DSP computing array hardware cost is relatively high, and has the defects of low computing efficiency, slow program loading, and inconvenient program version management. Summary of the invention
[0004] The technical problem to be solved by the present invention is that the existing DSP computing array hardware cost is high, and there are problems such as low computing efficiency, slow program loading and inconvenient program version management. The purpose is to provide a computing system based on a DSP chip array to solve the above problems.
[0005] The present invention is achieved through the following technical solutions:
[0006] A computing system based on a DSP chip array, comprising a DSP chip array and an array support unit connected via a hardware link;
[0007] The DSP chip array includes (m+1)*(n+1) DSP chips; wherein m and n are both integers greater than or equal to zero;
[0008] The hardware link includes a number of computing data flow channels equal to (m+1)*(n+1) DSP chips and a number of loading and maintaining data flow channels equal to (m+1)*(n+1) DSP chips; wherein each DSP chip in the DSP chip array is connected to the array support unit through the corresponding computing data flow channel and loading and maintaining data flow channel;
[0009] The array support unit includes a loading and maintaining data flow channel for completing program burning and program booting, and a computing data flow channel for completing data sharing and algorithm acceleration between (m+1)*(n+1) DSP chips.
[0010] Furthermore, it also includes a program burning module, a program boot module, a hardware driver and a storage body; the loading and maintenance data flow channel and the program burning module, the program boot module, the hardware driver and the storage body are interconnected;
[0011] The program burning module is used to write the DSP program data to be burned into the storage body by calling the hardware driver of the storage body and operating the storage body;
[0012] The program boot module is used to reset all DSP chip signals in the DSP chip array, then read the DSP program data burned in the storage body into the RAM through the hardware driver, and then distribute the burned DSP program data to each DSP chip to complete the parallel program boot of the DSP chip array.
[0013] The present invention is composed of a chip array composed of multiple DSP chips and an array support unit. Each DSP chip and the array support unit are connected through a hardware link, and the hardware link includes two channels: a computing data flow channel and a loading and maintenance data flow channel. The computing data flow channel is used to provide an information data transmission channel for application programs and computing programs; the loading and maintenance data flow channel is used for the boot program loading channel when the DSP chip is started, the program burning data distribution channel, the DSP chip running state control, the DSP chip computing program loading and program information reading, etc. The computing data flow channel and the loading and maintenance data flow channel of each DSP chip are directly connected to the array support unit, and the data exchange between DSP chips needs to be completed through the array support unit.
[0014] Furthermore, it also includes a data sharing module and a resource locking module; the computing data flow channel is interconnected with the data sharing module and the resource locking module; the data sharing module is used for sharing DSP program data between the same DSP chip or different DSP chips in the DSP chip array.
[0015] Furthermore, the data sharing module includes a data export and a data import; the data sharing method is: the data to be shared starts from the data export of the DSP chip and passes through the high-speed port and the resource lock to the DSP chip.
[0016] Furthermore, the computing data flow channel is also interconnected with the algorithm acceleration module; the data sharing method is: the data to be shared starts from the data outlet of the DSP chip, first passes through the algorithm acceleration, and then passes through the high-speed outlet and resource lock to the DSP chip.
[0017] Furthermore, the algorithm acceleration module includes a user-defined layer and a general computing power support layer; the general computing power support layer includes multiplication and addition, convolution, pooling, sigmoid function, ReLU function and sofmax regression algorithm.
[0018] Further, the resource lock module includes an ID latch and a data selector;
[0019] When the DSP chip generates an access request, the high-speed port module parses the access request into an access request ID and a data stream; the data stream enters the FIFO buffer, and the access request ID enters the ID latch, the ID latch internally constructs an ID register, and the ID latch determines whether there is an ID in the current ID register:
[0020] If there is an ID in the current ID register, it will wait until the ID register is cleared and then latched. During the waiting process, a flow control signal is sent to the high-speed port module, and the high-speed port module then controls the request source;
[0021] If there is no ID in the current ID register, the current ID is latched and a data select signal is sent to the data selector to allow access to the data source; after the data source completes the access operation, a data operation completion signal is sent to the ID latch to clear the ID register.
[0022] Furthermore, the resource lock module also includes a FIFO memory, and the FIFO memory is arranged between the high-speed port and the data selector.
[0023] Furthermore, the data export and data import realize the sending and receiving of data through the SRIO hardware interface.
[0024] Furthermore, the storage body is a non-volatile memory device.
[0025] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0026] 1. The present invention provides a computing system based on a DSP chip array, which is composed of a chip array composed of multiple DSP chips and an array support unit. It has the ability of parallel computing of multiple DSP chips and multiple cores, and the computing power is greatly improved compared with the current single-core computing module. The design adopts a special DSP array boot solution, which has more efficient program loading and lower hardware cost by reusing the storage body. The reused memory hardware cooperates with a special DSP array burning solution, so that the design has the advantages of simpler production and maintenance, and more convenient program version management. The DSP array data sharing solution is adopted to ensure that the computing array will not cause performance degradation due to the increase of computing power cores, so that the computing array has the advantage of breaking through the computing power upper limit of the single-core module computing unit; the algorithm acceleration module provides a user-defined layer and a general computing power support layer, so that the computing array has the advantage of accelerating specific algorithms.
[0027] 2. The present invention provides a computing system based on a DSP chip array. In the module computing unit, the hardware support design is adopted to realize the reuse of a dedicated DSP array guide channel, so that program loading is more efficient and hardware costs are reduced. On the basis of the array guide channel reuse, the DSP array program burning function is realized through software support, so that production and maintenance become simple and version management is more convenient.
[0028] 3. The present invention provides a computing system based on a DSP chip array, in which all DSP chips in the array realize interconnected data sharing through the SRIO hardware interface, so that the inter-core communication overhead of each DSP chip is minimized, and the growth relationship between the number of cores and computing power is close to a linear relationship, achieving the purpose of the computing array having the ability to break through the upper limit of the computing power of a single-core module computing unit; the SRIO data sharing hardware support is implemented using programmable logic devices, so that the top level of the data interaction method can be software-defined, achieving the purpose of accelerating certain specific algorithms.
[0029] 4. The present invention provides a computing system based on a DSP chip array, which supports parallel computing, greatly improves the computing power of module-level computing units, and has faster startup and loading, which effectively reduces hardware costs; it is simple to produce and maintain, and convenient to develop, design and manage. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:
[0031] Figure 1 Design a block diagram for the system of the present invention;
[0032] Figure 2 Loading and maintaining data flow link implementation block diagram for the present invention;
[0033] Figure 3 This is a block diagram of the data sharing module of the present invention;
[0034] Figure 4 It is a block diagram of data sharing and acceleration implementation of different DSPs of the present invention;
[0035] Figure 5 This is a block diagram of data sharing and acceleration implementation of the same DSP of the present invention;
[0036] Figure 6 This is a block diagram of resource lock implementation of the present invention;
[0037] Figure 7This is a block diagram of the acceleration module implementation of the present invention. DETAILED DESCRIPTION
[0038] The present invention realizes a computing array composed of DSPs in a single module computing unit to achieve the purpose of greatly improving the computing power of the single module computing unit; in the module computing unit, a dedicated DSP array guide channel reuse is realized by adopting hardware support design to achieve the purpose of efficient program loading and lower hardware cost; on the basis of array guide channel reuse, a program burning function for the DSP array is realized by software support to achieve the purpose of simple production and maintenance and convenient version management; a unique data sharing solution is adopted: all DSP processors in the array share data through SRIO interconnection, so that the inter-core communication overhead of each processor is minimized, and the growth relationship between the number of cores and the computing power is close to a linear relationship, so that the operation array has the purpose of breaking through the upper limit of the computing power of the single-core module computing unit; data sharing hardware support is realized by programmable logic devices, so that the top layer of the data interaction method can be software-defined, so as to achieve the purpose of accelerating certain specific algorithms.
[0039] The present invention supports parallel computing, greatly improves the computing power of module-level computing units, speeds up startup loading, reduces hardware costs, simplifies production and maintenance, facilitates development and design management, and can break through the computing power limit of single-core module computing units and provide additional acceleration for specific algorithms.
[0040] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.
[0041] Example
[0042] The imported DSP solution example of the present invention is the TMS320C6678 chip of Texas Instruments, and the domestic solution example is the 6678 chip of the National University of Defense Technology. The DSP solution is not limited to the above two types of chips.
[0043] Its advantages mainly lie in:
[0044] (1) Multiple DSPs form an array to achieve parallel computing, greatly improving the computing power of the module computing unit;
[0045] (2) Design a special DSP array program boot solution to enable the computing array to have the advantages of more efficient DSP program loading and lower hardware cost;
[0046] (3) Design a special DSP array program burning solution to make the operation array easier to produce and maintain, and more convenient to manage program versions;
[0047] (4) Design a special DSP array data sharing solution to prevent the performance of the computing array from degrading due to the increase in computing cores, so that the computing array has the advantage of breaking through the computing power limit of the single-core module computing unit.
[0048] (5) The data sharing scheme is implemented using programmable logic devices, giving the computing array the advantage of accelerating specific algorithms.
[0049] like Figure 1 As shown, this embodiment uses multiple DSP chips to form a chip array, and each chip is numbered (m, n). Each DSP chip and the array support unit are connected through a hardware link, and the hardware link includes two channels: a computing data flow channel and a loading and maintenance data flow channel. The computing data flow channel is used to provide an information data transmission channel for application programs and computing programs; the loading and maintenance data flow channel is used for the boot program loading channel when the DSP chip is started, the program burning data distribution channel, the DSP chip running status control, the DSP chip computing program loading and program information reading, etc. The computing data flow channel and the loading and maintenance data flow channel of each DSP chip are directly connected to the array support unit, and the data exchange between DSP chips needs to be completed through the array support unit.
[0050] The core of this embodiment is the array support unit, which is responsible for DSP program booting, burning, data interaction and sharing between DSP chips, and acceleration of specific algorithms. The array support unit includes four functional modules: program booting module, program burning module, data sharing module, algorithm acceleration module, and other support modules: hardware driver module, resource lock module, and storage body.
[0051] like Figure 2 As shown, the loading and maintenance data flow channel and the program burning module, program booting module, hardware driver, and storage body in the array support unit are interconnected to complete the program burning function and program booting function.
[0052] First, the user sends the DSP program data that needs to be burned to the program burning module through the input interface of the program burning module. The program burning module calls the hardware driver of the storage body to operate the storage body and writes the data into the storage body. The storage body needs to use a non-volatile storage device to ensure that the data is not lost when the power is off. When powered on and loaded, the program boot module enables all DSP reset signals in the DSP array, and reads the burned DSP program data in the storage body into the RAM through the hardware driver. The RAM needs to be implemented inside the program boot module. After the program data is read, all DSP reset signals in the DSP array are released, and all DSP loading programs are enabled. At this time, because all DSP program data are in the RAM of the program boot module, all DSPs can be loaded in parallel, and the program data can be distributed to each DSP chip to complete the parallel program booting of the DSP array.
[0053] The computing data flow channel in the array support unit is interconnected with the data sharing module, algorithm acceleration module and resource lock module, and the computing data flow channel completes the functions of data sharing and algorithm acceleration. Figure 3 As shown in the figure, the data sharing module consists of a data export and a data import. It is stipulated that each DSP chip in the array has a data export and a data import to ensure that the data sharing performance is not affected when the array scale increases. Each DSP chip in the array is designed with a data export and a data import interface, and the data is sent and received through the SRIO hardware interface at the data export and import.
[0054] There are two paths for data flow in the data export part: the first path is from the DSP chip directly to the high-speed port without passing through the algorithm acceleration unit, such as Figure 4 and Figure 5 The "path 1" in this design is suitable for algorithm data sharing that does not require computing power acceleration. The second is that the data that the DSP chip needs to share is transmitted to the high-speed port after passing through the algorithm acceleration module, and then sent to other DSP chips by the high-speed port, such as Figure 4 and Figure 5 In the "Path 2" in the design, this method uses the algorithm acceleration unit on the shared channel in the design to achieve the computing power acceleration of a specific algorithm. This computing power acceleration has certain limitations and needs to meet one or more of the following conditions: data that needs to be handed over to the next computing unit for calculation, or output data of the entire computing module, or data with small data volume but large computing volume during the calculation process. These data can be additionally accelerated in the transmission path without affecting the computing performance of the overall algorithm in the computing power module.
[0055] There are also two ways to construct the computational data flow channel:
[0056] The first is the data sharing and acceleration between different DSP chips, such as Figure 4,This implementation form is mainly convenient for the algorithm calculation layering and ,modular design. When the user has multiple time-consuming algorithm modules, multiple DSP chips are ,used to bear the calculation load of multiple time-consuming algorithm modules to ,achieve the purpose of computing acceleration.
[0057] The second is the data sharing and acceleration form of the same DSP chip, such as Figure 5 In this way, a single DSP chip can send data outward, and then return to its own data entry through the algorithm acceleration module in the external transmission channel, realizing algorithm acceleration with small data volume and high computing power requirements. In practical applications, if the DSP chip is a multi-core processor, the data sharing and acceleration form of the same DSP can be used to realize multi-core data sharing and acceleration within the chip.
[0058] The resource lock module mainly completes the critical resource competition generated when multiple access request sources access the same data source at the same time. The access request source refers to any processing core in any DSP device in the DSP array, and the data source refers to the data shared by any processing core of any DSP device in the DSP array. Figure 6 As shown, Figure 6 The block diagram of resource lock implementation is shown in the figure. The access request initiated by the access request source in the array contains two pieces of information, the request source ID and data. The ID of each request source must be unique and cannot be repeated. The access request passes through the high-speed port, and the high-speed port module parses the access request into the request source ID and data. The data stream enters the FIFO buffer, and the access request ID enters the ID latch at the same time. The ID latch builds an ID register inside. The ID latch will determine whether the ID exists in the current register. If the ID exists, it will not be latched immediately, but will wait until the ID register is cleared before latching. During the waiting process, a flow control signal will be sent to the high-speed port module. The high-speed port module will then control the request source. If there is no ID occupied in the register, the current ID will be latched, and a data selection signal will be sent to the data selector to allow the requested data to access the data source. After the data source completes the access operation, it will send a data operation completion signal to the ID latch to clear the ID register. A FIFO buffer is added between the high-speed port and the data selector to ensure that the data has not reached the data selector when the ID latch sends a selection enable signal. Adding an ID latch ensures the consistency of data access and ensures that access operations to the same data source are atomic operations without being interrupted midway.
[0059] like Figure 7 As shown, Figure 7This is a block diagram of the algorithm acceleration module. The algorithm acceleration module is implemented in a programmable logic device. The upper part is open to user development. Users define the acceleration method of the upper part of the algorithm acceleration module according to their own algorithms; the lower part of the algorithm acceleration module implements general computing power support IP: multiplication and addition, convolution, pooling, sigmoid, ReLU, sofmax, etc., and provides users with custom layer calls; the bottom layer implements the internal interface of the input and output data link, which is responsible for the input and output of data and interacts with other modules in the system.
[0060] This embodiment implements a computing array composed of DSPs in a single module computing unit, which greatly improves the computing power of a single module computing unit; in the module computing unit, by adopting hardware support design, the reuse of a dedicated DSP array guide channel is realized, so that program loading is more efficient and the purpose of reducing hardware costs is achieved; on the basis of the reuse of the array guide channel, this embodiment implements the program burning function for the DSP array through software support, and the purpose of simple production and maintenance and convenient version management is achieved; this embodiment adopts a unique data sharing solution: all DSP chips in the array share data through SRIO interconnection, so that the inter-core communication overhead of each DSP chip is minimized, and the growth relationship between the number of cores and the computing power is close to a linear relationship, so that the operation array has the ability to break through the upper limit of the computing power of a single-core module computing unit; data sharing hardware support is implemented using programmable logic devices, so that the top level of the data interaction method can be software-defined, and the purpose of accelerating certain specific algorithms is achieved.
[0061] This embodiment proposes and implements a DSP operation array, implements a program loading and program burning scheme for the operation array under the DSP operation array framework, and designs a data sharing scheme specifically for the DSP operation array. The design includes hardware design and software design.
[0062] By applying this embodiment, the system completes the following Figure 1 The functions of each module of the DSP computing array shown in the figure; because the design itself is aimed at the computing array composed of DSP, the system top-level design has the ability of multiple DSPs and multiple cores to perform parallel computing, and the computing power is greatly improved compared with the current single-core computing module; the design adopts a special DSP array boot solution, which has more efficient program loading and lower hardware cost by reusing the storage body; the reused memory hardware is combined with a special DSP array burning solution, which makes the design easier to produce and maintain, and more convenient for program version management; the DSP array data sharing solution is adopted to ensure that the computing array will not suffer performance degradation due to the increase in computing power cores, so that the computing array has the advantage of breaking through the computing power limit of the single-core module computing unit; the algorithm acceleration module provides a user-defined layer and a general computing power support layer, so that the computing array has the advantage of accelerating specific algorithms.
[0063] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A computing system based on a DSP chip array, characterized in that: including a DSP chip array and an array support unit connected via a hardware link; The DSP chip array includes (m+1)*(n+1) DSP chips; wherein m and n are both integers greater than or equal to zero; The hardware link includes a number of computing data flow channels equal to (m+1)*(n+1) DSP chips and a number of loading and maintaining data flow channels equal to (m+1)*(n+1) DSP chips; wherein each DSP chip in the DSP chip array is connected to the array support unit through the corresponding computing data flow channel and loading and maintaining data flow channel; The array support unit includes a loading and maintenance data flow channel for completing program burning and program booting, and a computing data flow channel for completing data sharing and algorithm acceleration between (m+1)*(n+1) DSP chips; The computing system based on the DSP chip array further includes a data sharing module and a resource lock module; the computing data flow channel is interconnected with the data sharing module and the resource lock module; the data sharing module is used for sharing DSP program data within the same DSP chip or between different DSP chips in the DSP chip array; The resource lock module includes an ID latch and a data selector; when the DSP chip generates an access request, the high-speed port module parses the access request into an access request ID and a data stream; the data stream enters the FIFO buffer, and the access request ID enters the ID latch at the same time, the ID latch internally constructs an ID register, and the ID latch determines whether there is an ID in the current ID register: if there is an ID in the current ID register, it waits until the ID register is cleared and then latches, and sends a flow control signal to the high-speed port module during the waiting process, and the high-speed port module then controls the flow request source; if there is no ID in the current ID register, the current ID is latched, and a data selection signal is sent to the data selector to allow access to the data source; after the data source completes the access operation, it sends a data operation completion signal to the ID latch to clear the ID register.
2. A computing system based on a DSP chip array according to claim 1, characterized in that: It also includes a program burning module, a program boot module, a hardware driver and a storage body; the loading and maintenance data flow channel and the program burning module, the program boot module, the hardware driver and the storage body are interconnected; The program burning module is used to write the DSP program data to be burned into the storage body by calling the hardware driver of the storage body and operating the storage body; The program boot module is used to reset all DSP chip signals in the DSP chip array, and then read the DSP program data written in the storage body into the RAM through the hardware driver, and then distribute the written DSP program data to each DSP chip.
3. The computing system based on DSP chip array according to claim 1, characterized in that: The data sharing module includes a data export and a data import; the data sharing mode is: the data to be shared starts from the data export of the DSP chip and passes through the high-speed port and the resource lock to the DSP chip.
4. A computing system based on a DSP chip array according to claim 3, characterized in that: The computing data flow channel is also interconnected with the algorithm acceleration module; the data sharing method is: the data to be shared starts from the data outlet of the DSP chip, first passes through the algorithm acceleration, and then passes through the high-speed outlet and resource lock to the DSP chip.
5. A computing system based on a DSP chip array according to claim 4, characterized in that: The algorithm acceleration module includes a user-defined layer and a general computing power support layer; the general computing power support layer includes multiplication and addition, convolution, pooling, sigmoid function, ReLU function and sofmax regression algorithm.
6. The computing system based on DSP chip array according to claim 1, characterized in that: The resource lock module further comprises a FIFO memory, and the FIFO memory is arranged between the high-speed port and the data selector.
7. The computing system based on DSP chip array according to claim 3, characterized in that: The data export and data import realize the sending and receiving of data through the SRIO hardware interface.
8. The computing system based on DSP chip array according to claim 2, characterized in that: The storage body is a non-volatile memory device.
Citation Information
Patent Citations
Universal array signal processing plate
CN102073346A
Storage device and fetching method for multilayered cooperation and sharing in GPDSP (General-Purpose Digital Signal Processor)
CN104699631A
Cited By
Intelligent layout and connection system and method for DSP (Digital Signal Processor) module
CN121279235A
Intelligent layout and wiring system and method for DSP module
CN121279235B