Digital signal processing device and method
By designing a digital signal processing device, using the combination of the main control unit, the data reading and writing unit and the standard vector computing unit, the problem that digital signal processor architecture in the prior art is difficult to take into account both efficiency and versatility, and efficient and general digital signal processing is achieved.
Patent Information
- Application Number
- CN202510607344.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
AI Technical Summary
The existing digital signal processor architecture is difficult to take into account both efficiency and versatility. Although customized hardware circuits are efficient, they have a solid function, and the computing efficiency of software programming is low and there are safety risks.
A digital signal processing device is designed, including a main control unit, a data reading and writing unit and a scalar vector calculation unit. Through functional configuration and startup configuration, flexible data reading and writing and scalar calculation and parallel execution of vector calculation are realized.
It realizes digital signal processing that takes into account both efficiency and versatility, improves computing efficiency, reduces hardware resources and power consumption, and supports the versatility of multiple digital signal processing algorithms.
Smart Images

Figure CN120123290A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital signal processing, and particularly relates to a digital signal processing device and method. Background Art
[0002] Currently, digital signal processing is a key part of modern intelligent information processing systems. Efficient digital signal processing can effectively improve the overall system's ability to process external information inputs such as images, voices, and communication signals. Against the backdrop of the rapid development of intelligent systems, the demand for the ability to process complex digital signals in various intelligent terminals is also increasing day by day.
[0003] Digital signal processing in intelligent terminals is generally achieved through customized hardware circuits or by invoking general-purpose processors. Among them, using software programming to invoke general-purpose processors for digital signal processing has low computational efficiency, and there may even be security risks due to long-term operation. While customized hardware circuits have high signal processing efficiency and can support more efficient parallel computing and data reading and writing. However, the hardware module form of customized hardware circuits is fixed. Once an Application-Specific Integrated Circuit (ASIC) chip is formed, it is basically unchangeable, and the digital signal processing functions it provides are also difficult to expand or optimize. Even when using a Field-Programmable Gate Array (FPGA) chip or other flexibly programmable digital processing chips, the modification, upgrade, and comprehensive implementation processes of the hardware circuit will significantly affect the development and debugging efficiency of digital signal processing algorithms. On the other hand, more advanced intelligent terminals often need to perform various signal processes. For this, the terminal chip will need to have hardware circuit modules that support multiple digital signal processing algorithms, which will bring challenges in many aspects such as chip area and power consumption in actual development and application.
[0004] To address the above problems, in the research and development of intelligent terminals, a digital signal processor design that is specifically optimized based on the digital signal processing functions required by the terminal is usually carried out. However, current digital signal processor architectures all have their own limitations and are difficult to balance efficiency and generality. Summary of the Invention
[0005] Based on this, this application provides a digital signal processing device and method to achieve digital signal processing that balances efficiency and generality.
[0006] According to one aspect of the present application, a digital signal processing device is provided, including: a main control unit, configured to generate a function configuration and a startup configuration in response to a signal processing command of a central control system, where the function configuration includes a data read / write configuration and / or a calculation mode configuration; a data read / write unit, configured to be started based on the startup configuration and, according to the data read / write configuration, read input data from a first target storage space and / or output an algorithm processing result to a second target storage space; and a scalar-vector calculation unit, configured to be started based on the startup configuration and, according to the calculation mode configuration, perform a preset scalar calculation and / or vector calculation on the input data and output the calculation result as the algorithm processing result.
[0007] According to some embodiments, the scalar-vector calculation unit includes: an input-end calculation sub-unit, configured to be started based on the startup configuration and, according to the calculation mode configuration, perform a first preset calculation on the input data, select a result of the first preset calculation, and output the selected result as a preprocessing result; a shift register sub-unit, configured to be started based on the startup configuration and store the preprocessing result in a preset vector storage manner as a storage result; a multiply-accumulate sub-unit, configured to be started based on the startup configuration and perform a second preset calculation on the storage result and a target coefficient to obtain a second preset calculation result; and an output-end calculation sub-unit, configured to be started based on the startup configuration and, according to the calculation mode configuration, perform a third preset calculation on the second preset calculation result and output the calculation result as the algorithm processing result.
[0008] According to some embodiments, the scalar-vector calculation unit further includes: a coefficient configuration sub-unit, configured to determine a target coefficient according to a preset coefficient configuration; where the preset coefficient configuration includes a configuration input by the central control system, a configuration read from a preset storage space, a default coefficient configuration, and / or a coefficient configuration generated by the main control unit according to the signal processing command.
[0009] According to some embodiments, the startup configuration includes: starting the data read / write unit to read input data from the first target storage space after the data read / write configuration is completed; starting the input-end calculation sub-unit when the reading of the input data is valid; starting the shift register sub-unit when the input-end calculation sub-unit outputs a valid preprocessing result; starting the multiply-accumulate sub-unit when the stored data amount of the shift register sub-unit meets a preset condition and / or the coefficient configuration sub-unit has determined the target coefficient; starting the output-end calculation sub-unit when the multiply-accumulate sub-unit outputs a valid second preset calculation result; and / or starting the data read / write unit to output the algorithm processing result to the second target storage space when the output-end calculation sub-unit calculates the algorithm processing result.
[0010] According to some embodiments, the computing mode configuration includes the computing mode of the input - end computing sub - unit, the selection method of the output path of the input - end computing sub - unit, the computing mode of the output - end computing sub - unit, and / or the selection method of the output path of the output - end computing sub - unit.
[0011] According to some embodiments, the data read - write configuration includes the start address of reading and writing the target data, the data read - write length, and / or the data read - write mode; the target data includes input data and algorithm processing results, and the data read - write mode includes single - point read - write, continuous read - write, skip read - write, cyclic read - write, and / or repeated read - write.
[0012] According to some embodiments, the selection method of the output path includes: multiplexing, reusing computing resources, configuring operands, enabling signal control, and / or switching computing modules.
[0013] According to one aspect of the present application, a digital signal processing method includes: in response to a signal processing command of a central control system, generating a function configuration and a start configuration based on a main control unit, where the function configuration includes a data read - write configuration and / or a computing mode configuration; starting a data read - write unit based on the start configuration; based on the data read - write unit, reading input data from a first target storage space according to the data read - write configuration; starting a scalar - vector calculation unit based on the start configuration; according to the computing mode configuration, performing a preset scalar calculation and / or vector calculation on the input data based on the scalar - vector calculation unit, and outputting the calculation result as an algorithm processing result; starting a data read - write unit based on the start configuration; based on the data read - write unit, outputting the algorithm processing result to a second target storage space according to the data read - write configuration.
[0014] According to one aspect of the present application, an electronic device is provided, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0015] According to one aspect of the present application, a computer - readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method as described above is implemented.
[0016] Through the above - mentioned embodiments provided by the present application, the main control unit is connected to the central control system, performs function configuration and start configuration according to the received signal processing command, the data read - write unit is started according to the start configuration, is flexible in reading and writing and has an unlimited length, the scalar - vector calculation unit can perform scalar calculation and vector calculation according to the start configuration. For general digital signal processing algorithms, they can all be decomposed into the coupled processing of scalar operations and vector operations, with good versatility; the units can run in parallel, improving the computing efficiency. Description of the Drawings
[0017] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit the present application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without exceeding the scope of protection required by the present application.
[0019] Figure 1 It is a block diagram of the digital signal processing device provided by the embodiment of the present application; Figure 2 It is a schematic diagram of the implementation principle of the pipelined multiply-accumulate sub-unit provided by the embodiment of the present application; Figure 3 It is a detailed framework diagram of the digital signal processing device provided by the embodiment of the present application; Figure 4 It is a schematic diagram of the principles of the first pre-designed calculation and the third pre-designed calculation provided by the embodiment of the present application; Figure 5 It is a flowchart of the digital signal processing method provided by the embodiment of the present application; Figure 6 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0021] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present application. However, those of ordinary skill in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the present application.
[0022] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0023] The flowcharts shown in the drawings are only illustrative and do not necessarily include all content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0024] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the concept of this application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.
[0025] General-purpose processors generally only support basic general-purpose computing instruction sets. Therefore, it takes a long time to complete more complex digital signal processing algorithms, which will affect the intelligent information processing ability of the terminal. General-purpose processors that focus on large-scale computing, such as Graphics Processing Unit (GPU), are generally only applicable to medium and large intelligent terminals or workstations and cannot meet more flexible application requirements such as chip implementation. If a Central Processing Unit (CPU) that generally only supports serial computing is used to process digital signal processing tasks for a long time, it may affect the scheduling control of the overall terminal or chip, and there are certain security risks.
[0026] For specific implementation manners, reference may be made to the following embodiments.
[0027] Figure 1 It is a block diagram of the digital signal processing device provided by the embodiment of this application. As Figure 1 shown, the device includes a main control unit 110, a data reading and writing unit 120, and a scalar-vector calculation unit 130.
[0028] The main control unit 110 is configured to generate a function configuration and a start configuration in response to a signal processing command of a central control system.
[0029] The main control unit 110 is responsible for docking with the central control system of the chip, receiving signal processing commands from the central control system, and starting to work in response to the signal processing commands. By parsing the specific command content of the signal processing commands, function configurations and startup configurations are generated to control subsequent units to carry out algorithm processing and data reading and writing.
[0030] According to the exemplary embodiments, the receiving manners of the signal processing commands include: receiving by using a register interface, receiving through an interrupt, receiving through a bus protocol, receiving by directly connecting through a hardware circuit, etc.
[0031] Among them, the function configuration includes a data reading and writing configuration and / or a calculation mode configuration, and the startup configuration includes the startup timing of other functional units inside the device. The data reading and writing configuration includes the specific control logic for the data reading and writing unit 120 to perform data reading and writing, and the calculation mode configuration includes the calculation mode and output path of the scalar-vector calculation unit 130.
[0032] According to the exemplary embodiments, the main control unit 110 communicates with other functional units inside the device through signal communication modes such as direct connection of hardware circuit signals, register reading and writing, signal transfer, and bus protocol, so as to complete the function configuration and startup configuration.
[0033] According to the exemplary embodiments, in the digital signal processing device, the main control unit 110 and other functional units are directly connected through hardware circuit signals, and the main control unit 110 configures all other functional units in parallel through signal transmission after responding to the commands.
[0034] According to the exemplary embodiments, the central control system is usually a CPU (Central Processing Unit, central processor).
[0035] Based on the above embodiments, in one embodiment, the main control unit 110 communicates with the chip CPU in a register reading and writing manner. The main control unit 110 responds to the signal processing commands from the CPU by checking the status of the interface register, and configures other functional units inside the device based on the task number and signal storage space index in the commands.
[0036] The data reading and writing unit 120 is used to start based on the startup configuration, and read input data from the first target storage space and / or output the algorithm processing result to the second target storage space according to the data reading and writing configuration.
[0037] The data reading and writing unit 120 interfaces with the chip data storage space and is used to perform two functions: First, the data reading and writing unit 120 is responsible for sending a data reading request to the corresponding chip data storage space according to the specific storage location of the input data (i.e., the first target storage space) in the protocol supported by the chip or intelligent terminal. After obtaining the corresponding data through reading, it sends the data to the scalar-vector calculation unit 130, and controls the order and length of the read data according to the data reading and writing configuration of the main control unit 110. Second, the data reading and writing unit 120 is responsible for outputting data to the corresponding chip data storage space in the protocol supported by the chip or intelligent terminal after obtaining the effective algorithm processing result from the scalar-vector calculation unit 130, and controls the order and length of the output data according to the data reading and writing configuration of the main control unit 110, and outputs the corresponding data to the second target storage space for storage.
[0038] Among them, the input data is the signal data to be processed included in the signal processing command.
[0039] It should be noted that both the first target storage space and the second target storage space can be storage spaces such as memory, external storage, and registers, and the present application does not limit this.
[0040] The protocols supported by the chip or intelligent terminal include: the data reading and writing protocols for RAM-type memory and data based on spatial addresses, as well as other communication protocols for data storage spaces such as memory, external storage, and registers. The present application does not limit this.
[0041] Furthermore, in some embodiments, the data reading and writing unit 120 can be split into an independent data reading control unit and a data output control unit. Among them, the data reading control unit performs the first function described above, and the data output control unit performs the second function described above.
[0042] When the data reading and writing unit 120 reads or outputs data, there is no limit on the quantity of data. It can read in or output multiple signal data simultaneously, improving the digital signal processing efficiency, especially applicable to chips or intelligent terminals with limited read-write interface bit width but a large amount of data to be processed.
[0043] When the data reading and writing unit 120 reads or outputs data, there is also no limit on the size of the data. The data bit width or the number of multi-data in parallel should depend on the data reading and writing protocol supported inside the chip or intelligent terminal and the data interface width.
[0044] Furthermore, it should be emphasized that the data reading and writing unit 120 can read the input data and output the algorithm processing result simultaneously.
[0045] According to the exemplary embodiment, one 32-bit complex number is read and written at a time.
[0046] The scalar-vector calculation unit 130 is configured to be started based on a startup configuration, and perform a preset scalar calculation and / or vector calculation on input data according to a calculation mode configuration, and output the calculation result as an algorithm processing result.
[0047] The scalar-vector calculation unit 130 is used to perform specific algorithm tasks included in a signal processing command on input data. The scalar-vector calculation unit 130 is preset with multiple calculation modes, including a preset scalar calculation and a vector calculation. When processing a specific algorithm task, specific sub-modules are called for calculation according to the calculation mode configuration of the main control unit 110, and the required calculation results are selected for output to obtain an algorithm processing result.
[0048] According to an exemplary embodiment, since the data reading and writing unit 120 has no limitation on the number of data read or output at the same moment, and can read in or output multiple signal data simultaneously, the scalar-vector calculation unit 130 can also process multiple input data in parallel, outputting the algorithm processing result corresponding to each input data, improving the digital signal processing efficiency, and is particularly applicable to chips or intelligent terminals with limited read-write interface bit widths but a large amount of data to be processed.
[0049] The digital signal processing device provided in the present application is connected to a central control system through the main control unit 110, performs function configuration and startup configuration according to a received signal processing command, the data reading and writing unit 120 is started according to the startup configuration, is flexible in reading and writing and has an unlimited length, the scalar-vector calculation unit 130 can perform scalar calculation and vector calculation according to the startup configuration, and for general digital signal processing algorithms, they can all be decomposed into a coupled processing of scalar operations and vector operations, with good versatility; the units can run in parallel, improving the calculation efficiency.
[0050] The present application has low requirements for computing resources and data interface bit widths, has strong versatility support capabilities for different digital signal microprocessor design requirements, and is widely applicable to various intelligent chips or intelligent terminals to improve their digital signal processing capabilities. The digital signal processing device provided in the present application can be used as a dedicated microprocessor architecture for processing digital signals in a chip or an intelligent terminal, and at the same time, through trimming, it can be used as an independent hardware digital signal processing acceleration module, which is mounted on the main processor peripheral bus for hardware acceleration of some specific digital signal processing operations.
[0051] According to some embodiments, the scalar-vector calculation unit 130 includes an input-end calculation sub-unit 131, a shift register sub-unit 132, a multiply-accumulate sub-unit 133, and an output-end calculation sub-unit 134.
[0052] The input-end calculation sub-unit 131 is configured to be started based on a startup configuration, perform a first preset calculation on input data according to a calculation mode configuration, and select the result of the first preset calculation, and output the selected result as a preprocessing result.
[0053] The input - end computing sub - unit 131 is mainly used to process the input data read by the data reading and writing unit 120 each time.
[0054] Furthermore, in order to provide a preset signal processing capability, the input - end computing sub - unit 131 can be designed with multiple computing modes, and when processing specific algorithm tasks, it selects the required computing result according to the computing mode configuration of the main control unit 110 for output.
[0055] According to the exemplary embodiment, the data reading and writing unit 120 reads in multiple input data simultaneously. Therefore, the input - end computing sub - unit 131 can be designed with the processing functions of mapping multiple input data calculations to a single result for output, and mapping multiple input data parallel calculations to multiple results for caching and output in sequence.
[0056] It should be emphasized that the first pre - designed calculation is a scalar or complex - number calculation.
[0057] The shift - register sub - unit 132 is used to start based on the start configuration and store the pre - processed result in a preset vector - type storage manner as the storage result.
[0058] The shift - register sub - unit 132 has a shift - storage function that supports the preset register - space scale and order rules. Each time a pre - processed result from the input - end computing sub - unit 131 is received, a shift operation is performed.
[0059] The shift operation specifically includes: storing the pre - processed result (scalar) of the input - end computing sub - unit 131 as a vector in a certain order, and the length of the vector is the same as the length of the target coefficient.
[0060] In the shift operation, the data is first - in - first - out. That is, the newly input scalar is stored in the storage space with the most backward order, the existing data in the storage space is shifted into the storage space with the previous data unit in the forward order, and the data located in the storage space with the most forward order is shifted out of the shift - register space.
[0061] Furthermore, in some embodiments, at the start of a signal - processing task, the storage space of the shift - register sub - unit 132 can be reset to a fixed state.
[0062] According to the exemplary embodiment, the preset vector - type storage manner includes linear shift - storage manner, multi - dimensional storage manner, sparse storage manner, and cyclic storage manner, etc.
[0063] The multiply - accumulate sub - unit 133 is used to start based on the start configuration and perform a second pre - designed calculation on the storage result and the target coefficient to obtain a second pre - designed calculation result.
[0064] The multiply-accumulate subunit 133 performs a second pre-designed calculation on the stored results in the shift register subunit 132 and the target coefficient sequence one by one according to the aforementioned data storage order, and outputs the second pre-designed calculation result.
[0065] In some embodiments, the second pre-designed calculation includes a multiply-accumulate calculation, that is, multiply and accumulate.
[0066] Further, based on the above embodiments, in some embodiments, considering the resource overhead and timing issues of the hardware circuit, the above second pre-designed calculation can be performed in a pipeline form. The first layer of the pipeline performs a one-to-one multiplication of the stored results and the target coefficients, the second layer of the pipeline processes partial summation, and the subsequent pipeline levels sequentially complete the remaining summation.
[0067] Figure 2 It is a schematic diagram of the implementation principle of the pipelined multiply-accumulate subunit 133 in the form of a one-dimensional vector shift register. It has certain advantages in terms of implementation difficulty, resource overhead, and lightweight development, and can meet most signal processing requirements. The data storage order in this implementation mode is the most basic linear order.
[0068] It can be understood that in the specific implementation process, a two-dimensional or more complex shift storage and multiply-accumulate order rule can also be adopted, or a more complex order rule can be mapped into a linear order in a certain way, and then implemented according to Figure 2 the shown method.
[0069] Further, the second pre-designed calculation can also adopt other implementation forms, including but not limited to designing different forms of pipeline summation, one-time all calculations, etc.
[0070] The output-end calculation subunit 134 is used to be started based on the start configuration, perform a third pre-designed calculation on the second pre-designed calculation result according to the calculation mode configuration, and output the calculated result as the algorithm processing result.
[0071] The output-end calculation subunit 134 processes the second pre-designed calculation result according to the specified calculation mode and sends the processing result to the data reading and writing unit 120 for output, that is, the calculation result of the output-end calculation subunit 134 belongs to the final algorithm processing result.
[0072] Further, in order to provide preset signal processing capabilities, multiple calculation modes can be designed in the output-end calculation subunit 134, and the required calculation result is selected for output according to the calculation mode configuration of the main control unit 110 when processing specific algorithm tasks.
[0073] It should be emphasized that the third pre-designed calculation is a scalar or complex calculation.
[0074] According to the exemplary embodiment, the data reading and writing unit 120 reads in multiple input data simultaneously. Therefore, the output end calculation sub-unit 134 can have the processing function of mapping a single second pre-designed calculation result to multiple output data for parallel output, and the processing function of parallelly calculating multiple second pre-designed calculation results to map to multiple results and caching and outputting them in sequence.
[0075] According to some embodiments, the scalar-vector calculation unit 130 further includes a coefficient configuration sub-unit 135.
[0076] The coefficient configuration sub-unit 135 is used to determine the target coefficient according to the preset coefficient configuration. The coefficient configuration sub-unit 135 responds to the preset coefficient configuration and generates a coefficient sequence with the same size and order rule as the storage result of the shift register sub-unit 132 as the target coefficient.
[0077] Among them, the preset coefficient configuration includes the configuration directly input by the central control system, the configuration read from the preset storage space, the default coefficient configuration, and / or the coefficient configuration generated based on the main control unit 110 according to the signal processing command.
[0078] According to the exemplary embodiment, the preset storage space includes the chip or terminal storage space.
[0079] The scalar-vector calculation unit 130 provides a relatively wide range of scalar calculation functions by designing various calculation functions in the input end calculation sub-unit 131 and the output end calculation sub-unit 134, and expands the signal vector multiplication and accumulation calculation function through the multiplication and accumulation sub-unit 133 with freely configurable coefficients, so as to provide a relatively general and flexible signal processing algorithm support ability. In practical applications, a specific signal processing algorithm can be deconstructed into a combination form of scalar calculation and vector multiplication and accumulation, and then implemented by the digital signal processing device provided in this application.
[0080] Since the multiplication and accumulation resources for implementing different algorithm functions are multiplexed, this microprocessor architecture has a certain lightweight feature. Combined with the flexible and length-unlimited data reading and writing mode, this architecture can support the processor to complete signal input, processing and output of various types and variable data lengths with limited resource overhead, especially suitable for digital chips or intelligent terminals with limited bus read and write interface bit widths but requiring support for large data volume throughput and calculation.
[0081] According to some embodiments, the startup configuration includes at least one of Configuration 1, Configuration 2, Configuration 3, Configuration 4, Configuration 5, and Configuration 6.
[0082] Configuration 1: After the data reading and writing configuration is completed, start the data reading and writing unit 120 to read the input data from the first target storage space.
[0083] Configuration 1 specifies the timing for starting the reading operation of the data reading and writing unit 120. According to the exemplary embodiment, the first target storage space is the RAM (Random Access Memory). After the data reading and writing unit 120 obtains the specific control logic in the data reading and writing configuration (such as the starting address of data reading and the length of data reading), the memory address controller calculates the subsequent read address sequence, that is, performs continuous memory space reading starting from the starting address; and after starting the data reading, sends a request to read the data corresponding to the address to the RAM (Random Access Memory) memory in sequence through the address interface, and then after obtaining the requested data, transfers the data to the input end calculation subunit 131.
[0084] It should be emphasized that after the data reading and writing unit 120 finishes reading the specified length of data, it automatically switches to transmitting all-zero data.
[0085] Furthermore, in some embodiments, in order to ensure the smoothness of the digital signal processing device in processing digital signals, Configuration 1 can also be: after both the function configuration and the calculation mode configuration are completed, start the data reading and writing unit 120 to read the input data from the second target storage space. That is, after all function unit configurations are completed, start the data reading.
[0086] Configuration 2: When the reading of the input data is valid, start the input end calculation subunit 131.
[0087] Configuration 2 specifies the timing for starting the input end calculation subunit 131. That is, in order to ensure the reliability and efficiency of the subsequent processes, when the read data is valid, start the input end calculation subunit 131 to perform preprocessing of the digital signal data.
[0088] Configuration 3: When the output of the input end calculation subunit 131 is a valid preprocessing result, start the shift register subunit 132.
[0089] Configuration 3 specifies the timing for starting the shift register subunit 132. That is, in order to ensure the reliability and efficiency of the subsequent processes, after obtaining a valid preprocessing result, start the shift storage of the shift register subunit 132.
[0090] Configuration 4: When the stored data volume of the shift register subunit 132 meets the preset conditions and / or the coefficient configuration subunit 135 has determined the target coefficient, start the multiply-accumulate subunit 133.
[0091] Configuration 4 specifies the timing for starting the multiply-accumulate subunit 133. That is, after the shift register subunit 132 stores sufficient data and the coefficient configuration subunit 135 completes the configuration, start the multiply-accumulate subunit 133.
[0092] Determine whether the shift register sub-unit 132 stores sufficient data through preset conditions, and the preset conditions may include a data volume threshold.
[0093] Configuration 5: When the multiply-accumulate sub-unit 133 outputs a valid second preset calculation result, start the output-end calculation sub-unit 134.
[0094] Configuration 5 specifies the startup timing of the output-end calculation sub-unit 134. To ensure the reliability and efficiency of subsequent processes, after the multiply-accumulate sub-unit 133 calculates a valid second preset calculation result, start the output-end calculation sub-unit 134.
[0095] Configuration 6: When the output-end calculation sub-unit 134 calculates the algorithm processing result, start the output of the data reading and writing unit 120, and output the algorithm processing result to the second target storage space.
[0096] Configuration 6 specifies the startup timing of the output of the data reading and writing unit 120. According to the exemplary embodiment, the second target storage space is the RAM (Random Access Memory) memory. After the output-end calculation sub-unit 134 calculates the required algorithm processing result, start the data output to the RAM (Random Access Memory) memory, and save the algorithm processing result to the RAM (Random Access Memory) memory.
[0097] Furthermore, to ensure the reliability and efficiency of subsequent processes, after the output-end calculation sub-unit 134 calculates a valid algorithm processing result, start the output of the data reading and writing unit 120.
[0098] According to some embodiments, the calculation mode configuration includes the calculation mode of the input-end calculation sub-unit 131, the selection method of the output path of the input-end calculation sub-unit 131, the calculation mode of the output-end calculation sub-unit 134, and / or the selection method of the output path of the output-end calculation sub-unit 134.
[0099] Specifically, the calculation mode configuration generated by the main control unit 110 configures the calculation mode of the input-end calculation sub-unit 131, so as to perform specified preprocessing (i.e., the first preset calculation) on each input data read. The calculation mode configuration generated by the main control unit 110 also configures the multiplexing mode of the input-end calculation sub-unit 131, that is, controls which path of the preprocessed data (preprocessing result) enters the shift register sub-unit 132 as a part of the input of the multiply-accumulate sub-unit 133.
[0100] The calculation mode configuration generated by the main control unit 110 also affects the calculation mode of the calculation sub-unit 134 at the output end, so as to perform a specified secondary processing (third preset calculation) on each valid second preset calculation result. The processing result will be transmitted to the data reading and writing unit 120 as the final algorithm processing result or a part of the result, and then saved to the second target storage space. The calculation mode configuration generated by the main control unit 110 also affects the multiplexing mode of the calculation sub-unit 134 at the output end.
[0101] Furthermore, the calculation mode configuration generated by the main control unit 110 also includes setting the working mode of the coefficient configuration sub-unit 135, that is, selecting from the preset coefficient configurations.
[0102] According to the exemplary embodiment, the corresponding coefficient mode is transmitted to the coefficient configuration sub-unit 135 according to the task number of the sub-task included in the signal processing command.
[0103] According to some embodiments, the selection methods of the output path include: multiplexing, multiplexing computing resources, configuring operands, enabling signal control, and / or switching computing modules.
[0104] Multiplexing can be implemented by a multiplexer controlled by the main control unit 110, or in other forms, such as multiplexing basic computing resources, configuring operands through the central configuration of the chip or intelligent terminal, enabling or disabling certain internal computing modules through signal control, etc.
[0105] According to some embodiments, the data reading and writing configuration includes the reading and writing start address, data reading and writing length, and / or data reading and writing mode of the target data; the target data includes input data and algorithm processing results.
[0106] Specifically, the data reading and writing configuration of the main control unit 110 configures the specific control logic of the data reading and writing unit 120 for reading and writing data, including but not limited to the reading position, length, and specific reading mode of the input data, and the output position, length, and specific output mode of the algorithm processing result.
[0107] The data reading and writing modes include single-point reading and writing, continuous reading and writing, skip reading and writing, cyclic reading and writing, and / or repeated reading and writing.
[0108] The data reading and writing modes can also include other various reading and writing modes, as well as combinations with the above reading and writing modes.
[0109] According to the exemplary embodiment, if the output length is 1, it is configured as a single-point output mode, and if the output length is greater than 1, it is configured as a continuous output mode.
[0110] It should be noted that in some embodiments, after the data reading and writing unit 120 obtains the configuration of the data output start address and the data output length, the memory address controller calculates the subsequent write address sequence. If the output mode is configured as single-point output, a result data is output to the start address. If it is continuous output, a segment of result data with a length determined by the command configuration is output to the continuous memory space starting from the start address; and after the algorithm processing result output is valid, a request to write data to the corresponding address is sent to the RAM memory of the chip in sequence through the address interface, and the data to be output is synchronously transmitted to the data interface of the chip's RAM memory.
[0111] To provide a more detailed description of the digital signal processing device provided in this application, a specific embodiment is given as Figure 3 shown Figure 3 which shows the specific framework diagram of the digital signal processing device of this embodiment.
[0112] In this embodiment, the main control unit 110 communicates with the chip CPU through register reading and writing. The CPU is the central control system of the overall chip system. The data reading and writing unit 120 reads the random access memory (RAM) of the chip through the DMA method. All the signal data to be processed in the chip is stored in a continuous space in the RAM memory, and the data bit width corresponding to one address is 32 bits. The data form is a complex number with the upper 16 bits being the imaginary part (Q channel) and the lower 16 bits being the real part (I channel). The reading and writing of the RAM memory space inside the chip are realized through the address indexing method.
[0113] The main control unit 110 will respond to the signal processing command from the CPU by checking the status of the interface register, and configure other functional units inside the digital signal processing device based on the task number and signal storage space index in the command. In the processor, the main control unit 110 and other functional units are directly connected through hardware circuit signals.
[0114] The function configuration is as described above, and this embodiment will not be elaborated here. The output modes in the data reading and writing configuration include: if the output length is 1, it is the single-point output mode; if the output length is greater than 1, it is the continuous output mode. The reading and writing addresses and lengths, and the reading mode in the data reading and writing configuration are as described above, and this embodiment will not be elaborated here.
[0115] The first preset calculation of the input end calculation sub-unit 131 is the calculation of the square of the modulus of the complex signal. The calculation result of the first preset calculation and the direct input data are connected to the multiplexer configured by the main control unit 110, and after selection, the preprocessing result required for the algorithm processing task is output. The Kth valid preprocessing output result is denoted as D K .
[0116] After obtaining valid input data and being started according to the startup configuration of the main control unit 110, the input terminal calculation sub-unit 131 performs the calculation of the squared modulus of the complex signal.
[0117] Figure 4 in Figure 4 (a), Figure 4 (b) sequentially explains the schematic diagrams of the direct pass and the calculation of the squared modulus of the complex signal of the input terminal calculation sub-unit 131 in this embodiment. Among them, the direct pass calculation does not introduce hardware processing delay, and the calculation of the squared modulus of the complex signal introduces hardware processing delay.
[0118] After obtaining the valid preprocessing result input and being started according to the startup configuration of the main control unit 110, the shift register sub-unit 132 stores them sequentially in the vectorized shift storage manner as a part of the input of the multiply-accumulate sub-unit 133. The storage space size is 64, that is, 64 data can be stored simultaneously. Before D K enters the shift register sub-unit 132, the data therein is denoted as D K-1 , D K-2 , ……, D K-64 .
[0119] After obtaining the working mode (such as the coefficient mode selection signal) in the calculation mode configuration from the main control unit 110, the coefficient configuration sub-unit 135 selects a preset coefficient as the other part of the input of the multiply-accumulate sub-unit 133, and the configured target coefficient is denoted as C 64 , C 63 , ……, C 1 (in the order from the back to the front).
[0120] After being started according to the startup configuration of the main control unit 110, the multiply-accumulate sub-unit 133 starts to calculate the multiply-accumulate results of the 64 data in the shift register sub-unit 132 and the 64 coefficients in the coefficient configuration sub-unit 135 in a pipeline form. Specifically, in the first layer, there are 64 multiplication calculations, and the results are the results of multiplying 64 data and 64 coefficients in one-to-one correspondence in the order from the front to the back. D K-1 , D K-2 , ……, D K-64 and C 64 , C 63 , ……, C 1 The multiplication result of is C 64 D K-1 , C 63 D K-2 , ……, C 1 D K-64 ; in the second layer, there are 32 addition operations, and the 64 multiplication results are subjected to a parallel addition once in the mode of adding adjacent pairs pairwise. C 64 D K-1 , C63 D K-2 ,..., C 1 D K-64 The parallel addition result of D is C 64 D K-1 +C 63 D K-2 , C 62 D K-3 +C 61 D K-4 ,..., C 2 D K-63 +C 1 D K-64 ; The subsequent third to seventh layers are similar to the second layer, with 16, 8, 4, 3, and 1 addition operations respectively. The calculation result of the seventh layer is D K-1 , D K-2 ,..., D K-64 and C 64 , C 63 ,..., C 1 The multiply-accumulate result .
[0121] The third pre-designed calculation of the output end calculation sub-unit 134 includes two calculations: maximum value search and shift calculation.
[0122] After the output end calculation sub-unit 134 obtains the effective multiply-accumulate result (i.e., the second pre-designed calculation result) and is started according to the start configuration of the main control unit 110, it synchronously performs two calculations: maximum value search and shift calculation. Among them, the maximum value search is to compare and record the maximum value among all multiply-accumulate results, and the shift calculation is to shift the input data to the right by 6 bits, which is equivalent to the result of dividing the input data by 64 and rounding down. The results of the two calculations and the direct multiply-accumulate result are connected to the multiplexer configured by the main control unit 110, and after selection, the required algorithm processing result is output to the data reading and writing unit 120.
[0123] Figure 4 (a), Figure 4 (c), Figure 4 (d) sequentially illustrate the schematic diagrams of the principle of the direct connection, 6-bit right shift calculation, and maximum value search calculation of the output end calculation sub-unit 134 in the embodiment. Among them, the direct connection calculation does not introduce hardware processing delay, and the other calculations will introduce hardware processing delay.
[0124] Based on the above functional units, this embodiment supports various common algorithms in digital signal processing, such as signal difference in the general sense, signal peak search, signal filtering, and signal smoothing. Among them, signal difference is widely applicable to various digital signal processing fields and aims to calculate the difference sequence of an input data sequence; signal peak search is commonly used in fields such as communication and speech processing and aims to detect the maximum energy value of a segment of input signal; signal filtering is also widely applicable to various digital signal processing fields and aims to calculate the output of a complex signal passing through a finite impulse response filter; signal smoothing aims to calculate the smoothed filtering result of signal power with a certain window length to reduce the influence of noise. The above four algorithms are very commonly used in various current digital signal processing chips.
[0125] Furthermore, based on the above embodiment, the calculation steps of the aforementioned signal difference algorithm supported by this embodiment are as follows.
[0126] The main control unit 110 receives and detects that the task number in the interface register is 01, then configures the data reading length as the length of the data to be processed, and the output data length as the length of the data to be processed minus 1; at the same time, configures the target coefficient as the difference vector in parallel, that is, C 64 =-1, C 63 =1, and the other coefficients are 0; configures the calculation mode of the input end calculation subunit 131 as direct pass-through in parallel, and the calculation mode of the output end calculation subunit 134 as direct pass-through.
[0127] After the above function configuration is completed, start the data reading and writing unit 120 to read the RAM memory data, and start the input end calculation subunit 131 and the shift register subunit 132 simultaneously after obtaining valid data.
[0128] When the shift register length is equal to 2, start the multiply-accumulate subunit 133. After the output of the multiply-accumulate subunit 133 is valid (the pipeline delay is 7), it passes through the output end calculation subunit 134 to the data reading and writing unit 120 directly and outputs. After the output data reaches the established length, the signal difference algorithm is completed.
[0129] Furthermore, based on the above embodiment, the calculation steps of the aforementioned signal peak search algorithm supported by this embodiment are as follows.
[0130] The main control unit 110 receives and detects that the task number in the interface register is 02, then configures the data reading length as the length of the search interval in parallel, and the output data length as 1; configures the target coefficient as C 64 =1, and the other coefficients are 0; configures the calculation mode of the input end calculation subunit 131 as the calculation of the squared modulus of the complex signal in parallel, and the output end calculation subunit 134 as the maximum value search.
[0131] After the above function configuration is completed, start the data reading and writing unit 120 to read the RAM memory data, and start the input end calculation sub-unit 131 after obtaining valid data; start the shift register sub-unit 132 after the preprocessing result is valid; when the shift register length is equal to 1, start the multiply-accumulate sub-unit 133; start the output end calculation sub-unit 134 after the output of the multiply-accumulate sub-unit 133 is valid; and set the maximum value search length to be equal to the length of the interval to be searched, and output the maximum value. This maximum value result is the signal peak value required by the algorithm.
[0132] Further, based on the above embodiment, the calculation steps of the foregoing signal filtering algorithm supported by this embodiment are as follows.
[0133] The main control unit 110 receives and detects that the task number in the interface register is 03, and then configures the data reading length to be the length of the data to be calculated in parallel, and the output data length to be the length of the data to be calculated minus 63; configures the target coefficient to be the preset finite impulse response filter coefficient (the maximum length is 64) in parallel; configures the calculation mode of the input end calculation sub-unit 131 to be direct pass, and the calculation mode of the output end calculation sub-unit 134 to be direct pass.
[0134] After the above function configuration is completed, start the data reading and writing unit 120 to read the RAM memory data, and start the input end calculation sub-unit 131 and the shift register sub-unit 132 simultaneously after obtaining valid data.
[0135] When the shift register length is equal to 64, start the multiply-accumulate sub-unit 133; after the output of the multiply-accumulate sub-unit 133 is valid, directly pass through the output end calculation sub-unit 134 to the data reading and writing unit 120 and output, and the algorithm is completed after the specified length is output. The output data is the signal filtering result of the data to be processed.
[0136] Further, based on the above embodiment, the calculation steps of the foregoing signal smoothing algorithm supported by this embodiment are as follows.
[0137] The main control unit 110 receives and detects that the task number in the interface register is 04, and then configures the data reading length to be the length of the data to be calculated in parallel, and the output data length to be the length of the data to be calculated minus 63; configures the target coefficient to be a vector of all 1s in parallel; configures the calculation mode of the input end calculation sub-unit 131 to be complex signal modulus square calculation, and the calculation mode of the output end calculation sub-unit 134 to be shift.
[0138] After the above function configuration is completed, start the data reading and writing unit 120 to read the RAM memory data, and start the input end calculation sub-unit 131 and the shift register sub-unit 132 simultaneously after obtaining valid data; when the shift register length is equal to 64, start the multiply-accumulate sub-unit 133; after the output of the multiply-accumulate sub-unit 133 is valid, start the output end calculation sub-unit 134; start data output after the shift calculation result is valid, and the algorithm is completed after outputting the specified length. The output data is the signal smoothing result of the data to be processed.
[0139] The digital signal processing device provided by this application can support the required algorithms through the lightweight customization design of the input end calculation sub-unit 131 and the output end calculation sub-unit 134, that is, only by designing the customized input end calculation sub-unit 131 and the output end calculation sub-unit 134 (the calculations therein), and multiple algorithms can reuse the scalar calculation resources and multiply-accumulate resources, which can effectively control the hardware circuit scale and power consumption of the digital signal processor. It realizes general support for including but not limited to the above four digital signal processing algorithms. Compared with separately designing hardware circuits for each algorithm, this embodiment can save a large amount of hardware resources and operating power consumption, and compared with using a general-purpose processor, this embodiment can provide better computing and data throughput efficiency.
[0140] The method embodiments of this application are described below, which can be used to control the device embodiments of this application. For details not disclosed in the method embodiments of this application, reference can be made to the device embodiments of this application.
[0141] Figure 5 The flowchart of a digital signal processing method according to an exemplary embodiment is shown.
[0142] As Figure 5 shown, the method includes step S510 - step S570.
[0143] In step S510, in response to the signal processing command of the central control system, based on the main control unit, generate a function configuration and a start configuration, where the function configuration includes a data reading and writing configuration and / or a calculation mode configuration.
[0144] In step S520, start the data reading and writing unit based on the start configuration.
[0145] In step S530, based on the data reading and writing unit, read the input data from the first target storage space according to the data reading and writing configuration.
[0146] In step S540, start the scalar-vector calculation unit based on the start configuration.
[0147] In step S550, according to the computing mode configuration, based on the scalar vector calculation unit, perform a preset scalar calculation and / or vector calculation on the input data, and output the calculation result as the algorithm processing result.
[0148] In step S560, start the data reading and writing unit based on the startup configuration.
[0149] In step S570, based on the data reading and writing unit, according to the data reading and writing configuration, output the algorithm processing result to the second target storage space.
[0150] The method performs similar functions to the device provided previously. For other functions, please refer to the previous description and will not be elaborated here.
[0151] Figure 6 An electronic device according to an exemplary embodiment of the present application is shown. The following will refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present application. Figure 6 The electronic device 600 shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0152] As Figure 6 shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0153] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the methods according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 610 can execute the method as described above.
[0154] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0155] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.
[0156] The bus 630 can represent one or more of several types of bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the various bus structures.
[0157] The electronic device 600 can also communicate with one or more external devices 300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or can communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0158] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. The technical solutions according to the embodiments of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above methods according to the embodiments of the present application.
[0159] The software product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0160] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0161] The program code for performing the operations of this application may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0162] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by a device, the computer-readable medium implements the foregoing functions.
[0163] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0164] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, server, mobile terminal, or network device, etc.) to execute the method according to the embodiments of this application.
[0165] The exemplary embodiments of the present application have been specifically shown and described above. It should be understood that the present application is not limited to the detailed structures, settings or implementation methods described herein; on the contrary, the present application is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A digital signal processing device, characterized in that: include: A main control unit, for generating a function configuration and a startup configuration in response to a signal processing command of a central control system, wherein the function configuration includes a data read and write configuration and / or a computing mode configuration; a data reading and writing unit, configured to start based on the startup configuration, and read input data from a first target storage space and / or output an algorithm processing result to a second target storage space according to the data reading and writing configuration; A scalar-vector calculation unit is used to start based on the startup configuration, perform preset scalar calculation and / or vector calculation on the input data according to the calculation mode configuration, and output the calculation result as the algorithm processing result.
2. The device according to claim 1, characterized in that The scalar vector calculation unit comprises: An input end calculation subunit, configured to start based on the startup configuration, perform a first preset calculation on the input data according to the calculation mode configuration, select a result of the first preset calculation, and output the selected result as a preprocessing result; A shift register subunit, configured to start based on the startup configuration, and store the preprocessing result in a preset vector storage manner as a storage result; a multiplication-accumulation subunit, configured to start based on the startup configuration, and perform a second preset calculation on the stored result and the target coefficient to obtain a second preset calculation result; The output end calculation subunit is used to start based on the startup configuration, perform a third preset calculation on the second preset calculation result according to the calculation mode configuration, and output the calculation result as an algorithm processing result.
3. The device according to claim 2, characterized in that The scalar vector calculation unit also includes: A coefficient configuration subunit, configured to determine the target coefficient according to a preset coefficient configuration; The preset coefficient configuration includes a configuration input by the central control system, a configuration read from a preset storage space, a default coefficient configuration and / or a coefficient configuration generated by the main control unit according to the signal processing command.
4. The device according to claim 3, characterized in that The startup configuration includes: After the data read and write configuration is completed, the data read and write unit is started to read input data from the first target storage space; In case that the reading of the input data is valid, starting the input end calculation subunit; In the case where the input-end calculation subunit outputs a valid preprocessing result, starting the shift register subunit; When the amount of stored data of the shift register subunit meets a preset condition and / or the coefficient configuration subunit has determined the target coefficient, starting the multiplication and accumulation subunit; In the case where the multiplication-accumulation subunit outputs a valid second preset calculation result, starting the output-end calculation subunit; and / or When the output-end calculation subunit calculates and obtains the algorithm processing result, the data reading and writing unit is started to output the algorithm processing result to the second target storage space.
5. The device according to claim 2, characterized in that The calculation mode configuration includes the calculation mode of the input-end calculation subunit, the selection method of the output path of the input-end calculation subunit, the calculation mode of the output-end calculation subunit and / or the selection method of the output path of the output-end calculation subunit.
6. The device according to claim 1, characterized in that The data read / write configuration includes a read / write start address of target data, a data read / write length and / or a data read / write mode; The target data includes the input data and the algorithm processing result, and the data reading and writing mode includes single-point reading and writing, continuous reading and writing, skipping reading and writing, cyclic reading and writing and / or repeated reading and writing.
7. The device according to claim 5, characterized in that The output path selection method includes: multi-path selection, multiplexing of computing resources, configuring operands, signal control enabling and / or switching computing modules.
8. A digital signal processing method, characterized in that: include: In response to a signal processing command from a central control system, generating a function configuration and a startup configuration based on a main control unit, wherein the function configuration includes a data read and write configuration and / or a computing mode configuration; Starting a data reading and writing unit based on the startup configuration; Based on the data reading and writing unit, and according to the data reading and writing configuration, reading input data from a first target storage space; Starting a scalar vector calculation unit based on the startup configuration; According to the calculation mode configuration, based on the scalar-vector calculation unit, a preset scalar calculation and / or vector calculation is performed on the input data, and a calculation result is output as an algorithm processing result; Starting a data reading and writing unit based on the startup configuration; Based on the data reading and writing unit, according to the data reading and writing configuration, the algorithm processing result is output to the second target storage space.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to claim 8.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the method according to claim 8 is implemented.
Citation Information
Patent Citations
Bit granularity oriented information processing system
CN107748674A
Universal digital signal processing device, method and system
CN114063977A
Vector processor having functional unit paths of differing pipeline lengths
US5598547A