Accelerated calculation method and device of filter, equipment and storage medium

By fixed-pointing the tap coefficient of the filter and configuring the filter in the accelerator, the problem of high computational complexity of higher-order filters is solved, the calculation efficiency is improved and the application scenario is expanded.

CN119995558APending Publication Date: 2025-05-13GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411894719.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The calculation amount of the filter is large, especially in higher-order filters, the calculation complexity is significantly increased, affecting the calculation efficiency.

Method used

By obtaining the configuration information of the target filter, including the order and tap coefficient, the number of bytes used for the tap coefficient is determined and fixed-pointed. Then, a target filter is configured in the accelerator and the input data is filtered through the accelerator running filter.

Benefits of technology

The calculation efficiency of the filter is improved, the impact caused by the large amount of calculation is reduced, and the application scenario of the filter is expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995558A_ABST
    Figure CN119995558A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an accelerated calculation method and device of a filter, equipment and a storage medium, and the method comprises the steps: obtaining the configuration information of a target filter, and the configuration information comprises an order and a tap coefficient; determining the number of bytes used by the tap coefficient; according to the byte number, carrying out fixed-point processing on the tap coefficient to obtain a fixed-point processed tap coefficient; configuring the target filter in an accelerator according to the order of the target filter and the fixed-point tap coefficient; and obtaining input data, running the target filter through the accelerator to filter the input data, and obtaining output data. The filter is configured in the accelerator, so that the calculation efficiency of the filter can be improved, the influence caused by large calculation amount of the filter is reduced, and the application scene of the filter is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of filter calculation technology, and in particular to a filter acceleration calculation method, a filter acceleration calculation device, an electronic device and a computer-readable storage medium. Background Art

[0002] FIR (Finite Impulse Response) filters are widely used in many fields (such as electrocardiogram signal processing, audio processing, communication systems, etc.) due to their simple design and linear phase characteristics. Their linear phase characteristics ensure that the signal will not be phase distorted during the filtering process, which is crucial for applications that need to maintain the integrity of the signal waveform. In addition, FIR filters have the characteristics of high stability and simple design, and can be directly designed through window function methods, frequency sampling methods, etc.

[0003] However, there are also some problems in the application of filters. The output of the filter is the weighted sum of the input signals. In order to achieve higher filtering accuracy, a longer filter order is usually required. The filter has a large amount of calculation, and each output requires multiple multiplication and addition operations. Especially in high-order filters, the calculation complexity increases significantly. As the filter order increases, the amount of calculation will also increase significantly, thus affecting the calculation efficiency of the filter. Summary of the invention

[0004] In view of the above problems, embodiments of the present invention are proposed to provide a method for accelerating calculation of a filter, an apparatus for accelerating calculation of a filter, an electronic device and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems.

[0005] In order to solve the above problem, a first aspect of an embodiment of the present invention provides a method for accelerating calculation of a filter, the method comprising:

[0006] Acquire configuration information of a target filter, wherein the configuration information includes an order and a tap coefficient;

[0007] Determining the number of bytes used by the tap coefficients;

[0008] According to the number of bytes, the tap coefficients are fixed-pointed to obtain fixed-point tap coefficients;

[0009] According to the order of the target filter and the fixed-point tap coefficients, configuring the target filter in an accelerator;

[0010] Input data is acquired, and the target filter is run through the accelerator to filter the input data to obtain output data.

[0011] Optionally, filtering the input data by running the target filter through the accelerator to obtain output data includes:

[0012] The target filter is run through the accelerator, and the target data to be calculated in each round is obtained from the input data according to the order, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain the output data.

[0013] Optionally, the accelerator includes a plurality of computing units and a buffer;

[0014] The target filter is run by the accelerator, target data to be calculated in each round is obtained from the input data according to the order, the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain output data, including:

[0015] The target filter is run through multiple computing units of the accelerator, and the target data to be calculated in each round is obtained from the input data according to the order through different computing units respectively, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation results of each round and write them into the buffer; after the calculation is completed, the calculation results of each round in the buffer are added to obtain the output data.

[0016] Optionally, the step of obtaining target data to be calculated in each round from the input data according to the order by using different calculation units respectively, accumulating and multiplying the target data with the fixed-point tap coefficients respectively to obtain the calculation result of each round and writing the result into the buffer includes:

[0017] Grouping the fixed-point tap coefficients to determine multiple groups of tap coefficients;

[0018] Allocating a plurality of computing units to the plurality of groups of tap coefficients;

[0019] Through the multiple computing units respectively, a round of target data that needs to be calculated is obtained from the input data in parallel according to the order; the target data is accumulated and multiplied with a group of tap coefficients corresponding to the multiple computing units respectively to obtain the calculation results of the multiple computing units; the calculation results of the multiple computing units are added to obtain the calculation results corresponding to the target data that needs to be calculated and write them into the buffer.

[0020] Optionally, the step of fixing the tap coefficients according to the number of bytes to obtain the fixed-point tap coefficients includes:

[0021] Determine a mapping multiple according to the number of bytes;

[0022] The tap coefficients are multiplied by the mapping multiples to obtain multiplication results, and the multiplication results are rounded to obtain fixed-point tap coefficients.

[0023] Optionally, the configuration information of the target filter also includes a delay period;

[0024] The step of configuring the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients includes:

[0025] According to the order of the target filter, the fixed-point tap coefficients and the delay period, configuring the target filter in the accelerator;

[0026] The target filter is run by the accelerator, target data to be calculated in each round is obtained from the input data according to the order, the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain output data, including:

[0027] The target filter is run through the accelerator, and the target data to be calculated in each round is obtained from the input data in sequence according to the delay period and the order, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain the output data.

[0028] According to a second aspect of the present invention, there is provided a filter acceleration calculation device, the device comprising:

[0029] An information acquisition module, used to acquire configuration information of a target filter, wherein the configuration information includes an order and a tap coefficient;

[0030] A byte number determination module, used to determine the number of bytes used by the tap coefficients;

[0031] A fixed-point processing module, used for performing fixed-point processing on the tap coefficients according to the number of bytes to obtain fixed-point tap coefficients;

[0032] A filter configuration module, configured to configure the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients;

[0033] The filter operation module is used to obtain input data, and filter the input data by running the target filter through the accelerator to obtain output data.

[0034] Optionally, the filtering operation module includes:

[0035] An operation submodule is used to run the target filter through the accelerator, obtain the target data to be calculated in each round from the input data according to the order, accumulate and multiply the target data with the fixed-point tap coefficients to obtain the operation result of each round; and add the operation results of each round to obtain output data.

[0036] Optionally, the accelerator includes a plurality of computing units and a buffer;

[0037] The operator module comprises:

[0038] A target data operation unit is used to run the target filter through multiple computing units of the accelerator, obtain the target data to be calculated in each round from the input data according to the order through different computing units, accumulate and multiply the target data with the fixed-point tap coefficients to obtain the operation result of each round and write it into the buffer; after the operation is completed, add the operation results of each round in the buffer to obtain output data.

[0039] Optionally, the target data computing unit includes:

[0040] A coefficient grouping subunit, used for grouping the fixed-point tap coefficients to determine multiple groups of tap coefficients;

[0041] A coefficient allocation subunit, used for allocating a plurality of calculation units to the plurality of groups of tap coefficients;

[0042] The coefficient operation subunit is used to obtain a round of target data that needs to be calculated from the input data in parallel according to the order through the multiple calculation units respectively; accumulate and multiply the target data with a group of tap coefficients corresponding to the multiple calculation units respectively to obtain the operation results of the multiple calculation units; add the operation results of the multiple calculation units to obtain the operation results corresponding to the target data that needs to be calculated in a round and write them into the buffer.

[0043] Optionally, the fixed-point processing module includes:

[0044] A multiple determination submodule, used for determining a mapping multiple according to the number of bytes;

[0045] The coefficient fixed-point submodule is used to multiply the tap coefficient and the mapping multiple to obtain a multiplication result, and round the multiplication result to obtain a fixed-point tap coefficient.

[0046] Optionally, the configuration information of the target filter also includes a delay period;

[0047] The filter configuration module comprises:

[0048] A target filter configuration submodule, configured to configure the target filter in the accelerator according to the order of the target filter, the fixed-point tap coefficients and the delay period;

[0049] The operator module comprises:

[0050] An output data acquisition unit is used to run the target filter through the accelerator, obtain the target data to be calculated in each round from the input data according to the delay period and the order in turn, and accumulate and multiply the target data with the fixed-point tap coefficients to obtain the calculation result of each round; and add the calculation results of each round to obtain the output data.

[0051] According to a third aspect of the present invention, there is provided an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of a method for accelerating calculation of a filter as described above.

[0052] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of a method for accelerating calculation of a filter as described above.

[0053] The technical solution provided by the embodiments of the present invention may have the following beneficial effects:

[0054] The embodiment of the present invention provides a method, device, equipment and storage medium for accelerating the calculation of a filter, the method comprising: obtaining configuration information of a target filter, the configuration information comprising an order and a tap coefficient; determining the number of bytes used for the tap coefficient; according to the number of bytes, fixing the tap coefficient to obtain the fixed-point tap coefficient; configuring the target filter in an accelerator according to the order of the target filter and the fixed-point tap coefficient; obtaining input data, filtering the input data by running the target filter through the accelerator, and obtaining output data. Configuring the filter in an accelerator can improve the computational efficiency of the filter, thereby reducing the impact of the large amount of computation on the filter and expanding the application scenarios of the filter. The accelerator has a highly parallel computing architecture that can process multiple computing tasks simultaneously, thereby significantly improving computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a flowchart of the steps of an accelerated calculation method of a filter provided by an embodiment of the present invention;

[0056] Figure 2It is a flowchart of an accelerated calculation method of a filter provided by an embodiment of the present invention;

[0057] Figure 3 It is a structural block diagram of an accelerated computing device for a filter provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] The computational complexity of the filter is large, and each output requires multiple multiplication and addition operations. Especially in high-order filters, the computational complexity increases significantly. As the filter order increases, the computational complexity will also increase significantly, thus affecting the computational efficiency of the filter.

[0060] One of the core concepts of the embodiment of the present invention is to obtain the configuration information of the target filter, the configuration information includes the order and tap coefficients; determine the number of bytes used for the tap coefficients; according to the number of bytes, fix the tap coefficients to obtain the fixed-point tap coefficients; configure the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients; obtain input data, and filter the input data by running the target filter through the accelerator to obtain output data. Configuring the filter in the accelerator can improve the computational efficiency of the filter, thereby reducing the impact of the filter caused by the large amount of calculation and expanding the application scenarios of the filter. The accelerator has a highly parallel computing architecture that can process multiple computing tasks at the same time, thereby significantly improving computing efficiency.

[0061] Reference Figure 1 , shows a flowchart of a method for accelerating calculation of a filter provided by an embodiment of the present invention, and the method may specifically include the following steps:

[0062] Step 101, obtaining configuration information of a target filter, wherein the configuration information includes an order and a tap coefficient;

[0063] A filter is an electronic or digital device used to process signals. Its main function is to selectively pass or suppress certain frequency components in the signal according to the frequency. Filters are widely used in communications, audio processing, image processing, control systems, biomedical engineering and other fields. Filters are the core tools in signal processing. Their design and selection need to be weighed according to specific application requirements (such as frequency selectivity, phase characteristics, computational complexity, etc.).

[0064] The target filter of the embodiment of the present invention mainly includes FIR filter, which is not limited by the embodiment of the present invention. FIR filter is a common digital filter, which is characterized by that the output depends only on the current and past input signals, but not on the past output signals. Due to its simple structure, good stability and linear phase, FIR filter has been widely used in many fields, such as the low-pass filter in electrocardiogram ECG, which often uses FIR low-pass filter and audio system.

[0065] The order of a filter refers to the number of poles contained in the filter. Poles are the complex roots in the filter transfer function and determine the frequency response characteristics of the filter. The higher the order of the filter, the better its frequency selectivity and transition band characteristics are usually, but it will also increase the computational complexity and implementation difficulty. Low-order filters (such as first-order and second-order): The frequency response is smoother, the transition band is wider, and the stopband attenuation is slower. High-order filters (such as third-order and above): The frequency response is steeper, the transition band is narrower, and the stopband attenuation is faster.

[0066] The order of the filter is usually determined based on the following factors: If a steeper transition band or higher stopband attenuation is required, a higher order is usually required. If you are sensitive to phase distortion, you may want to choose a lower order filter. In embedded systems or real-time processing, you may need to trade off filter performance and computational complexity.

[0067] The tap coefficients of a filter are the coefficients used to weight the input signal in the filter. For digital filters (such as FIR and IIR filters), the tap coefficients determine the frequency response characteristics of the filter. The tap coefficients are the core parameters in filter design and directly affect the performance of the filter. They can be generated by tools such as MATLAB. In digital filters, the tap coefficients are usually expressed as h(n), where n is the time index. For FIR filters, the tap coefficients are the impulse response of the filter.

[0068] In this embodiment, to obtain the configuration information of the target filter, including the order and tap coefficients, it is usually necessary to determine according to the type and design method of the filter. The order of the filter is usually related to the performance of the filter. For example, if a low-pass filter is to have better frequency selectivity (that is, the amplitude-frequency curve is steeper near the passband edge), it is usually necessary to increase the order.

[0069] The order of a filter is usually equal to the number of tap coefficients of the filter minus 1. For example, if the filter has N tap coefficients, then the order is N-1. The tap coefficients of a filter are the design parameters of the filter, which are usually designed by methods such as window function method, frequency sampling method or least squares method. The tap coefficients are usually expressed as h[n], where n is the index of the tap, ranging from 0 to N-1.

[0070] Step 102, determining the number of bytes used by the tap coefficients;

[0071] In digital filters, tap coefficients are usually expressed in the form of floating-point numbers or fixed-point numbers. The number of bytes of the tap coefficients depends on the numerical representation used and the accuracy requirements of the system. The number of bytes used for common tap coefficient precision is 1, 2, and 3 bytes. 1 byte (8 bits): suitable for scenarios with small dynamic range and low accuracy requirements; 2 bytes (16 bits): suitable for most filter designs and meet common accuracy requirements; 3 bytes (24 bits): suitable for filter designs with high accuracy or high dynamic range.

[0072] In this embodiment, the number of bytes used to determine the tap coefficients needs to be determined according to the accuracy requirements of the tap coefficients. 1 byte (8 bits): low accuracy, suitable for simple filters; 2 bytes (16 bits): high accuracy, suitable for most filter designs; 3 bytes (24 bits): very high accuracy, suitable for high dynamic range or high-precision filters. For example: if the filter design has high accuracy requirements, 2 bytes (16 bits) can be selected, and if higher accuracy is required, 3 bytes (24 bits) can be selected.

[0073] Step 103, according to the number of bytes, the tap coefficients are fixed-pointed to obtain fixed-point tap coefficients;

[0074] Fixed-Point Quantization is the process of converting floating-point numbers to fixed-point numbers, which is commonly used in digital signal processing (DSP), embedded systems, and other resource-constrained computing environments. Fixed-Point Quantization can significantly reduce computational complexity and storage requirements, but there is a trade-off between accuracy and dynamic range.

[0075] Fixed-point conversion is the process of converting floating-point numbers to fixed-point numbers, which is suitable for resource-constrained computing environments. Fixed-point conversion can significantly reduce computational complexity and storage requirements, but there is a trade-off between accuracy and dynamic range. The key steps of fixed-point conversion include determining the fixed-point format, scaling floating-point numbers, quantization, and dequantization.

[0076] In this embodiment, the tap coefficients are fixed-point processed according to the determined number of bytes, and can be converted into fixed-point representation. Fixed-point is the process of converting floating-point weights into fixed-point numbers, which are usually expressed as integers and retain the precision of floating-point numbers through scaling factors. According to the accuracy requirements, a suitable scaling factor (such as 127, 32767 or 8388607) is selected, each floating-point weight is multiplied by the scaling factor, and rounded to an integer, and the fixed-point weight is stored in hardware and used in the filter calculation.

[0077] Choose an appropriate scaling factor based on the number of bytes of fixed-point precision:

[0078] 1 byte (8 bits): scaling factor is 127 (signed integer range is -128 to 127); 2 bytes (16 bits): scaling factor is 32767 (signed integer range is -32768 to 32767); 3 bytes (24 bits): scaling factor is 8388607 (signed integer range is -8388608 to 8388607).

[0079] For example, the accuracy of the tap coefficients of an FIR filter depends on the usage requirements and the value range. For example, if there is a set of numbers, the smallest of which is 0.00123, then if you want to keep three decimal places, the accuracy is 0.001, and if you want to keep four decimal places, the accuracy is 0.0001. When doing fixed-point conversion, what is actually done is a mapping of the numerical range. The implementation process is performed on the accelerator, so the mapping is to map an integer range to a decimal range, such as mapping [-128,127] to [-1,1]. This is equivalent to using integers from -128 to 127 to represent the decimals between [-1,1]. However, this representation cannot represent all decimals, because at this time it can be considered that the division value between each of my integers is 1 / 127 (equivalent to using -128 to represent the decimal -1, and 127 to represent the decimal 1). In this case, 1 / 127 can be considered as the minimum precision that can be represented by my integer range. If this range can meet the precision requirements, one byte can be used to represent the value and perform calculations. If not, the integer range is expanded. For example, [-32768,32767] is used to represent [-1,1], so that the precision becomes 1 / 32767.

[0080] In some embodiments, step 103 includes the following sub-steps:

[0081] Sub-step S11, determining a mapping multiple according to the number of bytes;

[0082] When processing data, the relationship between the number of bytes and the mapping factor is often used to determine the scaling or mapping of the data. The mapping factor can be used to adjust the range, resolution, or scale of the data to suit specific application requirements.

[0083] Bytes: Indicates the size of data, usually in bytes. For example, 8-bit data occupies 1 byte, 16-bit data occupies 2 bytes, and 32-bit data occupies 4 bytes.

[0084] Mapping multiple: represents the scaling factor of data, which is used to map data from one range to another. For example, when mapping an 8-bit data (0-255) to a 16-bit data (0-65535), the mapping multiple is 256.

[0085] In this embodiment, it is first necessary to determine the number of bytes used for the filter tap coefficient accuracy, and then determine the mapping multiple according to the number of bytes. The 1-byte mapping multiple is 127; the 2-byte mapping multiple is 32767; and the 3-byte mapping multiple is 8388607.

[0086] Sub-step S12, respectively multiplying the tap coefficients by the mapping multiples to obtain multiplication results, and rounding the multiplication results to obtain fixed-point tap coefficients.

[0087] In this embodiment, multiplying the tap coefficient by the mapping multiple and rounding the multiplication result to obtain the fixed-point tap coefficient is a key step in achieving the fixed-point of the filter.

[0088] The process of converting a floating point number or a high-precision integer to a fixed point number is usually achieved by multiplying it by a mapping multiple and rounding it. Get the tap coefficient and mapping multiple: tap coefficient (floating point number or high-precision integer), mapping multiple (scale factor calculated based on the number of bytes). Multiply the tap coefficient by the mapping multiple, round off or truncate the result, and get the fixed-point tap coefficient.

[0089] Step 104, configuring the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients;

[0090] An accelerator is a hardware or software component used to accelerate specific computing tasks. It optimizes the use of computing resources, improves computing speed and efficiency, and is widely used in high-performance computing (HPC), artificial intelligence (AI), graphics processing, signal processing and other fields. An accelerator can include multiple computing units, which can operate in parallel to improve the computational efficiency of the filter. Common accelerators include: 8-bit accelerators, 16-bit accelerators, 32-bit accelerators, etc.

[0091] In this embodiment, the order of the filter determines the complexity of the filter and the required computing resources. The higher the order, the better the performance of the filter, but the greater the amount of calculation. The tap coefficient is a key parameter of the filter and determines the frequency response of the filter.

[0092] Set the order N of the filter in the accelerator to ensure that the accelerator can process filters of order N. The fixed-point tap coefficients need to be loaded in a form suitable for accelerator processing, and these coefficients are loaded into the registers or memory of the accelerator to ensure that the accelerator can access these coefficients for calculation. According to the order of the filter and the number of tap coefficients, configure the computing unit of the accelerator to ensure that the filtering operation can be performed efficiently, and set the accelerator's parallelism, pipeline depth and other parameters to optimize the filter's computing performance. Set the bit width, data format (such as fixed-point or floating-point), clock frequency and other parameters of the input and output interfaces. After completing the above configuration, start the accelerator to start performing filtering operations.

[0093] In some embodiments, the configuration information of the target filter further includes a delay period; and step 104 includes the following sub-steps:

[0094] Sub-step S21, configuring the target filter in the accelerator according to the order of the target filter, the fixed-point tap coefficients and the delay period;

[0095] In digital signal processing (DSP) or hardware accelerators, latency usually refers to the delay time or clock cycle introduced during signal processing. Latency can occur in filters, convolution operations, pipeline processing, and other scenarios.

[0096] In this embodiment, the order N of the filter is set in the accelerator to ensure that the accelerator can process an N-order filter. The fixed-point tap coefficients need to be loaded in a form suitable for accelerator processing, and the fixed-point tap coefficients are loaded into the registers or memory of the accelerator to ensure that the accelerator can access these coefficients for calculation. The delay period is the time delay introduced when the filter processes the signal, which is usually related to the order of the filter and the hardware architecture. The delay period is set in the accelerator to ensure that the accelerator can correctly handle the delay of the signal. According to the order of the filter and the number of tap coefficients, the computing unit of the accelerator is configured to ensure that the filtering operation can be performed efficiently.

[0097] Step 105 , obtaining input data, and filtering the input data by running the target filter through the accelerator to obtain output data.

[0098] In this embodiment, after the target filter is configured and runs in the accelerator, the next step is to obtain input data and filter the input data through the accelerator to finally obtain output data.

[0099] Input data is the processing object of the filter, usually a time domain signal (such as audio signal, image signal, etc.), or pre-stored data. Get input data from data sources (such as sensors, files, networks, etc.), convert input data into a format suitable for accelerator processing (such as fixed-point or floating-point numbers), and ensure that the bit width and data format of the input data match the input interface of the accelerator.

[0100] The input data is sent to the accelerator, and the input data is filtered using the configured target filter. The accelerator is started to start the filtering operation. The accelerator performs convolution operation (the core operation of filtering) on ​​the input data according to the configured filter order and tap coefficient. During the processing, the accelerator will perform weighted summation on the input data according to the characteristics of the filter (such as linear phase, nonlinear phase, etc.).

[0101] After filtering is completed, the accelerator generates output data, which is the filtered signal.

[0102] In some embodiments, step 105 includes the following sub-steps:

[0103] Sub-step S31, running the target filter through the accelerator, obtaining the target data to be calculated in each round from the input data according to the order, accumulating and multiplying the target data with the fixed-point tap coefficients to obtain the calculation results of each round; adding the calculation results of each round to obtain output data.

[0104] Accumulated multiplication is a commonly used operation in digital signal processing (DSP), especially in filters (such as FIR filters) and convolution operations. The core of the accumulated multiplication operation is to accumulate multiple product results to obtain the final output value.

[0105] In this embodiment, when the target filter is run in the accelerator, the input data is usually processed in segments, and the target data to be calculated in each round is obtained from the input data according to the order (number of taps) of the filter. Then, the target data is cumulatively multiplied with the fixed-point tap coefficients to obtain the calculation results of each round, and finally the calculation results of all rounds are added to obtain the output data.

[0106] The input signal of the filter is usually a time domain signal sequence. The target data of the current round is obtained from the input data. The length of the target data is equal to the order of the filter. The fixed-point tap coefficients are loaded, and the target data and tap coefficients of each round are accumulated and multiplied to obtain the operation results of each round. The operation results of all rounds are added to obtain the final output data.

[0107] In some embodiments, the accelerator includes a plurality of computing units and a buffer; the step S31 includes the following sub-steps:

[0108] Sub-step S311, running the target filter through multiple computing units of the accelerator, respectively obtaining the target data to be calculated in each round from the input data according to the order through different computing units, accumulating and multiplying the target data with the fixed-point tap coefficients to obtain the calculation results of each round and writing them into the buffer; after the calculation is completed, adding the calculation results of each round in the buffer to obtain the output data.

[0109] An accelerator usually includes multiple compute units and buffers. The compute unit is the core component of the accelerator that is responsible for performing the actual computing tasks. Each compute unit can process data independently or in parallel, thereby improving the overall computing efficiency. The buffer is a storage unit used to temporarily store data, usually located between the compute unit and the memory. The buffer can reduce data access latency and improve data transmission efficiency.

[0110] In this embodiment, when the target filter is run in parallel by multiple computing units in the accelerator, the computing efficiency can be significantly improved. Each computing unit independently processes a portion of the input data and writes the operation result into a buffer. Finally, all the operation results in the buffer are added to obtain the final output data.

[0111] According to the number of computing units of the accelerator, the input data is distributed to different computing units. Each computing unit obtains the target data of the current round from the input data, and the length of the target data is equal to the order of the filter. Each computing unit performs cumulative multiplication operations on the target data and the tap coefficients to obtain the operation results of the current round. Each computing unit writes the operation results of the current round into the buffer. After all computing units complete the operation, all the operation results in the buffer are added to obtain the final output data.

[0112] Sub-step S411, running the target filter through the accelerator, obtaining the target data to be calculated in each round from the input data according to the delay period and the order in turn, accumulating and multiplying the target data with the fixed-point tap coefficients to obtain the calculation result of each round; adding the calculation results of each round to obtain the output data.

[0113] In this embodiment, when the target filter is run in the accelerator, it is necessary to obtain each round of target data from the input data in sequence according to the delay period, and perform accumulation and multiplication operations, and finally add the operation results of all rounds to obtain output data.

[0114] For example: the number of taps of the filter is 2, the tap coefficient is two 2s, the input is [1, 2, 3, 4, 5, 6], and the first output is 1*2+2*2=6, which is equivalent to multiplying the first two values ​​of the input by their corresponding tap coefficients and then adding them together. Then the second output value is the input delayed by one cycle to become [2,3] multiplied by the respective tap coefficients and then added together.

[0115] In some embodiments, step S311 includes the following sub-steps:

[0116] The fixed-point tap coefficients are grouped to determine multiple groups of tap coefficients; multiple calculation units are allocated to the multiple groups of tap coefficients; a round of target data that needs to be calculated is obtained from the input data in parallel according to the order through the multiple calculation units respectively; the target data is accumulated and multiplied with a group of tap coefficients corresponding to the multiple calculation units respectively to obtain the calculation results of the multiple calculation units; the calculation results of the multiple calculation units are added to obtain the calculation results corresponding to the target data that needs to be calculated and write them into the buffer.

[0117] Usually the number of taps may be dozens or even hundreds, so the tap coefficients can be grouped and the accelerator can calculate each group of data in parallel to improve the computational efficiency of the filter.

[0118] In this embodiment, when the filter accumulation and multiplication operations are processed in parallel by multiple computing units in the accelerator, the fixed-point tap coefficients can be grouped and a group of tap coefficients can be allocated to each computing unit. Each computing unit independently processes a part of the input data and writes the operation result into a buffer. Finally, the operation results of all computing units are added together to obtain the operation result of one round of target data.

[0119] According to the number of computing units, the tap coefficients are divided into multiple groups, a group of tap coefficients is assigned to each computing unit, and each computing unit processes a group of tap coefficients. The target data of the current round is obtained from the input data. The length of the target data is equal to the order of the filter. Each computing unit performs cumulative multiplication operations on the target data and the corresponding group of tap coefficients to obtain the operation result of the current round. The operation results of all computing units are added together to obtain the operation result of the target data of the current round, and the operation result of the current round is written into the buffer.

[0120] The embodiment of the present invention determines the number of bytes used by the tap coefficients by obtaining the order and tap coefficients of the filter; the tap coefficients are fixed-pointed according to the number of bytes, and a target filter is configured in an accelerator according to the order of the filter and the fixed-point tap coefficients. The accelerator has a highly parallel computing architecture and can process multiple computing tasks at the same time, thereby improving the computing efficiency of the filter and reducing the impact of the large amount of calculation on the filter.

[0121] Reference Figure 2 , which shows a schematic flow chart of an accelerated calculation method for a filter provided by an embodiment of the present invention,

[0122] This method first determines the filter order and tap coefficients to ensure the filter performance. To obtain the configuration information of the target filter, including the order and tap coefficients, it is usually necessary to determine it according to the type and design method of the filter. Secondly, the number of bytes used to determine the tap coefficient accuracy needs to be determined according to the accuracy requirements of the tap coefficients. 1 byte (8 bits): low accuracy, suitable for simple filters; 2 bytes (16 bits): high accuracy, suitable for most filter designs; 3 bytes (24 bits): very high accuracy, suitable for high dynamic range or high precision filters.

[0123] Then, according to the determined number of bytes, the tap coefficients are fixed-point processed and converted into fixed-point representation. According to the accuracy requirements, a suitable mapping multiple (such as 127, 32767 or 8388607) is selected, and each floating-point weight is multiplied by the mapping multiple and rounded to an integer.

[0124] Set the order of the filter in the accelerator, load the fixed-point tap coefficients into the accelerator, ensure that the accelerator can access these coefficients for calculation, configure the accelerator's computing unit according to the filter order and the number of tap coefficients, set the accelerator's parallelism, pipeline depth and other parameters to optimize the filter's computing performance, and complete the configuration of the target filter in the accelerator.

[0125] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0126] Reference Figure 3 , shows a structural block diagram of an accelerated computing device for a filter provided by an embodiment of the present invention, which may specifically include the following modules:

[0127] An information acquisition module 301 is used to acquire configuration information of a target filter, wherein the configuration information includes an order and a tap coefficient;

[0128] A byte number determination module 302, used to determine the number of bytes used by the tap coefficients;

[0129] A fixed-point processing module 303 is used to perform fixed-point processing on the tap coefficients according to the number of bytes to obtain fixed-point tap coefficients;

[0130] A filter configuration module 304, configured to configure the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients;

[0131] The filtering operation module 305 is used to obtain input data, and filter the input data by running the target filter through the accelerator to obtain output data.

[0132] In some embodiments, the filtering operation module 305 includes:

[0133] An operation submodule is used to run the target filter through the accelerator, obtain the target data to be calculated in each round from the input data according to the order, accumulate and multiply the target data with the fixed-point tap coefficients to obtain the operation result of each round; and add the operation results of each round to obtain output data.

[0134] In some embodiments, the accelerator includes a plurality of computing units and a buffer; the operation submodule includes:

[0135] A target data operation unit is used to run the target filter through multiple computing units of the accelerator, obtain the target data to be calculated in each round from the input data according to the order through different computing units, accumulate and multiply the target data with the fixed-point tap coefficients to obtain the operation result of each round and write it into the buffer; after the operation is completed, add the operation results of each round in the buffer to obtain output data.

[0136] In some embodiments, the target data computing unit includes:

[0137] A coefficient grouping subunit, used for grouping the fixed-point tap coefficients to determine multiple groups of tap coefficients;

[0138] A coefficient allocation subunit, used for allocating a plurality of calculation units to the plurality of groups of tap coefficients;

[0139] The coefficient operation subunit is used to obtain a round of target data that needs to be calculated from the input data in parallel according to the order through the multiple calculation units respectively; accumulate and multiply the target data with a group of tap coefficients corresponding to the multiple calculation units respectively to obtain the operation results of the multiple calculation units; add the operation results of the multiple calculation units to obtain the operation results corresponding to the target data that needs to be calculated in a round and write them into the buffer.

[0140] In some embodiments, the fixed-point processing module 303 includes:

[0141] A multiple determination submodule, used for determining a mapping multiple according to the number of bytes;

[0142] The coefficient fixed-point submodule is used to multiply the tap coefficient and the mapping multiple to obtain a multiplication result, and round the multiplication result to obtain a fixed-point tap coefficient.

[0143] In some embodiments, the configuration information of the target filter also includes a delay period; the filter configuration module 304 includes:

[0144] A target filter configuration submodule, configured to configure the target filter in the accelerator according to the order of the target filter, the fixed-point tap coefficients and the delay period;

[0145] The operator module comprises:

[0146] An output data acquisition unit is used to run the target filter through the accelerator, obtain the target data to be calculated in each round from the input data according to the delay period and the order in turn, and accumulate and multiply the target data with the fixed-point tap coefficients to obtain the calculation result of each round; and add the calculation results of each round to obtain the output data.

[0147] In this embodiment, the order and tap coefficient of the filter are obtained through the information acquisition module, the byte number determination module determines the number of bytes used for the tap coefficient; the fixed-point processing module fixes the tap coefficient according to the number of bytes, and the filter configuration module configures the target filter in the accelerator according to the order of the filter and the fixed-point tap coefficient, so that the filter operation module 305 operates the filter to filter the input data to obtain output data. The calculation efficiency of the filter can be improved, thereby reducing the impact of the filter caused by the large amount of calculation.

[0148] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0149] An embodiment of the present invention also provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned embodiment of the accelerated calculation method of the filter are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0150] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned embodiment of the accelerated calculation method of a filter are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0151] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0152] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0153] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0154] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0156] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0157] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0158] The above is a detailed introduction to an accelerated calculation method for a filter and an accelerated calculation device for a filter provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for accelerating calculation of a filter, characterized in that: The method comprises: Acquire configuration information of a target filter, wherein the configuration information includes an order and a tap coefficient; Determining the number of bytes used by the tap coefficients; According to the number of bytes, the tap coefficients are fixed-pointed to obtain fixed-point tap coefficients; According to the order of the target filter and the fixed-point tap coefficients, configuring the target filter in an accelerator; Input data is acquired, and the target filter is run through the accelerator to filter the input data to obtain output data.

2. The filter acceleration calculation method according to claim 1, characterized in that: The step of filtering the input data by running the target filter through the accelerator to obtain output data includes: The target filter is run through the accelerator, and the target data to be calculated in each round is obtained from the input data according to the order, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain the output data.

3. The filter acceleration calculation method according to claim 2, characterized in that: The accelerator includes a plurality of computing units and a buffer zone; The target filter is run by the accelerator, target data to be calculated in each round is obtained from the input data according to the order, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; The results of each round of operations are added together to obtain the output data, including: The target filter is run through multiple computing units of the accelerator, and the target data to be calculated in each round is obtained from the input data according to the order through different computing units respectively, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation results of each round and write them into the buffer; after the calculation is completed, the calculation results of each round in the buffer are added to obtain the output data.

4. The accelerated calculation method of the filter according to claim 3, characterized in that: The method of obtaining target data to be calculated in each round from the input data according to the order through different calculation units respectively, accumulating and multiplying the target data with the fixed-point tap coefficients respectively to obtain the calculation result of each round and writing it into the buffer includes: Grouping the fixed-point tap coefficients to determine multiple groups of tap coefficients; Allocating a plurality of computing units to the plurality of groups of tap coefficients; Through the multiple computing units respectively, a round of target data that needs to be calculated is obtained from the input data in parallel according to the order; the target data is accumulated and multiplied with a group of tap coefficients corresponding to the multiple computing units respectively to obtain the calculation results of the multiple computing units; the calculation results of the multiple computing units are added to obtain the calculation results corresponding to the target data that needs to be calculated and write them into the buffer.

5. The filter acceleration calculation method according to claim 1, characterized in that: The step of fixing the tap coefficients according to the number of bytes to obtain the fixed-point tap coefficients comprises: Determine a mapping multiple according to the number of bytes; The tap coefficients are multiplied by the mapping multiples to obtain multiplication results, and the multiplication results are rounded to obtain fixed-point tap coefficients.

6. The accelerated calculation method of the filter according to claim 2, characterized in that: The configuration information of the target filter also includes a delay period; The step of configuring the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients includes: According to the order of the target filter, the fixed-point tap coefficients and the delay period, configuring the target filter in the accelerator; The target filter is run by the accelerator, target data to be calculated in each round is obtained from the input data according to the order, the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain output data, including: The target filter is run through the accelerator, and the target data to be calculated in each round is obtained from the input data in sequence according to the delay period and the order, and the target data is accumulated and multiplied with the fixed-point tap coefficients to obtain the calculation result of each round; and the calculation results of each round are added to obtain the output data.

7. A filter acceleration calculation device, characterized in that: The device comprises: An information acquisition module, used to acquire configuration information of a target filter, wherein the configuration information includes an order and a tap coefficient; A byte number determination module, used to determine the number of bytes used by the tap coefficients; A fixed-point processing module, used for performing fixed-point processing on the tap coefficients according to the number of bytes to obtain fixed-point tap coefficients; A filter configuration module, configured to configure the target filter in the accelerator according to the order of the target filter and the fixed-point tap coefficients; The filter operation module is used to obtain input data, and filter the input data by running the target filter through the accelerator to obtain output data.

8. The filter acceleration calculation device according to claim 7, characterized in that: The filtering operation module comprises: An operation submodule is used to run the target filter through the accelerator, obtain the target data to be calculated in each round from the input data according to the order, accumulate and multiply the target data with the fixed-point tap coefficients to obtain the operation result of each round; and add the operation results of each round to obtain output data.

9. The filter acceleration calculation device according to claim 8, characterized in that: The accelerator includes a plurality of computing units and a buffer zone; The operator module comprises: A target data operation unit is used to run the target filter through multiple computing units of the accelerator, obtain the target data to be calculated in each round from the input data according to the order through different computing units, accumulate and multiply the target data with the fixed-point tap coefficients to obtain the operation result of each round and write it into the buffer; after the operation is completed, add the operation results of each round in the buffer to obtain output data.

10. The filter acceleration computing device according to claim 9, characterized in that: The target data computing unit comprises: A coefficient grouping subunit, used for grouping the fixed-point tap coefficients to determine multiple groups of tap coefficients; A coefficient allocation subunit, used for allocating a plurality of calculation units to the plurality of groups of tap coefficients; The coefficient operation subunit is used to obtain a round of target data that needs to be calculated from the input data in parallel according to the order through the multiple calculation units respectively; accumulate and multiply the target data with a group of tap coefficients corresponding to the multiple calculation units respectively to obtain the operation results of the multiple calculation units; add the operation results of the multiple calculation units to obtain the operation results corresponding to the target data that needs to be calculated in a round and write them into the buffer.

11. The filter acceleration computing device according to claim 7, characterized in that: The fixed-point processing module includes: A multiple determination submodule, used for determining a mapping multiple according to the number of bytes; The coefficient fixed-point submodule is used to multiply the tap coefficient and the mapping multiple to obtain a multiplication result, and round the multiplication result to obtain a fixed-point tap coefficient.

12. The filter acceleration calculation device according to claim 8, characterized in that: The configuration information of the target filter also includes a delay period; The filter configuration module comprises: A target filter configuration submodule, configured to configure the target filter in the accelerator according to the order of the target filter, the fixed-point tap coefficients and the delay period; The operator module comprises: An output data acquisition unit is used to run the target filter through the accelerator, obtain the target data to be calculated in each round from the input data according to the delay period and the order in turn, and accumulate and multiply the target data with the fixed-point tap coefficients to obtain the calculation result of each round; and add the calculation results of each round to obtain the output data.

13. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of a method for accelerating calculation of a filter as described in any one of claims 1 to 6 are implemented.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a method for accelerating calculation of a filter according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Hybrid precision tap bit width optimization and hardware deployment method of adaptive MIMO equalizer

    CN121217242A

  • Hybrid precision tap bit-width optimization and hardware deployment method for adaptive MIMO equalizer

    CN121217242B