Hardware acceleration circuit and method, and integrated circuit
By designing a hardware acceleration circuit and utilizing pre-stored polynomial expansion coefficients and target precision constraint values to flexibly adjust the polynomial expansion order, the problem of low efficiency of existing hardware accelerators in double-precision floating-point function operations is solved, and efficient and flexible computing capabilities are achieved.
Patent Information
- Application Number
- CN202511331913.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing hardware accelerators cannot flexibly adjust the calculation accuracy and resource allocation when performing double-precision floating-point function operations, resulting in low efficiency when facing different tasks, and there are problems such as high calculation delay, high energy consumption, and insufficient parallelism.
A hardware acceleration circuit is designed, including an interface circuit, a control circuit, and an operation circuit. By pre-storing polynomial expansion coefficients and flexibly adjusting the polynomial expansion order in combination with the target precision constraint value, it supports multiple double-precision floating-point function operation types and reduces the number of memory accesses and operation delays through buffers and parallel processing.
It improves the efficiency of double-precision floating-point function operations, adapts to different precision requirements and operation types, reduces computing latency and energy consumption, and improves the applicability and performance of hardware accelerators.
Smart Images

Figure CN120832121A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiments of the present disclosure relate to the technical field of integrated circuits, and particularly, to a hardware acceleration circuit, method, and integrated circuit. BACKGROUND
[0002] In the fields of navigation algorithms, scientific simulations, and engineering analysis, a large number of computing tasks involving double-precision floating-point function operations need to be completed within a limited time. Double-precision floating-point numbers have high representation precision and calculation accuracy, but their complex operation process often requires a large amount of computing resources and time. Although general-purpose processors or graphics processing units can implement such operations through instruction sets, in the face of massive data and high real-time requirement application scenarios, there may still be problems such as high operation delay, high energy consumption, and insufficient parallelism. Therefore, the efficiency of double-precision floating-point function operations is worth attention. SUMMARY
[0003] Therefore, the embodiments of the present disclosure provide a hardware acceleration circuit, method, and integrated circuit to improve the efficiency of double-precision floating-point function operations. In a first aspect, a hardware acceleration circuit is provided, comprising: an interface circuit, a control circuit, and an operation circuit; the interface circuit is configured to read first data and a first parameter, and output second data and a second parameter to the outside, and is further configured to read third data, wherein the first data is double-precision floating-point input data, the second data is a double-precision floating-point calculation result, the first parameter includes a target precision constraint value, and the third data includes coefficients of polynomial expansion in a double-precision floating-point function operation; the control circuit is configured to read the first parameter through the interface circuit, instruct the interface circuit to read the first data and the third data according to the first parameter, enable the operation circuit to perform operations, and instruct the interface circuit to output the second data and the second parameter; and at least one operation circuit is configured to perform a double-precision floating-point function operation on the first data based on the first parameter and the third data to obtain the second data and the second parameter; wherein the double-precision floating-point function operation includes one or all of transcendental operations based on polynomial expansion and format processing operations, and the order of polynomial expansion is determined based on the target precision constraint value.
[0004] The above hardware acceleration circuit, by cooperatively designing the interface circuit, the control circuit, and the operation circuit, and combining the pre-stored polynomial expansion coefficient set, can flexibly adjust the order of polynomial expansion according to the target precision constraint value to balance the operation speed and precision, support multiple types of double-precision floating-point function operations, and effectively improve the efficiency of double-precision floating-point function operations while reducing the number of memory accesses and operation delays.
[0005] Optionally, the interface circuit comprises: a main bus interface module configured to read the first data and / or the third data and transmit the second data to the outside; and a slave bus interface module configured to read the first parameter and transmit the second parameter to the outside.
[0006] Optionally, the system further comprises: a buffer configured to buffer the first data read by the interface circuit and configured to buffer the second data generated by the operation circuit; the buffer comprises: a first storage unit and a second storage unit configured to buffer a first part and a second part of the first data and a first part and a second part of the second data respectively, the first part and the second part being divided based on a buffer segmentation manner; the first storage unit and the second storage unit are configured to provide the first data to the operation circuit and / or output the second data to the interface circuit when the operation circuit performs the double-precision floating-point function operation.
[0007] Optionally, the system further comprises: a register circuit configured to buffer the first parameter received by the interface circuit and buffer the second parameter written by the control circuit.
[0008] Optionally, the operation circuit comprises: a multiplication operation module configured to calculate a product of each order power term of a polynomial expansion of the first data based on the first parameter and the third data; and an addition operation module configured to perform an addition operation based on a result of the product of each order power term to determine a partial sum or a final sum of the polynomial expansion.
[0009] Optionally, the operation circuit further comprises: a preprocessing module configured to determine a calculation range of the first data before the polynomial expansion; and a result post-processing module configured to perform a format processing type operation on the partial sum or the final sum.
[0010] Optionally, the control circuit is further configured to determine a target operation flow of the operation circuit according to the first parameter, and control the multiplication operation module, the addition operation module, the preprocessing module and the result post-processing module to perform at least one operation link of the double-precision floating-point function operation in a preset operation sequence, the operation link comprising: operation interval determination, multiplication operation, addition operation and negation operation.
[0011] Optionally, the first parameter further comprises at least one of: an operation type, an operation item number and an input data number; the second parameter comprises at least one of: result overflow, result underflow, illegal operation and result inaccuracy; the transcendental operation comprises at least one of: a trigonometric function operation and a power function operation; and the format processing type operation comprises at least one of: a rounding operation, an integer operation and an absolute value operation.
[0012] In a second aspect, a hardware acceleration method is provided for performing a double-precision floating-point function operation by the hardware acceleration circuit provided in the first aspect, comprising: controlling the interface circuit to read first data, first parameters and third data, wherein the first data is double-precision floating-point input data, the first parameters include a target precision constraint value, and the third data includes coefficients of polynomial expansion in the double-precision floating-point function operation; enabling the operation circuit to perform a double-precision floating-point function operation on the first data based on the first parameters and the third data to determine second data and second parameters, the double-precision floating-point function operation including one or all of: a transcendental operation based on polynomial expansion or a format processing operation, wherein the order of the polynomial expansion is determined based on the target precision constraint value, and the second data is a double-precision floating-point calculation result; and outputting the second data and the second parameters to the outside.
[0013] In a third aspect, an integrated circuit is provided, comprising: a memory configured to pre-store third data, the third data including coefficients of polynomial expansion in a double-precision floating-point function operation; and the hardware acceleration circuit provided in the first aspect, configured to read first data, first parameters and the third data, perform a double-precision floating-point function operation on the first data based on the first parameters and the third data to determine second data and second parameters, the double-precision floating-point function operation including one or all of: a transcendental operation based on polynomial expansion or a format processing operation, wherein the order of the polynomial expansion is determined based on the target precision constraint value. BRIEF DESCRIPTION OF DRAWINGS
[0014] The drawings used in the description of the embodiments of the present disclosure are briefly described as follows: Figure 1 A structural schematic diagram of a hardware acceleration circuit provided in some embodiments of the present application is shown; Figure 2 A structural schematic diagram of an operation circuit provided in some embodiments of the present application is shown; Figure 3 A signal flow conversion schematic diagram of a hardware acceleration circuit provided in some embodiments of the present application is shown; Figure 4 A state conversion logic schematic diagram of an operation circuit provided in some embodiments of the present application is shown; Figure 5 A flow schematic diagram of a hardware acceleration method provided in some embodiments of the present application is shown. DETAILED DESCRIPTION
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will describe the embodiments of the present disclosure with reference to the drawings. The drawings in the following description are only some embodiments of the present disclosure, and for those skilled in the art, other drawings can be obtained from these drawings without creative labor, and other embodiments can be obtained, and the adjustments and improvements made without departing from the concept of the present disclosure are within the protection scope of the present disclosure.
[0016] In order to make the drawings simple, each drawing only schematically shows the part related to the embodiments, and it does not represent the actual structure of the product. In addition, in order to make the drawings simple and easy to understand, in some drawings, only some parts with the same structure or function are schematically shown, and there can be more or less parts with the same structure or function.
[0017] In the present disclosure, unless otherwise explicitly specified and limited, ordinal words such as "first", "second", etc. are only used to distinguish the description of the associated objects, and cannot be understood as indicating or implying the relative importance or order between the associated objects; in addition, it also does not represent the number of the associated objects. "Multiple" includes two or more, and other quantifiers are similar. " / " is used to describe the relationship between the associated objects, which represents the "or" relationship between the associated objects. "And / or" is used to describe the relationship between the associated objects, which includes any combination relationship between the associated objects, for example, "a and / or b" includes: "a alone", "b alone", or "a and b". "One or more" or "at least one" of multiple objects means any object or any combination of multiple objects, for example, "one or more of a1, a2, a3" or "at least one of a1, a2, a3" includes: "a1 alone", "a2 alone", "a3 alone", "a1 and a2", "a1 and a3", "a2 and a3", or "a1, a2 and a3".
[0018] Double-precision floating-point function operations are widely used in numerical calculation and scientific engineering fields, mainly involving transcendental operations such as trigonometric functions, exponential functions, power functions, and logarithmic functions, as well as format processing operations such as rounding, rounding, and absolute value. In implementation, double-precision floating-point transcendental operations usually rely on polynomial expansion methods (such as Taylor expansion, Chebyshev expansion, etc.), by calculating the power of each order of the polynomial and performing addition accumulation to obtain the approximate value. The higher the order of polynomial expansion, the higher the accuracy of the calculation result, but the required computing resources and time also increase. In different application scenarios, the requirement for operation accuracy is not the same, for example, navigation calculation may require high order to ensure minimal error, while some real-time control applications are more concerned about calculation speed than high accuracy. The existing hardware accelerator scheme often uses fixed polynomial order or fixed operation path when implementing double-precision floating-point function operations, resulting in the inability to flexibly adjust the operation accuracy and resource allocation when facing different tasks. In addition, the flow path of data inside the accelerator, the scheduling order of the operation module, and the organization method of the data cache also directly affect the overall calculation delay and energy efficiency. If the configurable and parallel management of the operation flow cannot be realized at the hardware level, it will be difficult to flexibly switch between different precision requirements and different operation types, thereby limiting the applicable range and performance potential of the accelerator. In view of this, the present application proposes a hardware acceleration circuit, method, and integrated circuit, which flexibly configures the polynomial expansion order based on the operation requirement, thereby balancing the calculation accuracy and efficiency, and supporting multiple double-precision floating-point function operation types, with efficient data reading, caching, and parallel processing capabilities, reducing memory access delay, and improving the operation efficiency of double-precision floating-point functions.
[0019] Figure 1A structural diagram of a hardware acceleration circuit provided in some embodiments of the present application is shown. The hardware acceleration circuit 100 is used for double-precision floating-point function operation, and includes an interface circuit 110, a control circuit 120, and an operation circuit 130. The interface circuit 110 is configured to read first data and a first parameter, and output second data and a second parameter to the outside, and is further configured to read third data. The first data is double-precision floating-point input data, the second data is a double-precision floating-point calculation result, the first parameter includes a target precision constraint value, and the third data includes coefficients used for polynomial expansion in double-precision floating-point function operation. The control circuit 120 is configured to read the first parameter through the interface circuit 110, instruct the interface circuit to read the first data and the third data according to the first parameter, enable the operation circuit 130 to perform operation, and instruct the interface circuit 110 to output the second data and the second parameter to the outside. The at least one operation circuit 130 is configured to perform double-precision floating-point function operation on the first data based on the first parameter and the third data, to obtain the second data and the second parameter. The double-precision floating-point function operation includes one or all of transcendental operation based on polynomial expansion and format processing operation. The order of polynomial expansion is determined based on the target precision constraint value.
[0020] The above embodiments of the hardware acceleration circuit can be used for double-precision floating-point function operation. The hardware acceleration circuit at least includes the interface circuit 110, the control circuit 120, and the at least one operation circuit 130. The first data is double-precision floating-point input data, and the second data is a calculation result of the first data after double-precision floating-point calculation. The first data and the first parameter can be stored in a memory outside the hardware acceleration circuit, such as the memory 10 or other memories. Therefore, the first data, the first parameter, and the third data can be stored in the same or different memories, which is not limited herein. Taking the third data stored in the memory 10 as an example, the third data can be stored in advance after the deployment of power-on initialization. The third data is a coefficient set used for polynomial expansion in double-precision floating-point function operation. For example, the third data can be a Taylor expansion coefficient set {1, -1 / 3!, 1 / 5!, -1 / 7!, …} of a sine function sin(x), or an exponential function exp(x) coefficient set {1, 1, 1 / 2!, 1 / 3!, 1 / 4!, …} the Taylor expansion coefficient set {1, 1 / 1!, 1 / 2!, 1 / 3!, …} of the Taylor expansion or the polynomial expansion coefficient set {1, -1 / 2, 1 / 3, -1 / 4, …} of the natural logarithm function ln(1+x). These coefficients can be stored in a double-precision floating-point format and can correspond to different expansion orders (such as 4th order, 6th order, 8th order, etc.), and the range of orders required to be called can be indicated by the target accuracy constraint value in the first parameter. By pre-storing the third data in the memory 10, dynamic calculation of factorials or coefficients during operation can be avoided, thereby reducing delay and reducing hardware overhead, and from the above coefficient set, based on the mathematical rules of Taylor expansion, although the functions of the expansion objects are different, there are a large number of repeated coefficients, and the same coefficients can not be stored repeatedly, so the third data does not excessively occupy the storage space of the memory 10.
[0021] The interface circuit 110 is configured to, under the instruction of the control circuit 120, first acquire the first parameter related to the calculation, for example, the first parameter can include the address of the first data to be read, the bit width, the operation logic to be performed (such as trigonometric function, exponential function), and the like, and accordingly the corresponding first data and third data can be read from the memory 10, and the second data obtained by the operation and the second parameter associated with the result can be output to the outside after the operation is completed. The second parameter can be used to indicate the state of the operation result (i.e., the second data). For example, the second parameter can include a state identifier for describing the operation state, such as result overflow, result underflow, illegal operation, or inaccurate result. Outputting the second data and the second parameter to the outside not only outputs the operation result, but also determines the precision, accuracy, and the like of the second data based on the second parameter. The reading of the first data and the third data can be synchronous or asynchronous, which is not limited here. The control circuit 120 responds to the first parameter (for example, the first parameter can be used to at least represent the operation type and the target accuracy constraint value), and sends control instructions for reading and outputting to the interface circuit 110, and enables the operation circuit 130 to perform operation. When the first parameter indicates to perform a transcendental operation, the control circuit 120 enables the operation circuit 130 to process the first data according to the polynomial expansion process, and the operation circuit 130 performs polynomial expansion calculation on the first data based on the polynomial coefficients provided by the third data, and the order of expansion is determined by the target accuracy constraint value in the first parameter (for example, the corresponding order is selected to complete the calculation under the condition of meeting the target accuracy), and the target accuracy constraint value can be determined before calculation. For example, when the hardware acceleration circuit is applied to the field of navigation positioning technology, the positioning accuracy of the navigation system can be flexibly adjusted by setting the target accuracy constraint value. For a high-precision requirement scenario, the target accuracy constraint value can be increased to make the double-precision floating-point function calculation have more expansion terms and obtain higher navigation accuracy. For a low-precision requirement scenario, the target accuracy constraint value is reduced, and the operation amount is correspondingly reduced.
[0022] Take the sine function, sin x, or "sin(x)" as an example. The hardware computation of this function is approximated by a Taylor series expansion, such as shown in Equation 1:
[0023] The more terms in the expansion, the more accurate the result, but the more computation is required. The number of terms in the series expansion can be configured by a target precision constraint, which is equivalent to telling the computing unit in hardware to expand to the nth term. For example, a target precision constraint of 7 indicates that the expansion is to be to the 7th term, x 15 / 15!, which has an error of about 7.65 x 10 -13 -4. A target precision constraint of 15 indicates that the expansion is to be to the 15th term, x 31 / 31!, which has an error of about 1.22 x 10 -34Accordingly, the calculation precision can be dynamically controlled as needed. When the first parameter indicates that a format processing type operation is to be performed, the control circuit 120 enables the operation circuit 130 to directly perform corresponding format processing on the first data according to the flow of format processing to obtain the second data, without calling the polynomial coefficients. For example, in the case of an integer function, the integer processing can be implemented by directly truncating or adjusting the decimal part of the floating-point number, or the correct integer processing can be performed on positive numbers and negative numbers. When processing an absolute value function, the sign bit of the floating-point number can be directly cleared (i.e., the positive number remains unchanged, and the negative number becomes positive), that is, the operation is implemented through bit operation or code shifting. The above two operation modes can exist at the same time, and no specific limitation is made herein. The output of the operation circuit 130 is provided to the interface circuit 110 as the second data, and the interface circuit 110 outputs the second data together with the second parameter derived from the result to the outside. The second parameter can at least reflect the state information related to the operation result. In an embodiment, the operation circuit 130 can be implemented in the form of a single operation unit, which is sequentially switched between different operation types by the control circuit 120 to complete different tasks. In another embodiment, the operation circuit 130 can include a plurality of parallelly configured operation units, and one or more operation units are selected by the control circuit 120 to participate in the current task according to one or more first parameters, so as to adapt to different throughput requirements. In addition, in some embodiments, the memory 10 can also be configured as part of the hardware acceleration circuit 100, so that the memory 10 is used as a dedicated memory for the coefficient set of the polynomial expansion in the double-precision floating-point function operation, and the coefficients are called as needed. No specific limitation is made herein. Through the above structure and flow, the pre-stored third data reduces the burden of coefficient generation or loading during operation, and the division of labor between the interface circuit 110 and the control circuit 120 makes the responsibilities of data transfer and operation triggering clear. The target precision constraint value can be used to determine the polynomial expansion order, which helps to make a controllable trade-off between precision and calculation overhead. The unified output of the result and the parameter can improve the adaptability of the device in different application scenarios and reduce unnecessary calculation and transmission overhead. The hardware acceleration circuit of the above embodiment can flexibly adjust the polynomial expansion order according to the target precision constraint value to balance the operation speed and precision while reducing the number of memory accesses and operation delays, support multiple double-precision floating-point function operation types, and thus improve the overall operation efficiency and applicability.
[0024] With reference to Figure 1 , the interface circuit 110 includes a main bus interface module 111 for reading the first data and / or the third data, and transmitting the second data to the outside; and a slave bus interface module 112 for reading the first parameter, and transmitting the second parameter to the outside.
[0025] Specifically, the master bus interface module 111 is configured to read the first data or the third data from the memory 10 under the scheduling of the control circuit 120, and transmit the second data to the external device through the master bus after the double-precision floating-point function operation is completed, so as to realize high-speed transmission of large data; the slave bus interface module 112 is configured to read the first parameter under the scheduling of the control circuit 120, and transmit the second parameter to the external device after the operation is completed, so as to realize fine transmission and control of different types of parameters.
[0026] In one specific example, the master bus interface module 111 can be implemented through a master interface (Master) of an Advanced High-performance Bus (AHB), and the slave bus interface module 112 can be implemented through a slave interface (Slave) of the AHB.
[0027] Through the design of the double-bus interface, the master bus interface module 111 focuses on high-throughput transmission of numerical data, and the slave bus interface module 112 focuses on fast access and update of operation parameters, so that the data flow and the control flow can be effectively separated at the hardware level, the delay caused by bus occupation conflict is avoided, and the overall data interaction efficiency and system response speed are improved.
[0028] With reference to the foregoing description, the interface circuit 110 is configured to read the first data from the memory 10 under the scheduling of the control circuit 120, and transmit the first data to the operation circuit 130 through the master bus interface module 111; the operation circuit 130 is configured to perform the double-precision floating-point function operation on the first data under the scheduling of the control circuit 120, and generate the second data; the operation circuit 130 is configured to transmit the second data to the interface circuit 110 through the slave bus interface module 112. Figure 1 Further comprising: a buffer 140 configured to buffer the first data read by the interface circuit 110, and configured to buffer the second data generated by the operation circuit 130; the buffer 140 comprises: a first storage unit and a second storage unit, configured to buffer a first part and a second part of the first data respectively, and configured to buffer a first part and a second part of the second data respectively, the first part and the second part being divided based on a buffer segmentation manner; the first storage unit and the second storage unit are configured to provide the first data to the operation circuit, or output the second data to the interface circuit, when the operation circuit performs the double-precision floating-point function operation.
[0029] In the above embodiment of the hardware acceleration circuit, the buffer 140 can be used to play the role of intermediate cache and parallel data supply in the process of data transmission and operation, on the one hand, to cache the first data read by the interface circuit 110 from the memory 10, and on the other hand, to cache the second data generated after the operation of the operation circuit 130. The first part and the second part are divided based on the segmentation of the buffer, for example, the data can be segmented according to the width of the internal transmission path, so that the originally continuous data is divided into two storage units that can be accessed in parallel or in series in physical structure, thereby supporting the reading or writing of two data segments. Taking the specific working process of parallel access as an example, the first storage unit and the second storage unit are configured to be able to provide the first data to the operation circuit in parallel when the operation circuit 130 performs the double-precision floating-point function operation, that is, two parts of the first data are transmitted to the operation unit at the same time in one clock cycle, so that each order power operation in the polynomial expansion calculation can quickly obtain the required complete input data; after the operation is completed, the first storage unit and the second storage unit can also output the second data to the interface circuit 110 in parallel, that is, the double parallel transmission of the result data is realized in one bus access, thereby reducing the bus occupation time and improving the data return efficiency.
[0030] In the first embodiment of the buffer 140, the buffer 140 can be implemented by using a hardware architecture of two memory banks (Bank), each Bank can be set as a static random access memory (SRAM), and the two Banks correspond to the first storage unit and the second storage unit respectively, and the access width of each Bank can match the bus width of the interface circuit (such as the first part corresponding to the data width of one bus access, and the second part corresponding to the next continuous data), which makes it possible to complete the serial or parallel scheduling of the double-part data internally through the buffer even if the bus width of the interface circuit itself is limited (for example, only the data width corresponding to the first part can be accessed at a time), thereby reducing the requirement for the bandwidth and access speed of the interface circuit.
[0031] In the second embodiment of the buffer 140, the buffer 140 adopts a hardware architecture of two Banks (each Bank is an SRAM). When the interface circuit 110 is a 32-bit bus interface, the buffer 140 can achieve complete access of 64-bit data by using each SRAM in a 32-bit manner, so as to sequentially transmit low 32-bit data and high 32-bit data under the condition that the bus bandwidth remains 32-bit, equivalently achieve the effect of 64-bit data processing, reduce the area overhead of the interface circuit 110 without expanding the bus width, and reduce the dependence on the single Bank capacity through the double Bank structure, thereby reducing storage redundancy and improving area efficiency, while ensuring efficient reading and writing of double-precision floating-point data. The above embodiments can still meet the supply demand of double-precision floating-point data without improving the bus bandwidth of the interface circuit, not only improve the data throughput rate of the operation stage and the result output stage, but also relieve the performance pressure of the interface circuit on the bus width and access frequency through the segmented caching and parallel access manner, so that the overall architecture realizes efficient data flow processing without increasing the complexity of the interface hardware.
[0032] With reference to the foregoing Figure 1 The register circuit 150 is further configured to cache the first parameter received by the interface circuit 110 and the second parameter written by the control circuit 120.
[0033] The register circuit 150 in the above embodiments can be used to cache the first parameter and the second parameter. The first parameter can include operation type, operation item number, input data quantity, target precision constraint value and the like, and the second parameter can include result overflow, result underflow, illegal operation or result inaccuracy and the like. The register circuit 150 can be implemented in a multi-register array structure, and each register is allocated to store parameters of a specific category, for example, operation control class parameters, precision class parameters and state identification class parameters are respectively stored in different register groups, so as to realize fast positioning and reading of parameters. For example, when the control circuit 120 receives the first parameter from the system through the interface circuit 110, for example, the first parameter can be issued by a central processor, and the first parameter is cached by the register circuit 150, and the first parameter can be cached in a configuration register in the register circuit 150. Or when the control circuit 120 outputs the second parameter through the interface circuit 110, the second parameter can be first cached in the register circuit 150, for example, in a state register. At this time, the register circuit 150 can support the control circuit 120 to update the operation state of the operation circuit 130 in the state register, for example, when it is determined that the current first data has result overflow after one operation, the corresponding state register of the register circuit 150 is set to “01”, and when result inaccuracy occurs, the state register is updated to “00”.
[0034] In another optional implementation, the register circuit 150 can also support batch writing and batch reading of parameters to reduce the number of bus interactions and improve data transmission efficiency. By setting the register circuit 150, a stable parameter cache layer can be formed between the interface circuit 110 and the control circuit 120, which not only avoids the loss or invalidation of parameters due to bus occupation or delay during operation, but also enables repeated use of unchanged parameter information during multiple operations, reduces the number of accesses to external memory, and thus reduces the overall system delay and improves the processing efficiency of the hardware acceleration circuit.
[0035] Figure 2 A structural diagram of an operation circuit provided in some embodiments of the present application is shown. The operation circuit 200 includes: a multiplication operation module 210 configured to calculate the product of each order power term of the polynomial expansion of the first data based on the first parameter and the third data; and an addition operation module 220 configured to perform accumulation operation based on the result of the product of each order power term to determine the partial sum or final sum of the polynomial expansion.
[0036] In the above application, it is assumed that the first data is sinx, the operation type is a sine function, the third data is a pre-stored polynomial expansion coefficient set, for example , and the target precision constraint value indicated by the first parameter determines the number of terms of the polynomial expansion (for example, 7 terms are taken). The operation circuit 130 first generates the power terms required for the polynomial expansion by the multiplication operation module 210: x is obtained in the initial period, and x 2 is calculated at the same time; the subsequent power terms are obtained in a recursive manner, for example, x 3 =x×x 2 , x 5 =x 3 ×x 2 , x 7 =x 5 ×x 2 , and so on, thereby avoiding repeated multiplication or sequentially multiplying, such as x 2 =x×x, x 3 =x×x×x, temporarily storing x and then repeatedly multiplying, which is not specifically limited here. Subsequently, the multiplication operation module 210 sequentially multiplies each power term with the corresponding coefficient to obtain the product value of the term, for example, the first term T1=1×x, the second term T2=(1 / 3!)×x 3 , the third term T3=(1 / 5!)×x 5 , and the fourth term T4=(1 / 7!)×x 7, until the last item limited by the target precision constraint value. Each product value obtained is immediately sent to the addition module 220, accumulated with the partial sum of the previous round, and the partial sum is continuously updated. For example, when x=1.0 and the target precision is 7 items, the result calculated by the multiplication module 210 and the addition module 220 is approximately equal to 0.8414709848. The difference between this final result and the mathematical library sin1≈0.841470984807 is within 10 -13 The multiplication module 210 of the hardware acceleration circuit is responsible for generating each value by "power term × coefficient", while the addition module 220 is responsible for accumulating each term to form the result. This division of labor can reduce the number of multiplications by recursively generating power terms, improve hardware computing efficiency, and flexibly control the balance between computing accuracy and computing latency based on the target accuracy constraint value given in the first parameter.
[0037] Continue to refer Figure 2 The operation circuit 130 also includes: a pre-processing module 230, configured to determine the calculation range of the first data before the polynomial is expanded; and a result post-processing module 240, configured to perform format processing operations on the partial sum or the final sum.
[0038] In the above embodiment, the pre-processing module 230 is used to determine the calculation range of the first data before the polynomial is expanded. For example, when performing trigonometric function operations, since trigonometric function operations have certain periodicity and symmetry, for example, in the calculation of the function sinx, only the function value in the interval [0, π] can be calculated. In the interval [π, 2π], the function value in the interval [0, π] is directly used, and the value of sinx is negated by the result post-processing module 240, so that the function value can be obtained without performing specific calculations on the function. The polynomial expansion operation is no longer repeated in the interval outside [0, π], thereby reducing the number of multiplication and addition calculations, reducing the overall operation delay and power consumption, and ensuring the calculation accuracy. The result post-processing module is used to perform format processing operations after the polynomial expansion obtains the partial sum or final sum, such as rounding operations (such as rounding to the nearest even number), rounding operations (such as rounding to zero, rounding up, or rounding down), or absolute value operations, to meet the requirements of different application scenarios for the result numerical format. For example, in some numerical computing environments, if the target output precision is IEEE754 double precision format, the result post-processing module can round the internal calculation result to make it conform to the representation range of 53 effective binary bits. When performing absolute value operations, the result post-processing module 240 can directly clear the sign bit without changing the mantissa, thereby efficiently obtaining the absolute value result. The pre-processing module 230 can also perform exponent alignment, such as when calculating 1×10 -5 +1×10 -4 When 1×10 -4Transform to 10x10 -5 Thus, the calculation is facilitated. In the above embodiment of the hardware acceleration circuit, the preprocessing module 230 can help determine the calculation interval, improve the calculation efficiency, and the result post-processing module 240 can ensure that the output result meets the specific numerical format or data processing requirement, thereby improving the adaptability and output consistency of the hardware acceleration circuit in different calculation tasks.
[0039] In some embodiments of the present application, the control circuit 120 is further configured to determine the target operation flow of the operation circuit 130 according to the first parameter, and control the multiplication operation module 210, the addition operation module 220, the preprocessing module 230, and the result post-processing module 240 to perform at least one operation link of the double-precision floating-point function operation according to the preset operation sequence, the operation link including operation interval determination, multiplication operation, addition operation, and negation operation.
[0040] In some embodiments of the present application, when the first parameter indicates that the operation type is a trigonometric function operation, the control circuit 120 first starts the preprocessing module to map the input first data to a predetermined principal value interval, for example, [0, π], to reduce the operation amount by utilizing the periodicity and symmetry of the function, and then controls the multiplication operation module 210 to calculate the product of each order power term of the polynomial expansion based on the third data (a set of polynomial expansion coefficients), and controls the addition operation module 220 to accumulate each order power term product to obtain a partial sum or a final sum of the polynomial expansion. In the Taylor expansion process, the control circuit 120 can coordinate the order of multiplication calculation and addition calculation between the multiplication operation module 210 and the addition operation module 220, so as to finally calculate and determine the second data. After obtaining the calculation result, the control circuit 120 starts the result post-processing module to perform format processing operation according to the original interval position of the first data, for example, when the input value is in the interval [π, 2π], the calculated result of sinx is negated, so as to directly obtain the function value in the target interval. In some embodiments, the control circuit 120 can also distribute different first data to different operation circuit 130 instances for parallel execution to improve the overall calculation throughput. In one operation circuit 130, the control circuit 120 can sequentially put the multiplication operation module 210, the addition operation module 220, the preprocessing module 230, and the result post-processing module 240 into work according to the preset operation sequence, form a pipeline processing flow, form the distribution of the total task amount in multiple operation circuits 130 and the calculation execution of a single task in one operation circuit 130, so as to complete the coordination of the entire calculation process and improve the calculation efficiency.
[0041] In some embodiments of the present application, the first data is double-precision floating-point input data, and the first parameters further include at least one of: an operation type, an operation item number, and an input data number; the second data is double-precision floating-point calculation result, and the second parameters include at least one of: result overflow, result underflow, illegal operation, and result inaccuracy. The transcendental operation includes at least one of: a trigonometric function operation and a power function operation; and the format processing operation includes at least one of: a rounding operation, an integer operation, and an absolute value operation.
[0042] The operation type can be used to distinguish the specific type of double-precision floating-point function operation, such as a trigonometric function operation or a power function operation. The operation item number can be used to specify the polynomial order used in the transcendental operation based on polynomial expansion, such as determining the use of the first several terms of Taylor expansion for calculating functions such as cosx, sinx, etc. The input data number parameter is used to indicate the number of double-precision floating-point data to be calculated in the current batch, so as to assist the control circuit 120 to schedule the operation circuit 130 to perform parallel or pipelined working mode. The target precision constraint value parameter is used to set the expected calculation precision, and the operation circuit can select an appropriate polynomial order or rounding strategy based on the precision constraint to reduce the calculation amount while ensuring the precision. In the second parameter, the result overflow indicates that the absolute value of the calculation result exceeds the maximum value range that can be represented by the double-precision floating-point number, such as when the exponential part exceeds the maximum representable exponent, usually returning positive or negative infinity. The result underflow means that the absolute value of the calculation result is smaller than the minimum non-zero value range that can be represented by the double-precision floating-point number, such as when the exponential part is lower than the minimum representable exponent, the result tends to zero, and can be represented as a non-normalized number or zero. The illegal operation means that there is a situation that does not conform to the mathematical definition in the operation process, such as square root operation on negative numbers, zero as a divisor, zero and infinity multiplication, etc. Such operations will cause the result to be undefined, and can return a "not a number" (NaN) flag. The result inaccuracy means that the calculation result is within the representation range of the double-precision floating-point number, but due to rounding or truncation, there is a deviation between the numerical value and the true mathematical result, such as rounding to the nearest representable numerical value, or the approximation error caused by insufficient polynomial order.
[0043] Figure 3 A signal flow conversion diagram of a hardware acceleration circuit provided in some embodiments of the present application is shown. Figure 3 The circuit structure of the present application is used to Figure 1Based on this, taking the interface circuit 110 having a 32-bit bit width as an example, the names of the various signals (data) in the figure include: m_rdata[31:0] is the data read from the external memory or peripheral device by the main bus interface module 111, which serves as the first data. m_wdata[31:0] is the result data written externally by the main bus interface module 111, which is derived from the second data in the buffer 140. m_haddr[31:0] is the address signal for the main bus interface module 111 to access the memory 10, used to locate the storage location of the first data or the third data (polynomial coefficients). m_hwrite is the write enable signal for the main bus interface module 111, which can be a high level, for example, to indicate a write operation. m_hsize[2:0] is the bit width information of the data accessed by the main bus interface module 111, such as 32 bits or 64 bits, which can be used to determine the reading method of the first and second parts of the first data. m_hready is the handshake signal indicating that data transmission is ready. wdata[31:0] is the control or parameter data written by the external host, i.e., the first parameter, received by the slave bus interface module 112. haddr[31:0] is the address signal used by bus interface module 112 to access the register circuit or control circuit. hwrite is the write enable signal from bus interface module 112, which can be a high level, for example, to indicate writing to register circuit 150. hsize[2:0] is the bit width information for data accessed from bus interface module 112. rdata[31:0] is the read data returned from bus interface module 112 to the external host, i.e., the second parameter. hready is the handshake signal from bus interface module 112 indicating that data transfer is ready. src_data is the first data provided by buffer 140 to arithmetic circuit 130. result is the second data output by the arithmetic circuit to buffer 140. bf_ctrl is a control signal for buffer 140, used to manage the writing and output order of the second data. data_cnt_i is the amount of input data received by control circuit 120, which can correspond to the input data amount in the first parameter. ctrl is a control signal issued by control circuit 120, which can enable interface circuit 110 to receive or send data. data_info is the first parameter passed from the bus interface module 112 or register circuit 150, and includes the operation type, number of operations, target precision constraint value, etc. sys_busy is the system busy signal; a high level indicates that the arithmetic circuit is currently executing an operation and cannot accept new tasks. intr is an interrupt signal used to notify the external host when an operation is completed or an error occurs. bf_data_info is the data information signal of buffer 140. It can be used by the control circuit 120 to instruct the main bus interface module 111 to store data in the buffer 140 and to instruct the arithmetic circuit 130 to read data from a specific location in the buffer 140.state is a state signal outputted by the operation circuit 120 to indicate the current operation stage state, such as "data fetching stage", "calculation stage", "result transmission stage". enable is an operation enable signal used to control the operation circuit 120 to start the operation circuit 130 to execute a specific operation. cal_type is an operation type signal, such as sin, cos, pow, etc., which can correspond to the "operation type" in the first parameter. Continue to refer to. Figure 2 The first data (double-precision floating-point input data) src_data inputted from outside, the constant data constant_data in the first parameter, such as 1 / 2!, 1 / 4!, 1 / 6! and other polynomial coefficients, are sent into the operation circuit 130 together. When the enable signal provided by the control circuit is valid, the first calculated part of the data src_A in the first data src_data is subjected to range control by the pre-processing module 230, and the A value and its sign bit sign_A are outputted. The multiplication operation module 210 receives A and calculates A 2 , A 4 , A 6 ... and the corresponding factorial reciprocal term (-1)^(n-1)A^(2n-1) / (2n-1)!, and the summation or subtraction operation is sequentially executed on each term by the addition operation module 220 to obtain the sum of the polynomial expansion. The result post-processing module 240 generates the final sign bit sign_C in combination with sign_A, and performs floating-point normalization processing on sum. frac_C represents the decimal part (i.e. the mantissa) of the floating-point number to accurately describe the effective bits of the value. exp_C represents the exponent part of the floating-point number, which is the exponent field in the double-precision floating-point number format, used to represent the magnitude of the value. The second data and the second parameter in the double-precision floating-point format are outputted (such as result overflow = 0, result underflow = 0, illegal operation = 0, and inaccuracy = 0).
[0044] The signal flow process of the above embodiment takes the execution of cosx operation as an example. The first parameter is written from the bus interface module 112 to the register circuit 150, including operation type = cos, input data number = 1, operation item number = 5 (polynomial order), and target precision constraint value = 1e-15. The control circuit 120 receives the first parameter and instructs the main bus interface module 111 to read the first data (double-precision floating point value of x, for example, x = 1.0471975512) and the third data (polynomial expansion coefficients of cosx, for example, 1, -1 / 2!, 1 / 4!, -1 / 6!, 1 / 8!) from the memory 10. The first data is divided into high 32 bits (first part) and low 32 bits (second part) and stored in the buffer 140. The control circuit 120 sends cal_type = cos, enable signal, and bf_data_info signal for reading data at specific position of the buffer to the operation circuit 130 and determines x in the calculation range of [0, π] by the preprocessing module; then the multiplication operation module 210 calculates x 2 , x 4 … in turn; then the addition operation module 220 performs polynomial accumulation according to the coefficients to obtain the approximate value of cosx. If x exceeds a certain range of π, the preprocessing module will reduce the actual operation amount by taking the inverse or sign transformation, for example, cos(π + θ) = -cosθ. The operation result (double-precision floating point format) is stored in the buffer 140, and the high 32 bits (first part) and the low 32 bits (second part) are stored respectively. The main bus interface module 111 writes the second data back to the outside and outputs the second parameter (including whether to overflow, whether to be inaccurate, etc.) from the bus interface module 112.
[0045] The operation circuit 130 in the above embodiment can form four state conversion logics in the working process, Figure 4 The operation circuit state conversion logic diagram provided in some embodiments of the present application is shown. When converting between different working states, the idle state idle indicates that the system has no task and waits for external start instruction (for example Figure 3The entering condition is transfer_done = 0 and sys_busy = 0. The triggering condition is enable = 1 (start operation instruction) is received, entering get_data, at which time the first data is read from the interface circuit 110. The execution flag sys_busy = 1 is set, and at the same time, intr_clr = 0 is cleared, hbus_req = 1 is set, the buffer 140 is prepared for writing, the buffer 140 address signal is haddr = start_addr + cnt_in * 4 (offset addressing by word), the data width is hsize = 3'b010 (word), and the data ready condition data_ready is indicated, entering calculating, calling the operation circuit 130, and executing the multiplication operation module 210, the addition operation module 220, the preprocessing module 230, and the like. The overflow, underflow, result illegal NaN, or illegal operation invalid is continuously monitored during the calculation. After the calculation is completed, cal_done = 1, entering transfer, the second data of the calculation result is written into the buffer 140, and is output to the outside through the interface circuit 110, at which time transfer_done = 1, cal_en = 0, and hbus_req = 1 is initiated, hwrite = 1, and the result storage address haddr = result_addr, hsize = 3'b010 (word) are indicated, and the data transmission is completed, returning to idle.
[0046] Figure 5 A flowchart of a hardware acceleration method provided in some embodiments of the present application is shown. The hardware acceleration method is used for double-precision floating-point function operation by the hardware acceleration circuit provided in the above embodiments, comprising: S510: controlling the interface circuit to read first data, first parameters, and third data, wherein the first data is double-precision floating-point input data, the first parameters include a target precision constraint value, and the third data includes coefficients of polynomial expansion in the double-precision floating-point function operation; S520: enabling the operation circuit to perform double-precision floating-point function operation on the first data based on the first parameters and the third data to determine second data and second parameters, the double-precision floating-point function operation including one or all of transcendental operation or format processing operation based on polynomial expansion, wherein the order of polynomial expansion is determined based on the target precision constraint value, and the second data is a double-precision floating-point calculation result; S530: outputting the second data and the second parameters to the outside.
[0047] The specific content of the above embodiments can refer to the schemes and beneficial effects related to the hardware acceleration circuit described above, and will not be repeated here.
[0048] Based on the same technical concept, the present application also provides an integrated circuit, comprising: a memory configured to pre-store third data, the third data being used for coefficients of polynomial expansion in double-precision floating-point function operation; and the hardware acceleration circuit provided in the above embodiments, configured to read the first data, the first parameter and the third data, perform double-precision floating-point function operation on the first data based on the first parameter and the third data, and determine the second data and the second parameter, the double-precision floating-point function operation including one or all of transcendental operation or format processing operation based on polynomial expansion, wherein the order of the polynomial expansion is determined based on a target precision constraint value. The hardware acceleration circuit can serve as a double-precision floating-point function calculation performer of a central processing unit (or other processing unit), and such calculation is performed by the central processing unit and then downgraded to the hardware acceleration circuit for calculation, thereby sharing the processing pressure of the central processing unit. Moreover, due to the structural design of the hardware acceleration circuit of the present application, such as directly reading the third data through the interface circuit, the hardware acceleration circuit avoids occupying the bus by frequently accessing the memory through the system bus, further reduces the processing pressure of the system, gives more resources to the central processing unit, and improves the processing efficiency of the entire processing system.
[0049] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments. In addition, the above embodiments can be freely combined as needed.
Claims
1. A hardware acceleration circuit, comprising: The application relates to a double-precision floating-point function operation circuit. The interface circuit is configured to read first data and a first parameter, and output second data and a second parameter to the outside, and is further configured to read third data, wherein the first data is double-precision floating-point input data, the second data is double-precision floating-point calculation results, the first parameter comprises a target precision constraint value, and the third data comprises coefficients of polynomial expansion in double-precision floating-point function operation. The control circuit is configured to read the first parameter through the interface circuit, instruct the interface circuit to read the first data and the third data according to the first parameter, enable the operation circuit to perform operation, and instruct the interface circuit to output the second data and the second parameter. At least one operation circuit is configured to perform double-precision floating-point function operation on the first data based on the first parameter and the third data, to obtain the second data and the second parameter. The double-precision floating-point function operation comprises one or all of transcendental operation based on polynomial expansion and format processing operation, wherein the order of the polynomial expansion is determined based on the target precision constraint value. The interface circuit comprises:
2. The hardware acceleration circuit of claim 1, wherein, a main bus interface module for reading the first data and / or the third data, and transmitting the second data to the outside; a slave bus interface module for reading the first parameter, and transmitting the second parameter to the outside. Further comprising:
3. The hardware acceleration circuit of claim 1, wherein, a buffer area configured to buffer the first data read by the interface circuit, and configured to buffer the second data generated by the operation circuit; The buffer area comprises a first storage unit and a second storage unit for buffering a first part and a second part of the first data respectively, and buffering the first part and the second part of the second data respectively, wherein the first part and the second part are divided based on a buffer area segmentation mode; The first storage unit and the second storage unit are configured to provide the first data to the operation circuit and / or output the second data to the interface circuit when the operation circuit performs the double-precision floating-point function operation. Further comprising:
4. The hardware acceleration circuit of claim 1, wherein, a register circuit for buffering the first parameter received through the interface circuit, and buffering the second parameter written through the control circuit. The operation circuit comprises:
5. The hardware acceleration circuit of any of claims 1 to 4, wherein, a multiplication operation module configured to calculate the product of each order power term of polynomial expansion of the first data based on the first parameter and the third data; an addition operation module configured to perform accumulation operation based on the result of the product of each order power term, to determine the partial sum or final sum of the polynomial expansion. The operation circuit further comprises:
6. The hardware acceleration circuit of claim 5, wherein, a preprocessing module configured to determine the calculation range of the first data before the polynomial expansion; a result post-processing module configured to perform the format processing operation on the partial sum or the final sum. 7. The hardware acceleration circuit of claim 6, wherein, The control circuit is further configured to determine a target operation flow of the operation circuit according to the first parameter, and control the multiplication operation module, the addition operation module, the preprocessing module and the result post-processing module to perform at least one operation link of the double-precision floating-point function operation in a preset operation sequence, the operation link including operation interval determination, multiplication operation, addition operation and negation operation.
8. The hardware acceleration circuit of any of claims 1 to 4, wherein, The first parameter further includes at least one of operation type, operation item number and input data number. The second parameter includes at least one of result overflow, result underflow, illegal operation and result inaccuracy. The transcendental operation includes at least one of trigonometric function operation and power function operation. The format processing operation includes at least one of rounding operation, integer operation and absolute value operation.
9. A hardware acceleration method, comprising: A method for performing a double-precision floating-point function operation by the hardware acceleration circuit of any one of claims 1 to 8, comprising: controlling the interface circuit to read first data, first parameters and third data, wherein the first data is double-precision floating-point input data, the first parameters include a target precision constraint value, and the third data includes coefficients of a polynomial expansion in the double-precision floating-point function operation; enabling the operation circuit to perform the double-precision floating-point function operation on the first data based on the first parameters and the third data to determine second data and second parameters, the double-precision floating-point function operation including one or all of a transcendental operation or a format processing operation based on a polynomial expansion, wherein an order of the polynomial expansion is determined based on the target precision constraint value, and the second data is a double-precision floating-point calculation result; outputting the second data and the second parameters to the outside.
10. An integrated circuit, characterized by comprising: a memory configured to pre-store third data including coefficients of a polynomial expansion in a double-precision floating-point function operation; the hardware acceleration circuit of any one of claims 1 to 8, configured to read first data, first parameters and the third data, perform the double-precision floating-point function operation on the first data based on the first parameters and the third data, and determine second data and second parameters, the double-precision floating-point function operation including one or all of a transcendental operation or a format processing operation based on a polynomial expansion, wherein an order of the polynomial expansion is determined based on a target precision constraint value.
Citation Information
Patent Citations
Hardware-friendly high-precision floating point transcendental function calculation system
CN119201036A
Floating point operation unit of sine and cosine functions, floating point operation acceleration method, equipment and medium
CN120631302A
Data processor, data processing method and arithmetic control program
JP2006323710A
Configurable function approximation based on switching mapping table content
US11423313B1
Processing unit
US20100332573A1