A Hardware Architecture and Method for Algorithm Acceleration

By designing register arrays, controllers and accessors in the hardware architecture, the feature extraction and matching of fingerprint algorithms is accelerated without increasing the chip area, solving the problems of large area and low efficiency of hardware acceleration solutions in the prior art, and improving the performance and flexibility of fingerprint recognition.

CN114327366BActive Publication Date: 2025-07-25SHANGHAI AISINOCHIP ELECTRONICS TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111548900.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-07-25
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

The existing fingerprint algorithm acceleration technology mainly relies on software implementation. The hardware acceleration solution is large in area and is accelerated only for specific steps, making it difficult to meet the needs of efficient authentication.

Method used

Design a hardware architecture, including register array, controller, operator and accessor. Through the control register, set the algorithms between two arrays, complete the multiplication and accumulation of different lengths, the distance between feature points and sum operations, and use 2 SRAM to complete data reading and cache in the same period to improve algorithm performance.

Benefits of technology

Without increasing the chip area, the performance of the fingerprint algorithm is improved, the feature extraction and matching process is accelerated, and flexible algorithm acceleration is realized, which is suitable for a variety of fingerprint algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114327366B_ABST
    Figure CN114327366B_ABST
Patent Text Reader

Abstract

The present invention discloses a hardware architecture and method for algorithm acceleration. The hardware architecture includes: a register array for storing control parameters, fixed parameters, and operation results; a controller for controlling the operation rules of an arithmetic unit, and for controlling the arithmetic unit to read values from the register array and values from an accessor; an arithmetic unit for performing corresponding multiplication, addition, subtraction, and mixed operations according to the operation rules given by the controller; and an accessor for storing operation data. In this way, by configuring the controller and the register array, multiplication-accumulation of different lengths, distances between feature points, and summation operations, etc. can be completed, accelerating various algorithms in fingerprint algorithm feature extraction and matching; by setting two SRAMs, the performance of the fingerprint algorithm can be improved without increasing the overall chip area. This design has flexible configuration, strong versatility, and small occupied area, and can accelerate different fingerprint algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of algorithm technologies, and in particular, to a hardware architecture and method for algorithm acceleration. Background Art

[0002] With the rapid development of the economy and the application of information technology, computing efficiency is crucial. Algorithm acceleration is applied in various fields such as the Internet, e-commerce, finance, and security. In the field of fingerprint algorithms, protecting personal privacy has become increasingly important, and the reliability requirements for identity verification and information security encryption are getting higher and higher. Traditional identity verification methods mainly include token-based or password-based. These traditional identity verification methods are easy to lose, forget, and crack, so that they can be misused by criminals, resulting in privacy leakage for users and potentially causing property losses. Therefore, traditional identity verification methods have been difficult to meet the current needs of information security and social development. Biometric technology identifies a person's identity based on physiological characteristics such as fingerprints, faces, and irises. Due to its convenience, that is, there is no need to carry or remember specifically; high security, that is, it is not easy to lose or forget; and differentiation, that is, each person has unique biometric characteristics different from others, it has gradually become an important supplement to traditional identity verification solutions and is widely used in fields such as finance and security.

[0003] Compared with other biometric methods such as faces and irises, fingerprint images have very good uniqueness, that is, the fingerprints of each finger are different; stability, that is, fingerprint patterns do not change easily; and convenience, that is, fingerprint collection is very convenient and the cost is relatively low. These excellent characteristics make fingerprint recognition the most widely used and mature biometric technology in the current market. As early as hundreds of years ago, fingerprint recognition was applied in people's production and life such as commercial trade and criminal sentencing. However, the fingerprint recognition systems at that time mainly judged whether it was the same finger through careful manual observation and analysis. This manual implementation method was easily interfered by human subjectivity, difficult to ensure its security and reliability, and had low efficiency, but it also greatly promoted the application of fingerprint recognition. With the development of human information technology, computers and mobile phones that can run programs emerged, and naturally fingerprint recognition algorithms appeared, so fingerprint recognition has gradually been converted into a more secure and efficient automatic fingerprint recognition system.

[0004] Currently, most fingerprint algorithms mainly include three parts: fingerprint feature extraction, fingerprint feature fusion, and fingerprint feature matching. Among them, fingerprint feature extraction is generally divided into the following steps: fingerprint image smoothing / normalization, image segmentation, orientation field calculation, image enhancement, binarization, thinning, and feature point extraction algorithms. A large number of multiply-accumulate, summation, and other operations are used in these algorithms. The fingerprint feature fusion and matching parts mainly use fingerprint feature matching algorithms, and calculating the distance between two feature points is essential in the matching algorithm. Of course, summation and mean value calculation operations are also used. In the existing technical solutions for accelerating fingerprint algorithms, most are still software. Even if there is hardware acceleration, it only accelerates specific step algorithms, and generally, the hardware area is relatively large. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a hardware architecture and method for algorithm acceleration.

[0006] A hardware architecture for algorithm acceleration provided by the present invention includes: a register array for storing control parameters, fixed parameters, and operation results; a controller for controlling the operation rules of the arithmetic unit and controlling the arithmetic unit to read the values in the register array and the values in the accessor; an arithmetic unit for performing corresponding multiplication, addition, subtraction, and mixed operations according to the operation rules given by the controller; an accessor for storing operation data; the register array, the controller, and the arithmetic unit are connected bidirectionally in sequence, and the accessor is respectively connected to the controller and the arithmetic unit.

[0007] The accessor includes 2 parallel static random access memories for performing read and cache operations in the same cycle. The register array includes a control register for configuring parameters such as the length of array operations, the operation rules between two arrays, the operation rules between operation results, and the data sign bit; an array address register for storing the address corresponding to the array; a fixed parameter register for storing fixed parameters during multiply-accumulate operations; and an operation result register for storing the final result of the operation.

[0008] The array address register includes an array 1 start address register, an array 1 offset address register, an array 2 start address register, and an array 2 offset address register.

[0009] The 0th bit of the control register is used to start the operation, the 1st bit is used to enable fixed parameters to participate in the operation, the 2nd bit is used to set the sign bit of the input data, the 4 bits from the 3rd to the 6th bits are used to set the length of the array operation, the 2 bits from the 7th to the 8th bits are used to set the operation rules for the corresponding data of arrays 1 and 2, the 2 bits from the 9th to the 10th bits are used to set the operation rules for the results after the operation of the corresponding data of arrays 1 and 2, the 11th bit sets the operation rules for the results between the first data of the array and the results between the second data of the array, the 12th bit sets the operation rules for the results between the third data of the array and the results between the fourth data of the array, the 13th bit sets the operation rules for the results between the fifth data of the array and the results between the sixth data of the array, and the 14th bit sets the operation rules for the results between the seventh data of the array and the results between the eighth data of the array.

[0010] There are 9 addresses corresponding to the data in Array 1. The address corresponding to the first data is the starting address register of Array 1. The address corresponding to the second data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 1 of Array 1. The address corresponding to the third data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 1 of Array 1. The address corresponding to the fourth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 2 of Array 1. The address corresponding to the fifth data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 2 of Array 1. The address corresponding to the sixth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 3 of Array 1. The address corresponding to the seventh data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 3 of Array 1. The address corresponding to the eighth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 4 of Array 1. The address corresponding to the ninth data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 4 of Array 1; there are also 9 addresses corresponding to the data in Array 2. The address corresponding to the first data is the starting address register of Array 2. The address corresponding to the second data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 1 of Array 2. The address corresponding to the third data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 1 of Array 2. The address corresponding to the fourth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 2 of Array 2. The address corresponding to the fifth data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 2 of Array 2. The address corresponding to the sixth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 3 of Array 2. The address corresponding to the seventh data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 3 of Array 2. The address corresponding to the eighth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 4 of Array 2. The address corresponding to the ninth data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 4 of Array 2.

[0011] There are 3 fixed-parameter registers, and the 3 fixed-parameter registers are used to store 9 fixed parameters, and each fixed parameter is 8-bit data.

[0012] According to the settings of the control register and other registers, the controller reads the data corresponding to the array address register and controls the arithmetic unit to complete operations such as multiplication and accumulation, distance, and summation of arrays of different lengths. When the first bit of the control register is set to 1, fixed parameters are used, and the controller only reads the data corresponding to the array 1 address to participate in the operation; otherwise, it reads the data corresponding to the array 2 address to participate in the operation.

[0013] An algorithm acceleration method based on a hardware architecture provided by the present invention includes the following steps:

[0014] Step S1, storing control parameters, fixed parameters, and operation results in a register array;

[0015] Step S2, controlling the operation rules of the arithmetic unit through the controller, and controlling the arithmetic unit to read the values in the register array and the accessor;

[0016] Step S3, performing corresponding multiplication, addition, subtraction, and mixed operations in the arithmetic unit according to the operation rules given by the controller;

[0017] Step S4, storing the operation data in the accessor.

[0018] The said step S4 includes: setting 2 parallel static random access memories to perform read and cache operations in the same cycle.

[0019] The present invention is mainly improved from two aspects. One is to set the operation rules between two arrays through the control register, which can complete various operations such as multiplication and accumulation of different lengths, distance between feature points, and summation operations, and can accelerate various algorithms in fingerprint algorithm feature extraction and matching; the other is to set 2 SRAMs, which can complete reading data from 2 SRAMs and completing an operation in the same cycle. After the fingerprint algorithm is operated, these 2 SRAMs can be used as ordinary SRAMs, and the performance of the fingerprint algorithm can be improved without increasing the overall chip area. Description of the Drawings

[0020] Figure 1 It is a schematic diagram of the hardware architecture for fingerprint algorithm acceleration according to the present invention;

[0021] Figure 2 It is a schematic diagram of the control register according to the present invention;

[0022] Figure 3 It is a diagram of the address composition of the array data according to the present invention;

[0023] Figure 4 This is the data operation control flow chart of the hardware architecture for fingerprint algorithm acceleration according to the present invention.

[0024] Among them:

[0025] 1 - Register array; 2 - Controller; 3 - Arithmetic unit; 4 - Accessor. Specific implementation mode

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment 1

[0028] As Figure 1 shown, a hardware architecture for algorithm acceleration provided by the present invention includes: a register array 1 for storing control parameters, fixed parameters, and operation results; a controller 2 for controlling the operation rules of the arithmetic unit 3 and controlling the arithmetic unit 3 to read the values in the register array 1 and the accessor 4; an arithmetic unit 3 for performing corresponding multiplication, addition, subtraction, and mixed operations according to the operation rules given by the controller 2; an accessor 4 for storing operation data; the register array 1, the controller 2, and the arithmetic unit 3 are connected bidirectionally in sequence, and the accessor 4 is respectively connected to the controller 2 and the arithmetic unit 3.

[0029] The accessor 4 includes two static random access memories SRAM. After the fingerprint algorithm calculation is completed, the two static random access memories SRAM can be used as ordinary static random access memories SRAM. By setting the operation rules between two arrays through the control register, various operations such as multiply-accumulate of different lengths, distance between feature points, and summation operation can be completed, which can accelerate various algorithms in fingerprint algorithm feature extraction and matching; by setting two static random access memories SRAM, the data of the two static random access memories SRAM can be read and an operation can be completed in the same cycle. After the fingerprint algorithm operation is completed, these two static random access memories SRAM can be used as ordinary static random access memories SRAM, which can improve the performance of the fingerprint algorithm without increasing the overall chip area.

[0030] Embodiment 2

[0031] As Figure 1As shown in the figure, a hardware architecture for algorithm acceleration provided by the present invention includes: a register array 1 for storing control parameters, fixed parameters, and operation results; a controller 2 for controlling the operation rules of an arithmetic unit 3, and for controlling the arithmetic unit 3 to read the values in the register array 1 and the values in an accessor 4; an arithmetic unit 3 for performing corresponding multiplication, addition, subtraction, and mixed operations according to the operation rules given by the controller 2; an accessor 4 for storing operation data; the register array 1, the controller 2, and the arithmetic unit 3 are sequentially connected bidirectionally electrically, and the accessor 4 is respectively connected to the controller 2 and the arithmetic unit 3.

[0032] Further, the accessor includes 2 parallel static random access memories for performing read and cache operations in the same cycle, and the 2 static random access memories SRAM can be used as ordinary static random access memories SRAM after the fingerprint algorithm calculation is completed.

[0033] Those skilled in the art can understand that by controlling the register to set the operation rules between two arrays, various operations such as multiplication and accumulation of different lengths, the distance between feature points, and summation operations can be completed, which can accelerate various algorithms in fingerprint algorithm feature extraction and matching; by setting 2 static random access memories SRAM, it is possible to complete reading the data of 2 static random access memories SRAM and completing an operation in the same cycle. After the fingerprint algorithm operation is completed, these 2 static random access memories SRAM can be used as ordinary static random access memories SRAM, improving the performance of the fingerprint algorithm without increasing the overall chip area.

[0034] Further, the register array 1 includes 1 control register, 10 array address registers, 3 fixed parameter registers, and 1 operation result register, a total of 15 registers;

[0035] The control register is used to configure parameters such as the length of array operations, the operation rules between two arrays, the operation rules between operation results, and the data sign bit;

[0036] The 10 array address registers include 1 starting address register of array 1, 4 offset address registers of array 1, 1 starting address register of array 2, and 4 offset address registers of array 2. These registers mainly store the corresponding addresses of the data of array 1 and array 2;

[0037] The 3 fixed parameter registers are used to store fixed parameters during multiplication and accumulation operations;

[0038] The operation result register is used to store the final result of the operation.

[0039] Embodiment III

[0040] A method for a hardware architecture for fingerprint algorithm acceleration, wherein the 0th bit of the control register is used to start the operation, the 1st bit is used to enable fixed parameters to participate in the operation, the 2nd bit is used to set the sign bit of the input data, the 4 bits from the 3rd to the 6th bits are used to set the array operation length, the 2 bits from the 7th to the 8th bits are used to set the operation rules for the corresponding data of array 1 and array 2, the 2 bits from the 9th to the 10th bits are used to set the operation rules for the results after the corresponding data operations of array 1 and array 2, the 11th bit sets the operation rules for the results between the first data of the array and the results between the second data of the array, the 12th bit sets the operation rules for the results between the third data of the array and the results between the fourth data of the array, the 13th bit sets the operation rules for the results between the fifth data of the array and the results between the sixth data of the array, and the 14th bit sets the operation rules for the results between the seventh data of the array and the results between the eighth data of the array.

[0041] Example 4

[0042] Such as Figure 3As shown in the figure, a hardware architecture method for fingerprint algorithm acceleration provided by the present invention. There are 9 addresses corresponding to the data in Array 1. The address corresponding to the first data is the starting address register of Array 1. The address corresponding to the second data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 1 of Array 1. The address corresponding to the third data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 1 of Array 1. The address corresponding to the fourth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 2 of Array 1. The address corresponding to the fifth data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 2 of Array 1. The address corresponding to the sixth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 3 of Array 1. The address corresponding to the seventh data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 3 of Array 1. The address corresponding to the eighth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 4 of Array 1. The address corresponding to the ninth data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 4 of Array 1. There are also 9 addresses corresponding to the data in Array 2. The address corresponding to the first data is the starting address register of Array 2. The address corresponding to the second data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 1 of Array 2. The address corresponding to the third data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 1 of Array 2. The address corresponding to the fourth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 2 of Array 2. The address corresponding to the fifth data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 2 of Array 2. The address corresponding to the sixth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 3 of Array 2. The address corresponding to the seventh data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 3 of Array 2. The address corresponding to the eighth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 4 of Array 2. The address corresponding to the ninth data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 4 of Array 2.

[0043] Example Five

[0044] A hardware architecture method for fingerprint algorithm acceleration provided by the present invention includes the following steps:

[0045] The register array 1 is used to store control parameters, fixed parameters, and operation results; the controller 2 is used to control the operation rules of the arithmetic unit 3, and to control the arithmetic unit 3 to read the values in the register array 1 and the values in 2 blocks of static random access memory SRAM; the arithmetic unit 3 is used to perform corresponding multiplication, addition, subtraction, and mixed operations according to the operation rules given by the controller 2; the 2 blocks of static random access memory SRAM are used to store operation data and can be used as ordinary static random access memory SRAM after the fingerprint algorithm calculation is completed. The register array 1 includes 1 control register, 10 array address registers, 3 fixed parameter registers, and 1 operation result register, a total of 15 registers; the control register is used to configure parameters such as the length of array operations, the operation rules between two arrays, the operation rules between operation results, and the data sign bit; 10 of the array address registers include 1 starting address register of array 1, 4 offset address registers of array 1, 1 starting address register of array 2, and 4 offset address registers of array 2, and these registers mainly store the corresponding addresses of the data of array 1 and array 2; 3 of the fixed parameter registers are used to store fixed parameters during multiplication and accumulation operations; the operation result register is used to store the final result of the operation.

[0046] Further, the 3 fixed parameter registers store 9 fixed parameters, and each fixed parameter is 8-bit data.

[0047] Embodiment Six

[0048] As Figure 4 shown, for a hardware architecture for algorithm acceleration provided by the present invention, the controller reads the data corresponding to the array address register according to the settings of the control register and other registers, and controls the arithmetic unit to complete operations such as multiplication and accumulation, distance, and summation of arrays of different lengths. When the first bit of the control register is set to 1, that is, fixed parameters are used, the controller only reads the data corresponding to the address of array 1 to participate in the operation, otherwise it is also necessary to read the data corresponding to the address of array 2 to participate in the operation. In practical applications, the data of array 1 and array 2 should be placed in different static random access memory SRAMs respectively, so that the controller can read the data in array 1 and array 2 in one cycle.

[0049] First, according to the first bit FIX_PARA_SEL of the control register, select where the data in array 2 comes from. If this bit is set to 1, the data comes from the fixed parameter register, otherwise the data comes from the static random access memory SRAM2; then the controller according to Figure 3The data corresponding to the middle address reading participates in the operation. According to the 2 bits of the 7th - 8th bits of the control register, the operation rules for the corresponding data of array 1 and array 2 are set. 00 means multiplying the corresponding data, 01 means adding the corresponding data, and 10 means subtracting the corresponding data; according to the 2 bits of the 9th - 10th bits of the control register, the operation rules for the result after data operation are set. 00 means no operation is performed on the result after operation, 01 means the result after operation multiplies itself, and 10 means the result after operation adds itself; according to the above settings, the operation result of a pair of data is obtained; then according to the 4 bits of the 3rd - 6th bits of the control register, the operation length of the array is set. If it is set to 0000, it means the array operation length is 1, and setting it to 0001 means the array operation length is 2, and so on. Setting 1000 means the array operation length is 9; when it is set to 1000, there are 9 result data accordingly. Then, according to the 11th - 14th bits of the control register, the operation rules between the corresponding two results are set, and finally the sum of the operation results after calculation is obtained to get the final operation result.

[0050] Embodiment Seven

[0051] A hardware - architecture - based algorithm acceleration method provided by the present invention includes the following steps:

[0052] Step S1, storing control parameters, fixed parameters, and operation results in a register array;

[0053] Step S2, controlling the operation rules of the arithmetic unit through a controller, and controlling the arithmetic unit to read the values in the register array and the values in the accessor;

[0054] Step S3, performing corresponding multiplication, addition, subtraction, and mixed operations in the arithmetic unit according to the operation rules given by the controller;

[0055] Step S4, storing the operation data in the accessor.

[0056] Further, the step S4 includes: setting 2 parallel static random - access memories, and performing read and cache operations in the same cycle.

[0057] Based on the above - mentioned fingerprint algorithm acceleration hardware architecture and method of the embodiment, various operations such as multiplying and accumulating with different lengths, calculating the distance between feature points, and summing operations can be completed. It can accelerate various algorithms in fingerprint algorithm feature extraction and matching. Moreover, it can complete reading the data of 2 static random - access memories SRAM and performing one operation in the same cycle. After the fingerprint algorithm operation is completed, these 2 static random - access memories SRAM can be used as ordinary static random - access memories SRAM. Without increasing the overall chip area, the performance of the fingerprint algorithm can be improved. The circuit structure is simple, the area is relatively small, it is convenient to be transplanted under various processes, and it is very easy to be added to the overall chip design.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A hardware system with algorithm acceleration, characterized in that Comprising: A register array for storing control parameters, fixed parameters, and operation results; a controller for controlling the operation rules of the arithmetic unit, as well as controlling the arithmetic unit to read the values in the register array and the values in the accessor; an arithmetic unit for performing corresponding multiplication, addition, subtraction, and mixed operations according to the operation rules given by the controller; an accessor for storing the data to be operated; the register array, the controller, and the arithmetic unit are electrically connected bidirectionally in sequence, and the accessor is respectively connected to the controller and the arithmetic unit; The register array includes a control register. The 0th bit of the control register is used to start the operation, the 1st bit is used to enable the fixed parameters to participate in the operation, the 2nd bit is used to set the sign bit of the input data, the 3rd - 6th bits, a total of 4 bits, are used to set the length of the array operation, the 7th - 8th bits, a total of 2 bits, are used to set the operation rules for the corresponding data of arrays 1 and 2, the 9th - 10th bits, a total of 2 bits, are used to set the operation rules for the results after the corresponding data operations of arrays 1 and 2, the 11th bit sets the operation rules for the results between the first data of the array and the results between the second data of the array, the 12th bit sets the operation rules for the results between the third data of the array and the results between the fourth data of the array, the 13th bit sets the operation rules for the results between the fifth data of the array and the results between the sixth data of the array, and the 14th bit sets the operation rules for the results between the seventh data of the array and the results between the eighth data of the array.

2. The hardware system for algorithm acceleration according to claim 1, characterized in that The accessor includes 2 parallel static random access memories for performing read and cache operations in the same cycle.

3. The hardware system for algorithm acceleration according to claim 2, wherein The register array includes a control register for configuring parameters such as the length of the array operation, the operation rules between two arrays, the operation rules between operation results, and the data sign bit; an array address register for storing the address corresponding to the array; a fixed parameter register for storing the fixed parameters during the multiply - accumulate operation; and an operation result register for storing the final result of the operation.

4. The hardware system for algorithm acceleration according to claim 3, characterized in that, The array address register includes an array 1 start address register, an array 1 offset address register, an array 2 start address register, and an array 2 offset address register.

5. The hardware system for algorithm acceleration according to claim 4, wherein There are 9 addresses corresponding to the data in Array 1. The address corresponding to the first data is the starting address register of Array 1. The address corresponding to the second data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 1 of Array 1. The address corresponding to the third data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 1 of Array 1. The address corresponding to the fourth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 2 of Array 1. The address corresponding to the fifth data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 2 of Array 1. The address corresponding to the sixth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 3 of Array 1. The address corresponding to the seventh data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 3 of Array 1. The address corresponding to the eighth data is the starting address register of Array 1 plus the lower 16 bits of the offset address register 4 of Array 1. The address corresponding to the ninth data is the starting address register of Array 1 plus the upper 16 bits of the offset address register 4 of Array 1. There are also 9 addresses corresponding to the data in Array 2. The address corresponding to the first data is the starting address register of Array 2. The address corresponding to the second data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 1 of Array 2. The address corresponding to the third data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 1 of Array 2. The address corresponding to the fourth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 2 of Array 2. The address corresponding to the fifth data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 2 of Array 2. The address corresponding to the sixth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 3 of Array 2. The address corresponding to the seventh data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 3 of Array 2. The address corresponding to the eighth data is the starting address register of Array 2 plus the lower 16 bits of the offset address register 4 of Array 2. The address corresponding to the ninth data is the starting address register of Array 2 plus the upper 16 bits of the offset address register 4 of Array 2.

6. The hardware system for algorithm acceleration according to claim 3 or 4, characterized in that, There are 3 fixed parameter registers, and the 3 fixed parameter registers are used to store 9 fixed parameters, and each fixed parameter is 8-bit data.

7. The hardware system for algorithm acceleration according to claim 4, wherein The controller reads the data corresponding to the array address register according to the settings of the control register and other registers, and controls the arithmetic unit to complete operations such as multiplication-accumulation, distance, and summation of arrays of different lengths. When the first bit of the control register is set to 1, that is, fixed parameters are used, the controller only reads the data corresponding to the address of Array 1 to participate in the operation, otherwise it reads the data corresponding to the address of Array 2 to participate in the operation.

8. An algorithm acceleration method for a hardware system accelerated by the algorithm as described in any one of claims 1 to 7, characterized in that, It includes the following steps: Step S1, store the control parameters, fixed parameters, and operation results in the register array; Step S2, control the operation rules of the arithmetic unit through the controller, and control the arithmetic unit to read the values in the register array and the accessor; Step S3, perform corresponding multiplication, addition, subtraction, and mixed operations in the arithmetic unit according to the operation rules given by the controller; Step S4, storing the operation data in the accessor.

9. The algorithm acceleration method of the hardware system based on algorithm acceleration according to claim 8, characterized in that The step S4 includes: setting 2 parallel static random access memories to perform read and cache operations in the same cycle.

Citation Information

Patent Citations

  • Montgomery analog multiplication algorithm and its analog multiplication and analog power operation circuit

    CN1492316A