Hardware acceleration method and device for data anomaly detection
By combining a pipelined adder with the Cordic algorithm, the computation speed and efficiency of the LOF algorithm are optimized, solving the problem of low computation speed under large data volumes and achieving efficient data anomaly detection.
Patent Information
- Application Number
- CN202510820850.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-28
AI Technical Summary
The LOF algorithm has a low computation speed when processing large amounts of data, and existing technologies have not been able to effectively solve the problems of computation time consumption and low efficiency.
A hardware acceleration method combining pipelined adders and the Cordic algorithm is adopted. The pipelined adder optimizes the time-parallel computation of Euclidean distance, and the hardware modulus unit enables the rapid calculation of data modulus.
It improves the computation speed and efficiency of the LOF algorithm in big data scenarios, reduces critical path latency, and enhances the performance of the adder in high-frequency applications.
Smart Images

Figure CN120848836A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data anomaly detection technology, and in particular to a hardware acceleration method and apparatus for data anomaly detection. Background Art
[0002] Anomaly detection uses data mining to find anomalous data in big data. By analyzing the distribution characteristics of anomalous data, it can extract anomalous information from massive amounts of data, find feature patterns, and discover potential and useful information.
[0003] The LOF (Local Outlier Factor) algorithm is a density-based algorithm for detecting local outliers. Outliers, also known as anomalies, are defined as data objects that are significantly different from other data distributions. The basic idea of this algorithm is to first calculate the reachability distance of each data point based on the data density around it, then calculate the local reachability density using the reachability distance, and finally further calculate the local outlier factor for each data point using the local reachability density. This outlier factor indicates the degree of outlier status of the data point; a larger factor value indicates a higher degree of anomaly, and a smaller factor value indicates a lower degree of anomaly.
[0004] The LOF algorithm requires calculating the reachability distance, reachability density, and outlier factor for each data point based on Euclidean distance. Therefore, when the amount of data is huge, running the LOF algorithm on general-purpose computing devices such as PCs (Personal Computers) is limited by several factors. First, the limited resources of the ALU (Arithmetic Logic Unit) in the CPU (Central Processing Unit) restrict the high-speed throughput of data. Second, the PC's data computation process involves operations such as fetching instructions and reading and writing data to memory. Actions unrelated to direct computation consume a lot of time, and the sequential execution method itself also consumes a lot of computation time.
[0005] In the prior art, two Chinese invention patent documents have been proposed: CN115827932A, published on March 21, 2023, and CN113049963A, published on June 29, 2021. Both patent documents disclose the processing of sample data, the storage of computer programs on storage media, and the execution of a local outlier factor algorithm program via computer equipment to perform anomaly detection. The prior art, represented by these two patent documents, still relies on the software-level implementation of the algorithm using a general-purpose computer system, and does not address the problems of computational time consumption and low efficiency during algorithm execution.
[0006] Among the prior art, Chinese invention patent document CN113630236A, published on July 21, 2021, replaces the adders on the critical path affecting the algorithm's calculation speed with carry-lookahead adders. The 32-bit carry-lookahead adder is composed of 11 cascaded groups of 3-bit carry-lookahead adders. Also among the prior art, Chinese invention patent document CN105335128A, published on February 17, 2016, focuses on the ALU circuit, which includes four parts: a decoding station, an inter-station register, a general-purpose register, and an execution station. The execution station contains three cascaded 64-bit carry-lookahead adders. The adder consists of four 4-bit adders cascaded together to form a first-stage 16-bit adder, two 16-bit adders cascaded together to form a second-stage 32-bit adder, and two 32-bit adders cascaded together to form a third-stage 64-bit adder.
[0007] The adders in the two patent documents mentioned above are combinational logic circuits, which will affect the delay of the critical path of the circuit and affect the high-speed characteristics. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention proposes a hardware acceleration method and apparatus for data anomaly detection, which can solve the problems of low Euclidean distance calculation speed and long detection time in the LOF anomaly detection algorithm under the condition of large data volume processing.
[0009] This invention is achieved by adopting the following technical solution:
[0010] A hardware acceleration method for data anomaly detection includes the following steps:
[0011] Step S1. Two sets of 16-bit wide data a(x) a ,y a ), b(x b ,y b The input is processed by a pipelined adder using two's complement addition to obtain x. a -x b ,y a -y b The calculation results;
[0012] The pipelined adder construction method includes: cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder; combining the two 8-bit carry-lookahead adders with three registers to form a 16-bit pipelined adder; of the two 8-bit carry-lookahead adders, one is used for low-bit addition calculation, and the other is used for high-bit addition calculation; the three registers are used to adjust the synchronization timing to form a pipelined structure.
[0013] Step S2. Using the vector pattern of the Cordic algorithm, the data modulus is solved by iterative calculation of the original data.
[0014] Step S1 specifically includes the following steps:
[0015] Step S 11 The generated signal and the propagated signal are generated respectively through a two-input AND gate and a two-input XOR gate;
[0016] Step S 12 Multiple carry signals are generated in parallel;
[0017] Step S 13 The propagation signal and the carry signal are XORed to generate the sum signal;
[0018] Step S 14 Cascade two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder;
[0019] Step S 15 An 8-bit carry-lookahead adder is used as the low 8-bit addition unit, and another 8-bit carry-lookahead adder is used as the high 8-bit addition unit. An 8-bit wide register reg1 is added after the low 8-bit carry-lookahead adder, and two 8-bit wide registers reg2 and reg3 are added before the high 8-bit carry-lookahead adder, ultimately forming a pipelined adder with a bit width of 16 bits.
[0020] Step S 16 Parallel numerical computation.
[0021] Step S 11 Specifically, this refers to: the i-th bit of the addend and augend is used to generate a signal G through a two-input AND logic gate. i The propagation signal P is generated through a two-input XOR logic gate. i , i≥0.
[0022] The step S 12 Specifically, it refers to: simultaneously generating multiple carry signals, where the carry signal C at any i-th bit... ifor:
[0023] C i =G i-1 +P i-1 G i-2 +P i-1 P i-2 G i-3 +…+P i-1 P i-2 …P0C0;
[0024] In the above formula, the subscripts of each letter are greater than or equal to 0.
[0025] The step S 16 Parallel computation specifically refers to the following: while the high 8-bit carry-lookahead adder is calculating the high-order bits of the current set of data, the low 8-bit carry-lookahead adder is simultaneously calculating the low-order bits of the next set of data.
[0026] Step S2 specifically includes the following steps:
[0027] Step S 21 Set the number of iterations i, and set register reg1 to store the scaling factor K. i ;
[0028] Step S 22 The input data is in the form of (x0, y0). Set the first parameter value x0 of the data stored in register reg2, and set the second parameter value y0 of the data stored in register reg3.
[0029] Step S 23 The parameter values in registers reg2 and reg3 are shifted i bits to the right using a shift operation before being stored.
[0030] Step S 24 The rotation sign d is determined based on the parameter value in register reg3. i ;
[0031] Step S 25 Iterative calculation to determine the rotation sign d i The value, and based on the rotation symbol d i The value of is updated to update the parameter values stored in registers reg2 and reg3;
[0032] Step S 26 Calculate the data modulus according to the following formula.
[0033]
[0034] In the formula, x i The modulus value after scaling is calculated after i iterations.
[0035] The step S 23 Specifically, this means: shifting the parameter value in register reg2 to the right by i bits through a shift operation, and storing the result in register reg4; shifting the parameter value in register reg3 to the right by i bits through a shift operation, and storing the result in register reg5.
[0036] The step S 24 Specifically, if the parameter value in register reg3 is greater than 0, then the rotation sign d... i =-1, if the parameter value in register reg3 is less than or equal to 0, then rotate the sign d. i =+1.
[0037] The step S 25 Specifically, if the rotation symbol d i If the result of adding register reg2 and register reg5 is -1, then the result is assigned to register reg2, and the parameter value stored in register reg2 is updated.
[0038] Subtracting register reg4 from register reg3 and assigning the result to register reg3 updates the parameter value stored in register reg3; if the rotation symbol d i =+1, then the result of subtracting register reg5 from register reg2 is assigned to register reg2, updating the parameter value stored in register reg2;
[0039] The result of adding register reg3 and register reg4 is assigned to register reg3, thus updating the parameter value stored in register reg3.
[0040] A hardware acceleration device for data anomaly detection includes a hardware fast addition unit and a hardware modulo unit; the hardware fast addition unit is used to construct a pipelined adder and process two sets of 16-bit wide data a(x) inputs. a ,y a ), b(x b ,y b Perform two's complement addition to obtain x. a -x b ,y a -y b The calculation results;
[0041] The pipelined adder construction method includes: cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder; combining the two 8-bit carry-lookahead adders with three registers to form a 16-bit pipelined adder; of the two 8-bit carry-lookahead adders, one is used for low-bit addition calculation, and the other is used for high-bit addition calculation; the three registers are used to adjust the synchronization timing to form a pipelined structure.
[0042] The hardware modulus unit is used to solve for the modulus of the data by iteratively calculating the original data using the vector pattern of the Cordic algorithm.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. In this invention, an adder based on a pipelined structure is formed by inserting inter-stage registers into a carry-lookahead adder. This design, as a sequential circuit, optimizes timing and is particularly suitable for high-frequency operating environments. Specifically, the pipelined adder decomposes the addition operation into multiple stages and inserts registers between each stage, allowing the computation of each stage to be executed in parallel within different clock cycles. This method not only reduces critical path latency but also improves overall throughput, thereby enhancing the adder's performance in high-frequency applications.
[0045] This invention also combines hardware fast addition and hardware modulo calculation, implementing the key calculation process (Euclidean distance calculation) that affects the detection time in the LOF anomaly detection algorithm through hardware parallel computing. This can effectively solve the problem of slow calculation speed in the software implementation of the algorithm and improve the calculation speed of the algorithm in big data scenarios.
[0046] 2. The pipelined adder constructed in this invention allows the carry calculation of the subsequent stage of the carry adder to not wait for the carry of the previous stage. Therefore, the critical path of the circuit is shortened, the circuit delay is reduced, and the calculation speed is improved. The pipelined structure can make full use of adder resources for parallel computing, while reducing the path length of combinational logic in the circuit and increasing the operating frequency of the adder, thus improving the calculation speed in two ways.
[0047] 3. In this invention, for multi-group data calculation under large data volume, while the high 8-bit carry-lookahead adder is calculating the high-order data of the current group of data, the low 8-bit carry-lookahead adder is simultaneously calculating the low-order data of the next group of data. Therefore, calculation results are continuously generated in each clock cycle, which greatly improves the calculation efficiency.
[0048] 4. This invention performs i iterations of calculation, y iThe value of x will approach 0, at which point x i The value is the modulus after scaling; therefore, it can be obtained through x. i and storage scaling factor K i The data modulus can be quickly calculated. Attached Figure Description
[0049] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments, wherein:
[0050] Figure 1 This is a schematic diagram of the process of the present invention;
[0051] Figure 2 This is a schematic diagram of the 4-bit carry-lookahead adder in this invention;
[0052] Figure 3 This is a schematic diagram of the 16-bit pipelined adder in this invention;
[0053] Figure 4 This is a schematic diagram of the data iterative shift calculation in this invention. Detailed Implementation
[0054] Example 1
[0055] As a basic embodiment of the present invention, the present invention includes a hardware acceleration method for data anomaly detection, comprising the following steps:
[0056] Step S1. Two sets of 16-bit wide data a(x) a ,y a ), b(x b ,y b The input is processed by a pipelined adder using two's complement addition to obtain x. a -x b ,y a -y b The calculation results.
[0057] The pipelined adder is constructed by cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder, and then combining these two 8-bit carry-lookahead adders with three registers to form a 16-bit adder. Of the two 8-bit carry-lookahead adders, one performs the lower 8-bit addition calculation, and the other performs the higher 8-bit addition calculation. The three registers are used to adjust the synchronization timing, forming the pipelined structure.
[0058] Step S2. Using the vector pattern of the Cordic algorithm, the data modulus is solved by iterative calculation of the original data.
[0059] Example 2
[0060] In a preferred embodiment of the present invention, the present invention includes a hardware acceleration method for data anomaly detection, comprising the following steps:
[0061] Step S1. Two sets of 16-bit wide data a(x) a ,y a ), b(x b ,y b The input is processed by a pipelined adder using two's complement addition to obtain x. a -x b ,y a -y b The calculation results.
[0062] The pipelined adder is constructed by cascading two 4-bit carry-lookahead adders to form the first-stage 8-bit carry-lookahead adder, and then combining these two 8-bit carry-lookahead adders with three registers to form a 16-bit adder. Of the two 8-bit carry-lookahead adders, one performs the lower 8-bit addition calculation, and the other performs the higher 8-bit addition calculation. The three registers are used to adjust the synchronization timing, forming the pipelined structure.
[0063] Specifically, step S1 includes the following steps:
[0064] Step S 11 The generated signal G is generated through a two-input AND gate and a two-input XOR gate, respectively. i With propagation signal G i .
[0065] Step S 12 Multiple carry signals are generated in parallel. Specifically, multiple carry signals are generated simultaneously, with the carry signal C for any i-th bit being... i for:
[0066] C i =G i-1 +P i-1 G i-2 +P i-1 P i-2 G i-3 +…+P i-1 P i-2 …P0C0.
[0067] Step S 13 The propagation signal and the carry signal are XORed to generate the sum signal.
[0068] Step S 14Two 4-bit carry-lookahead adders are cascaded to form an 8-bit carry-lookahead adder.
[0069] Step S 15 An 8-bit carry-lookahead adder is used as the low 8-bit addition unit, and another 8-bit carry-lookahead adder is used as the high 8-bit addition unit. An 8-bit wide register reg1 is added after the low 8-bit carry-lookahead adder, and two 8-bit wide registers reg2 and reg3 are added before the high 8-bit carry-lookahead adder, ultimately forming a pipelined adder with a bit width of 16 bits.
[0070] Step S 16 Parallel numerical computation.
[0071] Step S2. Using the vector pattern of the Cordic algorithm, the data modulus is solved by iterative calculation of the original data.
[0072] Example 3
[0073] In another preferred embodiment of the present invention, the present invention includes a hardware acceleration method for data anomaly detection, comprising the following steps:
[0074] Step S1. Two sets of 16-bit wide data a(x) a ,y a ), b(x b ,y b The input is processed by a pipelined adder using two's complement addition to obtain x. a -x b ,y a -y b The calculation results.
[0075] The pipelined adder is constructed by cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder, and then combining these two 8-bit carry-lookahead adders with three registers to form a 16-bit adder. Of the two 8-bit carry-lookahead adders, one performs the lower 8-bit addition calculation, and the other performs the higher 8-bit addition calculation. The three registers are used to adjust the synchronization timing, forming the pipelined structure.
[0076] Step S2. Utilize the vector pattern of the Cordic algorithm to solve for the data modulus through iterative calculations on the original data. This specifically includes the following steps:
[0077] Step S 21 Set the number of iterations i, and set register reg1 to store the scaling factor K. i .
[0078] Step S 22 The input data is in the form of (x0, y0). The first parameter value x0 is set in register reg2, and the second parameter value y0 is set in register reg3.
[0079] Step S 23 The parameter values in registers reg2 and reg3 are shifted i bits to the right using a shift operation before being stored.
[0080] Step S 24 The rotation sign d is determined based on the parameter value in register reg3. i .
[0081] Step S 25 Iterative calculation to determine the rotation sign d i The value, and based on the rotation symbol d i The value is updated to update the data stored in registers reg2 and reg3.
[0082] Step S 26 Calculate the data modulus according to the following formula.
[0083]
[0084] In the formula, x i The modulus value after scaling is calculated after i iterations.
[0085] Example 4
[0086] In another preferred embodiment of the present invention, the present invention includes a hardware acceleration method for data anomaly detection, as described in the appendix to the specification. Figure 1 This includes the following steps:
[0087] Step S1. Hardware fast addition. Specifically, two sets of 16-bit wide data a(x) a ,y a ), b(x b ,y b The input is processed by a pipelined adder using two's complement addition to obtain x. a -x b ,y a -y b The calculation results.
[0088] The pipelined adder construction method includes: using basic logic gate structures to form a 4-bit carry-lookahead adder; cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder; and combining the two 8-bit carry-lookahead adders with three registers to form a 16-bit pipelined adder. Of the two 8-bit carry-lookahead adders, one performs the lower 8-bit addition calculation, and the other performs the higher 8-bit addition calculation. The three registers are used to adjust the synchronization timing to form the pipelined structure.
[0089] Refer to the instruction manual appendix Figure 2 In a carry-lookahead adder, the carry calculation of the subsequent stage does not need to wait for the carry of the previous stage. Therefore, the critical path of the circuit is shortened, the circuit delay is reduced, and the calculation speed is increased. The pipelined structure can fully utilize adder resources for parallel computation, while simultaneously reducing the path length of combinational logic in the circuit and increasing the operating frequency of the adder. This dual approach improves calculation speed. The structure of the pipelined adder is shown in the attached manual. Figure 3 As shown.
[0090] Therefore, step S1 specifically includes the following steps:
[0091] Step S 11 The generation signal and propagation signal are generated respectively through a two-input AND gate and a two-input XOR gate. That is, the i-th bit of the addend and augend is generated into the generation signal G through the two-input AND gate. i The propagation signal P is generated through a two-input XOR logic gate. i , i≥0.
[0092] Specifically, Bit0 of the addend and augend is used to generate the generation signal G0 through a two-input AND gate, and the propagation signal P0 is generated through a two-input XOR gate. Similarly, Bit1 of the addend and augend is used to generate G1 and P1 through a logic gate, Bit2 of the addend and augend is used to generate G2 and P2 through a logic gate, and Bit3 of the addend and augend is used to generate G3 and P3 through a logic gate.
[0093] Step S 12 Multiple carry signals are generated in parallel, that is, multiple carry signals are generated simultaneously, and the carry signal C of any i-th bit is... i for:
[0094] C i =G i-1 +P i-1 G i-2 +P i-1 P i-2 G i-3 +…+P i-1 P i-2 …P0C0;
[0095] In the above formula, the subscripts of each letter are greater than or equal to 0.
[0096] Specifically, the output C0P0 is generated by a two-input AND logic gate, and the carry C1 is generated by a two-input OR logic gate.
[0097] C1 = G0 + C0P0.
[0098] The output is G0P1 through a two-input AND gate, P1P0C0 through a three-input AND gate, and carry C2 through a three-input OR gate: C2 = G1 + G0P1 + P1P0C0.
[0099] The outputs are P2P1P0C0 via a four-input AND gate, P2P1G0 via a three-input AND gate, P2G1 via a two-input AND gate, and carry C3 via a four-input OR gate: C3 = G2 + P2G1 + P2P1G0 + ...
[0100] P2P1P0C0.
[0101] The outputs are C0P3P2P1P0 through a five-input AND gate, G0P3P2P1 through a four-input AND gate, G1P3P2 through a three-input AND gate, G2P3 through a two-input AND gate, and finally carry C4 through a five-input OR gate: C4 = G3 + G2P3 + G1P3P2 + G0P3P2P1 + C0P3P2P1P0.
[0102] Among them, C0, C1, C2 and C3 are the four carry-in numbers generated simultaneously.
[0103] Step S 13 The propagation signal and the carry signal are XORed to generate the sum signal.
[0104] Specifically, the propagation signal P0 is XORed with the carry C0 to generate the sum S0; the propagation signal P1 is XORed with the carry C1 to generate the sum S1; the propagation signal P2 is XORed with the carry C2 to generate the sum S2; and the propagation signal P3 is XORed with the carry C3 to generate the sum S3.
[0105] Step S 14 Two 4-bit carry-lookahead adders are cascaded to form an 8-bit carry-lookahead adder. Specifically, two 4-bit carry-lookahead adders are cascaded one after the other. The final carry C4 of the output of the first 4-bit carry-lookahead adder is used as the input C0 of the second 4-bit carry-lookahead adder to produce an 8-bit carry-lookahead adder.
[0106] Step S 15 Refer to the attached instruction manual. Figure 3An 8-bit carry-lookahead adder is used as the low 8-bit addition unit, and another 8-bit carry-lookahead adder is used as the high 8-bit addition unit. An 8-bit wide register reg1 is added after the low 8-bit carry-lookahead adder, and two 8-bit wide registers, namely register reg2 and register reg3, are added before the high 8-bit carry-lookahead adder, ultimately forming a pipelined adder with a bit width of 16 bits.
[0107] Step S 16 Parallel computation of numerical values. Two sets of 16-bit wide data a(x) are processed. a ,y a ), b(x b ,y b The input pipeline adder performs two's complement addition to obtain x. a -x b ,y a -y b The calculation results are obtained by using a pipelined architecture. Based on the characteristics of the pipelined structure, for multi-group data calculations with large datasets, while the high 8-bit carry-lookahead adder is calculating the high-order bits of the current group of data, the low 8-bit carry-lookahead adder is simultaneously calculating the low-order bits of the next group of data. Therefore, calculation results are continuously generated each clock cycle, greatly improving computational efficiency.
[0108] Step S2. Hardware Modulus Calculation. Using the vector pattern of the Cordic algorithm, the modulus of the data is calculated iteratively on the original data. This includes the following steps:
[0109] Step S 21 Parameter preset: Set the number of iterations i, and set the scaling factor K to be stored in register reg1. i .
[0110] Step S 22 Data storage: The input data is in the form of (x0, y0). The first parameter value x0 is set in register reg2, and the second parameter value y0 is set in register reg3.
[0111] Step S 23 Data shifting: The parameter values in registers reg2 and reg3 are shifted i bits to the right and then stored.
[0112] For details, please refer to the instruction manual appendix. Figure 4In the hardware implementation, the ÷2 operation is performed by shifting data to the right. The data in register reg2 is shifted i bits to the right (calculated in the i-th iteration), and the result is stored in register reg4; the data in register reg3 is shifted i bits to the right, and the result is stored in register reg5 (calculated in the i-th iteration).
[0113] Step S 24 Determine the rotation sign: The rotation sign d is determined based on the parameter value in register reg3. i Set register reg6 to store the rotation symbol d. i The initial value is -1. If the parameter value in register reg3 is greater than 0, then d i =-1, if the parameter value in register reg3 is less than or equal to 0, then d i = +1. The mathematical expression is as follows:
[0114]
[0115] Step S 25 Numerical iterative update: Iterative calculation to determine the rotation sign d i The value, and based on the rotation symbol d i The value of is updated to update the parameter values stored in registers reg2 and reg3.
[0116] Specifically, if d i If the expression is -1, then the result of adding reg2 and reg5 is assigned to reg2, updating the parameter value stored in reg2; the result of subtracting reg3 and reg4 is assigned to reg3, updating the parameter value stored in reg3.
[0117] If d i If = +1, then the result of subtracting reg5 from reg2 is assigned to reg2, updating the parameter value stored in reg2; the result of adding reg3 to reg4 is assigned to reg3, updating the parameter value stored in reg3. The mathematical expression is as follows:
[0118] x i+1 =x i -d i y i 2 -i ,
[0119] y i+1 =y i +d i x i 2 -i .
[0120] Among them, addition and subtraction calculations can be performed quickly using the pipeline adder established in step S1.
[0121] Step S 26 Calculate the modulus: After i iterations, y i The value of x will approach 0, at which point x i The value is the modulus after scaling; therefore, through x... i ×K i The data modulus can be calculated: The specific calculation method is as follows:
[0122]
[0123] Example 5
[0124] In another preferred embodiment of the present invention, the present invention includes a hardware acceleration device for data anomaly detection, comprising a hardware fast addition unit and a hardware modulus calculation unit.
[0125] The hardware fast adder unit is used to construct a pipelined adder and to process the two sets of 16-bit wide input data a(x) a ,y a ), b(x b ,y b Perform two's complement addition to obtain x. a -x b ,y a -y b The calculation results are as follows. Specific calculation methods can be found in the methods described in Example 2 or Example 4.
[0126] The pipelined adder is constructed by cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder, and then combining these two 8-bit carry-lookahead adders with three registers to form a 16-bit adder. Of the two 8-bit carry-lookahead adders, one performs the lower 8-bit addition calculation, and the other performs the higher 8-bit addition calculation. The three registers are used to adjust the synchronization timing, forming the pipelined structure.
[0127] The hardware modulus calculation unit is used to solve for the modulus of the data by iteratively calculating the vector pattern of the Cordic algorithm on the original data. The specific solution method can be found in the methods described in Embodiment 3 or Embodiment 4.
[0128] In summary, any other corresponding modifications made by those skilled in the art after reading this invention document, without requiring creative mental effort, based on the technical solutions and concepts of this invention, are all within the scope of protection of this invention.
Claims
1. A hardware acceleration method for data anomaly detection, characterized in that: Includes the following steps: Step S1. Two sets of 16-bit wide data a(x) a ,y a ), b(x b ,y b The input is processed by a pipelined adder using two's complement addition to obtain x. a -x b ,y a -y b The calculation results; The pipelined adder construction method includes: cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder; combining the two 8-bit carry-lookahead adders with three registers to form a 16-bit pipelined adder; of the two 8-bit carry-lookahead adders, one is used for low-bit addition calculation, and the other is used for high-bit addition calculation; the three registers are used to adjust the synchronization timing to form a pipelined structure. Step S2. Using the vector pattern of the Cordic algorithm, the data modulus is solved by iterative calculation of the original data.
2. The hardware acceleration method for data anomaly detection according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S 11 The generated signal and the propagated signal are generated respectively through a two-input AND gate and a two-input XOR gate; Step S 12 Multiple carry signals are generated in parallel; Step S 13 The propagation signal and the carry signal are XORed to generate the sum signal; Step S 14 Cascade two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder; Step S 15 An 8-bit carry-lookahead adder is used as the low 8-bit addition unit, and another 8-bit carry-lookahead adder is used as the high 8-bit addition unit. An 8-bit wide register reg1 is added after the low 8-bit carry-lookahead adder, and two 8-bit wide registers reg2 and reg3 are added before the high 8-bit carry-lookahead adder, ultimately forming a pipelined adder with a bit width of 16 bits. Step S 16 Parallel numerical computation.
3. The hardware acceleration method for data anomaly detection according to claim 2, characterized in that: Step S 11 Specifically, this means that the i-th bit of the addend and augend is used to generate a signal G through a two-input AND logic gate. i The propagation signal P is generated through a two-input XOR logic gate. i , i≥0.
4. The hardware acceleration method for data anomaly detection according to claim 3, characterized in that: Step S 12 Specifically, it refers to: simultaneously generating multiple carry signals, where the carry signal C at any i-th bit... i for: C i =G i-1 +P i-1 G i-2 +P i-1 P i-2 G i-3 +…+P i-1 P i-2 …P0C0; In the above formula, the subscripts of each letter are greater than or equal to 0.
5. A hardware acceleration method for data anomaly detection according to claim 2, characterized in that: Step S 16 Parallel computation specifically refers to the following: while the high 8-bit carry-lookahead adder is calculating the high-order bits of the current set of data, the low 8-bit carry-lookahead adder is simultaneously calculating the low-order bits of the next set of data.
6. A hardware acceleration method for data anomaly detection according to claim 1 or 5, characterized in that: Step S2 specifically includes the following steps: Step S 21 Set the number of iterations i, and set register reg1 to store the scaling factor K. i ; Step S 22 The input data is in the form of (x0, y0). Set the first parameter value x0 of the data stored in register reg2, and set the second parameter value y0 of the data stored in register reg3. Step S 23 The parameter values in registers reg2 and reg3 are shifted i bits to the right using a shift operation before being stored. Step S 24 The rotation sign d is determined based on the parameter value in register reg3. i ; Step S 25 Iterative calculation to determine the rotation sign d i The value, and based on the rotation symbol d i The value of is updated to update the parameter values stored in registers reg2 and reg3; Step S 26 Calculate the data modulus according to the following formula. In the formula, x i The modulus value after scaling is calculated after i iterations.
7. A hardware acceleration method for data anomaly detection according to claim 6, characterized in that: Step S 23 Specifically, this means: shifting the parameter value in register reg2 to the right by i bits through a shift operation, and storing the result in register reg4; shifting the parameter value in register reg3 to the right by i bits through a shift operation, and storing the result in register reg5.
8. A hardware acceleration method for data anomaly detection according to claim 7, characterized in that: Step S 24 Specifically, if the parameter value in register reg3 is greater than 0, then the rotation sign d... i =-1, if the parameter value in register reg3 is less than or equal to 0, then rotate the sign d. i =+1.
9. A hardware acceleration method for data anomaly detection according to claim 8, characterized in that: Step S 25 Specifically, if the rotation symbol d i If the value is -1, then the result of adding register reg2 and register reg5 is assigned to register reg2, updating the parameter value stored in register reg2; the result of subtracting register reg4 from register reg3 is assigned to register reg3, updating the parameter value stored in register reg3; if the rotation sign d i =+1, then the result of subtracting register reg5 from register reg2 is assigned to register reg2, updating the parameter value stored in register reg2; the result of adding register reg3 to register reg4 is assigned to register reg3, updating the parameter value stored in register reg3.
10. A hardware acceleration device for data anomaly detection, characterized in that: It includes a hardware fast addition unit and a hardware modulo unit; the hardware fast addition unit is used to construct a pipelined adder and to process two sets of 16-bit wide input data a(x) a ,y a ), b(x b ,y b Perform two's complement addition to obtain x. a -x b ,y a -y b The calculation results; The pipelined adder construction method includes: cascading two 4-bit carry-lookahead adders to form an 8-bit carry-lookahead adder; combining the two 8-bit carry-lookahead adders with three registers to form a 16-bit pipelined adder; of the two 8-bit carry-lookahead adders, one is used for low-bit addition calculation, and the other is used for high-bit addition calculation; the three registers are used to adjust the synchronization timing to form a pipelined structure. The hardware modulus unit is used to solve for the modulus of the data by iteratively calculating the original data using the vector pattern of the Cordic algorithm.
Citation Information
Patent Citations
64-bit fixed-point ALU (arithmetic logical unit) circuit based on three-stage carry lookahead adder in GPDSP
CN105335128A
Lithium battery pack consistency detection method and device based on local outlier factor
CN113049963A
SM3 data encryption method and related device
CN113630236A
Data outlier isolated point detection method and system, computer equipment and storage medium
CN115827932A