Message decoding device and method and hardware acceleration system
Through hardware optimization of BDD algorithm, hybrid computing arrays and multi-stage pipeline technology are adopted to solve the problem of high computational complexity in high-dimensional grid spaces in traditional message decoding methods, and efficient and low-latency hardware accelerated decoding is achieved, which is suitable for quantum computing environments.
Patent Information
- Application Number
- CN202510480817.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional message decoding methods have high computational complexity, high resource consumption and slow speed in high-dimensional grid spaces, especially when it is difficult to achieve efficient information restoration under the threat of quantum computing.
The bounded distance decoding algorithm is optimized by hybrid computing arrays and expansion factor, and the BDD algorithm is optimized through hardware architecture, parallel layer and iterative layer are introduced, combined with multi-stage pipeline technology, square root operation is omitted, floating-point operation is reduced, quad-value addition and subtraction units are multiplexed, and hardware resources are dynamically adjusted.
It reduces the computational complexity and hardware resource overhead, improves decoding efficiency and adaptability, and realizes low latency and high parallel hardware acceleration, which is suitable for secure decoding of post-quantum cryptography algorithms.
Smart Images

Figure CN120342620A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security technology, and in particular, to a message decoding device, method, and hardware acceleration system. Background Art
[0002] Traditional decoding methods, such as maximum likelihood decoding, are difficult to be practical due to the computational complexity of high-dimensional lattice spaces and need to be combined with specific decoding strategies, such as bounded distance decoding (BDD), to balance between accuracy and efficiency. Such decoding needs to handle the complex geometric relationships between lattice points and cope with noise interference to ensure reliable and efficient information restoration under the threat of quantum computing.
[0003] BDD is closely related to intractable problems in lattice-based cryptography and provides a theoretical basis for the security of lattice-based cryptosystems. For example, BDD and the unique shortest vector problem (uSVP) are equivalent in computational complexity, which provides a basis for constructing secure encryption algorithms. In the lattice-based cryptosystem, message decoding is a key link in the decryption process of structured / unstructured lattices. Since lattice-based cryptography relies on high-dimensional mathematical lattice structures to resist quantum attacks, its decryption needs to recover the original information from lattice vectors with noise. The message decoding algorithm can accurately recover the original message from the lattice point matrix data with noise interference. By adopting bounded distance decoding technology, it can efficiently find the lattice point closest to the target vector in the high-dimensional lattice space, and then further extract the original binary message from the lattice point by de-tagging, and eliminate coding redundancy and noise through hierarchical decomposition and modular arithmetic.
[0004] However, in software implementation, the use of a recursive divide-and-conquer strategy, recursive calls, and stack operations will introduce additional time delays and generate high computational cycles, resulting in problems such as high computational complexity, large resource consumption, and slow speed. Summary of the Invention
[0005] This application provides a message decoding device, method, and hardware acceleration system to solve the problems of high computational complexity, large resource consumption, and slow speed existing in the message decoding method.
[0006] In a first aspect, this application provides a message decoding device, including:
[0007] A storage unit, including multiple registers, for decomposing and storing an input data matrix;
[0008] A preprocessing unit, for converting the input data matrix into a complex vector;
[0009] A decoding unit, comprising a hybrid computing array and a computing distance module. The hybrid computing array is configured to perform a recursive operation on the complex vector by means of a bounded distance decoding algorithm and based on an expansion factor to output a target vector, and the expansion factor is used to determine the number of iterations of the hybrid computing array; the computing distance module includes a computing module, a comparison module, and a tree adder. The computing module is configured to calculate the Euclidean distance based on the target vector, the comparison module is configured to determine a target lattice point based on the Euclidean distance, and the tree adder is used to assist the computing module in calculating the Euclidean distance;
[0010] A post-processing unit, configured to extract a decoding result based on the target lattice point.
[0011] In some feasible embodiments, the preprocessing unit is further configured to:
[0012] Obtain an input data matrix and extract a vector based on the input data matrix;
[0013] Perform a scaling process on the vector, and perform a block process on the scaled vector to obtain a complex vector.
[0014] In some feasible embodiments, the hybrid computing array includes a parallel layer and an iterative layer;
[0015] The parallel layer includes a first number of four-dimensional logic units, and the four-dimensional logic units are configured to perform a recursive operation on the complex vector by means of a bounded distance decoding algorithm to generate a first target vector, a second target vector, a third target vector, and a fourth target vector;
[0016] The iterative layer includes a preset dimension logic unit, and the preset dimension logic unit is configured to perform a recursive operation on the first target vector, the second target vector, the third target vector, and the fourth target vector by means of a bounded distance decoding algorithm to generate a target vector, and the target vector is a preset dimension lattice point vector.
[0017] In some feasible embodiments, the post-processing unit includes: a second number of three-subtraction modules, a second number of subtract-and-add modules, and an extraction module;
[0018] The three-subtraction module is configured to perform a subtraction operation on the imaginary part of the preset dimension lattice point vector to output an imaginary part result;
[0019] The subtract-and-add module is configured to perform two subtraction and one addition operations on the real part of the preset dimension lattice point vector to output a real part result;
[0020] The extraction module is configured to output a decoding result based on the imaginary part result and the real part result, and the decoding result is a binary message.
[0021] In some feasible embodiments, the expansion factor of the parallel layer is a third quantity, the number of rounds of recursive operations of the iterative layer is a fourth quantity, and the fourth quantity is twice the third quantity;
[0022] The hybrid computing array further includes a splitting layer, and the splitting layer is configured to split the computing processes of the parallel layer and the iterative layer into a fifth quantity of pipeline stages.
[0023] In some feasible embodiments, the four-dimensional logic unit is configured to perform recursive operations on the complex vector based on the bounded distance decoding algorithm to generate a first target vector, a second target vector, a third target vector, and a fourth target vector, and is specifically configured as follows:
[0024] Split the complex vector into a first complex vector and a second complex vector;
[0025] Recursively perform the bounded distance decoding algorithm on the first complex vector and the second complex vector to obtain dimension information, where the dimension information is the dimension information of the first complex vector and the second complex vector;
[0026] If the dimension in the dimension information reaches a preset rule, output a sub-vector result;
[0027] Output a first target vector based on the sub-vector result.
[0028] In some feasible embodiments, the post-processing unit further includes an arithmetic logic unit, and the arithmetic logic unit is configured to perform modulo operations and masking operations in the decoding unit.
[0029] In some feasible embodiments, the computing module is configured to perform Euclidean distance calculation based on the first target vector, the second target vector, the third target vector, the fourth target vector, and the target vector to obtain a distance result;
[0030] The comparison module is configured to generate a comparison result based on the distance result, and obtain a target lattice point based on the comparison result, where the target lattice point is the lattice point of the vector with the smallest Euclidean distance;
[0031] The number of tree adders is two, the tree adder has a binary tree structure, and the number of levels of the tree adder is the logarithm of the dimension of the target vector, and is used to perform addition operations in the Euclidean distance calculation hierarchically.
[0032] In a second aspect, the present application provides a message decoding method, including:
[0033] Obtain an input data matrix, where the input data matrix is stored after decomposition;
[0034] The input data matrix is converted into a complex vector by a preprocessing unit;
[0035] Based on the hybrid computing array, a recursive operation is performed on the complex vector based on the bounded distance decoding algorithm to output a target vector;
[0036] The Euclidean distance is calculated by a calculation module based on the target vector. The calculation distance module includes a calculation module, a comparison module, and a tree adder. The comparison module is used to determine a target lattice point based on the Euclidean distance, and the tree adder is used to assist the calculation module in calculating the Euclidean distance;
[0037] The target lattice point is extracted as a decoding result by a postprocessing unit.
[0038] In a third aspect, the present application provides a hardware acceleration system, including a housing and the message decoding device. The housing has a cavity, and the message decoding device is disposed in the cavity.
[0039] As can be seen from the above technical solutions, the present application provides a message decoding device, method, and hardware acceleration system. The device includes a storage unit, a preprocessing unit, a decoding unit, and a postprocessing unit. The storage unit includes a plurality of registers for decomposing and storing an input data matrix. The preprocessing unit is used to convert the input data matrix into a complex vector. The decoding unit includes a hybrid computing array and a calculation distance module. The hybrid computing array is used to perform a recursive operation on the complex vector based on the bounded distance decoding algorithm and based on an expansion factor to output a target vector. The expansion factor is used to determine the number of iterations of the hybrid computing array. The calculation distance module includes a calculation module, a comparison module, and a tree adder. The calculation module is used to calculate the Euclidean distance based on the target vector. The comparison module is used to determine a target lattice point, and the tree adder is used to assist the calculation module in calculating the Euclidean distance. The postprocessing unit is used to extract a decoding result based on the target lattice point. By optimizing the BDD algorithm in the message decoding process and introducing a hybrid computing array and an expansion factor, the device can reduce the computational complexity and hardware resource overhead. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a schematic diagram of the BDD algorithm provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic diagram of the de-tagging algorithm provided by an embodiment of the present application;
[0043] Figure 3 Schematic diagram of the message decoding algorithm provided by the embodiments of the present application;
[0044] Figure 4 Schematic diagram of the structure of the message decoder provided by the embodiments of the present application;
[0045] Figure 5 Schematic diagram of the structure of the hybrid computing array provided by the embodiments of the present application;
[0046] Figure 6 Schematic flow diagram of the message decoding method provided by the embodiments of the present application. Detailed implementation manners
[0047] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0048] Based on quantum computing, public key cryptosystems, such as RSA (Rivest-Shamir-Adleman) and ECC (Elliptic Curve Cryptography), may be threatened, while lattice cryptography has the ability to resist quantum computing. The post-quantum lattice cryptography algorithm Scloud + is a key encapsulation mechanism based on the unstructured learning with errors problem. This algorithm can significantly improve the computing and communication efficiency by using ternary secrets and lattice coding.
[0049] The sender encodes the message into an encryption matrix for secure communication between the two parties. The receiver uses the corresponding private key and the message decoder to calculate the original message from the encryption matrix. However, the message decoder of the lattice-based cryptography scheme occupies most of the computational complexity and resource cost of the message receiver. Therefore, the decoding process may become a performance bottleneck for the decryptor to decrypt quickly.
[0050] In the post-quantum cryptography algorithm, the unstructured lattice Scloud + The decryption process of the cryptography algorithm adopts an efficient message decoding algorithm to cope with the potential threats brought by quantum computers. By using the bounded distance decoding (BDD) technology, such as Figure 1As shown, Algorithm 1 can find the lattice point closest to the target vector in the high-dimensional lattice space, and then further extract the original binary message from the lattice point through the Delabeling Algorithm 2. The encoding redundancy and noise are eliminated through hierarchical decomposition and modular arithmetic. The message decoding Algorithm 3 accurately recovers the original message from the lattice matrix data disturbed by noise.
[0051] Specifically, for Algorithm 1, the input of Algorithm 1 is the target vector and the Barnes-Wall lattice BW n , where n is the dimension of the lattice. First, it is detected whether n is equal to 2, which is the termination condition of the recursion. Because in the low dimension, that is, when n = 2, the decoding calculation is relatively simple. Because in the two-dimensional case, the lattice point closest to the target vector can be obtained through a simple rounding operation. If n = 2, the rounded value of the target vector is directly returned.
[0052] When n > 2, the target vector is first split into a first vector and a second vector, each part with a dimension of n / 2. The BDD algorithm is recursively called for the first vector and the second vector respectively. When calculating in the Barnes-Wall lattice BW n / 2 with a dimension of n / 2, the corresponding first sub-lattice vector and second sub-lattice vector are obtained. This process is looped until the dimension is reduced to 2, triggering the termination condition.
[0053] The first correction term z1 and the second correction term z2 are calculated respectively, where z1 and z2 are calculated by the following formula:
[0054]
[0055] where, is a certain specific transformation operation, y1 is the first sub-lattice vector, and y2 is the second sub-lattice vector.
[0056] Compare the magnitudes of the two distance metrics ‖y1 - t1‖ 2 + 2‖z1‖ 2 and ‖y2 - t2‖ 2 + 2‖z2‖ 2 If ‖y1 - t1‖ 2 + 2‖z1‖ 2 < ‖y2 - t2‖ 2 + 2‖z2‖ 2 , then return the combined vector Otherwise, return
[0057] Through the above steps, this algorithm can find the lattice vector closest to the target vector in the Barnes-Wall lattice, realizing the function of bounded distance decoding.
[0058] For Algorithm 2, which is used to extract binary messages from lattice vectors, the input consists of three positive integers n, τ, and μ, and the following conditions are satisfied. These parameters are used to define the computational range and related constraints of the algorithm:
[0059] n = 2 k ≥ 4;
[0060]
[0061] The input also includes a lattice vector w ∈ C, where C is based on a nested lattice That is, the lattice vector comes from a specific nested lattice structure. The output is a binary message m ∈ {0, 1} μ .
[0062] The algorithm process includes an outer loop and an inner loop. The outer loop processes the vector dimension, and the inner loop processes the complex vector. The outer loop iterates over the variable l from k - 1 to 1. In each iteration, refer to Figure 2 Step 2 in, divide the lattice vector according to specific rules, and then update the lattice vector. In the way of Step 3, process the lattice vector step by step through this update method.
[0063] After the outer loop ends, enter the inner loop, transform the representation of the lattice vector, refer to Figure 2 Step 5 in, then iterate over the variable j in the way of Step 6, and then perform message extraction through Step 11, where is a specific inverse transformation operation. Then intercept the first μ bits of the vector m ′ as the final binary information m. Through the above steps, the de-tagging algorithm can successfully extract the required binary message from the given lattice vector.
[0064] For Algorithm 3, which is used to recover the original binary message from the encoded integer matrix, as Figure 3 in the input and input, the algorithm flattens the input encoded integer matrix M into a vector y, and then through vector truncation, vector scaling, block processing, and then cyclic decoding, and then combines all the extracted binary message blocks m y into the final binary message and returns it. Through the above steps, the message decoding algorithm can gradually recover the original binary message from the encoded integer matrix.
[0065] However, in software implementation, using the recursive divide-and-conquer strategy, recursive calls and stack operations will introduce additional time delays, and at the same time generate high computing cycles, and there are problems such as difficulty in fully utilizing multi-core parallelism, low throughput, resource waste, and high latency.
[0066] To solve the above problems, some embodiments of the present application provide a message decoding device. This device optimizes the hardware architecture of the BDD. Through a hybrid computing array instead of a single recursively expanded fully parallel architecture, and through a multi-stage pipelining technique, the throughput of the message decoder can be significantly improved. And by reconstructing the formula, the square root operation in the original algorithm is omitted, thereby avoiding more complex floating-point operations and reducing the consumption of required hardware resources. During the de-tagging process, the resource consumption is further reduced by multiplexing the four-value addition and subtraction unit multiple times. For the three different security levels of Scloud + , the device can be adapted by dynamically adjusting parameters through registers.
[0067] As Figure 4 shown, the device includes a storage unit, a preprocessing unit, a decoding unit, and a post-processing unit. The storage unit includes multiple registers. For example, the storage unit is a 32-bit register bank REG_0. The register bank is used to store the input data matrix M. q is a modulus parameter for scaling operations. The input data matrix is decomposed and stored in multiple registers, storing the high-order data M_high with a width of m×π / 16, where π represents the number of real parts after splitting the complex dimension. M_high and the m parameter are combined into the data stream to be decoded. Specifically, the input data matrix is split into sub-blocks according to the dimension and allocated to different registers for storage. In some embodiments, a data acquisition module is further included, and the decomposed input data matrix is extracted by the data acquisition module and prepared for processing.
[0068] The preprocessing unit is used to convert the input data matrix into a complex vector, including flattening, scaling, and partitioning into negative pairs. Among them, the input data matrix is a two-dimensional matrix. In some embodiments, the preprocessing unit is also applied to obtain the input data matrix, extract vectors based on the input data matrix, perform scaling processing on the vectors, and then perform partitioning processing on the scaled vectors to obtain complex vectors.
[0069] Specifically, the input data matrix is expanded into a one-dimensional vector by rows or columns, and then the elements in the vector are numerically adjusted to map the input data matrix to a preset numerical range, that is, each element is multiplied by a scaling factor according to the preset security parameter and modulus. The preset numerical range is a numerical range suitable for lattice cryptography decoding. Then the scaled vector is divided into multiple sub-blocks of a fixed length. In each group of sub-blocks, two adjacent elements are respectively used as the real part and the imaginary part of the complex number to generate a complex vector. By processing the input data matrix, the geometric relationship of the data can be retained, the lattice point matching calculation can be simplified, and floating-point operations can be avoided, reducing the consumption of hardware resources.
[0070] The decoding unit is used to locate the closest lattice point from a complex vector, i.e., the target vector or target lattice point. The decoding unit includes a hybrid computing array and a distance calculation module. The hybrid computing array is used to perform recursive operations on the complex vector by the bounded distance decoding algorithm and based on the expansion factor to output the target vector. The distance calculation module includes a calculation module, a comparison module, and a tree adder. The calculation module is used to calculate the Euclidean distance based on the target vector. The comparison module is used to determine the target lattice point. The tree adder is used to assist the calculation module in calculating the Euclidean distance.
[0071] In some embodiments, the number of tree adders is two. The tree adder has a binary tree structure. The number of levels of the tree adder is the logarithm of the dimension of the target vector and is used to perform the addition operation in the Euclidean distance calculation hierarchically.
[0072] As Figure 5 shown, in some embodiments, the hybrid computing array includes a parallel layer and an iterative layer. The parallel layer includes a first number of four-dimensional logic units BDD4, which are used to process the bounded distance decoding tasks of four-dimensional complex vectors. The input four-dimensional complex vector is split into lower-dimensional sub-vectors, the nearest lattice points are calculated in parallel, and the recursive algorithm is implemented through pure combinational logic to eliminate the stack overhead and latency of software recursion, and the candidate lattice points and correction terms are output for iterative integration by the iterative layer.
[0073] Among them, the expansion factor is used to determine the number of iterations of the hybrid computing array, indicating the hierarchical depth at which the BDD algorithm is expanded in hardware implementation. In some embodiments, the expansion factor of the parallel layer is a third number, and the number of rounds of recursive operations of the iterative layer is a fourth number. The fourth number is twice the third number. The hybrid computing array further includes a split layer, and the split layer is used to split the calculation processes of the parallel layer and the iterative layer into a fifth number of pipeline stages.
[0074] The third number is 8, and the fifth number is 6. That is to say, the number of iterations is 16 times, which can balance resources and speed. The number of pipeline stages is six. In the parallel layer, 4 BW4 units at the bottom layer calculate in parallel to generate 4 groups of sub-results. In the iterative layer, the upper layer completes the BW8 / BW 16 calculation through 16 iterations and reuses the BW4 units. The number of BW2 units is reduced from 256 to 16, a reduction of 94%. The total area is reduced to 25% of the fully parallel case. Inserting a six-stage pipeline reduces the critical path from 10 ns, i.e., combinational logic, to 2 ns, i.e., single cycle, and the frequency is increased to 500 MHz.
[0075] According to the path splitting of BDD 32 such as BW2→BW4→BW8→EdC2→EdC4→EdC8. It can be understood that different expansion factors can also be set as shown in the following table:
[0076]
[0077] When the expansion factor is 8, the product A×T of the area A and the number of cycles T is minimized, reaching Pareto optimality. The A×T values of fully parallel (A = 100, T = 1) and fully serial (A = 10, T = 256) are 100 and 2560 respectively, while the A×T of the expansion factor 8 is 50×16 = 800, which is better.
[0078] It can also be set to support multiple security levels, such as a flexible decoder of NIST Level 1 - 5. When at a high security level, the expansion factor 16 is used to increase the parallelism. When at a low security level, it switches to the expansion factor 8 to save power, enabling dynamic resource allocation.
[0079] The four - dimensional logic unit is used to perform recursive operations on the complex vector based on the bounded - distance decoding algorithm. Among them, the first quantity is four, and the four four - dimensional logic units process different four - dimensional inputs simultaneously. Each unit independently completes the splitting and decoding tasks, and generates the first target vector, the second target vector, the third target vector, and the fourth target vector respectively.
[0080] In some embodiments, the complex vector is split into a first complex vector and a second complex vector, and then the bounded - distance decoding algorithm is recursively executed on the first complex vector and the second complex vector to obtain dimension information, where the dimension information is the dimension information of the first complex vector and the second complex vector; if the dimension in the dimension information reaches a preset rule, a sub - vector result is output; and the first target vector is output based on the sub - vector result.
[0081] Taking the first target vector as an example, BDD4 first horizontally splits the four - dimensional vector t into two two - dimensional sub - vectors t1 and t2. For t1 and t2, the two - dimensional lattice points y1 and y2 are respectively output through the BDD algorithm, and then the closest points in the two - dimensional lattice space are calculated. Linear transformations are performed on the two - dimensional lattice points y1 and y2 to generate z1 and z2, that is, through operations. At this time, The operation represents scaling and rotation in the complex domain. Calculate the squared distances of two paths. The first path distance is ‖y1 - t1‖ 2 +2‖z1‖ 2 , and the second path distance is ‖y2 - t2‖ 2 +2‖z2‖ 2 . In this process, the squared differences of each dimension can be hierarchically accumulated through a tree - shaped adder, and the path with a smaller distance is selected and merged into a four - dimensional lattice point vector, that is, the first target vector.
[0082] In some embodiments, the four-dimensional logic unit is further configured to output a correction term, where the correction term is used to represent the adjustment amount between word vectors. In the iterative layer, if the first path is selected, the correction term passes through z1 to perform recursive correction on higher-dimensional vectors. BDD4 can achieve low-latency and high-parallel hardware acceleration while ensuring accuracy.
[0083] The iterative layer includes a preset-dimension logic unit, which is configured to perform recursive operations on the first target vector, the second target vector, the third target vector, and the fourth target vector based on the bounded distance decoding algorithm to generate a target vector, where the target vector is a preset-dimension lattice point vector.
[0084] In this embodiment, the preset-dimension logic unit is BDD 32 , that is to say, the preset dimension is 32. It can be understood that if the preset dimension is 32 dimensions, by iteratively reusing BDD4, BDD8 and BDD 16 are constructed. For BDD8, two four-dimensional sub-blocks are merged into an eight-dimensional vector. In the first round of iteration, four BDD4 processing sub-blocks 1-4 output an intermediate result temp[1]. The temporary register group REG_2 is used to store the intermediate result. The width of the intermediate result is the same as that of the register group, and the depth is min(m,n) / 32, where min(m,n) represents the smaller dimension of the input matrix. In the second round of iteration, four BDD4 processing sub-blocks 5-8 output an intermediate result.
[0085] Temporary registers (temp[0]-temp
[31] ), real part: temp[2×i+0] (such as temp[0], temp[2], temp[4], etc.). Imaginary part: temp[2×i+1] (such as temp[1], temp[3], temp[5], etc.), which are used to store the complex vector intermediate results at different recursive levels and support separate real and imaginary part calculations.
[0086] Through the post-processing unit, a decoding result is generated based on the target vector, and a binary message is extracted from the target lattice. The input of the post-processing unit is the target lattice w j ∈BW 32 , and the post-processing unit includes a three-subtraction module, a subtract-add module, and an extraction module.
[0087] To merge two intermediate results, in some embodiments, through a second number of subtract-add modules, the subtract-add module is configured to perform two subtractions and one addition operation on the real part of the preset-dimension lattice point vector to output a real part result.
[0088] Exemplarily, the intermediate result is a four-dimensional lattice vector. For example, the real part components of the intermediate result temp[1] are [a1, a2, a3, a4], and the real part components of another intermediate result are [b1, b2, b3, b4]. Then, the subtraction / addition module calculates the difference vector through the following formula:
[0089] Δ Re = [a1 - b1, a2 - b2, a3 - b3, a4 - b4];
[0090] Then, the difference vector is superimposed with the correction value through the following formula:
[0091]
[0092] where, is the real part after the linear transformation of the correction value of the intermediate result A.
[0093] The output of the subtraction / addition module is the combined four-dimensional real part vector, that is, the real part result.
[0094] For the output imaginary part result, in some embodiments, the triple subtraction module is used to perform the subtraction operation of the imaginary part of the dimensional lattice vector. The triple subtraction module corresponds to the complex vector update operation in step three of algorithm 2, that is, update where, the imaginary part calculation of the complex subtraction needs to be decomposed into three subtractions.
[0095] Specifically, the input of the triple subtraction module is the adjacent complex pair w j = a + bi and w j+1 = c + di (real parts a, c, imaginary parts b, d). The imaginary part result is calculated through the following formula:
[0096] [(d - b) - (c - a)] / 2;
[0097] The corresponding three subtraction operations can be calculated by three stages of subtractors, that is, the first-stage subtractor, the second-stage subtractor, and the third-stage subtractor. The triple subtraction module is multiplexed by a pipeline, and each module processes adjacent complex numbers. Among them, d - b is the imaginary part difference between w j+1 and w j , c - a is the real part difference between w j+1 and w j , and (d - b) - (c - a) is the subtraction operation between the imaginary part difference and the real part difference. Then, fixed-point processing is performed by shifting one bit to the right, that is, by the ratio with 2.
[0098] It can be expanded through calculation operations to obtain the following formula:
[0099]
[0100] That is to say, the real part calculation section includes two subtractions and one addition, and the imaginary part calculation section includes three subtractions. In this embodiment, eight three-subtraction modules and eight subtract-and-add modules are reused. For different security levels, in this embodiment, the total number of cycles consumed by the message decoding device is
[0101] The distance calculation module reduces the computational complexity by omitting the square root and only comparing the squared distances, while ensuring accuracy. By reusing the subtractor chain, independent calculation of the imaginary part for each complex number pair is avoided, reducing the number of subtractors by 50%. By replacing division with fixed-point right shift, floating-point operations are avoided while retaining the rounding of the least significant bit.
[0102] See again Figure 1 The Euclidean distance calculation and lattice point selection are gradually completed from temp
[29] to temp
[31] .
[0103] In some embodiments, the post-processing unit includes an arithmetic logic unit (ALU), and the arithmetic logic unit is used to perform modular operations and masking operations in the decoding unit.
[0104] For the modular operation, the inputs are the numerical values a (real part) and b (imaginary part), and the modulus Calculate the mask value mask = m - 1 according to m, and then perform a bitwise AND operation to output the truncated a' and b'.
[0105] Then, a binary message is generated through a masking operation and a bit selector. For the masking operation, the input is the complex number that has been processed by the modular operation and the target bit width, and then bit screenshot and splicing operations are performed, that is, the first μ bits m of the binary message block temp
[31] are extracted from w through Algorithm 3 j ∈ {0, 1} j as the decryption result, all the message block splicing bits are finally output as m = (m1,..., m μ ) lm / μ )
[0106] In some embodiments, the recursive algorithm can also be fully expanded into combinational logic, and all computing tasks are executed in parallel without iterative reuse. That is, at each recursive level, such as BDD4 and BDD8, hardware modules are independently deployed without pipeline registers and are implemented only through combinational logic.
[0107] The message decoding device provided in this embodiment can achieve the best balance between hardware resources and computing efficiency while ensuring post-quantum security. And while improving the decoding accuracy, this device also provides the possibility for an efficient and low-cost hardware implementation, has broad application prospects, and particularly has important technical advantages in fields such as communication and data transmission.
[0108] Such as Figure 6As shown in the figure, based on the above-mentioned message decoding device, some embodiments of the present application further provide a message decoding method, including:
[0109] S100: Obtain an input data matrix.
[0110] The input data matrix is stored after being decomposed.
[0111] S200: Convert the input data matrix into a complex vector through a preprocessing unit.
[0112] S300: Perform a recursive operation on the complex vector based on the bounded distance decoding algorithm by a hybrid computing array to output a target vector;
[0113] S400: Calculate the Euclidean distance based on the target vector through a calculation module.
[0114] S500: Determine a target lattice point based on the Euclidean distance through a comparison module.
[0115] S600: Extract the target lattice point as a decoding result through a post-processing unit.
[0116] For the effects during the operation of this embodiment, reference can be made to the effects of the above-mentioned device embodiment, which will not be elaborated here.
[0117] Based on the above-mentioned message decoding device, some embodiments of the present application further provide a hardware acceleration system, including a housing and the message decoding device, the housing has a cavity, and the message decoding device is arranged in the cavity.
[0118] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application, and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other embodiments extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.
Claims
1. A message decoding device, characterized in that, Including: A storage unit including multiple registers for decomposing and storing an input data matrix; A preprocessing unit for converting the input data matrix into a complex vector; A decoding unit including a hybrid computing array and a distance calculation module. The hybrid computing array is used to perform recursive operations on the complex vector based on a bounded distance decoding algorithm and an expansion factor to output a target vector. The expansion factor is used to determine the number of iterations of the hybrid computing array. The distance calculation module includes a calculation module, a comparison module, and a tree adder. The calculation module is used to calculate the Euclidean distance based on the target vector. The comparison module is used to determine a target lattice point based on the Euclidean distance. The tree adder is used to assist the calculation module in calculating the Euclidean distance; A postprocessing unit for extracting a decoding result based on the target lattice point.
2. The message decoding device according to claim 1, characterized in that, The preprocessing unit is further used for: Obtaining an input data matrix and extracting vectors based on the input data matrix; Performing a scaling process on the vectors, and then performing a block process on the scaled vectors to obtain a complex vector.
3. The message decoding device according to claim 1, wherein The hybrid computing array includes a parallel layer and an iterative layer; The parallel layer includes a first number of four-dimensional logic units for performing recursive operations on the complex vector based on a bounded distance decoding algorithm to generate a first target vector, a second target vector, a third target vector, and a fourth target vector; The iterative layer includes a preset dimension logic unit for performing recursive operations on the first target vector, the second target vector, the third target vector, and the fourth target vector based on a bounded distance decoding algorithm to generate a target vector, and the target vector is a preset dimension lattice point vector.
4. The message decoding device according to claim 3, wherein The postprocessing unit includes: a second number of three-minus modules, a second number of subtract-and-add modules, and an extraction module; The three-minus module is used to perform a subtraction operation on the imaginary part of the preset dimension lattice point vector to output an imaginary part result; The subtract-and-add module is used to perform two subtraction and one addition operations on the real part of the preset dimension lattice point vector to output a real part result; The extraction module is used to output a decoding result based on the imaginary part result and the real part result, and the decoding result is a binary message.
5. The message decoding device according to claim 3, wherein The expansion factor of the parallel layer is a third number, and the number of rounds of recursive operations of the iterative layer is a fourth number, and the fourth number is twice the third number; The hybrid computing array further includes a split layer for splitting the calculation processes of the parallel layer and the iterative layer into a fifth number of pipeline stages.
6. The message decoding device according to claim 3, wherein The four-dimensional logic unit is used to perform recursive operations on the complex vector based on a bounded distance decoding algorithm to generate a first target vector, a second target vector, a third target vector, and a fourth target vector, and is specifically configured to: Split the complex vector into a first complex vector and a second complex vector; Recursively perform a bounded distance decoding algorithm on the first complex vector and the second complex vector to obtain dimension information, and the dimension information is the dimension information of the first complex vector and the second complex vector; If the dimension in the dimension information reaches a preset rule, output a sub-vector result; Output a first target vector based on the sub-vector result.
7. The message decoding device according to claim 1, wherein The post-processing unit further includes an arithmetic logic unit, which is used to perform modulo operations and masking operations in the decoding unit.
8. The message decoding device according to claim 1, wherein The calculation module is used to perform Euclidean distance calculation based on the first target vector, the second target vector, the third target vector, the fourth target vector, and the target vector to obtain a distance result; The comparison module is used to generate a comparison result based on the distance result, and obtain a target lattice point based on the comparison result, where the target lattice point is the lattice point of the vector with the smallest Euclidean distance; The number of the tree adders is two, the tree adder is a binary tree structure, and the number of levels of the tree adder is the logarithm of the dimension of the target vector, which is used to perform the addition operation in the Euclidean distance calculation in a hierarchical manner.
9. A message decoding method, characterized in that, Comprising: Obtain an input data matrix, which is stored after decomposition; Convert the input data matrix into a complex vector through a preprocessing unit; Perform a recursive operation on the complex vector based on a bounded distance decoding algorithm by a hybrid computing array to output a target vector; Calculate the Euclidean distance based on the target vector through a calculation module, the calculation distance module includes a calculation module, a comparison module, and a tree adder, the comparison module is used to determine a target lattice point based on the Euclidean distance, and the tree adder is used to assist the calculation module in calculating the Euclidean distance; Extract the target lattice point as a decoding result through a post-processing unit.
10. A hardware acceleration system, characterized in that Comprising a housing and the message decoding device according to any one of claims 1-8, the housing has a cavity, and the message decoding device is arranged in the cavity.