Message coding method and device

By processing the addition and subtraction units and addition units of the chunked vector in parallel, tensor product operations are performed, grid vectors are generated and encoding matrix is constructed, which solves the calculation efficiency and resource consumption problems of the structureless LWE scheme and realizes an efficient message encoding process.

CN120342619APending Publication Date: 2025-07-18NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510480798.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing structureless LWE schemes have shortcomings in computing efficiency and resource consumption, which are difficult to meet the needs of lightweight. The complexity of high-dimensional grid structure leads to low coding efficiency and high hardware implementation difficulty.

Method used

By obtaining chunked vectors, tensor product operations are performed using the parallel architecture of addition and subtraction units and addition units, grid vectors are generated, and target vectors are output, and encoding matrix is finally constructed. The parallel architecture is used to process real and imaginary computing units to improve computing efficiency and reduce hardware resource overhead.

Benefits of technology

It significantly improves computing efficiency, reduces hardware resource consumption, and meets the requirements of post-quantum cryptography system for real-time and resource efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342619A_ABST
    Figure CN120342619A_ABST
Patent Text Reader

Abstract

The invention provides a message coding method and device, and the method comprises the steps: firstly obtaining a block vector after block processing, and then carrying out the coding operation to generate a lattice vector. And on the basis of the lattice vector, by means of a first preset group number operation module comprising an addition and subtraction unit and an addition unit, executing tensor product parallel calculation, outputting a target vector, and outputting a coding matrix according to the target vector. According to the method, real part and imaginary part operation units of tensor product operation in a message coding process are innovatively multiplexed, and the two operation units are constructed by adopting a parallel architecture, so that the problems that high-dimensional complex number tensor product operation cannot be effectively parallelized, hardware resources are redundant and occupied due to real part and imaginary part operation and the like in an existing software implementation scheme are solved; the computing efficiency can be remarkably improved, the hardware resource overhead is greatly reduced, and powerful support is provided for efficient operation of the post quantum cryptography system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information security technology, and in particular, to a message encoding method and apparatus. Background Art

[0002] With the rapid development of quantum computing technology, traditional public-key cryptography algorithms (such as RSA and Elliptic Curve Cryptography ECC) are facing the potential threat of being cracked by quantum algorithms (such as Shor's algorithm). Therefore, researching cryptographic solutions that can resist quantum computing attacks has become an important task in the current field of cryptography. For this reason, cryptographic solutions based on the Learning with Errors (LWE) problem have emerged. The security of the LWE problem is based on the approximate shortest vector problem (SVP) and closest vector problem (CVP) on a lattice, and these problems are still considered difficult to solve in the quantum computing model. Therefore, LWE-based cryptographic schemes are regarded as key candidate schemes in post-quantum cryptography.

[0003] Existing LWE-based cryptographic schemes are divided into two categories: unstructured LWE and structured LWE. Compared with structured LWE, although unstructured LWE schemes have higher security, their computational and communication efficiency is lower, which limits their application in practical systems. Therefore, how to improve the computational efficiency while ensuring high security has become an important research direction for unstructured LWE schemes.

[0004] Message Encoding (MsgEnc) is a key step in the lattice-based cryptographic system. In the Scloud+ scheme of unstructured LWE, message encoding is used to map ciphertexts into vectors suitable for communication transmission. Since the lattice-based cryptographic system resists quantum attacks through high-dimensional lattice structures, the data needs to be mapped into these complex lattice structures during the message encoding process. Due to the high dimension and complexity of the lattice, traditional encoding methods may lead to low computational efficiency. Especially in unstructured LWE schemes, the encoding process involves complex linear operations and noise control, and has the following deficiencies: (1) The mapping in the high-dimensional lattice space needs to rely on the processing of complex geometric relationships, and the algorithm complexity is high. (2) The adaptability between message blocks and lattice structures is insufficient, resulting in encoding redundancy. (3) When implemented in hardware, it occupies a large amount of resources, and the hardware implementation difficulty is high, making it difficult to meet the lightweight requirements.

[0005] Under this background, there is an urgent need for an efficient, compact and unstructured LWE framework-adapted message encoding method and apparatus, which can significantly reduce the computational complexity and resource consumption while ensuring quantum-resistant security. Summary of the Invention

[0006] The present application provides a message encoding method and apparatus to solve the problems of high computational complexity, large resource consumption, and slow computational speed.

[0007] In a first aspect, the present application provides a message encoding method, including:

[0008] Obtain a segmented vector, where the segmented vector is a vector after segmentation;

[0009] Perform encoding on the segmented vector to generate a lattice vector;

[0010] Based on the lattice vector, perform tensor product parallel computing through a first preset number of operation modules to output a target vector, where the operation module includes an addition and subtraction unit and an addition and addition unit;

[0011] Output an encoding matrix through the target vector.

[0012] In some feasible embodiments, the performing encoding on the segmented vector to generate a lattice vector includes:

[0013] Pad the segmented vector to generate a segmented vector of a preset length;

[0014] Perform a recursive grouping operation on the segmented vector of the preset length through a tagging method, and generate a lattice vector of a preset dimension through modulo operation, where the lattice vector includes a real part and an imaginary part.

[0015] In some feasible embodiments, the obtaining a segmented vector includes:

[0016] Obtain the information to be encoded and the data width of the register;

[0017] Store the message vector in the register based on the data width;

[0018] Calculate a second preset number, where the second preset number is the ratio of the data width to the preset width;

[0019] Clip the message vector into a second preset number of segments with a preset width to output a segmented vector.

[0020] In some feasible embodiments, the performing tensor product parallel computing through a first preset number of operation modules based on the lattice vector to output a target vector includes:

[0021] Perform real part calculation of the first group of lattice vectors in the second preset number through the addition and subtraction unit to output a first real part result, and perform imaginary part calculation of the first group of lattice vectors in the preset number through the addition and addition unit to output a first imaginary part result, where the number of the first group of lattice vectors is a first number;

[0022] Based on the first real part result and the first imaginary part result, a first updated vector is obtained;

[0023] Store the first updated vector in a register to perform the next iteration with the updated vector.

[0024] In some feasible embodiments, the number of the addition-subtraction units and the addition-addition units is a second number;

[0025] The performing tensor product parallel calculation on the lattice vector through a first preset number of operation modules to output a target vector includes:

[0026] Regroup the first updated vector to obtain a third number of vectors, where the first number is a multiple of the third number;

[0027] Perform real part calculation on the third number of vectors through the addition-subtraction units to output a second real part result, and perform imaginary part calculation on the third number of vectors in the set of preset numbers through the addition-addition units to output a second imaginary part result;

[0028] Based on the second real part result and the second imaginary part result, a second updated vector is obtained.

[0029] In some feasible embodiments, the performing tensor product parallel calculation on the lattice vector through a first preset number of operation modules to output a target vector includes:

[0030] Regroup the second updated vector to obtain a fourth number of vectors, where the third number is a multiple of the fourth number;

[0031] Perform real part calculation on the fourth number of vectors through the addition-subtraction units to output a third real part result, and perform imaginary part calculation on the fourth number of vectors in the set of preset numbers through the addition-addition units to output a third imaginary part result;

[0032] Based on the third real part result and the third imaginary part result, a third updated vector is obtained.

[0033] In some feasible embodiments, the performing tensor product parallel calculation on the lattice vector through a first preset number of operation modules to output a target vector includes:

[0034] Regroup the third updated vector to obtain a fifth number of vectors, where the fourth number is a multiple of the fifth number;

[0035] Perform a modulus operation on the fifth number of vectors, and the output is the target vector.

[0036] In some feasible embodiments, the outputting an encoding matrix through the target vector includes:

[0037] Concatenate the target vectors to construct the original vector;

[0038] Pad the original vector to a preset dimension and perform a scaling operation to output the encoding matrix.

[0039] In a second aspect, the present application provides a message encoding device, including:

[0040] An acquisition module is configured to acquire block vectors, where the block vectors are vectors after being blocked;

[0041] A label processing module is configured to perform encoding on the block vectors to generate lattice vectors;

[0042] An operation module is configured to perform tensor product parallel computing based on the lattice vectors through a first preset number of operation modules to output updated vectors, where the operation module includes an addition and subtraction unit and an addition and addition unit;

[0043] A calculation module is configured to output an encoding matrix through the updated vectors.

[0044] In some feasible embodiments, the operation module further includes a multiplexer and a timing control unit;

[0045] The timing control unit is configured to: acquire vector processing information, where the vector processing information includes real part processing information and imaginary part processing information;

[0046] The multiplexer is configured to: acquire the vector processing information sent by the timing control unit, generate control signals, where the control signals include a start subtraction signal and a disable subtraction signal; and select input sources based on the control signals, where the input sources include an addition and subtraction unit and an addition and addition unit.

[0047] As can be seen from the above technical solutions, the present application provides a message encoding method and device. The method includes: acquiring block vectors, where the block vectors are vectors after being blocked, then performing encoding on the block vectors to generate lattice vectors, performing tensor product parallel computing based on the lattice vectors through a first preset number of operation modules to output target vectors, where the operation module includes an addition and subtraction unit and an addition and addition unit, and outputting an encoding matrix through the target vectors. By reusing the real part and imaginary part operation units of the tensor product operation in the message encoding process, and the real part and imaginary part operation units adopt a parallel architecture, the calculation efficiency can be improved and the hardware resource overhead can be reduced. Description of the Drawings

[0048] To more clearly illustrate the technical solution of the present application, the accompanying drawings required in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a schematic flowchart of the message encoding method provided by the embodiment of the present application;

[0050] Figure 2 It is a schematic flowchart of the iterative process provided by the embodiment of the present application;

[0051] Figure 3 It is a schematic structural diagram of the message encoding device provided by the embodiment of the present application. Specific embodiments

[0052] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0053] In the current field of information security, the development of quantum computing technology poses a severe challenge to traditional public-key cryptography algorithms. Classical algorithms represented by RSA and Elliptic Curve Cryptography (ECC) face the potential risk of being successfully cracked by quantum algorithms such as Shor's algorithm. Against this background, a cryptographic system based on the Learning with Errors (LWE) problem has emerged. It provides a solid guarantee for the security of cryptographic schemes by solving linear equations with accompanying noise.

[0054] The security of the LWE problem closely relies on the approximate Shortest Vector Problem (SVP) and Closest Vector Problem (CVP) on the lattice. These problems are also considered extremely challenging and difficult to effectively overcome in the quantum computing model. Cryptographic schemes based on LWE are mainly divided into two categories: unstructured-LWE and structured-LWE.

[0055] Compared with structured-LWE, although the unstructured-LWE scheme has higher security, it performs poorly in terms of computing and communication efficiency, severely restricting its wide application in practical systems.

[0056] Post-quantum lattice cryptography algorithm Scloud+ is a key encapsulation mechanism based on unstructured lattice LWE. By innovatively introducing ternary secrets and lattice coding techniques, it opens up a new path for improving the efficiency of unstructured lattice LWE schemes. The application of ternary secrets can effectively reduce the noise interference in the decryption process, while lattice coding further optimizes the parameter selection by enhancing the error correction ability. Through these technologies, Scloud + while maintaining high security, can significantly improve the computing and communication efficiency, and strongly promotes the application process of post-quantum cryptography in practical scenarios.

[0057] Message Encoding (MsgEnc) is a key step in lattice-based cryptographic systems, such as the unstructured lattice LWE-based Scloud+. In lattice-based cryptographic systems, the core goal is to accurately convert the original message into a lattice vector suitable for encryption requirements. Given that lattice-based cryptographic systems rely on high-dimensional mathematical lattice structures to resist quantum attacks, the message must be precisely mapped into these complex lattice structures during the encoding process.

[0058] Some encoding methods are limited by the complexity of high-dimensional lattice structures, resulting in low encoding efficiency.

[0059] Scloud + The message encoding process in Scloud combines ternary secrets and lattice coding techniques. By mapping the message into the Barnes-Wall lattice (BW32) for encoding. This method makes full use of the characteristics of tensor product operations. In the high-dimensional lattice space, the message can be efficiently converted into lattice point elements. In addition, to handle the complex geometric relationships in the high-dimensional lattice space, Scloud + adopts a specially optimized algorithm to reduce the computational complexity during the encoding process, thereby improving the efficiency and reducing the demand for hardware resources.

[0060] In this process, the message is split into multiple groups and encoded using the labeling method. By dividing the message data into blocks of different widths, mapping them into suitable lattice structures, and using the characteristics of ternary secrets to simplify the encoding operation. That is to say, message encoding not only needs to ensure the accuracy and security of encoding, but also needs to optimize the computational efficiency to meet the high requirements for the performance of the cryptographic system under the threat of quantum computing, enabling Scloud + while maintaining high security, can efficiently execute the message encoding process.

[0061] In post-quantum cryptographic algorithms, the unstructured lattice Scloud + The encryption process of the cryptographic algorithm depends on an efficient message encoding algorithm to cope with the potential threats brought by quantum computers. Scloud+ In the cryptographic algorithm, the core encoding process that maps a binary message to the Barnes-Wall lattice code space is the tagging method. The tagging method can improve the encoding efficiency and fault tolerance, achieve effective compression of parameters, reduce communication overhead, enhance quantum-resistant security, and demonstrate excellent advantages in improving efficiency, ensuring security, and reducing implementation costs.

[0062] Scloud + The message encoding method for the core encoding that realizes the mapping from the plaintext message to the lattice code space in the cryptographic algorithm can convert a binary message into a formatted matrix that meets the requirements of LWE encryption. This method unifies the design of error-correcting coding and the LWE encryption layer, avoids the efficiency loss of the discrete architecture of traditional schemes, and realizes the on-demand balance of security and efficiency through the flexible configuration of μ and τ. It shows significant advantages in improving noise robustness, parameter optimization and communication efficiency improvement, and quantum-resistant security design.

[0063] However, there are many limitations in the existing message encoding schemes implemented by software. The high-dimensional complex tensor product operations involved in the algorithm, such as the iterative operations in the Barnes-Wall lattice update, need to be processed recursively layer by layer. Since software is difficult to effectively parallelize the computing units, the computing speed is severely limited. At the same time, the architecture of separate real and imaginary part operations needs to repeatedly call independent hardware resources, resulting in redundant occupation of computing units and significantly increasing the hardware area overhead. These problems lead to low computing efficiency and huge resource consumption, and it is difficult to meet the strict requirements of the post-quantum cryptosystem for real-time performance and resource efficiency.

[0064] Some embodiments of the present application propose a message encoding method. This method can significantly improve the computing efficiency and at the same time greatly reduce the hardware resource overhead by reusing the real and imaginary part operation units of the tensor product operation in the message encoding process and constructing the real and imaginary part operation units with a parallel architecture. The method is expected to provide strong support for the efficient operation of the post-quantum cryptosystem, effectively solve the key technical problems faced in the current message encoding process, and promote the continuous development of the information security field in the era of quantum computing.

[0065] As Figure 1 shown, the message encoding method includes the following steps:

[0066] S100: Obtain a block vector.

[0067] The block vector is a vector after being blocked, which is a sub-vector obtained by cutting the original message vector m according to a preset width. To adapt to the register bit width, that is, the data width, and realize parallel processing, the message vector is cut into a block vector.

[0068] In some embodiments, the information to be encoded and the data width of the register are obtained, and then the message vector is stored in the register based on the data width, and the second preset number of groups is calculated, where the second preset number of groups is the ratio of the data width to the preset width; the message vector is cut into the second preset number of groups with the preset width to output a segmented vector, and the segmented vector after cutting is defined as m j , where

[0069] According to the register bit width l m Determine the block width μ. The block width is determined by the hardware resources and needs to be a power of 2 (such as 16, 32) to adapt to the grouping level. Then calculate the number of blocks. Exemplarily, if the bit width is 64 bits and the block width is 16 bits, then the number of block groups is 64 / 16 = 4 groups. The message vector is sliced into 4 groups of segmented vectors according to the width and stored in the register.

[0070] In some embodiments, the vector can be split by a bitmask or a shift operation to reduce the probability of information loss. Dynamically partitioning based on the register bit width can maximize the parallel computing efficiency.

[0071] Among them, the original message vector has three formats, as shown in the following table:

[0072] Specification 0 1 2 Effective data bit width τ 3 4 3 Number of data blocks subm 2 2 4 Size of data block μ 64 96 64 <![CDATA[Message length l m > 128 192 256

[0073] Among them,

[0074] S200: Perform encoding on the segmented vector to generate a lattice vector.

[0075] Perform encoding on the segmented vector through a tagging method. In some embodiments, fill the segmented vector to generate a segmented vector of a preset length, and then perform a recursive grouping operation on the segmented vector of the preset length through the tagging method, and generate a lattice vector of a preset dimension through modulo operation. The lattice vector includes a real part and an imaginary part.

[0076] Among them, to store the intermediate processing results, set An intermediate register with a bit width can store the filled vector and the scaled vector.

[0077] Filling the segmented vector is to ensure that the length of the segmented vector meets the hierarchical requirements of the recursive grouping, such as a length of 2^n. Specifically, if the length of the original segmented vector is L and the target length is N = 2 k (k is the number of recursive levels), then the filling method is: zero-padding and repeated filling. Zero-padding is to append N - L zeros at the end of the segmented vector, and repeated filling is to repeat the first N - L elements of the segmented vector.

[0078] Exemplarily, the input block vector: 1, 2, 3 (length L = 3), the target length N = 4 (k = 2), and the result of zero-padding: 1, 2, 3, 0.

[0079] The padded block vector is recursively grouped hierarchically to provide a structured input for modular arithmetic. The tagging method uses binary tags to identify the grouping levels, and the input to the tagging method is the message vector m ∈ {0, 1} μ , positive integers n, τ, such that n = 2 k ≥ 4, The output is the lattice vector x ∈ C, where C is based on a nested lattice

[0080] The specific process is as follows:

[0081] 1. Pad m with 0 to obtain a vector m′ of length ;

[0082] 2. Denote m′ = (u0, u1, …, u n / 2-1 ), such that u j ∈ {0, 1} 2τ-ωH(j) ;

[0083] 3. Calculate v = (v0, v1, …, v n / 2-1 ), where v j = f 2τ-ωH(j) (u j ), 0 ≤ j < n / 2, where ωH(j) is the number of 1s in the binary representation of j, and the sub-block length is dynamically adjusted by ωH(j) to adapt to the recursive level requirements;

[0084] 4. For l starting from 1 to k - 1, loop through and group and update the vector layer by layer. The layer l determines the granularity of the current block;

[0085] 5. Write where

[0086] 6. Update where the parity of the tag determines whether to add the φw term. For example, when the last bit of the tag is 0, the original block is retained, and when the last bit is 1, adjacent blocks are combined; the number of bits of the tag at the layer determines the grouping depth of the current operation.

[0087] 7. End the loop;

[0088] 8. Calculate

[0089] 9. Return w.

[0090] Exemplarily, if l = 2, the input vector v = (w1, w2, w3, w4), and the labels are 00, 01, 10, 11 respectively. Group by the first 2 bits of the labels (00 and 01 as one group, 10 and 11 as another group). Combine w1 and w2: w1 + φw2; combine w3 and w4: w3 + φw4. The output vector v = (w1, w1 + φw2, w3, w3 + φw4).

[0091] Then take the modulo 2 of the output vector v τ , to make the numerical range controllable, map the grouped values to the complex space, generate the real and imaginary parts of the lattice vector, and output the encoded complex vector w.

[0092] The labeling method hierarchically subdivides the blocks through binary labels to form a tree structure, adapts to the hardware pipeline, and through the imaginary unit φ, can enhance the multi-dimensional characteristics of the encoding, improve the anti-noise ability, and can also reduce the complexity from O(n 2 ) to O(n log n).

[0093] S300: Based on the lattice vector, perform tensor product parallel calculation through the operation modules of the first preset number of groups to output the target vector.

[0094] The operation module includes an addition and subtraction unit and an addition and addition unit. The number of the addition and subtraction unit and the addition and addition unit is the second number, and the second number is 8. That is to say, the number of both the addition and subtraction unit and the addition and addition unit is 8.

[0095] The addition and subtraction unit processes the real parts of the first group of 4 block vectors, such as performing addition and subtraction alternately, and outputs the real part result. The addition and addition unit processes the imaginary parts, such as performing continuous addition, and outputs the imaginary part result. Combine the real part and the imaginary part results to obtain the first updated vector, and the length is reduced to 1 / 4 of the original. In each subsequent round, the vector is regrouped, for example, into 4 groups, 2 groups, 1 group, and the separation calculation of the real part and the imaginary part is repeated. Through the modulo operation, for example, modulo 256 to constrain the numerical range, and output the target vector.

[0096] As Figure 2 shown, for the first round of iteration, in some embodiments, the addition and subtraction unit is used to perform the real part calculation of the first group of lattice vectors in the second preset number of groups to output the first real part result, and, the addition and addition unit is used to perform the imaginary part calculation of the first group of lattice vectors in the set number of groups to output the first imaginary part result. The number of the first group of lattice vectors is the first number, and the first number is 32, corresponding to 32 temp registers, that is, REG_2; then based on the first real part result and the first imaginary part result, obtain the first updated vector; store the first updated vector in the register to perform the next iteration through the updated vector.

[0097] Among them, the register is a 32-channel real and imaginary component register REG_2, which is used to store complex vectors and supports parallel processing.

[0098] Store all tagged signals into the REG_2 register as follows. abc is the tagged signal, i.e., the first update vector, temp is the REG_2 register, temp[0*2] <= a[0], temp[0*2 + 1] <= a[1], temp[1*2] <= a[2], temp[1*2 + 1] <= b[0], temp[2*2] <= a[3], temp[2*2 + 1] <= b[1], temp[3*2] <= b[2], temp[3*2 + 1] <= b[3], temp[4*2] <= a[4], temp[4*2 + 1] <= b[4], temp[5*2] <= b[5], temp[5*2 + 1] <= b[6], temp[6*2] <= b[7], temp[6*2 + 1] <= b[8], temp[7*2] <= b[9], temp[7*2 + 1] <= c[0], temp[8*2] <= a[5], temp[8*2 + 1] <= b

[10] , temp[9*2] <= b

[11] , temp[9*2 + 1] <= b

[12] , temp[10*2] <= b

[13] , temp[10*2 + 1] <= b

[14] , temp[11*2] <= b

[15] , temp[11*2 + 1] <= c[1], temp[12*2] <= b

[16] , temp[12*2 + 1] <= b

[17] , temp[13*2] <= b

[18] , temp[13*2 + 1] <= c[2], temp[14*2] <= b

[19] , temp[14*2 + 1] <= c[3], temp[15*2] <= c[4], temp[15*2 + 1] <= c[5].

[0099] Among them, the addition and subtraction unit is an adder and a subtractor, and the addition and addition unit is two adders aa0_0. Several sub-signals constitute the addition and addition unit (PPU) and the addition and subtraction unit (PMU). Among them, for the addition and addition unit: / / AAArray: a + b + c, assign aa0 = aa0_0 + aa0_1 + aa0_2, assign aa1 = aa1_0 + aa1_1 + aa1_2, assign aa2 = aa2_0 + aa2_1 + aa2_2, assign aa3 = aa3_0 + aa3_1 + aa3_2, assign aa4 = aa4_0 + aa4_1 + aa4_2, assign aa5 = aa5_0 + aa5_1 + aa5_2, assign aa6 = aa6_0 + aa6_1 + aa6_2, assign aa7 = aa7_0 + aa7_1 + aa7_2.

[0100] Addition and subtraction unit: / / AS Array: a + b - c: assign as0 = as0_0 + as0_1 - as0_2, assign as1 = as1_0 + as1_1 - as1_2, assign as2 = as2_0 + as2_1 - as2_2, assign as3 = as3_0 + as3_1 - as3_2, assign as4 = as4_0 + as4_1 - as4_2, assign as5 = as5_0 + as5_1 - as5_2, assign as6 = as6_0 + as6_1 - as6_2, assign as7 = as7_0 + as7_1 - as7_2.

[0101] Combine the real and imaginary parts of the results to obtain the first updated vector, write it into the register, and prepare for the next iteration.

[0102] For the second iteration, in some embodiments, the first updated vector is regrouped. Among them, the regrouping is to sequentially assign the first updated vector to the aa sub-signals and as sub-signals respectively. There are a total of 24 aa sub-signals and 24 as sub-signals. The sub-signals pass through the PPU and PMU to obtain 8 aa signals and 8 as signals, that is, to obtain a third quantity of vectors. The first quantity is a multiple of the third quantity, and the third quantity is 8; the real part calculation of the third quantity of vectors is performed by the addition and subtraction unit to output the second real part result, and the imaginary part calculation of the third quantity of vectors in the set number is performed by the addition unit to output the second imaginary part result; based on the second real part result and the second imaginary part result, a second updated vector is obtained.

[0103] Transfer the labeled signal stored in REG_2 to the PPU and PMU, and continue to store the aa and as signals respectively output by the PPU and PMU in REG_2.

[0104] For example, at this time, aa0_0 = a[0], aa0_1 = a[2], aa0_2 = b[0], as0_0 = a[1], as0_1 = a[2], as0_2 = b[0], etc. The following is the signals of the PPU and PMU stored in the register REG_2, temp[1*2] <= as0, temp[1*2 + 1] <= aa0, temp[3*2] <= as1, temp[3*2 + 1] <= aa1, temp[5*2] <= as2, temp[5*2 + 1] <= aa2, temp[7*2] <= as3, temp[7*2 + 1] <= aa3, temp[9*2] <= as4, temp[9*2 + 1] <= aa4, temp[11*2] <= as5, temp[11*2 + 1] <= aa5, temp[13*2] <= as6, temp[13*2 + 1] <= aa6, temp[15*2] <= as7, temp[15*2 + 1] <= aa7.

[0105] For the third round of iteration, in some embodiments, the second update vector is regrouped. Specifically, the regrouping is to sequentially assign the second update vector to the aa sub-signals and the as sub-signals respectively. There are a total of 24 aa sub-signals and 24 as sub-signals. The sub-signals pass through the PPU and PMU to obtain 4 aa signals and 4 as signals, that is, a fourth quantity of vectors is obtained. The third quantity is a multiple of the fourth quantity, and the fourth quantity is 4. The real part calculation of the fourth quantity of vectors is performed by the addition and subtraction unit to output the third real part result, and the imaginary part calculation of the fourth quantity of vectors in the set number is performed by the addition unit to output the third imaginary part result. Based on the third real part result and the third imaginary part result, the third update vector is obtained.

[0106] For the fourth round of iteration, the third update vector is regrouped. Specifically, the regrouping is to sequentially assign the second update vector to the aa sub-signals and the as sub-signals respectively. There are a total of 24 aa sub-signals and 24 as sub-signals. The sub-signals pass through the PPU and PMU to obtain 2 aa signals and 2 as signals, that is, 2 sets of vectors are obtained. The real part calculation of the 2 sets of vectors is performed by the addition and subtraction unit to output the fourth real part result. The imaginary part calculation of the 2 sets of vectors is performed by the addition unit to output the fourth imaginary part result. Based on the fourth real part result and the fourth imaginary part result, the fourth update vector is obtained.

[0107] The aa and as signals stored in REG_2 are passed to the PPU and PMU, and the aa and as signals respectively output by the PPU and PMU are continuously stored in REG_2. This can be regarded as a loop iteration. The following is the continuous storage of the output aa and as signals in REG_2.

[0108] 3'h2: begin / / computer w, step1, temp[1*2] <= as0, temp[1*2 + 1] <= aa0, temp[3*2] <= as1, temp[3*2 + 1] <= aa1, temp[5*2] <= as2, temp[5*2 + 1] <= aa2, temp[7*2] <= as3, temp[7*2 + 1] <= aa3, temp[9*2] <= as4, temp[9*2 + 1] <= aa4, temp[11*2] <= as5, temp[11*2 + 1] <= aa5, temp[13*2] <= as6, temp[13*2 + 1] <= aa6, temp[15*2] <= as7, temp[15*2 + 1] <= aa7, end.

[0109] 3'h3: begin / / computer w, step2, temp[2*2] <= as0, temp[2*2 + 1] <= aa0, temp[3*2] <= as1, temp[3*2 + 1] <= aa1, temp[6*2] <= as2, temp[6*2 + 1] <= aa2, temp[7*2] <= as3, temp[7*2 + 1] <= aa3, temp[10*2] <= as4, temp[10*2 + 1] <= aa4, temp[11*2] <= as5, temp[11*2 + 1] <= aa5, temp[14*2] <= as6, temp[14*2 + 1] <= aa6, temp[15*2] <= as7, temp[15*2 + 1] <= aa7, end.

[0110] 3'h4: begin / / computer w, step3, temp[4*2] <= as0, temp[4*2 + 1] <= aa0, temp[5*2] <= as1, temp[5*2 + 1] <= aa1, temp[6*2] <= as2, temp[6*2 + 1] <= aa2, temp[7*2] <= as3, temp[7*2 + 1] <= aa3, temp[12*2] <= as4, temp[12*2 + 1] <= aa4, temp[13*2] <= as5, temp[13*2 + 1] <= aa5, temp[14*2] <= as6, temp[14*2 + 1] <= aa6, temp[15*2] <= as7, temp[15*2 + 1] <= aa7, end.

[0111] 3'h5: begin / / computer w, step4, temp[8 * 2] <= as0, temp[8 * 2 + 1] <= aa0, temp[9 * 2] <= as1, temp[9 * 2 + 1] <= aa1, temp[10 * 2] <= as2, temp[10 * 2 + 1] <= aa2, temp[11 * 2] <= as3, temp[11 * 2 + 1] <= aa3, temp[12 * 2] <= as4, temp[12 * 2 + 1] <= aa4, temp[13 * 2] <= as5, temp[13 * 2 + 1] <= aa5, temp[14 * 2] <= as6, temp[14 * 2 + 1] <= aa6, temp[15 * 2] <= as7, temp[15 * 2 + 1] <= aa7, end.

[0112] The operation of passing the aa and as signals stored in REG_2 to the PPU and PMU is rather cumbersome. That is, the elements in temp are respectively passed to the plus-minus signals such as aa0_0, aa0_1, aa0_2, aa1_0, aa1_1, aa1_2, etc., and the plus-plus signals such as as0_0, as0_1, as0_2, as1_0, as1_1, as1_2, etc., and then calculated by the PPU and PMU.

[0113] For the final iteration, the fourth update vector is regrouped to obtain a fifth number of vectors, where the fourth number is a multiple of the fifth number and the fifth number is 1; a modulo operation is performed on the vectors of the fifth number, and the output is the target vector.

[0114] Perform a modulo operation on the final vector to constrain the numerical range, prevent numerical overflow, and adapt to the integer requirements of the subsequent encoding matrix. Output the target vector for constructing the encoding matrix M.

[0115] In the tensor product operation stage, the plus-minus unit and the plus-plus unit adopt a time-division multiplexing mechanism. The same set of hardware units process the real part and the imaginary part respectively in different clock cycles, and dynamically switch the operation mode through a multiplexer. When the first sub-block enters the tensor product operation stage, the second sub-block has already started label encoding. In the tensor product operation, 8 plus-minus units and 8 plus-plus units simultaneously process different data segments, and 16-element real-imaginary operations (8 groups × 2 elements) can be completed in each clock cycle.

[0116] S400: Output the encoding matrix through the target vector.

[0117] In some embodiments, the target vectors are concatenated to construct the original vector, the original vector is filled to a preset dimension, and a scaling operation is performed to output the encoding matrix. Exemplarily, multiple target vectors after iteration are sequentially concatenated into the original vector, filled to the preset dimension (such as a 1024×1024 matrix), and padded with zeros or interpolated. Scaled proportionally (such as normalized to 0-255) to adapt to the transmission protocol.

[0118] The encoding method decomposes the high-dimensional tensor product into multi-level grouped iterations, uses addition-subtraction units and addition-addition units for parallel processing to reduce complexity. The result of each round of iteration is directly written into the register, and the next round reads it, hiding the calculation latency. The vector is dynamically segmented according to the register bit width to maximize the utilization of the hardware bit width. The addition-subtraction units and addition-addition units are reused in different iteration stages, reducing the number of logic units.

[0119] Traditional software implementation requires O(n 2 ) operations for each sub-block. The method compresses the single sub-block processing time to a fixed 7 cycles through hardware parallelism, that is, 1 cycle for preprocessing + 5 cycles for core calculation + 1 cycle for postprocessing.

[0120] Based on the above message encoding method, some embodiments of the present application further provide a message encoding device, including: an acquisition module, a tag processing module, an operation module, and a calculation module.

[0121] The acquisition module is configured to acquire the segmented vector, and the segmented vector is the vector after being segmented. The tag processing module is configured to perform encoding on the segmented vector to generate a lattice vector. The operation module is configured to perform tensor product parallel calculation based on the lattice vector through a first preset number of operation modules to output an updated vector, and the operation module includes an addition-subtraction unit and an addition-addition unit. The calculation unit is configured to output the encoding matrix through the updated vector.

[0122] As Figure 3 shown, the original input data acquired by the acquisition module is input to the input register REG_0, with different widths, supporting multiple segmentation modes. l is the length of the original message, m is the number of segments, and it can dynamically adapt to different data volumes. According to the register bit width, for example, 16 bits, the data in REG_0 is segmented into fixed-width segmented vectors. If the width of REG_0 is 64 bits, each block is segmented into 4 groups of 16-bit data.

[0123] The intermediate register REG_1, with a width of 4 / 16 bits, adapts to different precision requirements, a depth of 32, and supports batch data caching. Stores the intermediate data after clipping, performs real part calculation through the PMU (addition-subtraction unit), supports shift operations, and is used for numerical alignment or scaling.

[0124] The result register REG_2 supports multi-channel parallel output, such as 32-channel synchronous transmission. It contains 32 elements (temp[0] to temp

[31] ). Each element stores the real and imaginary components. The indexing rule is as follows: temp[2i + 0] is the real part of the i-th block. For example, temp[0] is the real part of the 0-th block. temp[2i + 1] is the imaginary part of the i-th block. For example, temp[1] is the imaginary part of the 0-th block.

[0125] In some embodiments, the operation module further includes a multiplexer and a timing control unit;

[0126] The timing control unit is configured to: obtain vector processing information, where the vector processing information includes real part processing information and imaginary part processing information;

[0127] The multiplexer is configured to: obtain the vector processing information sent by the timing control unit, generate control signals, where the control signals include a start subtraction signal and a disable subtraction signal; and select an input source based on the control signals, where the input source includes an addition-subtraction unit and an addition-addition unit.

[0128] Exemplarily, the timing control unit generates a mode_sel signal according to the current processing requirement (real part / imaginary part). When processing the real part, mode_sel = 0 (enable subtraction), and when processing the imaginary part, mode_sel = 1 (disable subtraction). The multiplexer selects the input source according to mode_sel. If mode_sel = 0, it is connected to the addition-subtraction unit. If mode_sel = 1, it is connected to the addition-addition unit.

[0129] For the effects during the operation of this embodiment, reference can be made to the effects of the above method embodiment, which will not be elaborated here.

[0130] As can be seen from the above technical solutions, the present application provides a message encoding method and apparatus. The method includes: obtaining a segmented vector, where the segmented vector is a vector after being segmented, and then performing encoding on the segmented vector to generate a lattice vector. Based on the lattice vector, performing tensor product parallel calculation through a first preset number of operation modules to output a target vector. The operation module includes an addition-subtraction unit and an addition-addition unit, and outputting an encoding matrix through the target vector. By reusing the real part and imaginary part operation units of the tensor product operation in the message encoding process, and the real part and imaginary part operation units adopting a parallel architecture, the calculation efficiency can be improved, and the hardware resource overhead can be reduced.

[0131] For the similar parts among the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts fall within the protection scope of this application.

Claims

1. A message encoding method, characterized in that, Including: Obtain a block vector, where the block vector is a vector after being blocked; Perform encoding on the block vector to generate a lattice vector; Based on the lattice vector, perform tensor product parallel calculation through a first preset number of operation modules to output a target vector, where the operation module includes an addition and subtraction unit and an addition and addition unit; Output an encoding matrix through the target vector.

2. The message encoding method according to claim 1, wherein The performing encoding on the block vector to generate a lattice vector includes: Pad the block vector to generate a block vector with a preset length; Perform a recursive grouping operation on the block vector with the preset length through a labeling method, and generate a lattice vector with a preset dimension through modulo operation, where the lattice vector includes a real part and an imaginary part.

3. The message encoding method according to claim 2, wherein The obtaining a block vector includes: Obtain information to be encoded and the data width of a register; Store the message vector into the register based on the data width; Calculate a second preset number, where the second preset number is the ratio of the data width to a preset width; Clip the message vector into the second preset number of groups with a preset width to output a block vector.

4. The message encoding method according to claim 3, wherein The performing tensor product parallel calculation through a first preset number of operation modules based on the lattice vector to output a target vector includes: Perform real part calculation of the first group of lattice vectors in the second preset number of groups through the addition and subtraction unit to output a first real part result, and perform imaginary part calculation of the first group of lattice vectors in the set number of groups through the addition and addition unit to output a first imaginary part result, where the number of the first group of lattice vectors is a first number; Obtain a first updated vector based on the first real part result and the first imaginary part result; Store the first updated vector into the register to perform the next iteration through the updated vector.

5. The message encoding method according to claim 4, wherein The number of the addition and subtraction unit and the addition and addition unit is a second number; The performing tensor product parallel calculation through a first preset number of operation modules based on the lattice vector to output a target vector includes: Regroup the first updated vector to obtain a vector with a third number, where the first number is a multiple of the third number; Perform real part calculation of the vector with the third number through the addition and subtraction unit to output a second real part result, and perform imaginary part calculation of the vector with the third number in the set number of groups through the addition and addition unit to output a second imaginary part result; Obtain a second updated vector based on the second real part result and the second imaginary part result.

6. The message encoding method according to claim 5, wherein The performing tensor product parallel calculation through a first preset number of operation modules based on the lattice vector to output a target vector includes: Regroup the second updated vector to obtain a vector with a fourth number, where the third number is a multiple of the fourth number; Perform real part calculation of the vector with the fourth number through the addition and subtraction unit to output a third real part result, and perform imaginary part calculation of the vector with the fourth number in the set number of groups through the addition and addition unit to output a third imaginary part result; Obtain a third updated vector based on the third real part result and the third imaginary part result.

7. The message encoding method according to claim 6, wherein The performing tensor product parallel calculation through a first preset number of operation modules based on the lattice vector to output a target vector includes: Regroup the third update vector to obtain a fifth quantity of vectors, where the fourth quantity is a multiple of the fifth quantity; Perform a modulo operation on the fifth quantity of vectors, and the output is the target vector.

8. The message encoding method according to claim 1, wherein The output of the encoding matrix through the target vector includes: Concatenate the target vectors to construct an original vector; Pad the original vector to a preset dimension and perform a scaling operation to output an encoding matrix.

9. A message encoding device, characterized in that, Includes: An acquisition module is configured to acquire a segmented vector, where the segmented vector is a segmented vector; A label processing module is configured to perform encoding on the segmented vector to generate a lattice vector; An operation module is configured to perform tensor product parallel calculations through a first preset number of operation modules based on the lattice vector to output an update vector, where the operation module includes an addition / subtraction unit and an addition / addition unit; A calculation module is configured to output an encoding matrix through the update vector.

10. The message encoding device according to claim 9, wherein The operation module further includes a multiplexer and a timing control unit; The timing control unit is configured to: acquire vector processing information, where the vector processing information includes real part processing information and imaginary part processing information; The multiplexer is configured to: acquire the vector processing information sent by the timing control unit, generate control signals, where the control signals include a start subtraction signal and a disable subtraction signal; and select an input source based on the control signals, where the input sources include an addition / subtraction unit and an addition / addition unit.