Data processing device, method, equipment and chip
By dynamically configuring the data processing device of the hardware processing module, the problems of polynomial multiplication computing efficiency and high storage resource occupation are solved, and efficient hardware acceleration and resource optimization of polynomial computing tasks are realized.
Patent Information
- Application Number
- CN202510459687.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In the prior art, the time complexity of polynomial multiplication calculation is O(N²), and the storage resource occupies a high level, which cannot meet the technical requirements of large-length polynomial multiplication.
It provides a data processing device, which dynamically configures the hardware processing module through the controller module to build a target task path suitable for polynomial computing tasks, and supports the individual or combination execution of number theory transformation, multiplication and inverse number theory transformation tasks.
Hardware acceleration of polynomial computing tasks is realized, the scalability and resource utilization of the hardware structure are improved, and the calculation waiting time and resource waste are reduced.
Smart Images

Figure CN119987855A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a data processing device, method, equipment and chip. Background Art
[0002] With the continuous development of artificial intelligence and big data algorithms, using encryption algorithms to encrypt data is an important way to protect data security. Mainstream privacy computing technologies such as Fully Homomorphic Encryption (FHE) and Zero Knowledge Proof (ZKP) require a large number of polynomial multiplication calculations.
[0003] In the related technology, in the process of polynomial multiplication, if two polynomials of length N are directly multiplied, the time complexity is O(N²). Number Theoretic Transform (NTT) and Inverse Number Theoretic Transform (INTT) can optimize the polynomial multiplication calculation and reduce its time complexity to O(N²). The calculation speed of NTT / INTT is an important factor affecting the efficiency of privacy computing. However, the storage resource usage in the related technology is high and the task process is rigid, which cannot meet the technical requirements of large-length polynomial multiplication. Summary of the invention
[0004] In response to the technical problems existing in the prior art, the present application provides a data processing device, method, equipment and chip for dynamically configuring hardware processing modules to adapt to the computing requirements of various polynomial multiplication tasks, realize hardware acceleration of polynomial computing tasks, greatly improve the scalability of the hardware structure, and improve resource utilization.
[0005] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to a data processing device, the device comprising: a controller module, a task execution unit, and a reading result module; the data processing method comprises: The data processing device receives a task selection signal from a task request end, and determines a target task to be performed by the task execution unit based on the task selection signal; the target task is one or more polynomial calculation tasks; the polynomial calculation tasks include at least: number theory change tasks, multiplication tasks, and inverse number theory change tasks; The data processing device selects a corresponding hardware processing module based on the target task to construct a corresponding target task path; the hardware processing modules in the task execution unit include: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module; the number, type, and connection method of the hardware processing modules in the target task path are related to the task type to which the target task belongs; The data processing device receives the data to be processed, and inputs the data to be processed into the target task path to start the target task and obtain the task execution result; The data processing device outputs the task execution result to the task request end.
[0006] In a second aspect, an embodiment of the present application provides a data processing device, the device comprising: a controller module, a task execution unit, and a reading result module; The controller module is used to receive a task selection signal from a task request end, and determine a target task to be performed by the task execution unit based on the task selection signal; the target task is one or more polynomial calculation tasks; the polynomial calculation tasks at least include: number theory change tasks, multiplication tasks, and inverse number theory change tasks; The task execution unit is used to select a corresponding hardware processing module based on the target task to construct a corresponding target task path; the hardware processing modules in the task execution unit include: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module; the number, type, and connection method of the hardware processing modules in the target task path are related to the task type to which the target task belongs; The task execution unit is further used to receive the data to be processed, and input the data to be processed into the target task path to start the target task and obtain the task execution result; The reading result module is used to output the task execution result to the task request end.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising: The memory is used to store computer programs; the processor is used to read and execute the computer programs to implement the data processing device of the second aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a chip, in which a computer program is loaded, and the computer program is used to implement the data processing device as in the second aspect.
[0009] In a fifth aspect, a computer-readable storage medium is provided, which includes instructions, and when the instructions are executed on a computer, the computer executes the data processing device of the second aspect.
[0010] In the technical solution of the present application, the data processing device includes a controller module, a task execution unit, and a reading result module. The hardware processing modules in the task execution unit include: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module. In the above device, the controller module receives a task selection signal from the task request end, and determines the target task to be executed by the task execution unit based on the task selection signal; the task execution unit selects the corresponding hardware processing module based on the target task to construct the corresponding target task path; the task execution unit receives the data to be processed, and inputs the data to be processed into the target task path to start the target task and obtain the task execution result; the reading result module outputs the task execution result to the task request end. In the technical solution of the present application, the data processing device can determine the target task according to the task selection signal by the controller module, support the individual or combined execution of tasks such as number theory transformation, multiplication, and inverse number theory transformation in polynomial calculation tasks, meet diversified needs, and is easy to adapt to new computing needs by expanding task selection signals and hardware processing modules, and has strong scalability. In terms of task execution mechanism, the task execution unit can dynamically construct the optimal task path according to the target task, configure the hardware processing module on demand, reduce resource waste and inefficiency under the fixed architecture, and quickly build the path to reduce computing waiting time and improve overall computing efficiency. In terms of resource optimization and utilization, the device allocates hardware resources on demand, and only enables necessary modules when executing different tasks, saving resources and energy consumption. At the same time, it flexibly constructs paths to achieve hardware module reuse, improve hardware utilization, and reduce hardware deployment costs. In summary, the technical solution of this application adapts to the computing needs of various polynomial multiplication tasks through task selection signals and dynamic construction of hardware processing modules, improves the scalability of the hardware structure, and improves resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a structural schematic diagram of a data processing device according to an embodiment of the present application; Figure 2 is a schematic diagram of the architecture of a data processing device according to an embodiment of the present application; Figure 3 It is a schematic diagram of the principle of a data processing device according to an embodiment of the present application; Figure 4 It is a flowchart of a data processing method according to an embodiment of the present application; Figure 5 is a schematic diagram of the structure of an electronic device according to an embodiment of the present application; Figure 6 It is a structural schematic diagram of a medium in an embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0013] With the continuous development of artificial intelligence and big data algorithms, more and more people realize that data has become a valuable resource with rich value. While constantly exploring and utilizing the value of data, protecting data security and privacy has become an important issue. Using encryption algorithms to encrypt data is an important way to protect data security. Mainstream privacy computing technologies such as FHE and ZKP require a large number of polynomial multiplication calculations.
[0014] Currently, in the process of polynomial multiplication, if two polynomials of length N are directly multiplied, the time complexity is O(N²). NTT and INTT can optimize polynomial multiplication and reduce its time complexity to O(Nlog2N).
[0015] In the related art, number theoretic transformation (NTT), multiplication, and inverse number theoretic transformation (INTT) are performed as a whole. When calculating the multiplication of two polynomials, it is necessary to first perform NTT on P1 and P2 to obtain P1' and P2', then multiply them to obtain M', and finally perform INTT to obtain the result M. The time complexity of this process is optimized to O (N log²N). However, the above solution requires a large storage space. Taking the above calculation requirements as an example, the amount of data that needs to be stored is two polynomial coefficient data of length N and two rotation factors of length N, a total of 4N data. For example, when the data bit width is 255 bits and the polynomial length is 2×18, the amount of data that needs to be stored is huge, which will exceed the on-board storage resources of most FPGAs on the market.
[0016] Secondly, the above scheme is not flexible enough in the stage of executing calculations. It is necessary to perform a complete set of data processing procedures including number theory transformation, multiplication, and inverse number theory transformation. If you want to support polynomial multiplication of a larger length such as 2×30, it is limited by the existing chip structure and hardware circuit design. It is impossible to directly implement ordinary number theory transformation and inverse number theory transformation on existing chips or other hardware circuits. Matrix number theory transformation and its inverse transformation are even more difficult to support on existing chips.
[0017] It can be seen that the relevant technologies have the following technical problems: high storage resource usage, rigid task process, and inability to meet the technical requirements of large-length polynomial multiplication.
[0018] To solve at least one technical problem in the related art, the embodiments of the present application provide a data processing device, method, equipment and chip.
[0019] In the technical solution provided by the present application, the data processing device includes: a controller module, a task execution unit, and a reading result module. The controller module is used to receive a task selection signal from the task request end, and determine the target task that the task execution unit needs to execute based on the task selection signal; the target task is one or more of the polynomial calculation tasks; the polynomial calculation tasks at least include: number theory change tasks, multiplication tasks, and inverse number theory change tasks. The task execution unit is used to select the corresponding hardware processing module based on the target task to construct the corresponding target task path; the number, type, and connection method of the hardware processing modules in the target task path are related to the task type to which the target task belongs. The task execution unit is also used to receive data to be processed, and input the data to be processed into the target task path to start the target task and obtain the task execution result. The reading result module is used to output the task execution result to the task request end.
[0020] The technical solution of the present application, first, the controller module can determine the target task according to the task selection signal, and these target tasks can be one or more of the polynomial calculation tasks. This means that the device can handle the individual execution of number theory transformation tasks, multiplication tasks, and inverse number theory transformation tasks, and can also handle their combined execution, which greatly meets the diverse needs of different users or application scenarios. For example, in some complex cryptographic calculations, it may be necessary to perform number theory transformation first, then multiplication operations, and finally inverse number theory transformation. The device can easily realize such a task combination. With the development of technology and the emergence of new computing needs, it is only necessary to expand the task selection signal and the corresponding hardware processing module to allow the device to support more types of polynomial calculation tasks or other related tasks. This makes the device highly scalable and able to adapt to future developments and changes.
[0021] Furthermore, the task execution unit selects the corresponding hardware processing module according to the target task to construct the target task path. The number, type and connection method of the hardware processing modules are related to the task type to which the target task belongs. This dynamic construction method can be optimally configured according to the specific task. For example, for simple multiplication tasks, only a few multiplication-related hardware processing modules may need to be connected; while for complex number theory transformation combination tasks, multiple number theory transformations and multiplication modules can be flexibly combined to form an efficient task execution path, avoiding resource waste and inefficiency under a fixed architecture. Since the appropriate task path can be quickly constructed according to the target task, the data to be processed can be directly input into the path to execute the task, reducing unnecessary intermediate links and waiting time. Compared with the traditional fixed process calculation method, the device can complete tasks more efficiently and improve the overall computing efficiency.
[0022] Then, the data flow of the entire device is clear and definite. It receives the task selection signal and the data to be processed from the task request end, determines the task through the controller module, and the task execution unit executes the task. Finally, the reading result module outputs the task execution result to the task request end. This clear data flow makes the data processing process easy to monitor and manage, and is also convenient for troubleshooting and optimization. For the user at the task request end, they only need to send the task selection signal and the data to be processed to get the required task execution result, without having to worry about the complex task execution process inside the device. This lowers the user's usage threshold and improves the user experience.
[0023] Finally, by dynamically selecting hardware processing modules according to the target tasks, the device can realize on-demand allocation of hardware resources. When executing different tasks, only the necessary hardware processing modules are enabled, avoiding the waste of resources and increased energy consumption caused by the simultaneous operation of all hardware modules. For example, when executing a multiplication task, it is not necessary to enable hardware modules related to number theory transformation, thereby saving hardware resources and power consumption. Since the task path can be flexibly constructed, the hardware processing modules can be reused between different tasks, improving the overall utilization of hardware resources. This reduces the hardware cost of the device to a certain extent, and also improves the cost performance of the device.
[0024] In summary, in this technical solution, in terms of flexibility and scalability, the data processing device can determine the target task based on the task selection signal by means of the controller module, support the individual or combined execution of tasks such as number theory transformation, multiplication, and inverse number theory transformation in polynomial calculation tasks, meet diversified needs, and facilitate the adaptation to new calculation needs by expanding the task selection signal and hardware processing module, and has strong scalability. In terms of task execution mechanism, the task execution unit can dynamically construct the optimal task path according to the target task, configure the hardware processing module on demand, reduce the waste of resources and inefficiency under the fixed architecture, and quickly build the path to reduce the calculation waiting time and improve the overall calculation efficiency. In terms of data interactivity, its data flow is clear, from receiving the task selection signal and the data to be processed, to determining the task, executing the task, and then outputting the result, which is convenient for monitoring, management and troubleshooting, and the user only needs to send the signal and data to obtain the result, which reduces the threshold of use and improves the user experience. In terms of resource optimization and utilization, the device allocates hardware resources on demand, and only enables necessary modules when executing different tasks, saving resources and energy consumption, and flexibly constructs the path to realize hardware module reuse, improve hardware utilization, and reduce hardware deployment costs. In summary, this technical solution can adapt to the computing requirements of various polynomial multiplication tasks through task selection signals and dynamically combined hardware processing modules, and directly execute polynomial calculation tasks on hardware, greatly improving the scalability of the hardware structure and improving resource utilization.
[0025] The data processing solution provided in the embodiment of the present application can also be executed by a chip, or by a chip equipped with other electronic components in cooperation, or by a service program installed in an electronic device to implement the data processing solution.
[0026] Figure 1 A flow chart of a data processing device provided in an embodiment of the present application is as follows: Figure 1 As shown, the device includes the following units: The controller module is used to receive a task selection signal from the task request end, and determine a target task to be performed by the task execution unit based on the task selection signal; A task execution unit, used to select a corresponding hardware processing module based on a target task to construct a corresponding target task path; The task execution unit is further used to receive the data to be processed and input the data to be processed into the target task path to start the target task and obtain the task execution result; The result reading module is used to output the task execution results to the task request end.
[0027] In the embodiment of the present application, the target task is one or more of the polynomial calculation tasks. The polynomial calculation tasks include at least: number theory change tasks, multiplication tasks, and reverse number theory change tasks. See the examples below for details, which will not be expanded here.
[0028] In the embodiment of the present application, the number, type, and connection mode of the hardware processing modules in the target task path are related to the task type to which the target task belongs. It is understandable that when executing the target task, the data processing device will build a target task path that is adapted thereto. The construction of this path is highly targeted, and the number, type, and connection mode of the hardware processing modules in the path all depend on the type of the target task. When the target task is a single number theory transformation task, the task path will contain necessary modules such as the initialization module, the storage module, and the butterfly calculation module, and the link will be built according to the steps of initialization, data storage, and butterfly operation; if it is a multiplication task, the initialization module, the storage module, and the multiplication module will be enabled to build a path around data input and multiplication operation. When performing complex number theory transformation and multiplication combination tasks, the device will connect the butterfly calculation module, the multiplication module, etc. in an orderly manner based on the task flow to meet the task requirements. Thus, the hardware processing module can be dynamically configured according to the target task, which can not only avoid resource waste and improve hardware utilization efficiency, but also greatly enhance the flexibility and adaptability of the device, so that the data processing device can operate efficiently in a variety of scenarios to meet the computing needs of different users.
[0029] In the embodiment of the present application, the controller module, as a key component of the data processing device, plays a core control and coordination role in the entire system. The primary function of the controller module is to receive a task selection signal from the task request end. This signal is a key input used to indicate the target task that the system needs to perform. The target task can be one or more polynomial calculation tasks, including number theory transformation tasks, multiplication tasks, inverse number theory transformation tasks, etc. For example, when the task selection signal is a specific code (such as 2'b00), the controller can identify and determine that the number theory transformation task is to be executed; when it is 2'b01, it determines to execute the multiplication task, etc. By accurately interpreting the task selection signal, the controller provides a clear direction for subsequent task execution.
[0030] Then, based on the determined target task, the controller module is responsible for coordinating and scheduling the work of each hardware processing module in the task execution unit. It determines which hardware processing modules need to participate in the task execution, as well as the working order and coordination mode between these modules according to the requirements of the target task. For example, when executing the number theory transformation task, the controller will arrange the initialization module to receive the data to be processed and temporarily store it in the storage module, and then trigger the butterfly calculation module, reversal module, etc. in turn to operate according to the predetermined process to ensure that the entire task can be carried out in an orderly manner.
[0031] In addition, the controller module monitors and manages the various states of the entire data processing process in real time. It pays attention to the working status of the hardware processing module, such as whether the initialization module has completed data storage, whether the butterfly calculation module has completed calculation, etc. By receiving the status signals fed back by each module (such as ntt_btf_finish_in, reverse_finish_in, etc.), the controller can accurately understand the progress of task execution and make corresponding decisions based on the status information, such as whether to continue to advance the task process or perform error handling.
[0032] It is worth mentioning that the controller module uses a finite state machine (FSM) to implement state transitions and control logic. The finite state machine contains a series of states (such as kIdle, kInitMem, kWaitNTTTF, kWaitINTTTF, kReverse, kPolyNTT, kPolyINTT, kPolyMult, kReadResult, and kFinish, etc.), as well as the transition conditions and actions between states. According to the task selection signal and the state signals fed back by each module, the finite state machine jumps between different states to achieve precise control of task execution. For example, in the initial state kIdle, if the external start signal poly_start is 1, the state changes from kIdle to kInitMem, entering the initialization memory state.
[0033] In addition to the finite state machine, the controller module also contains signal processing and logic control units. These units are responsible for processing and performing logical operations on task selection signals, feedback signals from various modules, and other related signals. For example, the task selection signal is decoded to determine the specific target task; the module feedback signal is analyzed to determine whether the task execution is normal, etc. Through these signal processing and logic control, the controller can make accurate decisions to ensure the stable operation of the system.
[0034] The controller module receives key information such as task selection signal and start signal (such as poly_start) from the task request end. The task request end sends these signals to indicate the task to be performed by the controller and the start time of the task. The controller starts the corresponding operation according to the received signal.
[0035] The controller module is connected to each hardware processing module (initialization module, storage module, reversal module, butterfly calculation module, multiplication module, modular multiplication module, etc.) in the task execution unit. It sends control signals to these modules to instruct them to perform corresponding operations, and receives status signals and data processing results fed back by each module. For example, the controller sends instructions to the initialization module to request it to receive and store the data to be processed; at the same time, it receives the calculation completion signal ntt_btf_finish_in fed back by the butterfly calculation module to decide whether to continue the subsequent operation.
[0036] When the task is completed, the controller module will control the reading result module to output the task execution result to the task request end. It sends output instructions to the reading result module to ensure that the results can be returned to the user accurately and in a timely manner.
[0037] Thus, the controller module can flexibly determine and execute the target task according to different task selection signals, so that the data processing device can adapt to diverse application scenarios and user needs. Whether it is executing a certain polynomial calculation task alone or executing multiple tasks in combination, the controller can effectively schedule and manage, greatly improving the flexibility of the system. By accurately coordinating and scheduling the hardware processing modules in the task execution unit, the controller module avoids the waste of resources and invalid operations, and reduces the waiting time and intermediate links in the task execution process. For example, dynamically selecting and configuring hardware processing modules according to task requirements enables the system to execute tasks in the best way, thereby improving the overall task execution efficiency. The controller module monitors and manages the system status in real time, and can promptly detect and handle abnormal situations that occur during task execution. When a module fails or an error occurs in task execution, the controller can take corresponding measures, such as pausing the task, performing error recovery or restarting, etc., to ensure the stability and reliability of the system.
[0038] In summary, the controller module plays a vital role in the data processing device. It provides strong support for the performance and functions of the data processing device through precise task control, efficient module coordination and comprehensive status management. As an optional embodiment, the hardware processing module in the task execution unit includes: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module.
[0039] Wherein, the initialization module is connected to the storage module; the initialization module is used to receive the data to be processed and temporarily store the data to be processed in the storage module.
[0040] Among them, the storage module is respectively connected to the reversal module, the butterfly calculation module, and the multiplication module, and the butterfly calculation module and the multiplication module are respectively connected to the modular multiplication module; the data processing results of the reversal module, the butterfly calculation module, the multiplication module, and the modular multiplication module are output to the reading result module through the storage module.
[0041] In some optional embodiments, the hardware processing module of the task execution unit adopts a modular design, each module has a clear division of labor and cooperates with each other, thereby achieving task decoupling and flexible dynamic path construction.
[0042] The initialization module is responsible for receiving the data to be processed from the external task request end. These data can be polynomial coefficients, rotation factors and other data related to polynomial calculations. The received data is converted to the necessary format or preprocessed, such as padding the data with zeros to meet specific data format requirements, or performing operations such as sign extension. The preprocessed data is accurately written to the storage module to prepare for subsequent calculations. On the input side, it is connected to the task request end to receive the data to be processed. On the output side, it is directly connected to the storage module to temporarily store the processed data.
[0043] The storage module plays an important role as a data hub. During the entire data processing process, it caches the intermediate results and input data of each stage to ensure that data can flow smoothly between different modules. It has the characteristics of time-sharing multiplexing. In the number theory transformation (NTT) and inverse number theory transformation (INTT) stages, it stores polynomial coefficients and corresponding rotation factors; in the multiplication task stage, it stores the two polynomial data involved in the multiplication operation. Through a reasonable address mapping mechanism, fast data reading and writing operations are achieved, data access efficiency is improved, and the operating speed of the entire device is improved. When inputting, it receives the data written by the initialization module, and also receives the data fed back from the reversal module, butterfly calculation module, multiplication module and modular multiplication module after processing. On the output, it provides a data input source for the reversal module, butterfly calculation module, and multiplication module, and finally passes the task execution results processed by each module to the reading result module.
[0044] The reversal module is mainly used to perform bit reversal and permutation operations on data during the bit reversal stage, that is, during the pre-processing or post-processing of NTT or INTT. Specifically, the binary bit of the data index i is reversed to obtain r, and then the values of P[i] and P[r] are exchanged to meet the specific data format requirements of number theory transformation. The input reads the data that needs to be bit reversed from the storage module. The output writes the reversed data back to the storage module for further processing by other modules.
[0045] The butterfly calculation module is responsible for executing the core butterfly operations of NTT and INTT, and is a key module for realizing polynomial transformation. It supports parallel computing mode, can make full use of hardware resources, greatly improve the computing speed, and effectively improve the efficiency of the device in processing polynomial computing tasks. The input reads the polynomial coefficients and the corresponding rotation factors from the storage module as the operands of the butterfly calculation. The output writes the results of the butterfly calculation into the storage module, and at the same time transmits the results to the modular multiplication module for modular operation processing.
[0046] The multiplication module performs point-by-point multiplication operations on the two input polynomials to obtain new polynomial product data. It can support fixed-point or floating-point multiplication operations, and flexibly select the appropriate operation method according to different accuracy requirements to adapt to various computing scenarios. The input reads the two polynomial coefficient data involved in the multiplication operation from the storage module. The output writes the calculated product result to the storage module, and sends the result to the modular multiplication module for modular operation.
[0047] The modular multiplication module performs modular operations on the results output by the butterfly calculation module or the multiplication module to ensure that the calculation results are within a finite field and meet the requirements of polynomial calculations in a specific mathematical domain. For example, the input receives the output results from the butterfly calculation module and the multiplication module. The output writes the results after modular operation back to the storage module for subsequent result output or further processing.
[0048] Taking NTT task execution as an example, the initialization module receives the polynomial P and NTT rotation factor W data, writes them into the storage module after preprocessing. The butterfly calculation module obtains P and W from the storage module, performs the core butterfly calculation operation, and stores the calculation result back to the storage module. The modular multiplication module performs modular operation on the butterfly calculation result to ensure that the result is within the finite field. The reversal module performs bit reverse order processing on the data in the storage module. The storage module outputs the final processed result through the reading result module.
[0049] Taking the multiplication task as an example, the initialization module receives two polynomials P1 and P2 data, writes them into the storage module after preprocessing. The multiplication module reads P1 and P2 from the storage module, performs point-by-point multiplication, and writes the product result back to the storage module. The modular multiplication module performs modular operation on the result obtained by multiplication. The storage module outputs the final result through the reading result module.
[0050] In this way, modular decoupling of each hardware processing module is independently designed with good independence and encapsulation, which makes it more convenient for maintenance and upgrading. For example, the multiplication module can be easily upgraded to a higher precision version without affecting the normal operation of other modules. Storage resource reuse The time-sharing cache data mechanism of the storage module effectively reduces the occupation of hardware resources, improves the utilization of storage resources, and achieves more efficient data processing under limited hardware resource conditions. Parallel processing capability The butterfly computing module and the multiplication module support parallel work, give full play to the parallel computing capability of the hardware, significantly improve the data processing throughput of the entire device, and speed up the task execution speed. In particular, dynamic path adaptation can flexibly combine various modules according to different task types, such as NTT, INTT or multiplication tasks, without large-scale modification of the hardware circuit, which enhances the adaptability and flexibility of the device. It can efficiently realize polynomial multiplication acceleration on hardware platforms such as FPGA, while taking into account flexibility and resource utilization, providing strong support for applications in related fields. In the embodiment of the present application, the reading result module is the key part responsible for outputting the final data in the data processing device, and plays an important role in the data flow and task completion process of the whole system. The core function of the reading result module is to read the final data after the task execution unit is processed from the storage module. These data are the task execution results obtained after a series of processing such as the initialization module, the reversal module, the butterfly calculation module, the multiplication module, the modular multiplication module, etc., such as the polynomial after the number theory transformation, the product result after the multiplication operation, or the final polynomial after the inverse number theory transformation. The module can accurately locate the location of the storage result data in the storage module, and read the data according to the predetermined rules and format. The task execution results read from the storage module are output to the task request end. It is responsible for converting the data in the internal processing format into a format suitable for the task request end to receive and process, and ensuring that the data can be correctly received and used by the task request end. For example, converting the data in binary format into a specific encoding format or encapsulating the data according to the protocol required by the task request end.
[0051] Further optionally, before outputting the data, the reading result module may perform certain verification operations on the task execution results. By checking the integrity and correctness of the data, it is ensured that the data output to the task request end is reliable. For example, the checksum of the data is calculated and compared with a preset value, or the data is checked to see if it complies with a specific format specification. If a problem is found in the data, the reading result module can take corresponding measures, such as re-reading the data, feeding back error information to the controller module, etc.
[0052] The read result module first waits for the output instruction of the controller module. When the task execution unit completes the task processing and the controller module determines that the result can be output, it will send an output instruction to the read result module. The instruction contains some key information about the result data, such as the storage location of the data, the length of the data, the format of the data, etc. The read result module prepares to perform data reading operations based on this information. According to the information provided by the controller module, the read result module accesses the storage module and reads the task execution result data from the specified location. During the reading process, it will operate according to the access rules and interface protocol of the storage module to ensure that the data can be read accurately. If the amount of data is large, the read result module may adopt a batch reading method to avoid reading too much data at one time and causing system resource shortage. After reading the data, the read result module performs necessary processing on the data, such as format conversion, data parsing, etc., to make it meet the requirements of the task request end. At the same time, the data is verified to check the correctness and integrity of the data. If the verification passes, continue to the next step of output operation; if the verification fails, take corresponding processing measures according to the specific situation, such as re-reading the data or reporting errors to the controller module. The processed and verified data is output to the task request end. The read result module sends the data according to the communication protocol and interface specifications agreed with the task requester. During the output process, it monitors the data transmission status to ensure that the data can be successfully delivered to the task requester. If there is a problem during the transmission process, such as data loss, communication interruption, etc., the read result module will try to resend the data or feedback the abnormal situation to the controller module.
[0053] From the connection relationship, the read result module is directly connected to the storage module, and reads the task execution result data from the storage module. The storage module stores the intermediate results and final results of each stage during the task execution process. The read result module performs data access and reading operations through the interface with the storage module. The read result module receives the control instructions from the controller module. The controller module decides when to let the read result module read and output the results according to the progress and status of the task execution, and provides relevant parameters and information to the read result module. At the same time, the read result module will also feedback the status information of the data reading and output process to the controller module, such as whether the data is successfully read and whether the data is successfully output. The read result module outputs the task execution result to the task request end. It transmits data with the task request end through a specific communication interface and protocol to ensure that the data can be accurately received and processed by the task request end.
[0054] The result reading module is a key link in the data interaction between the data processing device and the task request end. It ensures that the data processed inside the device can be smoothly returned to the task request end, so that the task request end can obtain the required calculation results and realize the effective output and interaction of data. By verifying and processing the task execution results, the result reading module improves the reliability and accuracy of the output data. This is very important for application scenarios that rely on data processing results, ensuring that users can get the correct calculation results and avoiding problems and losses caused by data errors. As the last link of the data processing device, the existence of the result reading module improves the function of the entire system. It works closely with other modules to complete the complete process from data input, processing to output, so that the data processing device can operate efficiently and reliably to meet the needs of different users and application scenarios.
[0055] For example, Figure 2 Taking the large bit width long sequence polynomial multiplication acceleration circuit architecture shown as an example, multiple modules work together to achieve polynomial multiplication acceleration. The controller module receives task selection signals and other parameter signals, contains a finite state machine, determines the target task according to the signal, coordinates and schedules the work of other modules, and is the control core of the entire architecture. The initialization module receives input data, performs preprocessing, and temporarily stores the data in the storage module to prepare for subsequent calculations. The two storage modules are used to store input data and intermediate calculation results, and play a role in data caching, so that each module can read the required data at any time. In tasks such as number theory transformation, the reversal module performs bit reverse order processing on the data to meet specific calculation format requirements. The butterfly calculation module performs butterfly operations in number theory transformation and is a key operation module for realizing polynomial transformation. The multiplication module performs multiplication operations on the input polynomial data and calculates the polynomial product. The modular multiplication module performs modular operations on the butterfly calculation or multiplication results to ensure that the results are within the finite field. The reading result module reads the final calculation result from the storage module and outputs the calculation result. exist Figure 2 In the process, input data enters the initialization module, which processes and stores it in the storage module. Each computing module (reversal, butterfly calculation, multiplication, modular multiplication) reads data from the storage module for processing, and stores the processing results back to the storage module. The modules interact with each other through the storage module. The reading result module obtains the final result from the storage module and outputs it, completing the entire data transmission and processing flow. Through this architecture and data transmission method, the accelerated calculation of large-bit width long sequence polynomial multiplication is achieved.
[0056] As an optional embodiment, the storage module includes a first storage module and a second storage module; the first storage module and the second storage module respectively occupy a unit storage space. When the task execution unit selects the corresponding hardware processing module based on the target task to construct the corresponding target task path, it is specifically used to: the task execution unit selects the corresponding target processing unit from the first storage module, the second storage module, the reversal module, the butterfly calculation module, the multiplication module, and the modular multiplication module based on the target task; the selected target processing units are combined into corresponding data processing links, and the combined data processing links are used as the target task path.
[0057] For example, in the number theory transformation task, the polynomial data to be processed is read from the first storage module or the second storage module, the reversal module is selected to pre-process the data by reversing the bit order, and then the butterfly calculation module performs the core butterfly operation, and the operation result is processed by the modular multiplication module, and finally the result is stored back to the storage module. This process selects the storage module, the reversal module, the butterfly calculation module, and the modular multiplication module to form the data processing link of the number theory transformation task.
[0058] For example, in a multiplication task, two polynomial data involved in the multiplication operation are read from two storage modules respectively, and the data is input into the multiplication module for point-by-point multiplication. The multiplication result is subjected to modular operation by the modular multiplication module, and the final result is stored back into the storage module. Here, the storage module, multiplication module, and modular multiplication module are selected to construct the target task path of the multiplication task and efficiently complete the multiplication calculation.
[0059] Thus, by setting up two storage modules and independently occupying unit storage space, the task execution unit can read data from different storage modules as needed, and build task paths in combination with other hardware processing modules to achieve flexible allocation of resources. For example, in different polynomial calculation tasks, different types of data can be stored in two modules respectively to avoid resource conflicts and improve resource utilization. Selecting target processing units from multiple modules to build data processing links can customize exclusive execution paths for different target tasks. For simple multiplication tasks, only multiplication modules and storage modules can be selected to build links; for complex number theory transformation combination tasks, butterfly calculation modules, reversal modules, etc. can be included in the link to accurately match task requirements and improve execution efficiency. This modular selection and combination method makes it easy to add or modify modules when new task requirements arise. In the future, if there are new polynomial calculation algorithms or task types, they can adapt to new requirements by simply expanding on the basis of existing modules, reducing the cost of hardware deployment and upgrades.
[0060] As an optional embodiment, the controller module is also used to control the data processing status of the data to be processed in the target task path; when the task execution unit inputs the data to be processed into the target task path to execute the target task, it is specifically used to: under the control and scheduling of the controller module, the task execution unit loads the data to be processed into the target task path, and switches the hardware processing unit where the data to be processed is located in the target task path to execute the corresponding data processing flow in the target task.
[0061] For example, when the target task is a number theory transformation task, the controller module issues an instruction to control the data to be processed to first enter the initialization module, and after completing the data preprocessing, it is loaded into the first storage module or the second storage module for temporary storage. Next, the controller schedules the data to enter the reversal module for bit reverse order. Subsequently, the data is switched to the butterfly calculation module to perform the butterfly operation, and the result after the operation enters the modular multiplication module for modular operation processing. During the whole process, the controller module continuously monitors and controls the data processing status to ensure that the data is accurately switched in sequence between the hardware processing units to successfully complete the number theory transformation task.
[0062] For example, for the combined task of multiplication and number theory transformation, the controller first guides the data to be processed into the initialization module, and then stores it in the corresponding storage module. First, according to the multiplication task process, the data is switched to the multiplication module for multiplication operation. After being processed by the modular multiplication module, according to the requirements of the number theory transformation task, the control data is sequentially entered into the reversal module, butterfly calculation module, modular multiplication module, etc. for processing, so as to realize the coherent execution of the combined task.
[0063] In this way, the controller module's control of the data processing status can accurately control the processing timing and order of the data to be processed in each hardware processing unit, avoid data processing confusion, and ensure that the task is executed according to the expected process. By flexibly switching the hardware processing unit in the target task path, the functions of each module can be fully utilized, the data waiting time can be reduced, the task execution can be accelerated, and the overall computing efficiency can be improved. Regardless of simple or complex combination tasks, this mechanism can adapt to different task requirements through the controller's scheduling and data processing unit switching, enhance the device's processing capabilities for various polynomial computing tasks, and improve the system's versatility and flexibility.
[0064] As an optional embodiment, the controller module includes a first state machine; the first state machine is used to receive a task start signal before determining the target task to be performed by the task execution unit based on the task selection signal; based on the task start signal, the initialization module is switched to the initialization memory state to trigger the initialization module to perform the initialization operation on the data to be processed; after receiving the initialization end signal, it continues to stay in the initialization memory state and triggers the step of determining the target task to be performed by the task execution unit based on the task selection signal. The initialization module is also used to perform initialization analysis on the data to be processed, and after the initialization is completed, the initialization end signal is fed back to the first state machine.
[0065] In this data processing device, the first state machine is in the initial kIdle (idle) state. At this time, it can receive the start signal poly_start from the outside. If poly_start is 0, it means that the task has not started, and the first state machine continues to stay in the kIdle state, waiting for the start instruction. When poly_start becomes 1, the first state machine receives the start signal, and the state changes from kIdle to kInitMem (initialization memory state). After entering this state, the first state machine controls the initialization module to start working and performs initialization operations on the input data to be processed, such as data format conversion, storage location allocation, etc. During the initialization process, the initialization module will continue to process data. As long as it has not completed the initialization, it will feedback poly_finish_in=0 to the first state machine, and the first state machine will continue to stay in the kInitMem state. Once the initialization is completed, the initialization module will feedback poly_finish_in=1 to the first state machine. At this time, although the first state machine is still in the kInitMem state, it will determine the target task to be executed by the task execution unit next according to the received task selection signal control_signal. For example, if control_signal represents a number theory transformation task, the first state machine will instruct subsequent modules to perform corresponding number theory transformation operations.
[0066] In this way, the first state machine decides whether to start the task based on the poly_start signal, avoiding false triggering of the task. Only after receiving a clear start signal will the system enter the initialization process, ensuring the accuracy and reliability of task startup. During the initialization process, the first state machine determines whether the initialization is completed based on the poly_finish_in signal. Only when the initialization is truly completed will the target task be determined based on the control_signal, avoiding subsequent calculation errors due to incomplete initialization and improving the stability of the system. Determining the target task based on the control_signal allows the system to flexibly configure tasks according to different needs. This enhances the adaptability of the system, and can cope with a variety of different types of polynomial calculation tasks to meet a variety of application scenarios.
[0067] As an optional embodiment, when the controller module determines the target task to be performed by the task execution unit based on the task selection signal, it is specifically used to: if the task selection signal is a first task signal, the controller module determines that the target task is a number theory transformation task NTT.
[0068] The controller module includes a second state machine; the second state machine is used to switch the task execution unit from the initialization memory state to the rotation factor loading state, butterfly calculation state, reversal state, and result output state in sequence; in the number theory conversion task, the data to be processed includes polynomial coefficients and NTT rotation factors.
[0069] In NTT, the target task path consists of the initialization module, the butterfly calculation module, and the reversal module; in the rotation factor loading state, the initialization module is used to obtain the NTT rotation factor from the second storage module and perform initialization analysis on the NTT rotation factor; if the analysis is completed, the NTT rotation factor analysis results are returned to the second storage module respectively, and a rotation factor analysis completion signal is sent to the second state machine to switch the second state machine to the butterfly calculation state; in the butterfly calculation state, the butterfly calculation module is used to perform butterfly calculation on the polynomial coefficients in the first storage module and the NTT rotation factor analysis results in the second storage module; the obtained butterfly calculation intermediate result is stored in the first storage module, and a butterfly calculation completion signal is sent to the second state machine to switch the second state machine to the reversal state; in the reversal state, the reversal module is used to perform bit reversal permutation on the butterfly calculation intermediate result in the first storage module to obtain the NTT calculation result; a reversal completion signal is sent to the second state machine to switch the second state machine to the result output state, and the NTT calculation result is output through the reading result module in the result output state.
[0070] In the above embodiment, the controller module determines the target task of the task execution unit by means of the task selection signal. When the task selection signal is the first task signal, it is determined to execute the number theory transformation task (NTT). This process is precisely controlled by the second state machine in the controller module, which guides the task execution unit in an orderly manner from the initialization memory state to the rotation factor loading state, the butterfly calculation state, the reversal state and the result output state.
[0071] Exemplarily, when control_signal=2'b00, the circuit needs to perform the number theory transformation task, and the state machine will enter the kWaitNTTTF state from the KInitMem state. This state is waiting for the rotation factor of the number theory transformation to complete initialization. When ntt_tf_is_avlb=0 sent by the initialization module, it means that the rotation factor of the number theory transformation has not completed initialization, and the state machine stays in the kWaitNTTTF state. When ntt_tf_is_avlb=1, the rotation factor of the number theory transformation has completed initialization, and the state machine will enter the kPolyNTT state. This state is the butterfly calculation of the number theory transformation. When ntt_btf_finish_in=0 sent by the butterfly calculation module, it means that the butterfly calculation module has not completed the calculation, and the state machine stays in the kPolyNTT state. When ntt_btf_finish_in=1, the butterfly calculation module completes the calculation, and the state machine stays in the kPolyNTT state. The state machine will enter the kReverse state, which is to reverse the data written back to the first storage module by the butterfly calculation module. When reverse_finish_in=0 sent by the reversal module, it means that the reversal module has not completed the reversal operation, and the state machine stays in the kReverse state. When reverse_finish_in=1, the reversal module has completed the reversal operation, and the state machine will enter the kReadResult state, which is to output the calculation result of the number theory transformation. When read_finish_in=0 sent by the result reading module, it means that the output is not completed, and the state machine stays in the kReadResult state. When read_finish_in=1, the calculation result has been completed, and the state machine will enter the kFinish state. Next, the state machine will automatically jump to the kIdle state and wait for the next poly_start signal from the outside world. For the above state switching process and related data processing flow, see Figure 3 shown.
[0072] Based on the above example, in the initialization memory state, when the system receives the task start signal, it enters the initialization memory state, and the initialization module stores the polynomial coefficients and NTT rotation factors to be processed in the first and second storage modules respectively. In the rotation factor loading state, after the second state machine switches to the rotation factor loading state, the initialization module obtains the NTT rotation factor from the second storage module and parses it. After the parsing is completed, the parsing result is returned to the second storage module, and the rotation factor parsing completion signal is sent to the second state machine to push the state machine into the butterfly calculation state. This is similar to waiting for the rotation factor to be initialized in the example. In the example, the ntt_tf_is_avlb signal is used to determine whether the rotation factor is initialized, and this embodiment implements state switching through the parsing completion signal. In the butterfly calculation state, the butterfly calculation module reads the polynomial coefficients from the first storage module, reads the NTT rotation factor parsing result from the second storage module, and performs butterfly calculation. The intermediate result obtained by calculation is stored in the first storage module, and a butterfly calculation completion signal is sent to the second state machine, prompting the state machine to enter the reversal state. This is similar to the mechanism in the example where the ntt_btf_finish_in signal is used to determine whether the butterfly calculation is completed and then switch the state. In the reverse state, the reversal module performs bit-reversal permutation on the intermediate result of the butterfly calculation in the first storage module to obtain the NTT calculation result, and sends a reversal completion signal to the second state machine to make the state machine enter the result output state. This is similar to the example where the reverse_finish_in signal is used to determine whether the reversal operation is completed to switch the state. In the result output state, the reading result module outputs the NTT calculation result. For the above state switching process and related data processing flow, see Figure 3 shown.
[0073] In this way, the second state machine performs phased and refined control over the task execution process. The switching of each state is based on the completion signal of the previous stage, ensuring the orderly execution of tasks and reducing the probability of errors. By clarifying the operations of each module in different states, efficient data flow between hardware processing units is achieved, avoiding idle resources and improving the utilization of hardware resources. In this way, the state switching and signal feedback mechanism enable the system to respond to the state changes of each module in a timely manner when executing NTT tasks, effectively improving the stability and reliability of the system.
[0074] As an optional embodiment, if there are multiple NTT rotation factors, the initialization module is further used to load the multiple NTT rotation factors into the second storage module at different times, so as to respectively calculate and obtain multiple NTT calculation results corresponding to the multiple NTT rotation factors.
[0075] As an optional embodiment, when the controller module determines the target task to be executed by the task execution unit based on the task selection signal, it is specifically used to: if the task selection signal is a second task signal, the controller module determines that the target task is a multiplication task.
[0076] The controller module includes a third state machine; the third state machine is used to switch the task execution unit from the initialization memory state to the multiplication calculation state and the result output state in sequence; in the multiplication task, the data to be processed includes polynomial coefficients.
[0077] In the multiplication task, the target task path consists of the initialization module and the multiplication module; in the multiplication calculation state, the multiplication module is used to perform multiplication calculation on the first polynomial coefficients stored in the first storage module and the second polynomial coefficients stored in the second storage module; the product result is written into the first storage module, and a multiplication calculation completion signal is sent to the third state machine, so that the second state machine switches to the result output state, and outputs the product result through the reading result module in the result output state.
[0078] For example, assuming that the task selection signal control_signal is used as an external input, when control_signal is assigned to the second task signal, which is equivalent to l2'b01 in the circuit logic, the system starts to execute the multiplication task. Initialize the memory state. After the system receives the task start signal, it enters the initialize memory state. The initialization module receives two sets of polynomial coefficients from the outside, and stores the first polynomial coefficient in the first storage module and the second polynomial coefficient in the second storage module. In the multiplication calculation state, the third state machine switches the task execution unit to the multiplication calculation state. The multiplication module reads the first polynomial coefficient from the first storage module and the second polynomial coefficient from the second storage module to perform multiplication operations. During the operation, the multiplication module continues to process data. At this time, if the multiplication module feeds back poly_mult_finish_in = 0 to the third state machine, the third state machine keeps the task execution unit in the multiplication calculation state. After the multiplication module completes the calculation, it can send a poly_mult_finish_in = 1 signal to the third state machine, which is just like the poly_mult_finish_in signal feedback mechanism in the example. The third state machine switches the task execution unit to the result output state based on the signal. In the result output state, the result reading module obtains the product result obtained by the multiplication calculation from the first storage module, and outputs the result to the outside, completing the execution of the entire multiplication task. The subsequent state jump is the same as when control_signal = 2'b00, that is, when the result reading module sends read_finish_in = 1, the entire process enters the end state, and then automatically returns to the idle state, waiting for the next task start signal.
[0079] Therefore, for multiplication tasks, the complex tasks are broken down into three stages: initialization, multiplication calculation, and result output. The third state machine performs orderly switching. Compared with the solution without state machine control, the execution process of the multiplication task is greatly simplified and the complexity of the design is reduced. Through signal feedback and state machine state switching, the controller module can monitor the execution status of the multiplication task in real time and realize precise control of task execution. When an exception occurs in task execution, it can also respond and process it in time, which improves the stability and reliability of the system. For the above state switching process and related data processing flow, see Figure 3 shown.
[0080] As an optional embodiment, when the controller module determines the target task to be executed by the task execution unit based on the task selection signal, it is specifically used to: if the task selection signal is a third task signal, the task execution unit determines that the target task is an inverse number theory transformation task INTT.
[0081] The controller module includes a fourth state machine; the fourth state machine is used to switch the task execution unit from the initialization memory state to the reversal state, the rotation factor loading state, the butterfly calculation state, and the result output state in sequence; in the number theory conversion task, the data to be processed includes polynomial coefficients and INTT rotation factors.
[0082] In the number theory conversion task, the target task path is composed of the initialization module, the butterfly calculation module, and the reversal module; in the reversal state, the reversal module is used to bit-reverse and permute the polynomial coefficients in the first storage module, and write the obtained polynomial reversal result back to the first storage module; a reversal completion signal is sent to the fourth state machine to switch the fourth state machine to the rotation factor loading state; in the rotation factor loading state, the butterfly calculation module is also used to butterfly calculate the polynomial reversal result in the first storage module and the INTT rotation factor result in the second storage module to obtain the INTT calculation result; a butterfly calculation completion signal is sent to the fourth state machine to switch the fourth state machine to the result output state, and output the INTT calculation result through the reading result module in the result output state.
[0083] Exemplarily, assuming that the task selection signal control_signal is used as an external input, when the control_signal is assigned to the third task signal, which is equivalent to 2'b10 in the circuit logic, the system starts the inverse number theory transformation task. The controller module receives the task start signal and enters the initialization memory state. The initialization module stores the polynomial coefficients and INTT rotation factors to be processed in the first and second storage modules respectively. The fourth state machine switches the task execution unit to the reversal state, and the reversal module reads the polynomial coefficients from the first storage module and performs a bit reversal permutation operation. During the operation, if the reversal module feeds back reverse_finish_in = 0 to the fourth state machine, the fourth state machine maintains the task execution unit in the reversal state; when the reversal operation is completed, the reversal module sends a reverse_finish_in = 1 signal to the fourth state machine, which is consistent with the feedback mechanism of the reverse_finish_in signal in the given example, and the fourth state machine switches the task execution unit to the rotation factor loading state accordingly. After entering the rotation factor loading state, the butterfly calculation module reads the polynomial reversal result from the first storage module, reads the INTT rotation factor from the second storage module, and performs butterfly calculation. During calculation, if the butterfly calculation module feeds back intt_btf_finish_in = 0 to the fourth state machine, the fourth state machine keeps the task execution unit in this state; when the butterfly calculation module completes the calculation, it sends a signal intt_btf_finish_in = 1 to the fourth state machine, and the fourth state machine switches the task execution unit to the result output state. This part of the signal feedback and state switching mechanism is similar to the intt_btf_finish_in signal controlling the butterfly calculation state switching in the example. In the result output state, the result reading module obtains the INTT calculation result from the first storage module and outputs it to the outside to complete the inverse number theory transformation task. The subsequent state jump is the same as when control_signal = 2'b00. When the result reading module sends read_finish_in = 1, the process enters the end state, and then automatically returns to the idle state, waiting for the next task start signal. For the above state switching process and related data processing flow, see Figure 3 shown.
[0084] Therefore, with the help of the fourth state machine, the inverse number theory transformation task is decomposed into multiple stages such as initialization, reversal, rotation factor loading and result output. The function of each stage is clear and the state switching is orderly. Compared with the complex out-of-order execution scheme, the task execution logic is greatly simplified, and the readability and maintainability of the design are improved. Through the feedback signal and the state switching of the state machine, the controller module can monitor the execution status of the inverse number theory transformation task in real time and realize precise control of the task. When an exception occurs during the task execution, the system can respond and handle it in time, significantly improving the stability and reliability of the system. The target task path is composed of the initialization module, the butterfly calculation module and the reversal module, which avoids unnecessary module participation and reduces resource consumption. The modules cooperate with each other during the task execution process, effectively improving the efficiency of hardware resource utilization.
[0085] Figure 3 The working principle of a finite state machine for at least one of number theoretic transformation, inverse number theoretic transformation, and multiplication is shown. Figure 3 In the kldle (idle) state, the initial state of the state machine, waiting for the start signal poly_start. When poly_start is 0, it remains in this state; when poly_start becomes 1, it switches to the klnitMem (initialization memory) state, marking the start of the task. In the klnitMem state, the initialization module processes and stores the input data, and enters the subsequent state according to different conditions after completion. During the initialization process, the initialization module will continue to process data. If the initialization is not completed, the initialization module will feedback poly_finish_in=0 to the state machine, and the state machine will continue to stay in the kInitMem state. Once the initialization is completed, the initialization module will feedback poly_finish_in=1 to the state machine. At this time, although the state machine is still in the kInitMem state, it will determine the target task to be executed by the task execution unit next based on the received task selection signal control_signal.
[0086] exist Figure 3In the kWaitNTTTF state, if the conditions for entering kWaitNTTTF are met (such as waiting for the relevant signals for initialization of the number-theoretic transformation twiddle factors to be met), switch to this state. If the conditions for entering kPolyMult are met (such as the state switching conditions for triggering the multiplication task), switch to this state. If the conditions for entering kReverse are met (such as the completion of some data preprocessing), switch to the kReverse state. In the kWaitNTTTF (waiting for the number-theoretic transformation twiddle factors) state, wait for the initialization of the number-theoretic transformation twiddle factors to be completed. When the relevant signals (such as ntt_tf_is_avlb=1) are met, indicating that the twiddle factors have been initialized, switch to the kPolyNTT state. In the kPolyNTT (performing number-theoretic transformation butterfly calculation) state, the butterfly calculation module performs butterfly calculations. When the calculation completion signal (such as ntt_btf_finish_in=1) appears, switch to the kReverse state. In the kReverse (reverse) state, the reversal module performs a bit reversal permutation operation. When the operation is completed (such as reverse_finish_in=1), enter different states according to the conditions. If the conditions for entering kWaitINTTTF are met (such as preparing for the inverse number theory transform and waiting for the inverse rotation factor to be initialized), switch to this state. If the conditions for entering kReadResult are met (such as preparing to output the result), switch to the kReadResult state. kWaitINTTTF (waiting for the inverse number theory transform rotation factor) state, wait for the inverse number theory transform rotation factor to be initialized, and switch to the kPolyINTT state after the conditions are met. kPolyINTT (performing the inverse number theory transform butterfly calculation) state, execute the inverse number theory transform butterfly calculation, and switch to the kReadResult state after completion. kReadResult (output result) state, read the result module to output the calculation result, and the result output is completed (such as read_finish_in=1), and switch to the kFinish state. In the kFinish (finished) state, the task is executed, and then it automatically jumps to the kldle state, waiting for the next task start signal. For the specific state switching conditions and switching process, please refer to the relevant description in the previous embodiment, which will not be repeated here.
[0087] Figure 3 The state machine shown coordinates the work of each hardware processing unit in the number theory transformation task through orderly state switching and control of each module, accurately controls the data processing flow and state, ensures the accurate and efficient execution of the number theory transformation task, and realizes complete process control from data initialization, rotation factor preparation, butterfly calculation, data reversal to result output.
[0088] In the prior art solution, a data processing device is used to perform number theory transformation, multiplication and inverse number theory transformation as a whole. When calculating the multiplication of two polynomials of length N, the amount of data that needs to be stored is two polynomial coefficient data of length N, and two rotation factors of length N, a total of 4N data. When the data bit width is large (such as 255 bits) and the polynomial length is long (such as 2×18), the amount of data required to be stored is huge, which will exceed the on-board storage resources of most FPGAs on the market. In the embodiment of the present application, by decoupling the tasks, one task is executed at a time. When performing number theory transformation or inverse number theory transformation, only one polynomial coefficient data of length N and one rotation factor of length N need to be stored; when performing multiplication, two polynomial coefficient data of length N are stored, and the total amount of data is reduced to 2N. This saves 50% of storage space compared to the prior art. Under the same hardware resource conditions, due to the saving of storage space, this solution can support the calculation of 2 times the sequence length. For example, under specific hardware resources, only calculations with shorter sequence lengths could be supported. However, with this solution, longer sequence lengths can be supported, thus increasing the scale of calculations.
[0089] The existing technical solutions are not flexible enough when performing calculations, and must be executed according to a complete set of number theory transformation, multiplication, and inverse number theory transformation. For polynomial multiplication of larger lengths (such as 2×30), it is difficult to implement ordinary number theory transformation and inverse number theory transformation on FPGA, and the existing technical solutions cannot support matrix number theory transformation and its inverse transformation.
[0090] In the embodiment of the present application, the state jump process of the finite state machine is controlled by the task selection signal to realize any one or more tasks in number theory transformation, multiplication or inverse number theory transformation, and the calculation process can be flexibly changed according to the needs. It can not only support ordinary number theory transformation and its inverse transformation, but also support matrix number theory transformation and its inverse transformation. Under similar hardware resource conditions, the sequence length of number theory transformation and its inverse transformation supported by this scheme has achieved a square increase. For example, the maximum sequence length that can be supported in the related technology is 2×15. After adopting the technical solution adopted in the embodiment of the present application, the maximum sequence length that can be supported can be increased to 2×30, which greatly expands the range of polynomial sequence length that can be processed and enhances the applicability and computing power of the scheme. In general, in the embodiment of the present application, by decoupling and flexible control of tasks, the computing efficiency and utilization rate of hardware resources are improved, which can better meet the needs of polynomial calculation in practical applications.
[0091] Figure 4 The flowchart of a data processing method provided in the embodiment of the present application is shown in FIG. The method is applied to the data processing device described in the above embodiment. Figure 4 As shown, the method comprises the following steps: 401, a data processing device receives a task selection signal from a task requesting end, and determines a target task to be executed by a task executing unit based on the task selection signal; 402, the data processing device selects a corresponding hardware processing module based on the target task to construct a corresponding target task path; 403, the data processing device receives the data to be processed, and inputs the data to be processed into the target task path to start the target task and obtain the task execution result; 404, the data processing device outputs the task execution result to the task requesting end.
[0092] In the embodiment of the present application, the target task is one or more polynomial calculation tasks. The polynomial calculation tasks at least include: number theory change tasks, multiplication tasks, and inverse number theory change tasks.
[0093] In the embodiment of the present application, the number, type and connection mode of the hardware processing modules in the target task path are related to the task type to which the target task belongs. The hardware processing modules in the task execution unit include: initialization module, storage module, reversal module, butterfly calculation module, multiplication module and modular multiplication module.
[0094] Further optionally, the hardware processing module in the task execution unit includes: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module; wherein the initialization module is connected to the storage module; the initialization module is used to receive the data to be processed and temporarily store the data to be processed in the storage module; wherein the storage module is respectively connected to the reversal module, the butterfly calculation module, and the multiplication module, and the butterfly calculation module and the multiplication module are respectively connected to the modular multiplication module; the data processing results of the reversal module, the butterfly calculation module, the multiplication module, and the modular multiplication module are output to the reading result module through the storage module.
[0095] Further optionally, the storage module includes a first storage module and a second storage module; the first storage module and the second storage module respectively occupy a unit storage space; the selecting a corresponding hardware processing module based on the target task to construct a corresponding target task path includes: Based on the target task, a corresponding target processing unit is selected from the first storage module, the second storage module, the reversal module, the butterfly calculation module, the multiplication module, and the modular multiplication module; the selected target processing units are combined into corresponding data processing links, and the combined data processing links are used as the target task path.
[0096] Further optionally, the method also includes: controlling the data processing status of the data to be processed in the target task path; inputting the data to be processed into the target task path to execute the target task, including: under the control and scheduling of the controller module, loading the data to be processed into the target task path, and switching the hardware processing unit where the data to be processed is located in the target task path to execute the corresponding data processing flow in the target task.
[0097] Further optionally, the controller module includes a first state machine; before determining the target task to be performed by the task execution unit based on the task selection signal, the method further includes: receiving a task start signal through the first state machine; based on the task start signal, switching the initialization module to the initialization memory state to trigger the initialization module to perform an initialization operation on the data to be processed; after receiving the initialization end signal, continuing to stay in the initialization memory state, and triggering the step of determining the target task to be performed by the task execution unit based on the task selection signal. After the initialization parsing of the data to be processed is completed, it also includes: feeding back the initialization end signal to the first state machine through the initialization module.
[0098] Further optionally, determining the target task to be executed by the task execution unit based on the task selection signal includes: if the task selection signal is a first task signal, determining the target task to be a number theory transformation task NTT.
[0099] The controller module includes a second state machine; the control of the data processing state of the data to be processed in the target task path includes: switching the task execution unit from the initialization memory state to the rotation factor loading state, butterfly calculation state, reversal state, and result output state in sequence through the second state machine; in the number theory conversion task, the data to be processed includes polynomial coefficients and NTT rotation factors.
[0100] In NTT, the target task path is composed of the initialization module, the butterfly calculation module, and the reversal module; under the control and scheduling of the controller module, the data to be processed is loaded into the target task path, and the hardware processing unit where the data to be processed is located is switched in the target task path, including: in the rotation factor loading state, the NTT rotation factor is obtained from the second storage module through the initialization module, and the NTT rotation factor is initialized and parsed; if the parsing is completed, the NTT rotation factor parsing results are returned to the second storage module respectively, and a rotation factor parsing completion signal is sent to the second state machine, so that the second state machine switches to the butterfly calculation state; in In the butterfly calculation state, the polynomial coefficients in the first storage module and the NTT rotation factor analysis results in the second storage module are subjected to butterfly calculation through the butterfly calculation module; the obtained intermediate result of the butterfly calculation is stored in the first storage module, and a butterfly calculation completion signal is sent to the second state machine to switch the second state machine to the reversal state; in the reversal state, the intermediate result of the butterfly calculation in the first storage module is bit-reversed and replaced through the reversal module to obtain the NTT calculation result; a reversal completion signal is sent to the second state machine to switch the second state machine to the result output state, and the NTT calculation result is output through the reading result module in the result output state.
[0101] Further optionally, the method also includes: if there are multiple NTT rotation factors, loading the multiple NTT rotation factors into the second storage module at different times respectively, so as to respectively calculate and obtain multiple NTT calculation results corresponding to the multiple NTT rotation factors.
[0102] Further optionally, determining the target task to be executed by the task execution unit based on the task selection signal includes: if the task selection signal is a second task signal, determining that the target task is a multiplication task.
[0103] The controller module includes a third state machine; the control of the data processing state of the data to be processed in the target task path includes: switching the task execution unit from the initialization memory state to the multiplication calculation state and the result output state in sequence through the third state machine; in the multiplication task, the data to be processed includes polynomial coefficients.
[0104] In a multiplication task, the target task path is composed of the initialization module and the multiplication module; under the control and scheduling of the controller module, the data to be processed is loaded into the target task path, and the hardware processing unit where the data to be processed is located is switched in the target task path, including: in the multiplication calculation state, the first polynomial coefficients stored in the first storage module are multiplied by the second polynomial coefficients stored in the second storage module through the multiplication module; the product result is written into the first storage module, and a multiplication calculation completion signal is sent to the third state machine, so that the second state machine switches to the result output state, and the product result is output through the reading result module in the result output state.
[0105] Further optionally, determining the target task to be executed by the task execution unit based on the task selection signal includes: if the task selection signal is a third task signal, determining the target task to be an inverse number theory transformation task INTT.
[0106] The controller module includes a fourth state machine; the control of the data processing state of the data to be processed in the target task path includes: switching the task execution unit from the initialization memory state to the reversal state, the rotation factor loading state, the butterfly calculation state, and the result output state in sequence through the fourth state machine; in the number theory conversion task, the data to be processed includes polynomial coefficients and INTT rotation factors.
[0107] In the number theory conversion task, the target task path is composed of the initialization module, the butterfly calculation module, and the reversal module; under the control and scheduling of the controller module, the data to be processed is loaded into the target task path, and the hardware processing unit where the data to be processed is located is switched in the target task path, including: in the reversal state, the polynomial coefficients in the first storage module are bit-reversed and permuted by the reversal module, and the obtained polynomial reversal result is written back to the first storage module; a reversal completion signal is sent to the fourth state machine to switch the fourth state machine to the rotation factor loading state; in the rotation factor loading state, the polynomial reversal result in the first storage module and the INTT rotation factor result in the second storage module are butterfly calculated by the butterfly calculation module; the obtained butterfly calculation intermediate result is stored in the first storage module, and a butterfly calculation completion signal is sent to the fourth state machine to switch the fourth state machine to the result output state, and the INTT calculation result is output by the reading result module in the result output state.
[0108] The above method can realize various functions in the above device embodiment, which will not be expanded here.
[0109] In the data processing method provided in the embodiment of the present application, the target task is determined according to the task selection signal, and the individual or combined execution of tasks such as number theoretic transformation, multiplication, and inverse number theoretic transformation in the polynomial calculation task is supported. This facilitates the adaptation to the calculation requirements of various polynomial multiplication tasks through the task selection signal and the dynamically combined hardware processing module, thereby improving the scalability of the hardware structure and improving resource utilization.
[0110] See also Figure 5 , Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of the present application. Figure 5 As shown, an embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the functions of the above-mentioned device embodiment are realized.
[0111] See also Figure 6 , Figure 6 A schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present application. Figure 6 As shown, this embodiment provides a computer-readable storage medium 600 on which instructions 611 are stored. When the instructions 611 are executed on a computer, the computer executes the functions of the above-mentioned device embodiment.
[0112] An embodiment of the present application further provides a chip, in which a computer program is loaded, and the computer program is used to implement the functions of the above-mentioned device embodiment.
[0113] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. It should be understood by those skilled in the art that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. Although the preferred embodiments of the present application have been described, once the technical personnel in the field know the basic creative concept, they can make additional changes and modifications to these embodiments. Therefore, the attached claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. In this way, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A data processing device, characterized in that: The device comprises: a controller module, a task execution unit, and a reading result module; The controller module is used to receive a task selection signal from a task request end, and determine a target task to be performed by the task execution unit based on the task selection signal; the target task is one or more polynomial calculation tasks; the polynomial calculation tasks at least include: number theory change tasks, multiplication tasks, and inverse number theory change tasks; The task execution unit is used to select a corresponding hardware processing module based on the target task to construct a corresponding target task path; the hardware processing modules in the task execution unit include: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module; the number, type, and connection method of the hardware processing modules in the target task path are related to the task type to which the target task belongs; The task execution unit is further used to receive the data to be processed, and input the data to be processed into the target task path to start the target task and obtain the task execution result; The reading result module is used to output the task execution result to the task request end.
2. The data processing device according to claim 1, characterized in that: The initialization module is connected to the storage module; the initialization module is used to receive the data to be processed and temporarily store the data to be processed in the storage module; Among them, the storage module is respectively connected to the reversal module, the butterfly calculation module, and the multiplication module, and the butterfly calculation module and the multiplication module are respectively connected to the modular multiplication module; the data processing results of the reversal module, the butterfly calculation module, the multiplication module, and the modular multiplication module are output to the reading result module through the storage module.
3. The data processing device according to claim 2, characterized in that: The storage module includes a first storage module and a second storage module; the first storage module and the second storage module respectively occupy a unit storage space; The task execution unit, when selecting a corresponding hardware processing module based on the target task to construct a corresponding target task path, is specifically used to: The task execution unit selects a corresponding target processing unit from the first storage module, the second storage module, the reversal module, the butterfly calculation module, the multiplication module, and the modular multiplication module based on the target task; combines the selected target processing units into corresponding data processing links, and uses the combined data processing links as the target task path.
4. The data processing device according to claim 2, characterized in that: The controller module is also used to control the data processing state of the to-be-processed data in the target task path; The task execution unit, when inputting the to-be-processed data into the target task path to execute the target task, is specifically used to: Under the control and scheduling of the controller module, the task execution unit loads the data to be processed into the target task path, and switches the hardware processing unit where the data to be processed is located in the target task path to execute the corresponding data processing flow in the target task.
5. The data processing device according to claim 4, characterized in that: The controller module includes a first state machine; the first state machine is used to receive a task start signal before determining the target task to be performed by the task execution unit based on the task selection signal; based on the task start signal, switch the initialization module to the initialization memory state to trigger the initialization module to perform the initialization operation on the data to be processed; after receiving the initialization end signal, continue to stay in the initialization memory state and trigger the step of determining the target task to be performed by the task execution unit based on the task selection signal; The initialization module is further used to perform initialization analysis on the data to be processed, and after the initialization is completed, feed back the initialization completion signal to the first state machine.
6. The data processing device according to claim 5, characterized in that: The controller module, when determining the target task to be performed by the task execution unit based on the task selection signal, is specifically used to: if the task selection signal is a first task signal, the controller module determines that the target task is a number theory transformation task NTT; The controller module includes a second state machine; the second state machine is used to switch the task execution unit from the initialization memory state to the rotation factor loading state, the butterfly calculation state, the reversal state, and the result output state in sequence; in the number theory conversion task, the data to be processed includes polynomial coefficients and NTT rotation factors; In NTT, the target task path consists of the initialization module, the butterfly calculation module, and the reversal module; In the rotation factor loading state, the initialization module is used to obtain the NTT rotation factor from the second storage module and perform initialization analysis on the NTT rotation factor; if the analysis is completed, the NTT rotation factor analysis results are returned to the second storage module respectively, and a rotation factor analysis completion signal is sent to the second state machine, so that the second state machine switches to the butterfly calculation state; In the butterfly calculation state, the butterfly calculation module is used to perform butterfly calculation on the polynomial coefficients in the first storage module and the NTT rotation factor analysis results in the second storage module; store the obtained butterfly calculation intermediate results in the first storage module, and send a butterfly calculation completion signal to the second state machine, so that the second state machine switches to the reversal state; In the reversal state, the reversal module is used to perform bit reversal permutation on the intermediate result of the butterfly calculation in the first storage module to obtain the NTT calculation result; A reversal completion signal is sent to the second state machine, so that the second state machine switches to the result output state, and outputs the NTT calculation result through the reading result module in the result output state.
7. The data processing device according to claim 6, characterized in that: If there are multiple NTT rotation factors, the initialization module is also used to load the multiple NTT rotation factors into the second storage module at different times, so as to respectively calculate and obtain multiple NTT calculation results corresponding to the multiple NTT rotation factors.
8. The data processing device according to claim 6, characterized in that: The controller module, when determining the target task to be executed by the task execution unit based on the task selection signal, is specifically configured to: if the task selection signal is a second task signal, the controller module determines that the target task is a multiplication task; The controller module includes a third state machine; the third state machine is used to switch the task execution unit from the initialization memory state to the multiplication calculation state and the result output state in sequence; In a multiplication task, the data to be processed includes polynomial coefficients; In the multiplication task, the target task path is composed of the initialization module and the multiplication module; In the multiplication calculation state, the multiplication module is used to perform multiplication calculation on the first polynomial coefficients stored in the first storage module and the second polynomial coefficients stored in the second storage module; write the product result into the first storage module, and send a multiplication calculation completion signal to the third state machine, so that the second state machine switches to the result output state, and outputs the product result through the reading result module in the result output state.
9. The data processing device according to claim 6, characterized in that: The controller module, when determining the target task to be performed by the task execution unit based on the task selection signal, is specifically used to: if the task selection signal is a third task signal, the controller module determines that the target task is an inverse number theory transformation task INTT; The controller module includes a fourth state machine; the fourth state machine is used to switch the task execution unit from the initialization memory state to the reversal state, the rotation factor loading state, the butterfly calculation state, and the result output state in sequence; In the number theory conversion task, the data to be processed includes polynomial coefficients and INTT rotation factors; In the number theory conversion task, the target task path is composed of the initialization module, the butterfly calculation module, and the reversal module; In the reversal state, the reversal module is used to perform bit reversal permutation on the polynomial coefficients in the first storage module, and write the obtained polynomial reversal result back to the first storage module; Sending a reversal completion signal to the fourth state machine to switch the fourth state machine to the rotation factor loading state; In the rotation factor loading state, the butterfly calculation module is further used to perform butterfly calculation on the polynomial reversal result in the first storage module and the INTT rotation factor result in the second storage module to obtain the INTT calculation result; A butterfly calculation completion signal is sent to the fourth state machine, so that the fourth state machine switches to the result output state, and outputs the INTT calculation result through the reading result module in the result output state.
10. A data processing method, applied to the data processing device according to any one of claims 1 to 9, the data processing method comprising: The data processing device receives a task selection signal from a task request end, and determines a target task to be executed by the task execution unit based on the task selection signal; The target task is one or more polynomial calculation tasks; the polynomial calculation tasks include at least: number theory change tasks, multiplication tasks, and reverse number theory change tasks; The data processing device selects a corresponding hardware processing module based on the target task to construct a corresponding target task path; the hardware processing modules in the task execution unit include: an initialization module, a storage module, a reversal module, a butterfly calculation module, a multiplication module, and a modular multiplication module; the number, type, and connection method of the hardware processing modules in the target task path are related to the task type to which the target task belongs; The data processing device receives the data to be processed, and inputs the data to be processed into the target task path to start the target task and obtain the task execution result; The data processing device outputs the task execution result to the task request end.
11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to read and execute the computer program to implement the data processing device according to any one of claims 1 to 9.
12. A chip, characterized in that: The chip is loaded with a computer program, and the computer program is used to implement the data processing device according to any one of claims 1 to 9.
Citation Information
Patent Citations
Point-adding and point-multiplying circuit based on binary extended domains and control method of point-adding and point-multiplying circuit
CN111198672A
Hardware resource scheduling method and device, computer equipment and storage medium
CN117112201A
Data processing method, system and device, computer equipment and storage medium
CN118426835A
High-performance polynomial multiplication hardware acceleration architecture for lattice cryptographic chip
CN118963703A
Acceleration unit and related apparatus and method
US20220255721A1