Method and apparatus for processing data, device, medium, and program product

By reordering the initial rotation factor and optimizing the offset, the load imbalance of ECNTT calculations in multi-core processors is solved, and a more efficient data processing process is achieved and computing efficiency is improved.

WO2025137881A1PCT designated stage expired Publication Date: 2025-07-03SUNLUNE (SINGAPORE) PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/142087
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In multi-core processors, the load imbalance of butterfly computing unit load of elliptic curve number theory transform (ECNTT) leads to a long calculation time, and the traditional processing method fails to effectively balance the load of each core, affecting the calculation efficiency.

Method used

By reordering the initial rotation factor, generating new rotation factor and recording offsets, optimizing the data processing flow, using parallel computing and load balancing strategies, reducing the amount of Montgomery multiplication operations and improving calculation efficiency.

Benefits of technology

It effectively solves the load imbalance problem of ECNTT calculation in multi-core processors, shortens the calculation time, and improves the calculation efficiency of ECNTT.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023142087_03072025_PF_FP_ABST
    Figure CN2023142087_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for processing data, a device, a medium, and a program product. The method comprises: acquiring an initial rotation factor of each input point among the number of input points corresponding to a butterfly computation unit; reordering the acquired initial rotation factors to obtain a reordered new rotation factor corresponding to each input point, and recording the offset in input data of an input value and an output value corresponding to each new rotation factor; and performing computation on data to be processed in rounds, performing, on the basis of the new rotation factors and the offset, butterfly computation processing on the input data of a corresponding number of input points in a current round to obtain an output result of each round, and determining an output result of the final round as a computation result corresponding to the data to be processed. An output result of a current round among non-final rounds is used as the input data of the corresponding number of input points in the next round, and the input data of the corresponding number of input points in a starting round is acquired on the basis of the data to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device, equipment, medium and program product Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, medium, and program product. Background Art

[0002] The Elliptic Curve Number-Theoretic Transform (ECNTT) is a special form of computation over finite fields whose computation can be accelerated using the parallel computing techniques of the Fast Fourier Transform (FFT). In multi-core processors, the butterfly computation units of the FFT are highly symmetric, allowing each core in the processor to process a single butterfly computation unit. Although FFT parallel computing acceleration techniques can be applied to ECNTT, load imbalance still exists in multi-core processors.

[0003] Summary of the Invention

[0004] Embodiments of the present application provide a data processing method, apparatus, device, medium, and program product.

[0005] In one aspect, an embodiment of the present application provides a data processing method, the method comprising:

[0006] Obtain the initial rotation factor of each input point in the number of input points corresponding to the butterfly calculation unit;

[0007] Reordering the initial rotation factors of each input point in the input points to obtain a reordered new rotation factor corresponding to each input point, and recording the offset of the input value and output value corresponding to each new rotation factor in the input data;

[0008] The data to be processed is calculated according to the rounds, and the input data of the input points corresponding to the current round are subjected to butterfly calculation processing based on the new rotation factor and the offset to obtain the output result of each round, and the output result of the last round is determined as the calculation result corresponding to the data to be processed, wherein the output result of the current round among the non-last rounds is used as the input data of the input points corresponding to the next round, and the input data of the input points corresponding to the starting round is obtained based on the data to be processed.

[0009] On the other hand, an embodiment of the present application provides a data processing device, comprising:

[0010] An acquisition unit, used to obtain an initial rotation factor of each input point in the number of input points corresponding to the butterfly calculation unit;

[0011] a sorting unit, configured to reorder the initial rotation factors of each input point in the input points to obtain a reordered new rotation factor corresponding to each input point, and record an offset of an input value and an output value corresponding to each new rotation factor in the input data;

[0012] A calculation unit is used to calculate the data to be processed according to the rounds, perform butterfly calculation processing on the input data of the corresponding input points of the current round based on the new rotation factor and the offset, obtain the output result of each round, and determine the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round among the non-last rounds is used as the input data of the corresponding input points of the next round, and the input data of the corresponding input points of the starting round is obtained based on the data to be processed.

[0013] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to execute the data processing method described in any of the above embodiments by calling the computer program stored in the memory.

[0014] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for loading by a processor to execute the data processing method described in any of the above embodiments.

[0015] On the other hand, an embodiment of the present application provides a computer program product, including computer instructions, which, when executed by a processor, implement the data processing method described in any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 is a butterfly calculation diagram.

[0018] Figure 2 is a computational diagram for scalar multiplication.

[0019] FIG3 is a characteristic diagram of the arrangement of calculation quantities generated by calculating the initial rotation factors of the standard FFT.

[0020] FIG4 is a flow chart of a data processing method according to an embodiment of the present application.

[0021] FIG5 is a schematic structural diagram of the GS structure provided in an embodiment of the present application.

[0022] FIG6 is a schematic structural diagram of a CT structure provided in an embodiment of the present application.

[0023] FIG7 is a schematic diagram of a two-dimensional array provided in an embodiment of the present application.

[0024] FIG8 is a characteristic diagram of the calculation quantity arrangement generated by calculating according to the reordered new rotation factors.

[0025] FIG9 is a diagram of the system architecture provided in an embodiment of the present application.

[0026] FIG10 is a schematic diagram of a data storage method provided in an embodiment of the present application.

[0027] FIG11 is a schematic diagram of the calculation process provided in an embodiment of the present application.

[0028] FIG12 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application.

[0029] FIG13 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0031] The embodiments of the present application provide a data processing method, apparatus, device, medium and program product. Specifically, the data processing method of the embodiments of the present application can be executed by a computer device, wherein the computer device can be a terminal or a server. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, a smart speaker, a wearable smart device, a smart car terminal and other devices. The terminal can also include a client, which can be a financial client, a browser client or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution network services, and basic cloud computing services such as big data and artificial intelligence platforms, but is not limited to this.

[0032] First, some nouns or terms that appear in the description of the embodiments of this application are explained as follows:

[0033] Elliptic Curve Cryptography (ECC) is a public key cryptography algorithm based on elliptic curve mathematics. Its security relies on the difficulty of the elliptic curve discrete logarithm problem.

[0034] Number Theoretic Transform (NTT) is a number theoretic transform based on Discrete Fourier Transform (DFT), which converts an integer sequence of length N into another integer sequence of length N.

[0035] Elliptic Curve number-theoretic transform (ECNTT) is a special form over finite fields, and its calculation process can be accelerated by parallel computing using the fast Fourier transform (FFT).

[0036] Fast Fourier transform (FFT), also known as FFT for short, is a general term for efficient and fast methods of calculating the discrete Fourier transform (DFT) using computers. The DFT is the Fourier transform's discrete form in both the time and frequency domains.

[0037] The twiddle factor (W) refers to the complex constant multiplied in the butterfly operation of the FFT algorithm.

[0038] DDR (double date rate), double data rate, DDR SDRAM is double data rate synchronous dynamic random access memory, people are accustomed to calling it DDR, where SDRAM is the abbreviation of Synchronous Dynamic Random Access Memory, that is, synchronous dynamic random access memory.

[0039] The radix-2 algorithm decomposes a DFT sequence of length N into a linear combination of two DFT subsequences of length N / 2. If N / 2 is still divisible by 2, the parity decomposition of the two subsequences of length N / 2 can be continued. If the length of the sequence N is an exponential multiple of 2, that is, N=2 n , n∈{0,1,2,…}, the radix-2 algorithm can be applied to gradually decompose the sequence of length N until the length of the sequence becomes 1.

[0040] ECNTT is a specialized form of FFT over finite fields, and its computational process can leverage the parallel computing acceleration techniques of FFT. In multi-core processors, as shown in the butterfly computation diagram in Figure 1, the FFT butterfly units are highly symmetrical, allowing each core in the multi-core processor to process a single butterfly unit. However, the multiplication calculations performed by ECNTT butterfly units involve finite field scalar multiplication, so each butterfly unit must perform the double() and add() operations used in ECC.

[0041] The calculation diagram of scalar multiplication, as shown in Figure 2, includes the point addition (PADD) operation add() and the point doubling (PDBL) operation double(). When the bit is equal to 1, both double() and add() are calculated; when the bit is equal to 0, only double() is calculated. As shown in Figure 2, seven double() calculations and three add() calculations are performed.

[0042] In the ECCNTT operation, the twiddle factor W corresponding to each butterfly unit has a bit length of 253 bits. Therefore, the number of double() calculations is fixed, while the number of add() calculations varies. Specifically, the number of add() calculations depends on the number of occurrences of bit 1. In this example, the number of Montgomery multiplications for ECC add() is (number of occurrences of bit 1) * 21, while the number of Montgomery multiplications for ECC double() is 253 * 8.

[0043] Figure 3 shows a characteristic diagram of the number of calculations generated using the standard FFT's rotation factors. The horizontal axis represents the input point number, and the vertical axis represents the number of add() operations. The points in Figure 3 represent points where bits are equal to 1. By statistically analyzing the distribution of bit 1s, we found that, with the exception of the first and N / 4th points, where the number of additions is relatively small, the number of additions for all other bit 1s is distributed between 100 and 160. This leads to load imbalance between the various computing cores. When multi-core processors have limited computing resources, the traditional approach is to use the maximum number of ECC add() operations as a benchmark, which results in longer ECNTT calculation times.

[0044] As shown in Table 1 below, the maximum number of addition operations (max ECC add() number) and the minimum number of addition operations (min ECC add() number) corresponding to each of the four batches calculated according to the rotation factor W of the standard FFT are illustrated.

[0045] Table 1

[0046] To address this load imbalance between computing cores, the present invention proposes a new data processing method that more effectively balances the load between computing cores, thereby shortening the ECNTT calculation time. This method effectively reduces the total number of Montgomery multiplications, improves the efficiency of ECNTT calculations, and solves the load imbalance problem of ECNTT calculations in multi-core processors.

[0047] Since ECNTT is a computationally intensive task, with a computational percentage exceeding 90%, the issue of input / output (IO) memory access can be temporarily ignored. Since scalar multiplication consumes a large amount of register resources within the chip, the radix-2 algorithm can be used as the basic computational unit for ECNTT in this embodiment.

[0048] The embodiment of the present application mainly reorders the generated initial rotation factors W according to the number of addition operations ECC add() corresponding to the initial rotation factors W, and reads the offset (index) of the input value and output value corresponding to each reordered new rotation factor in each batch in the input data. Since the generation of the reordered new rotation factors W' and the offset is performed in the preprocessing, it does not occupy the ECNTT calculation time. Since the length of the input data satisfies N=2 n , so the entire iterative process can be controlled by n.

[0049] It should be noted that the order of description of the following embodiments does not limit the priority order of the embodiments.

[0050] Please refer to Figures 4 to 11. Figure 4 is a flow chart of the data processing method provided in an embodiment of the present application. Figures 5 to 11 are schematic diagrams of application scenarios provided in an embodiment of the present application. For the specific schematic diagrams represented by each application scenario, please refer to the description of the above figures. The method may include the following steps 11 to 130:

[0051] Step 110: Obtain an initial rotation factor for each input point in the input points corresponding to the butterfly computing unit.

[0052] In some embodiments, obtaining an initial rotation factor for each input point in the number of input points corresponding to the butterfly computing unit includes:

[0053] Get the number of input points corresponding to the butterfly calculation unit;

[0054] An initial rotation factor of each input point in the input points is calculated according to the input points and a calculation formula of the butterfly calculation unit.

[0055] For example, the initial twiddle factor W corresponding to each input point can be pre-generated based on the calculation formula of the specific structure of the butterfly computation unit and the known number of input points. The number of input points is fixed for each butterfly computation unit and can be pre-set when the butterfly computation unit structure is established, or it can be set based on actual computational requirements. The number of input points represents the number of input data points used for the calculation and can also be used to determine the number of twiddle factors.

[0056] For example, the butterfly computation unit can use a GS (Gentleman-Sande) architecture. Based on the GS structure's calculation formula and the known number of input points, the initial twiddle factors W corresponding to each input point can be pre-generated. This step can be completed during the preprocessing phase to prepare for subsequent calculations. A specific implementation method may include first determining the number of initial twiddle factors W to be generated (n / 2+1) based on the length n of the input points. Then, each initial twiddle factor W is calculated using the GS structure and stored in a register.

[0057] Figure 5 shows a schematic diagram of the GS structure. This diagram illustrates the process of calculating two output values ​​(u+v, (uv)*w) using the GS structure's butterfly computation unit, given two input values ​​(u and v). u and v represent the input values, W represents the initial twiddle factor, mod represents the modulus, and mod q represents the modulus operation. A modulus operation by q (i.e., mod q) is performed on the two output values. When the number of input points is known, the initial twiddle factor W corresponding to each input point can be calculated based on the GS structure's calculation formula and the known number of input points.

[0058] For example, the structure of the butterfly computing unit can also use a CT (Cooley-Tukey) structure, and the initial rotation factor W corresponding to each input point can be pre-generated according to the calculation formula of the CT structure and the known number of input points, and stored in a register.

[0059] Figure 6 shows a schematic diagram of the CT structure, illustrating the process of calculating two output values ​​(u+v*w, uv*w) using the butterfly computation unit of the CT structure given two input values ​​(u and v). A modulo operation by q (i.e., mod q) is performed on the two output values. When the number of input points is known, the initial twiddle factor W corresponding to each input point can be calculated based on the calculation formula of the CT structure and the known number of input points.

[0060] Among them, when using the GS structure, the butterfly computing unit includes a modular addition module (+), a modular subtraction module (-), and a Barrett modular multiplication module (×); the modular addition module and the modular subtraction module respectively perform addition and subtraction operations on the two data to be processed, and the addition result is directly output. The subtraction result is multiplied with the corresponding rotation factor through the Barrett modular multiplication unit and then output. When using the CT structure, the butterfly computing unit includes a modular addition module (+), a modular subtraction module (-), and a Barrett modular multiplication module (×); the Barrett modular multiplication module multiplies one of the data to be processed with the corresponding rotation factor. The modular addition module and the modular subtraction module respectively perform addition and subtraction operations on the multiplication result output by the Barrett modular multiplication module and the other data to be processed, and then output the operation result.

[0061] Step 120: reorder the initial rotation factors of each input point in the input points to obtain a reordered new rotation factor corresponding to each input point, and record the offset of the input value and output value corresponding to each new rotation factor in the input data.

[0062] In some embodiments, reordering the initial rotation factors of each input point in the input points to obtain the reordered new rotation factors corresponding to each input point includes:

[0063] The initial rotation factors are reordered according to the number of addition operations corresponding to each of the initial rotation factors to obtain a reordered new rotation factor corresponding to each input point.

[0064] In some embodiments, reordering the initial rotation factors according to the number of addition operations corresponding to each of the initial rotation factors to obtain a reordered new rotation factor corresponding to each input point includes:

[0065] The initial rotation factors are reordered in ascending order according to the number of addition operations corresponding to each of the initial rotation factors to obtain a reordered new rotation factor corresponding to each input point.

[0066] In some embodiments, recording the offset of the input value and the output value corresponding to each of the new rotation factors in the input data includes:

[0067] A two-dimensional array is created, where the two-dimensional array is used to record the offset of the input value and the output value corresponding to each of the new rotation factors in the input data.

[0068] For example, the initial rotation factors W can be reordered according to the number of addition operations (i.e., the number of ECC add() operations) corresponding to each initial rotation factor W, and a two-dimensional array can be created to record the offset (index) of the input value and output value corresponding to each initial rotation factor W in the input data. This step can also be completed in the preprocessing stage to prepare for the subsequent data reading stage. Specifically, all initial rotation factors W can be traversed and the number of ECC add() operations for each initial rotation factor W can be counted. Then, each initial rotation factor W can be reordered in ascending order according to the number of ECC add() operations to obtain the reordered new rotation factor W' corresponding to each input point; at the same time, a two-dimensional array can be created to record the offset (index) of the input value and output value corresponding to each initial rotation factor W in the input data. The offset (index) represents the offset of the reordered new rotation factor W' relative to the input value and output value corresponding to the initial rotation factor W in the input data after the initial rotation factor W is reordered. For example, an initial rotation factor W3 originally corresponds to the third input data. After the initial rotation factors are reordered, W3 is ranked fourth. However, when reading the data required for the reordered W3, it will still be read according to the original third input data based on the corresponding offset.

[0069] For example, as shown in Figure 7, for W0, since it has the fewest ECCadd() calls, the offset (index) of its corresponding point in batch 0 is 0, and this offset (index) is placed at the first position in the array. In Figure 7, n represents the batch number, and N represents the number of input data points.

[0070] As shown in Figure 8, the calculation quantity arrangement feature diagram is generated by calculating according to the reordered new rotation factors, where the horizontal axis represents the input point sequence number, the vertical axis represents the number of addition operations add(), and the points in Figure 8 represent points where the bit is equal to 1.

[0071] As shown in Table 2 below, the maximum number of addition operations (max ECC add() number) and the minimum number of addition operations (min ECC add() number) corresponding to each of the four batches calculated according to the reordered rotation factors W' are illustrated.

[0072] Table 2

[0073] As shown in Table 2, due to the reordering of the initial twiddle factors, new twiddle factors corresponding to each input point are obtained. When applying these new twiddle factors to data processing, the maximum number of ECC add() calls is 126 when batch size is 0, which is 32 fewer than the sequential read method in Table 1. The maximum number of ECC add() calls is 126 when batch size is 1, which is 28 fewer than the sequential read method in Table 1. The maximum number of ECC add() calls is 131 when batch size is 2, which is 22 fewer than the sequential read method in Table 1. Since this example has four batches, this method reduces 82 × 21 Montgomery multiplications compared to the traditional sequential read method.

[0074] In some embodiments, the initial rotation factors of each input point in the input point count are reordered based on a hash algorithm to obtain a reordered new rotation factor corresponding to each input point, and the offset of the input value and output value corresponding to each new rotation factor in the input data is recorded.

[0075] A hash algorithm maps binary values ​​of arbitrary length to fixed-length binary values, the result typically represented by a hash code. The initial twiddle factors for each input point can be reordered using a hash algorithm. Specifically, for each input point, the initial twiddle factors are used as input and then hashed to obtain a hash code for the corresponding new twiddle factor. Based on the hash code sorting results, the initial twiddle factors for all input points are reordered to obtain the new twiddle factors corresponding to each input point. During this process, the offsets of the input and output values ​​corresponding to each new twiddle factor in the input data can also be recorded. This process can be achieved by traversing the reordered twiddle factor list and recording the correspondence between each new twiddle factor and the original input data. These offsets can be used in subsequent data processing to further improve data processing efficiency and accuracy. This sorting method leverages the deterministic and efficient nature of the hash algorithm, making the data processing process more concise and accurate.

[0076] In some embodiments, the initial rotation factors of each input point in the input points are reordered based on the number of occurrences of the initial rotation factors to obtain reordered new rotation factors corresponding to each input point, and the offsets of the input values ​​and output values ​​corresponding to each new rotation factor in the input data are recorded.

[0077] For example, for each input point, the number of occurrences of each initial rotation factor can be calculated, and then the initial rotation factors can be reordered from highest to lowest by occurrence to obtain the reordered new rotation factors corresponding to each input point. This sorting method allows initial rotation factors with high occurrences to be processed earlier in the calculation process. Initial rotation factors with high occurrences generally mean that they will be used more frequently in the calculation process. Therefore, processing these initial rotation factors early can reduce waiting time during the calculation process, thereby improving the computational efficiency of ECNTT.

[0078] In some embodiments, the initial rotation factors of each input point in the input points are reordered based on the core load condition to obtain a reordered new rotation factor corresponding to each input point, and the offset of the input value and output value corresponding to each new rotation factor in the input data is recorded.

[0079] For example, the initial twiddle factors can be reordered based on the load of each core. If a core is lightly loaded, more initial twiddle factors can be allocated to it, fully utilizing the core's computing power. This balances the load across multiple cores and improves ECNTT computational efficiency. By balancing the load across cores, we can avoid overloading some cores while leaving others idle, thereby improving ECNTT computational efficiency. This can be achieved through the use of load balancing algorithms or dynamic scheduling strategies.

[0080] In some embodiments, the initial rotation factors of each input point in the input points are reordered based on the modulus of the initial rotation factors to obtain reordered new rotation factors corresponding to each input point, and the offsets of the input values ​​and output values ​​corresponding to each new rotation factor in the input data are recorded.

[0081] For example, the initial twiddle factors can be reordered according to their modulus. This sorting method allows data points with fewer modulo operations to be processed earlier during the calculation process, potentially improving the computational efficiency of ECNTT. For example, this sorting method primarily sorts based on the modulus of the twiddle factors W. Modulo operations are typically relatively time-consuming in computers. Therefore, if the modulus of W is known in advance, data points with fewer modulo operations can be prioritized, thereby reducing overall computation time. It is important to note that this sorting method does not directly affect the order in which data points are processed, but rather optimizes the use of the initial twiddle factors during the calculation process. For example, the moduli of all initial twiddle factors can be calculated first, and then sorted according to their moduli. During the actual calculation process, initial twiddle factors with smaller moduli can be prioritized, thereby reducing the number of modulo operations.

[0082] In some embodiments, the method further comprises:

[0083] The new rotation factor and the offset are stored in a register.

[0084] As shown in Figure 9, the system architecture provided by the embodiment of the present application can use DDR memory as an external storage resource. The CPU chip has a multi-core structure, with multiple cores placed on the same CPU chip. Each core has register resources. The CPU chip also includes physical hardware such as the first-level cache (L1 cache) and the second-level cache (L2 cache). Among them, a single core accesses the L2 cache through the L1 cache.

[0085] The first-level cache (L1 cache), integrated into the CPU, temporarily stores data while the CPU is processing it. Because cached instructions and data operate at the same frequency as the CPU, a larger L1 cache capacity allows for more information to be stored, reducing the number of data exchanges between the CPU and memory and improving CPU computing efficiency. However, because cache memory is composed of static RAM and has a complex structure, the L1 cache capacity cannot be increased within the limited CPU chip area.

[0086] Due to the limited capacity of the L1 cache, a high-speed memory device is placed outside the CPU to further increase CPU speed. This L2 cache operates at a flexible frequency, allowing it to be synchronized with or different from the CPU. When reading data, the CPU first searches the L1 cache, then the L2, then the main memory, and finally the external memory. Therefore, the L2 cache's impact on the system cannot be ignored.

[0087] In an embodiment of the present application, when reading data, each thread reads the corresponding target data from the DDR memory into a register based on the read offset (index). Specifically, when reading data, the target data is first searched in the register (register). If the target data required by the thread is not found in the register (register), the target data is then searched from the L1 cache; if the target data required by the thread is also not found in the L1 cache, the target data is then searched from the L2 cache; if the target data required by the thread is also not found in the L2 cache, the target data is then searched from the DDR until the target data is found and the read target data is stored in the register (register). This is then used by the butterfly computing unit in the core to calculate the target data and obtain the calculation result. The calculation result is then stored in the DDR memory through the L1 cache and the L2 cache in sequence.

[0088] For example, the new rotation factor and the offset may be stored in any one of a register, a first-level cache (L1 Cache), a second-level cache (L2 Cache), and a DDR memory.

[0089] Among them, because in the subsequent steps, when the core calculates the data, the calculation-related data, parameters and instructions are immediately called from the registers for calculation. In order to be able to quickly read the relevant data based on the new rotation factor and offset in the subsequent steps, the new rotation factor and offset can be stored in the register to improve data acquisition efficiency and calculation efficiency.

[0090] Step 130: Calculate the data to be processed according to the rounds, perform butterfly calculation on the input data of the input points corresponding to the current round based on the new rotation factor and the offset, obtain the output result of each round, and determine the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round among the non-last rounds is used as the input data of the input points corresponding to the next round, and the input data of the input points corresponding to the starting round is obtained based on the data to be processed.

[0091] In some embodiments, the calculating the data to be processed according to rounds, performing butterfly calculation processing on the input data corresponding to the number of input points in the current round based on the new rotation factor and the offset, obtaining the output result of each round, and determining the output result of the last round as the calculation result corresponding to the data to be processed, includes:

[0092] For each of the multiple threads of the current round in the non-last round, read input data corresponding to each input point from the memory into a register according to the offset of each input point in the input points, and perform butterfly calculation processing on the read input data based on the new rotation factor of each input point, when each thread in the current round completes the calculation, obtain an output result of the current round in the non-last round, and store the output result of the current round in the non-last round in the memory as input data for the corresponding input point in the next round; and

[0093] For each of the multiple threads in the last round, input data corresponding to each input point is read from the memory into the register according to the sequence number from small to large, and butterfly calculation processing is performed on the read input data based on the new rotation factor of each input point. When each thread in the last round completes the calculation, the output result of the last round is obtained, and the output result of the last round is determined as the calculation result corresponding to the data to be processed.

[0094] In some embodiments, the method further comprises:

[0095] For the multiple threads in each round, during the butterfly calculation process, the multiple threads perform parallel calculations on the read input data.

[0096] This parallel computing approach can significantly improve computational efficiency. Butterfly computing is a common computing method that involves multiple steps and operations, including data reading, calculation, and result storage. By using multiple threads for parallel computing, multiple data points can be processed simultaneously, reducing computation time and improving efficiency.

[0097] For example, a certain number of threads can be created for each round, with each thread responsible for reading a portion of the input data and performing calculations. These threads can execute in parallel without conflict or competition. This approach can speed up the computation and reduce waiting time during the computation.

[0098] When using multiple threads for parallel computing, synchronization and data sharing between threads must be considered. To avoid race conditions and data inconsistencies, appropriate synchronization mechanisms are required to ensure data correctness and consistency. Furthermore, threads must be properly scheduled and managed to fully utilize overall resources and improve computing efficiency.

[0099] In some embodiments, the method further comprises:

[0100] Storing the data to be processed in the memory in rows, so that input data corresponding to the number of input points in the starting round is obtained based on the data to be processed;

[0101] The output results of the current round are stored in the memory in a row manner to serve as input data corresponding to the input points of the next round.

[0102] For example, as shown in the data storage diagram in Figure 10, the data to be stored is stored in rows within the DDR memory. Each point stores twelve 32-bit words horizontally, where P0 through PN represent data points numbered 0 through N, respectively. For example, the data to be stored can be the previously described data to be processed, or the output results of the current round to be stored.

[0103] For example, data to be stored is stored row-wise in global memory within DDR memory. Global memory is a key concept in computing and plays a crucial role in program execution. Global memory refers to the memory space that exists at all times during program execution and is accessible and usable by all threads or processes. Global memory is the memory space shared by all threads or processes. In multi-threaded or multi-process programs, each thread or process has its own private memory space for storing local variables, etc. Global memory, on the other hand, is a memory space shared by all threads or processes, storing global data such as global variables and static variables. Global memory plays a crucial role in programs. First, it enables data sharing. In multi-threaded or multi-process programs, different threads or processes need to interact and share data, which is where global memory comes in. By storing shared data in global memory, threads or processes can easily access and modify this data, enabling data sharing and communication. Global memory can also be used to store global state. In some cases, a program state needs to be shared between multiple threads or processes, and this state can be stored in global memory. For example, if the value of a counter needs to be accumulated in multiple threads, the counter can be stored in the global memory, and each thread can modify the value of the counter by accessing the global memory.

[0104] For example, the calculation process diagram shown in Figure 11 represents the specific situation of processing data in each thread TID based on the reordered new rotation factors and corresponding offsets, where the cross symbol represents the butterfly computing unit in each core.

[0105] For the computation process in non-final rounds, i.e., batches 0 to n-1 shown in Figure 11, each thread reads the corresponding input data from the DDR memory into a register based on the new twiddle factor and corresponding offset (index) for each input point, and then calculates the result according to the butterfly computation unit structure. For example, in batch 1, when thread tid0 reads the new twiddle factor W0 (referring to Figures 7 and 10), since the offset (index) corresponding to the new twiddle factor W0 is 0, thread tid0 takes the first point P0 (x, y, z, zz) in the input data and reads it into the register (register) in the order in which P0 is stored in memory (i.e., from x1 to x12, and from x12 to zz12), and then performs the relevant calculations according to the butterfly computation unit. Because the data processed by each thread is inconsistent between different batches, when each core in each batch completes the computation, the computation results of each core in each round are rewritten to the DDR memory, and this computation result serves as the input for the next batch. Among them, each thread performs related calculation processes based on the butterfly computing unit of each core.

[0106] For the last round of calculation, that is, batch n shown in Figure 11, since the rotation factors of each butterfly computing unit in batch n are consistent, each thread in batch n can read the data from the DDR memory in the order of sequence numbers from 0 to N.

[0107] By computing the data to be processed in batches, the computation can be made more efficient. At the same time, by optimizing the data reading method, the number of unnecessary computations can be reduced.

[0108] Through the above steps, the embodiment of the present application reorders the initial rotation factors according to the number of ECC add() additions corresponding to each initial rotation factor, obtains reordered new rotation factors, and creates a two-dimensional array to record the offset (index) of the input value and output value corresponding to each new rotation factor in the input data. Then, the processed data is calculated in batches, which can effectively reduce the number of Montgomery multiplication operations, effectively improve the computational efficiency of ECNTT, and solve the problem of load imbalance of ECNTT calculation in multi-core processors.

[0109] For example, before the calculation begins, the input data can be preprocessed, such as sorting and deduplication, to reduce redundant operations during the calculation process and improve calculation efficiency.

[0110] For example, ECNTT can perform parallel computing on multi-core processors, which can significantly shorten the computing time by assigning computing tasks to multiple cores for simultaneous processing.

[0111] For example, on a multi-core processor, parallel communication can be used to improve data transmission efficiency and reduce communication overhead.

[0112] For example, to address the load imbalance problem that may occur during the ECNTT calculation process, computing tasks can be dynamically allocated to different cores for processing based on the characteristics of the computing tasks.

[0113] All of the above technical solutions can be combined in any way to form the embodiments of the present application, and will not be described in detail here.

[0114] The embodiment of the present application obtains the initial rotation factor of each input point in the input points corresponding to the butterfly calculation unit; reorders the initial rotation factor of each input point in the input points to obtain the reordered new rotation factor corresponding to each input point, and records the input value and output value offset of each new rotation factor in the input data; calculates the data to be processed according to the round, and based on the new rotation factor and the offset, performs butterfly calculation processing on the input data corresponding to the input point of the current round to obtain the output result of each round, and determines the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round in the non-last round is used as the input data of the input point corresponding to the next round, and the input data of the input point corresponding to the starting round is obtained based on the data to be processed. The embodiment of the present application can more effectively balance the load between the various computing cores and improve computing efficiency by reordering the initial rotation factors and then performing corresponding calculation processing on the data to be processed based on the new rotation factor and the offset.

[0115] To facilitate better implementation of the data processing method of the present embodiment, the present embodiment further provides a client. Please refer to FIG12 , which is a schematic diagram of the structure of the data processing device provided in the present embodiment. The data processing device 200 may include:

[0116] An acquisition unit 210 is configured to acquire an initial rotation factor of each input point in the number of input points corresponding to the butterfly calculation unit;

[0117] a sorting unit 220 configured to reorder the initial rotation factors of each input point in the input points to obtain a reordered new rotation factor corresponding to each input point, and record an offset of an input value and an output value corresponding to each new rotation factor in the input data;

[0118] The calculation unit 230 is used to calculate the data to be processed according to the rounds, perform butterfly calculation processing on the input data of the input points corresponding to the current round based on the new rotation factor and the offset, obtain the output result of each round, and determine the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round among the non-last rounds is used as the input data of the input points corresponding to the next round, and the input data of the input points corresponding to the starting round is obtained based on the data to be processed.

[0119] In some embodiments, the sorting unit 220 may be configured to:

[0120] The initial rotation factors are reordered according to the number of addition operations corresponding to each of the initial rotation factors to obtain a reordered new rotation factor corresponding to each input point.

[0121] In some embodiments, the sorting unit 220 may be configured to:

[0122] The initial rotation factors are reordered in ascending order according to the number of addition operations corresponding to each of the initial rotation factors to obtain a reordered new rotation factor corresponding to each input point.

[0123] In some embodiments, when recording the input value and the offset of the output value corresponding to each new rotation factor in the input data, the acquisition unit 210 may be configured to:

[0124] A two-dimensional array is created, where the two-dimensional array is used to record the offset of the input value and the output value corresponding to each of the new rotation factors in the input data.

[0125] In some embodiments, the computing unit 230 may be configured to:

[0126] For each of the multiple threads of the current round in the non-last round, read input data corresponding to each input point from the memory into a register according to the offset of each input point in the input points, and perform butterfly calculation processing on the read input data based on the new rotation factor of each input point, when each thread in the current round completes the calculation, obtain an output result of the current round in the non-last round, and store the output result of the current round in the non-last round in the memory as input data for the corresponding input point in the next round; and

[0127] For each of the multiple threads in the last round, input data corresponding to each input point is read from the memory into the register according to the sequence number from small to large, and butterfly calculation processing is performed on the read input data based on the new rotation factor of each input point. When each thread in the last round completes the calculation, the output result of the last round is obtained, and the output result of the last round is determined as the calculation result corresponding to the data to be processed.

[0128] In some embodiments, the computing unit 230 may also be configured to: for multiple threads in each round, perform parallel computing on the read input data through the multiple threads during the butterfly computing process.

[0129] In some embodiments, the data processing device 200 further includes:

[0130] A storage unit is used to store the data to be processed in the memory in a row manner, so that the input data corresponding to the input point number of the starting round is obtained based on the data to be processed; and store the output result of the current round in the memory in a row manner, so as to serve as the input data corresponding to the input point number of the next round.

[0131] All of the above technical solutions can be combined in any way to form the embodiments of the present application, and will not be described in detail here.

[0132] It should be understood that the data processing device embodiment and the method embodiment may correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, the data processing device shown in FIG8 can execute the above-mentioned data processing method embodiment, and the aforementioned and other operations and / or functions of each unit in the data processing device respectively implement the corresponding processes of the above-mentioned method embodiment. For the sake of brevity, no further description is given here.

[0133] In one embodiment, the present application further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0134] Figure 13 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device can be a terminal or a server. As shown in Figure 13, the computer device 300 may include: a communication interface 301, a memory 302, a processor 303, and a communication bus 304. The communication interface 301, the memory 302, and the processor 303 communicate with each other via the communication bus 304. The communication interface 301 is used for data communication between the computer device 300 and an external device. The memory 302 can be used to store software programs and modules. The processor 303 runs the software programs and modules stored in the memory 302, such as the software programs for the corresponding operations in the aforementioned method embodiments.

[0135] In one embodiment, the processor 303 can call the software program and module stored in the memory 302 to perform the following operations: obtain the initial rotation factor of each input point in the input points corresponding to the butterfly calculation unit; reorder the initial rotation factor of each input point in the input points to obtain the reordered new rotation factor corresponding to each input point, and record the offset of the input value and output value corresponding to each new rotation factor in the input data; calculate the data to be processed according to rounds, and based on the new rotation factor and the offset, perform butterfly calculation processing on the input data of the input points corresponding to the current round to obtain the output result of each round, and determine the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round in the non-last round is used as the input data of the input points corresponding to the next round, and the input data of the input points corresponding to the starting round is obtained based on the data to be processed.

[0136] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0137] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of computer programs, which can be loaded by a processor to execute the steps of any of the data processing methods provided in the embodiments of the present application. The specific implementation of each of the above operations can be found in the previous embodiments and will not be repeated here.

[0138] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0139] Since the computer program stored in the storage medium can execute the steps in any data processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any data processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0140] The present application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding process of any data processing method in the present application. For the sake of brevity, these instructions are not further described here.

[0141] The present application also provides a computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding process of any data processing method in the present application. For the sake of brevity, the details are not further described here.

[0142] The above is a detailed introduction to a data processing method, client, server, equity incentive system and storage medium provided in the embodiments of the present application. Specific examples are used in the embodiments of the present application to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A data processing method, characterized in that, The method includes: obtaining an initial rotation factor of each input point in the number of input points corresponding to a butterfly calculation unit; reordering the initial rotation factors of each input point in the number of input points to obtain a newly reordered rotation factor corresponding to each input point, and recording the offsets of the input values and output values corresponding to each of the newly reordered rotation factors in the input data; calculating the data to be processed by rounds, and performing butterfly calculation processing on the input data corresponding to the number of input points in the current round based on the newly reordered rotation factor and the offset, to obtain the output result of each round, and determining the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round in non-last rounds is used as the input data corresponding to the number of input points in the next round, and the input data corresponding to the number of input points in the starting round is obtained based on the data to be processed.

2. The data processing method according to claim 1, wherein The reordering the initial rotation factors of each input point in the number of input points to obtain a newly reordered rotation factor corresponding to each input point includes: reordering the initial rotation factors according to the number of addition operations corresponding to each of the initial rotation factors to obtain a newly reordered rotation factor corresponding to each input point.

3. The data processing method according to claim 2, wherein The reordering the initial rotation factors according to the number of addition operations corresponding to each of the initial rotation factors to obtain a newly reordered rotation factor corresponding to each input point includes: reordering the initial rotation factors in ascending order of the number of addition operations corresponding to each of the initial rotation factors to obtain a newly reordered rotation factor corresponding to each input point.

4. The data processing method according to claim 1, characterized in that The recording the offsets of the input values and output values corresponding to each of the newly reordered rotation factors in the input data includes: creating a two-dimensional array for recording the offsets of the input values and output values corresponding to each of the newly reordered rotation factors in the input data.

5. The data processing method according to any one of claims 1-4, characterized in that, The calculating the data to be processed by rounds, and performing butterfly calculation processing on the input data corresponding to the number of input points in the current round based on the newly reordered rotation factor and the offset, to obtain the output result of each round, and determining the output result of the last round as the calculation result corresponding to the data to be processed includes: For each thread in multiple threads in the current round of non-last rounds, reading the input data corresponding to each input point from the memory to the register according to the offset of each input point in the number of input points, and performing butterfly calculation processing on the read input data based on the newly reordered rotation factor of each input point. When each thread in the current round completes the calculation, obtaining the output result of the current round in non-last rounds, and storing the output result of the current round in non-last rounds in the memory as the input data corresponding to the number of input points in the next round; and For each thread in the multiple threads of the last round, the input data corresponding to each input point is read from the memory into the register according to the serial numbers from small to large, and the butterfly calculation process is performed on the read input data based on the new rotation factor of each input point. When each thread in the last round completes the calculation, the output result of the last round is obtained, and the output result of the last round is determined as the calculation result corresponding to the data to be processed.

6. The data processing method according to claim 5, characterized in that, The method further includes: For multiple threads in each round, during the butterfly calculation process, the multiple threads perform parallel calculations on the read input data.

7. The data processing method according to claim 5, characterized in that The method further includes: The data to be processed is stored in the memory in a row manner as the input data corresponding to the number of input points in the starting round Input data; The output result of the current round is stored in the memory in a row manner as the input data corresponding to the number of input points in the next round.

8. A data processing device, characterized in that, The device includes: An acquisition unit, configured to acquire the initial rotation factor of each input point in the number of input points corresponding to the butterfly calculation unit; A sorting unit, configured to re-sort the initial rotation factors of each input point in the number of input points to obtain a new rotation factor after re-sorting corresponding to each input point, and record the offsets of the input values and output values corresponding to each new rotation factor in the input data; A calculation unit, configured to calculate the data to be processed according to rounds, perform butterfly calculation processing on the input data corresponding to the number of input points in the current round based on the new rotation factor and the offset, obtain the output result of each round, and determine the output result of the last round as the calculation result corresponding to the data to be processed, wherein the output result of the current round in non-last rounds is used as the input data corresponding to the number of input points in the next round, and the input data corresponding to the number of input points in the starting round is obtained based on the data to be processed.

9. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The processor is configured to execute the data processing method according to any one of claims 1-7 by calling the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the data processing method according to any one of claims 1-9.

11. A computer program product, comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the data processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Discrete Fourier transform (DFT) / inverse discrete Fourier transform (IDFT) rotation factor control method, device and DFT / IDFT arithmetic device for long term evolution (LTE) system

    CN102111365A

  • High-capacity reconfigurable FFT operation IP core based on FPGA

    CN113157637A

  • Iterative NTT staggered storage system based on BRAM

    CN116679905A

  • Method and apparatus for performing a FFT computation

    US20150006604A1