Data processing method, zero-knowledge proof accelerator, and storage medium

By directly adding the modulus N to the Montgomery multiplication of NTT, the butterfly operation is optimized, and the processor inefficiency caused by judgment operations in NTT is solved, and more efficient calculation is achieved.

WO2025138133A1PCT designated stage expired Publication Date: 2025-07-03SUNLUNE (SINGAPORE) PTE LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/143335
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the number theory transformation (NTT) in zero-knowledge proof calculations frequently judged operations in the butterfly operation of the Montgomery domain leads to inefficiency of the processor, especially the irregular branch selection and prediction, resulting in emptying of instructions in the pipeline.

Method used

By directly adding the modulus N in the subtraction process, omitting the judgment operation, optimizing the steps of Montgomery multiplication, reducing the number of judgments, and improving calculation efficiency.

Benefits of technology

It effectively reduces the number of judgment operations, reduces the computational complexity, and improves the performance of NTT and the computing efficiency of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023143335_03072025_PF_FP_ABST
    Figure CN2023143335_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method, a zero-knowledge proof accelerator, and a storage medium. The method comprises: obtaining first input data, second input data, and a twiddle factor (110); according to the first input data and the second input data, generating a subtraction result (120); subjecting a first addition result of the subtraction result and a modulus to a multiplication operation with the twiddle factor so as to obtain a multiplication result (130); according to the multiplication result, determining a calculation target; and according to the calculation target, determining an output result.
Need to check novelty before this filing date? Find Prior Art

Description

Data Processing Method, Zero-Knowledge Proof Accelerator and Storage Medium Technical Field The present application relates to the field of computer technologies, and particularly to a data processing method, a zero-knowledge proof accelerator and a storage medium. Background Art Zero-Knowledge Proof (ZKP) is a highly information-secure encryption technology, which is widely used in fields such as data transmission and identity authentication. In the future information transmission field, ZKP has broad application prospects. Among them, the Number Theoretic Transform (NTT) is one of the most core parts in the ZKP calculation field, and the butterfly operations involved in it are all completed in the Montgomery Domain. However, this calculation method has multiple reduce operations. How to optimize this part of the content is the key point and difficulty in determining the performance of the NTT. This kind of reduce operation is very unfriendly to the processor because the selection and prediction of branches have no rules, which will frequently cause the emptying of instructions in the pipeline and reduce the efficiency of the processor. Summary of the Invention Embodiments of the present application provide a data processing method, a zero-knowledge proof accelerator and a storage medium. On the one hand, embodiments of the present application provide a data processing method, and the method includes: Obtain first input data, second input data and a rotation factor; Generate a subtraction result according to the first input data and the second input data; Perform a multiplication operation on the first addition result of the subtraction result and a modulus and the rotation factor to obtain a multiplication result; Determine a calculation target according to the multiplication result; Determine an output result according to the calculation target. On the other hand, embodiments of the present application provide a zero-knowledge proof accelerator, and the zero-knowledge proof accelerator includes a number theoretic transform butterfly calculation module. Among them, the number theoretic transform butterfly calculation module includes a first calculation unit, a second calculation unit and a third calculation unit; The first calculation unit is used to perform a multiplication operation on the first addition result of the subtraction result and a modulus and the rotation factor to obtain a multiplication result, where the subtraction result is generated by subtracting the second input data from the first input data; The second calculation unit is used to determine a calculation target according to the multiplication result; The third calculation unit is used to determine an output result according to the calculation target. On the other hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded by a processor to execute the data processing method according to any one of the foregoing embodiments. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings. FIG. 1 is a schematic flowchart of the data processing method provided by an embodiment of the present application. FIG. 2 is a schematic structural diagram of the GS structure provided by an embodiment of the present application. FIG. 3 is a schematic diagram of an application scenario provided by an embodiment of the present application. FIG. 4 is a first schematic structural diagram of the zero-knowledge proof accelerator provided by an embodiment of the present application. FIG. 5 is a second schematic structural diagram of the zero-knowledge proof accelerator provided by an embodiment of the present application. FIG. 6 is a schematic structural diagram of the computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. An embodiment of the present application provides a data processing method, a zero-knowledge proof accelerator, and a storage medium. Specifically, the data processing method in the embodiment of the present application can be executed by a computer device, where the computer device can be a terminal or a server, etc. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart TV, a smart speaker, a wearable smart device, a smart vehicle terminal, etc. The terminal can also include a client, which can be a financial client, a browser client, or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution network services, and big data and artificial intelligence platforms, but is not limited thereto. First, some nouns or terms that appear in the process of describing the embodiments of the present application are explained as follows: Zero-Knowledge Proof (ZKP) is a highly information-secure encryption technology, which means that the prover can make the verifier believe that a certain assertion is correct without providing any useful information to the verifier. Zero-Knowledge Proof is essentially a protocol involving two or more parties, that is, a series of steps required for two or more parties to complete a task. The prover proves to the verifier and makes it believe that it knows or has a certain message, but the proof process cannot leak any information about the message being proved to the verifier. Number Theoretic Transform (NTT) is a number theoretic transform based on the Discrete Fourier Transform (DFT). It transforms an integer sequence of length N into another integer sequence of length N. Fast Fourier Transform (FFT) is a general term for efficient and fast calculation methods that use a computer to calculate the Discrete Fourier Transform (DFT), abbreviated as FFT. Among them, DFT is a form in which the Fourier transform is discrete in both the time domain and the frequency domain. In the ZKP algorithm, the polynomial (poly) calculation that NTT needs to solve can be expressed as the following formula (1): Among them, A(x) represents the polynomial function value; ∑ represents the summation symbol; i represents the number of the input data; N represents the number of input data, that is, the number of points used in the Fourier transform; A i represents the rotation factor; x i represents the input data. Among them, the scale of N is generally between 14 and 26, and this calculation step will be repeatedly called in the overall calculation process. Therefore, the performance of NTT is very critical to the performance of the overall protocol. Compared with the butterfly operation in the traditional FFT, the butterfly operation in NTT is different in that all addition, subtraction, and multiplication calculations are completed in the Montgomery domain, which will result in many more judgment (reduce) operations, such as judging whether to subtract the modulus. This kind of judgment operation is very unfriendly to the processor because the selection and prediction of branches have no rules, which will frequently cause the instructions in the pipeline to be emptied, reducing the efficiency of the processor. In the embodiments of the present application, taking the 256-bit GS (Gentleman-Sande Butterfly) structure as an example, the method for reducing this kind of judgment operation will be described in detail. The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments. Please refer to FIGS. 1 to 4. FIG. 1 is a schematic flowchart of a data processing method provided by an embodiment of the present application. FIG. 2 is a schematic structural diagram of a GS structure provided by an embodiment of the present application. FIG. 3 is a schematic diagram of an application scenario provided by an embodiment of the present application. FIG. 4 is a schematic hardware structure diagram of a zero-knowledge proof acceleration calculation unit provided by an embodiment of the present application. The method may include the following steps 11 to 150: Step 110, obtain first input data, second input data, and a rotation factor. For example, the butterfly calculation unit adopts a GS structure. As shown in the schematic diagram of the GS butterfly transformation process in FIG. 2, in the GS butterfly transformation, modulo addition / subtraction operations are first performed, and then modulo multiplication operations are performed. Among them, x1 is the first input data, x1 + n / 2 is the second input data, is the rotation factor, and y1 and z1 are output data. For example, according to actual needs, the bit widths of the first input data, the second input data, and the rotation factor can be configured. Among them, the output data of each stage of the butterfly calculation unit is the input data of the next stage. Step 120, generate a subtraction result according to the first input data and the second input data. As shown in FIG. 2, the first input data x1 and the second input data x1 + n / 2 are subtracted to generate a subtraction result a. Step 130, perform a multiplication operation on the subtraction result and a first addition result of a modulus with the rotation factor to obtain a multiplication result. The following gives an exemplary description of the calculation of Montgomery multiplication. Traditional Montgomery multiplication is as follows: Among them, in the traditional GS butterfly transformation, the GS structure needs to determine whether the first input data x1 (the minuend) is less than the modulus N, and then decide whether to add N to perform subtraction. If it is determined that the first input data x1 is less than the modulus N, then N needs to be added before subtraction; if it is determined that the first input data x1 is not less than the modulus N, then N does not need to be added before subtraction. In the embodiment of the present application, no judgment operation is required before subtraction. The subtraction result a is directly added with the modulus N, which will not affect the multiplication result r. As shown in the following equation (1), the result of multiplying the subtraction result a plus the modulus N by the rotation factor b and taking the modulus remains unchanged. (a * b) mod N = ((a + N) * b) mod N (1); Among them, a represents the subtraction result; b represents the rotation factor, and b is the same as the rotation factor described above Consistent; N represents the modulus. In step 130, the first step (Step1) of Montgomery multiplication is adjusted to: In this process, no judgment operation is performed. After subtraction, the modulus N is added. In the subtraction process, there is no need to determine whether to add the modulus N by judging whether the first input data x1 is less than the modulus N, effectively reducing the number of judgment operations, reducing the computational complexity, improving the performance of NTT, and reducing the computational burden. Step 140, determining the calculation target according to the multiplication result. In step 140, based on the second step (Step2) of Montgomery multiplication, the calculation target u can be determined according to the multiplication result r. In some embodiments, the multiplication result includes a high-order multiplication result and a low-order multiplication result. The determining the calculation target according to the multiplication result includes: Determining an intermediate result according to the low-order multiplication result, where the intermediate result includes a high-order intermediate result and a low-order intermediate result; Determining a low-order carry value according to the low-order multiplication result or the low-order intermediate result; Determining the high-order calculation target of the calculation target according to the high-order multiplication result, the high-order intermediate result, and the low-order carry value. In some embodiments, the determining the high-order calculation target of the calculation target according to the high-order multiplication result, the high-order intermediate result, and the low-order carry value includes: determining the high-order calculation target of the calculation target according to the second addition result of the high-order multiplication result, the high-order intermediate result, and the low-order carry value. Illustrated in combination with the above calculation method of Montgomery multiplication, the multiplication result r includes the high-order multiplication result r high and the low-order multiplication result r low . In the process of determining the calculation target u according to the multiplication result r, the specific implementation steps are as follows: 1) Determine the intermediate result k according to the low-order multiplication result r low , where the intermediate result k includes the high-order intermediate result k high and the low-order intermediate result k low . Specifically, multiply the low-order multiplication result r low by the inverse element of the modulus N, and then multiply by the modulus N to obtain the intermediate result k, and the expression is k = (r low *n')*N. 2) Determine the low-order carry value c low according to the low-order multiplication result r low or the low-order intermediate result k low . Due to the characteristics of Montgomery multiplication, the lower bit calculation target u of the calculation target low (i.e., the lower 256 bits) must be 0, so the lower bit carry value c low is only divided into 0 and 1, and u low = 0. Based on this characteristic, when the lower bit carry value c low is 1, there must be at least one of the highest bits of the lower bit multiplication result r low and the lower bit intermediate result k low that is 1. Therefore, in the embodiments of the present application, the second addition operation in Step2 can be omitted, that is, the operation of (u low , c low ) = add(r low , k low ) can be omitted, and the lower bit carry value c low can be directly determined according to the lower bit multiplication result r low or the lower bit intermediate result k low , and the expression is c low = r low

[0255] or (or) k low

[0255] . In this way, a 256-bit addition operation is saved, and it becomes two single-bit bit logical operations, simplifying the calculation of the lower bit carry in the 256-bit Montgomery multiplication. 3) Determine the upper bit calculation target u of the calculation target u according to the upper bit multiplication result r high , the upper bit intermediate result k high and the lower bit carry value c low . Specifically, according to the second addition result of the upper bit multiplication result r high , the upper bit intermediate result k high and the lower bit carry value c high , determine the upper bit calculation target u of the calculation target u low , and the expression is u high = adc(r high , k high , c high , c low ). In some embodiments, when determining the upper bit calculation target of the calculation target according to the second addition result of the upper bit multiplication result, the upper bit intermediate result and the lower bit carry value, it further includes: Obtain a compensation number, and the sum of the compensation number and the modulus is the (n - 1)th power of 2, where n represents the bit width; Based on the carry-save adder, process the upper bit multiplication result, the upper bit intermediate result and the compensation number to obtain an addition output value and a carry output value; Shift the carry output value one bit to the left, and fill the least significant bit of the carry output value with the low-order carry value to obtain an updated carry output value; Sum the updated carry output value and the addition output value to obtain a sum value; Return a temporary value according to a first judgment result of determining whether there is a carry from the (n - 1)-th bit to the n-th bit of the sum value. In some embodiments, the returning a temporary value according to a first judgment result of determining whether there is a carry from the (n - 1)-th bit to the n-th bit of the sum value includes: If the first judgment result is that there is a carry from the (n - 1)-th bit to the n-th bit of the sum value, the returned temporary value is the modulus; or If the first judgment result is that there is no carry from the (n - 1)-th bit to the n-th bit of the sum value, the returned temporary value is 0. In some embodiments, the determining whether there is a carry from the (n - 1)-th bit to the n-th bit of the sum value includes: If the sum value from the (n - 1)-th bit to the n-th bit is not 0, it is determined that there is a carry from the (n - 1)-th bit to the n-th bit of the sum value; or If the sum value from the (n - 1)-th bit to the n-th bit is 0, it is determined that there is no carry from the (n - 1)-th bit to the n-th bit of the sum value. In some embodiments, the method further includes: if the first judgment result is that there is no carry from the (n - 1)-th bit to the n-th bit of the sum value, determining whether there is a carry from the 0-th bit to the (n - 2)-th bit of the sum value, and if there is a carry, the returned temporary value is the modulus. Among them, for the third step (Step3) of the traditional Montgomery multiplication, it is necessary to wait until the high-order calculation target u is calculated in the second step (Step2) high before the subsequent step of determining whether the calculation target is greater than or equal to the modulus mod can be executed. There are two problems: there is a serious data dependence, and the subsequent step of determining whether the calculation target is greater than or equal to the modulus mod can only be executed after the carry addition instruction adc() ends; a 256-bit comparison circuit is also required. For these two problems, the high-order calculation target u in the Montgomery multiplication high has a theoretical result range in [0, 2*mod), and finally it is necessary to determine whether the high-order calculation target u high is greater than the modulus mod. Since the modulus mod is usually a large prime number, such as the BLS-381 curve, the value is 0x73eda753299d7d483339d80809a1d80553bda402fffe5bfeffffffff00000001. It is troublesome to compare directly with this value, but after adding a compensation number δ to this value, the following formula can be satisfied: mod+δ=2 n-1 , where n represents the bit width. For example, if the bit width is 256, then the following formula is satisfied: mod + δ = 2 255 ; For example, taking n as 256, the r in the carry addition instruction adc() is high +k high +c low The result (high calculation target u high ) and the modulus mod can be interpreted as follows: if the former is greater than or equal to the latter, that is, r high +k high +c low ≥mod, then it must be equivalent to: r high +k high +c low +δ≥mod+δ=2 255 ; Therefore, we only need to check whether there is a carry in the sum of the four items on the left side from the n-1th bit to the nth bit (for example, bit [256:255]). low Only the lowest bit is valid, so we can calculate the sum of the other three items first, that is, r high +k high +δ, in the actual process, the carry-save adder (CSA) structure can be used for 3-2 compression, and the following formula is obtained: (us,uc)=csa(r high ,k high ,δ); Among them, us represents the addition output value, and uc represents the carry output value. As shown in Figure 3, the obtained carry output value uc[255:0] is shifted left by one bit, and the low-order carry value c low The lowest bit of the carry output value uc is filled in to obtain the updated carry output value. The sum of the updated carry output value and the addition output value us[255:0] is the sum of the three, and the obtained sum value is the data of [256:0]. Among them, if the sum value in bit [256:255] is not 0, it is determined that the sum value in bit [256:255] has a carry, indicating that the high-order calculation target u highIf it is greater than or equal to the modulus mod, the returned temporary value temp is the modulus mod. If the sum value at bit[256:255] is 0, it is determined that there is no carry at bit[256:255] of the sum value, indicating the high-bit calculation target u high If it is less than the modulus mod, the returned temporary value temp is 0. Among them, if the first judgment result is that there is no carry at bit[256:255] of the sum value, the addition can be reused to judge whether there is a carry at bit[254:0]. If there is a carry, the returned temporary value temp is the modulus mod; if there is no carry, the returned temporary value temp is 0. When reusing the addition to judge whether there is a carry at bit[254:0], specifically judge whether the sum value m

[0255] at the 255th bit is 1, where m

[0255] = (us + uc[255:0], c_low). In the embodiment of the present application, when executing the carry addition instruction adc() in the second step (Step2) of Montgomery multiplication, it is also possible to calculate in parallel whether there is a carry. After the multiplication, it is equivalent to advancing the judgment step (reduction) in the third step (Step3) of Montgomery multiplication, reducing the data dependence and large-bit-width comparison in Step3. Step 150, determine the output result according to the calculation target. In some embodiments, the determining the output result according to the calculation target includes: Judge whether the calculation target is greater than or equal to the modulus; If the calculation target is greater than or equal to the modulus, determine the difference between the calculation target and the modulus as the output result; or If the calculation target is less than the modulus, determine the calculation target as the output result. In some embodiments, the determining the output result according to the calculation target includes: Judge whether the high-bit calculation target is greater than or equal to the modulus; If the high-bit calculation target is greater than or equal to the modulus, determine the difference between the high-bit calculation target and the modulus as the output result; or If the high-bit calculation target is less than the modulus, determine the high-bit calculation target as the output result. In step 150, based on the third step (Step3) of Montgomery multiplication, the output result can be determined according to the calculation target u. That is, judge the high-bit calculation target u high Whether it is greater than or equal to the modulus mod; If the high-bit calculation target u highIf it is greater than or equal to the modulus mod, then the high-order calculation target u high The difference from the modulus mod is determined as the output result. If the high-order calculation target u high is less than the modulus mod, then the high-order calculation target u high is determined as the output result. In some embodiments, determining the output result according to the calculation target includes: Determining the difference between the high-order calculation target and the temporary value as the output result. In the embodiments of the present application, in the second step (Step2) of Montgomery multiplication, while executing the carry addition instruction adc(), it is also possible to calculate in parallel whether there is a carry. After multiplication, it is equivalent to advancing the judgment step in the third step (Step3) of Montgomery multiplication, reducing the data dependence and large-bit-width comparison in Step3. In this way, Step3 not only reduces the data dependence and large-bit-width comparison, but also can calculate the temporary value temp in parallel in Step2. In the embodiments of the present application, the high-order calculation target u high The difference from the temporary value temp can be determined as the output result, and only u high –temp is required, and the output result can be obtained without judgment. The data processing method provided by the embodiments of the present application can reduce the dependence between Step2 and Step3 and quickly execute the butterfly operation. All the above technical solutions can be combined arbitrarily to form the embodiments of the present application, which will not be elaborated one by one here. The embodiments of the present application obtain the first input data, the second input data and the rotation factor; generate a subtraction result according to the first input data and the second input data; perform a multiplication operation on the first addition result of the subtraction result and the modulus and the rotation factor to obtain a multiplication result; determine a calculation target according to the multiplication result; and determine an output result according to the calculation target. By adding the modulus after subtraction, the embodiments of the present application do not need to determine whether to add the modulus to do subtraction by judging whether the first input data is less than the modulus during the subtraction process, effectively reducing the number of judgment operations, reducing the calculation complexity, improving the performance of the NTT, and reducing the calculation burden. To facilitate better implementation of the data processing method of the embodiments of the present application, the embodiments of the present application also provide a zero-knowledge proof accelerator. Please refer to FIG. 4. FIG. 4 is a first structural schematic diagram of the zero-knowledge proof accelerator provided by the embodiments of the present application. Among them, the zero-knowledge proof accelerator 200 includes a number theory transform butterfly calculation module 210. Among them, the number theory transform butterfly calculation module 210 includes a first calculation unit 211, a second calculation unit 212 and a third calculation unit 213; The first calculation unit 211 is configured to perform a multiplication operation on the first addition result of the subtraction result and the modulus with a rotation factor to obtain a multiplication result, where the subtraction result is generated by subtracting the second input data from the first input data; The second calculation unit 212 is configured to determine a calculation target according to the multiplication result; The third calculation unit 213 is configured to determine an output result according to the calculation target. For example, the first calculation unit 211 may include a data input unit, a first subtractor, and a first multiplier; the data input unit is configured to input the first input data, the second input data, and the rotation factor; the first subtractor may be configured to generate a subtraction result according to the first input data and the second input data; the first multiplier may be configured to perform a multiplication operation on the first addition result of the subtraction result and the modulus with the rotation factor to obtain a multiplication result. In some embodiments, the multiplication result includes a high-order multiplication result and a low-order multiplication result, and the second calculation unit 212 is configured to: determine an intermediate result according to the low-order multiplication result, where the intermediate result includes a high-order intermediate result and a low-order intermediate result; determine a low-order carry value according to the low-order multiplication result or the low-order intermediate result; and determine a high-order calculation target of the calculation target according to the high-order multiplication result, the high-order intermediate result, and the low-order carry value. In some embodiments, the third calculation unit 213 is configured to: determine whether the calculation target is greater than or equal to the modulus; if the calculation target is greater than or equal to the modulus, determine the difference between the calculation target and the modulus as the output result; or if the calculation target is less than the modulus, determine the calculation target as the output result. In some embodiments, the third calculation unit 213 is configured to: determine whether the high-order calculation target is greater than or equal to the modulus; if the high-order calculation target is greater than or equal to the modulus, determine the difference between the high-order calculation target and the modulus as the output result; or if the high-order calculation target is less than the modulus, determine the high-order calculation target as the output result. In some embodiments, when determining the high-order calculation target of the calculation target according to the high-order multiplication result, the high-order intermediate result, and the low-order carry value, the second calculation unit 212 is configured to: determine the high-order calculation target of the calculation target according to the second addition result of the high-order multiplication result, the high-order intermediate result, and the low-order carry value. In some embodiments, the second computing unit 212 may further be configured to: obtain a compensation number, where the sum of the compensation number and the modulus is 2 to the power of n - 1, and n represents the bit width; process the high-order multiplication result, the high-order intermediate result, and the compensation number based on a carry-save adder to obtain an addition output value and a carry output value; shift the carry output value one bit to the left and fill the least significant bit of the carry output value with the low-order carry value to obtain an updated carry output value; sum the updated carry output value and the addition output value to obtain a sum value; and return a temporary value according to a first judgment result of determining whether there is a carry in the (n - 1)-th bit to the n-th bit of the sum value. For example, the second computing unit 212 may include a carry-save adder (CSA), and the carry-save adder is configured to process the high-order multiplication result, the high-order intermediate result, and the compensation number to obtain an addition output value and a carry output value. In some embodiments, when the second computing unit 212 returns a temporary value according to a first judgment result of determining whether there is a carry in the (n - 1)-th bit to the n-th bit of the sum value, it may be configured to: if the first judgment result is that there is a carry in the (n - 1)-th bit to the n-th bit of the sum value, the returned temporary value is the modulus; or if the first judgment result is that there is no carry in the (n - 1)-th bit to the n-th bit of the sum value, the returned temporary value is 0. In some embodiments, when the second computing unit 212 determines whether there is a carry in the (n - 1)-th bit to the n-th bit of the sum value, it may be configured to: if the (n - 1)-th bit to the n-th bit of the sum value is not 0, determine that there is a carry in the (n - 1)-th bit to the n-th bit of the sum value; or if the (n - 1)-th bit to the n-th bit of the sum value is 0, determine that there is no carry in the (n - 1)-th bit to the n-th bit of the sum value. In some embodiments, the second computing unit 212 may further be configured to: if the first judgment result is that there is no carry in the (n - 1)-th bit to the n-th bit of the sum value, determine whether there is a carry in the 0-th bit to the (n - 2)-th bit of the sum value. If there is a carry, the returned temporary value is the modulus. In some embodiments, the third computing unit 213 is configured to: determine the difference between the high-order computing target and the temporary value as the output result. As shown in FIG. 5, FIG. 5 is a second structural schematic diagram of a zero-knowledge proof accelerator provided by an embodiment of the present application. The zero-knowledge proof accelerator 200 may further include a multi-scalar multiplication (MSM) module 220. The multi-scalar dot product calculation module 220 can be used to process the output result of the third calculation unit 213 to obtain the multi-scalar dot product calculation result. Among them, the number theory transform butterfly calculation module 210 may further include a reduction unit 214, and the reduction unit 214 can be used to perform reduction calculations. Among them, the first calculation unit 211 can simultaneously input multiple groups of data. Each group of input data may include first input data, second input data, and corresponding rotation factors, and perform pipelined parallel calculations on the multiple groups of input data input simultaneously to improve the calculation efficiency. The zero-knowledge proof accelerator 200 provided by the embodiments of the present application can reduce the dependencies of Step 2 executed in the second calculation unit 212 and Step 3 executed in the third unit 213. Among them, due to the reduction of the dependencies of Step 2 and Step 3, while the second calculation unit 212 determines the calculation target of the next round based on the multiplication result of the next round, the third unit 213 can parallelly process determining the output result of the current round based on the calculation target of the current round, realizing fast execution of the butterfly operation. Among them, the zero-knowledge proof accelerator 200 can also be connected to a host computer, and the zero-knowledge proof accelerator 200 can also be used to receive relevant parameters configured by the host computer. All the above technical solutions can be combined arbitrarily to form the embodiments of the present application, which will not be elaborated here one by one. It should be understood that the embodiments of the zero-knowledge proof accelerator and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, the zero-knowledge proof accelerator can execute the above data processing method embodiments, and the foregoing and other operations and / or functions of each unit in the zero-knowledge proof accelerator respectively implement the corresponding processes of the above method embodiments. For the sake of brevity, it will not be elaborated here. In one embodiment, the present application further provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented. FIG. 6 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device may be a terminal or a server. As shown in FIG. 6, the computer device 300 may include: a communication interface 301, a memory 302, a processor 303, and a communication bus 304. The communication interface 301, the memory 302, and the processor 303 communicate with each other through the communication bus 304. The communication interface 301 is used for the computer device 300 to perform data communication with external devices. The memory 302 may be used to store software programs and modules. The processor 303 runs the software programs and modules stored in the memory 302, such as the software programs for the corresponding operations in the foregoing method embodiments. In one embodiment, the processor 303 may call the software programs and modules stored in the memory 302 to perform the following operations: obtain first input data, second input data, and a rotation factor; generate a subtraction result according to the first input data and the second input data; perform a multiplication operation on the first addition result of the subtraction result and a modulus with the rotation factor to obtain a multiplication result; determine a calculation target according to the multiplication result; and determine an output result according to the calculation target. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the foregoing embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this reason, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute the steps in any one of the data processing methods provided by the embodiments of the present application. For the specific implementation of each of the above operations, reference can be made to the foregoing embodiments, which will not be elaborated herein. Among them, the storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, or the like. Since the computer programs stored in the storage medium can execute the steps in any one of the data processing methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any one of the data processing methods provided by the embodiments of the present application can be realized. For details, reference can be made to the foregoing embodiments, which will not be elaborated herein. An embodiment of the present application further provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding processes in any one of the data processing methods in the embodiments of the present application. For the sake of brevity, it will not be elaborated herein. The embodiments of the present application also provide a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the corresponding processes in any one of the data processing methods in the embodiments of the present application. For the sake of brevity, details are not described herein again. The embodiments of the present application also provide a chip, which includes a processor for implementing the functions involved in any one or more of the above embodiments, such as obtaining or processing the information or messages involved in the above methods. Optionally, the chip further includes a memory for storing the necessary program instructions and data for the processor to execute. The chip may be composed of chips or may include chips and other discrete devices. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining first input data, second input data, and a rotation factor; Generating a subtraction result based on the first input data and the second input data; Performing a multiplication operation on the subtraction result and a first addition result of a modulus with the rotation factor to obtain a multiplication result; Determining a calculation target based on the multiplication result; Determining an output result based on the calculation target.

2. The data processing method according to claim 1, wherein The multiplication result includes a high-order multiplication result and a low-order multiplication result. The determining the calculation target based on the multiplication result includes: Determining an intermediate result based on the low-order multiplication result, where the intermediate result includes a high-order intermediate result and a low-order intermediate result; Determining a low-order carry value based on the low-order multiplication result or the low-order intermediate result; Determining a high-order calculation target of the calculation target based on the high-order multiplication result, the high-order intermediate result, and the low-order carry value.

3. The data processing method according to claim 2, wherein The determining the output result based on the calculation target includes: Judging whether the calculation target is greater than or equal to the modulus; If the calculation target is greater than or equal to the modulus, then determining the difference between the calculation target and the modulus as the output result; or If the calculation target is less than the modulus, then determining the calculation target as the output result.

4. The data processing method according to claim 3, wherein The determining the output result based on the calculation target includes: Judging whether the high-order calculation target is greater than or equal to the modulus; If the high-order calculation target is greater than or equal to the modulus, then determining the difference between the high-order calculation target and the modulus as the output result; or If the high-order calculation target is less than the modulus, then determining the high-order calculation target as the output result.

5. The data processing method according to claim 2, wherein The determining the high-order calculation target of the calculation target based on the high-order multiplication result, the high-order intermediate result, and the low-order carry value includes: Determining the high-order calculation target of the calculation target based on a second addition result of the high-order multiplication result, the high-order intermediate result, and the low-order carry value.

6. The data processing method according to claim 5, characterized in that, When determining the high-order calculation target of the calculation target based on the second addition result of the high-order multiplication result, the high-order intermediate result, and the low-order carry value, it further includes: Obtaining a compensation number, where the sum of the compensation number and the modulus is \(2^{n - 1}\), and \(n\) represents the bit width; Processing the high-order multiplication result, the high-order intermediate result, and the compensation number based on a carry-save adder to obtain an addition output value and a carry output value; Shifting the carry output value one bit to the left and filling the lowest bit of the carry output value with the low-order carry value to obtain an updated carry output value; Summing the updated carry output value and the addition output value to obtain a sum value; Returning a temporary value according to a first judgment result of judging whether there is a carry in the \((n - 1)\)-th bit to the \(n\)-th bit of the sum value.

7. The data processing method according to claim 6, wherein The returning a temporary value according to a first judgment result of judging whether there is a carry in the \((n - 1)\)-th bit to the \(n\)-th bit of the sum value includes: If the first judgment result is that there is a carry in the \((n - 1)\)-th bit to the \(n\)-th bit of the sum value, then the returned temporary value is the modulus; or If the first judgment result is that there is no carry in the (n - 1)-th to n-th bits of the sum value, the returned temporary value is 0.

8. The data processing method according to claim 6, wherein The judgment of whether there is a carry in the (n - 1)-th to n-th bits of the sum value includes: If the sum value in the (n - 1)-th to n-th bits is not 0, it is determined that there is a carry in the (n - 1)-th to n-th bits of the sum value; or If the sum value in the (n - 1)-th to n-th bits is 0, it is determined that there is no carry in the (n - 1)-th to n-th bits of the sum value.

9. A zero-knowledge proof accelerator, characterized in that, The zero-knowledge proof accelerator includes a number-theoretic transform butterfly calculation module, where the number-theoretic transform butterfly calculation module includes a first calculation unit, a second calculation unit, and a third calculation unit; The first calculation unit is configured to perform a multiplication operation on the first addition result of the subtraction result and the modulus with a rotation factor to obtain a multiplication result, where the subtraction result is generated by subtracting the second input data from the first input data; The second calculation unit is configured to determine a calculation target according to the multiplication result; The third calculation unit is configured to determine an output result according to the calculation target.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the data processing method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Method for long-number division or modular reduction

    CN103339665A

  • Modular operation circuit and modular operation method

    CN116627385A

  • Iterative NTT staggered storage system based on BRAM

    CN116679905A

  • Iterative NTT system based on FIFO storage

    CN116893797A

  • Data processing method and device, equipment and storage medium

    CN117254902A