Dynamic reconfigurable zero-knowledge proof accelerator and acceleration method
Through a dynamically reconfigurable zero-knowledge proof accelerator, utilizing fine-grained finite field arithmetic units and custom instruction sets, the problems of poor adaptability and low resource utilization of existing accelerator algorithms are solved, and efficient zero-knowledge proof generation acceleration is achieved.
Patent Information
- Application Number
- CN202510792671.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-05
AI Technical Summary
Existing zero-knowledge proof accelerators use coarse-grained dedicated hardware acceleration units, resulting in poor algorithm adaptability and low resource utilization. They are difficult to adapt to protocol switching and updates, and the computing operation time is uneven.
It adopts a dynamically reconfigurable zero-knowledge proof accelerator, through a fine-grained finite field arithmetic unit array and on-chip interconnect device, combined with a custom instruction set, to achieve dynamic pipeline combination and memory access scheduling of different computing operations, thereby improving protocol adaptability and resource utilization efficiency.
It improves the protocol adaptability and resource utilization of the accelerator, realizes efficient acceleration of the zero-knowledge proof generation process, is applicable to multiple protocols and reduces the complexity of software control.
Smart Images

Figure CN120602098A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of zero-knowledge proof hardware acceleration, and relates to a dynamically reconfigurable zero-knowledge proof accelerator and an acceleration method. Background Art
[0002] Zero-knowledge proof (ZKP) is an important privacy-preserving computing protocol that enables one party (the prover) to prove the validity of a computational problem to another party (the verifier) without revealing any private data. This protocol is widely used in privacy-preserving blockchains, secure databases, and verifiable machine learning. However, generating ZKPs involves a large number of finite field arithmetic operations, resulting in significant computational and storage overhead. Therefore, designing an efficient ZKP accelerator is a key issue.
[0003] Existing zero-knowledge proof accelerators typically design dedicated hardware acceleration units for different computational operations in proof generation, leveraging coarse-grained parallelism to accelerate proof generation calculations. This coarse-grained acceleration approach has two drawbacks: First, dedicated hardware acceleration units struggle to adapt to algorithm switching and updates, limiting the types of protocols supported by the accelerator and reducing long-term usability; second, the time consumption of different computational operations in proof generation varies significantly, resulting in unbalanced load on the dedicated hardware acceleration units and low resource utilization. Summary of the Invention
[0004] To address the problems of poor algorithm adaptability and low resource utilization in the coarse-grained, dedicated hardware acceleration units widely used in existing zero-knowledge proof accelerators, the present invention provides a dynamically reconfigurable zero-knowledge proof accelerator and acceleration method. The accelerator constructs a unified hardware unit based on fine-grained finite field arithmetic operations. It dynamically combines different computational operation pipelines through on-chip interconnects, and uses a custom instruction set for pipeline configuration and memory access scheduling. This improves the accelerator's protocol adaptability, resource utilization efficiency, and computational performance, achieving efficient acceleration of the complete zero-knowledge proof generation process.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A dynamically reconfigurable zero-knowledge proof accelerator, comprising:
[0007] a finite field arithmetic operation unit array for performing addition, subtraction, and multiplication operations within a finite field;
[0008] an on-chip network, configured to interconnect a plurality of finite field arithmetic operation units in the finite field arithmetic operation unit array to construct data paths corresponding to different computing operations;
[0009] The instruction controller supports customized accelerator instruction sets, is used to read configuration instructions, memory access instructions, and calculation instructions, and controls the execution process of the accelerator based on the instruction type.
[0010] Furthermore, it also includes a memory access controller and an on-chip memory;
[0011] The memory access controller is used to perform data loading or writing operations between the off-chip memory and the on-chip memory according to the memory access instruction;
[0012] The on-chip memory is used to temporarily store intermediate data and result data required for calculation.
[0013] Furthermore, the finite field arithmetic operation unit array is divided into several operation clusters, each operation cluster is composed of multiple finite field arithmetic operation units interconnected by an on-chip network router, and is used to perform computing operations with high correlation in data streams; different operation clusters are used to carry computing tasks that have small dependencies or are resource-constrained and cannot coexist in the same cluster, and the on-chip network router is responsible for completing cross-cluster data forwarding and communication.
[0014] Furthermore, the finite field arithmetic operation unit includes three basic operation units: a finite field addition unit, a finite field subtraction unit and a finite field multiplication unit, which are used to perform corresponding arithmetic operations in different operation modes.
[0015] Furthermore, the finite field arithmetic operation unit further includes a calculation mode register, a decoder and a multiplexer;
[0016] The calculation mode register is used to store the current calculation mode value;
[0017] The decoder is used to decode the calculation mode value into a control signal;
[0018] The multiplexer is used to select the input path and output path of each basic operation unit in the finite field arithmetic operation unit according to the control signal to realize different operation modes.
[0019] Furthermore, the finite field arithmetic operation unit supports finite field operations of multiple bit widths, including 768 bits, 384 bits, 256 bits, 128 bits and 64 bits;
[0020] When performing 768-bit operations, the finite field arithmetic operation unit performs large bit width operations in a single task mode;
[0021] When executing 384-bit, 256-bit, 128-bit, and 64-bit operations, the finite field arithmetic operation unit adopts a bit width decomposition method to support parallel execution of multiple operation tasks.
[0022] Furthermore, the interconnection state of the on-chip network is controlled by an interconnection mode register, which is used to dynamically configure the connection state of each crossbar switch in the on-chip network to adapt to different data flow modes.
[0023] Furthermore, the instruction controller includes an instruction memory, a decoder, a data distributor, a mode control register table and an instruction engine;
[0024] The instruction memory is used to store various instruction sequences in the zero-knowledge proof generation process, including configuration instructions, memory access instructions and calculation instructions;
[0025] The decoder is used to parse the instruction content in the instruction memory, identify the instruction type and extract the corresponding operation parameters;
[0026] The data distributor is used to send the operation parameters of the instruction to the corresponding instruction engine according to the instruction type;
[0027] The mode control register table is used to store a plurality of preset calculation mode values and interconnection mode values, and is used to dynamically configure the calculation mode register of the finite field arithmetic operation unit and the interconnection mode register of the on-chip network;
[0028] The instruction engine is used to call the corresponding hardware module to perform corresponding configuration, memory access or calculation operations according to the instruction type and parameters output by the decoder.
[0029] Furthermore, the instruction engine includes a configuration engine, a memory access engine and a calculation engine;
[0030] The configuration engine is used to execute configuration instructions to complete dynamic configuration of hardware;
[0031] The memory access engine is used to execute memory access instructions to control the transmission of data between the off-chip memory and the on-chip memory;
[0032] The computing engine is used to execute computing instructions to schedule the finite field arithmetic operation unit array to perform specific computing operations on the data in the on-chip memory.
[0033] A zero-knowledge proof acceleration method based on a dynamically reconfigurable architecture is executed based on a dynamically reconfigurable zero-knowledge proof accelerator and includes the following steps:
[0034] 1) The instruction controller executes configuration instructions, calls the computing mode values and interconnection mode values in the mode control register table, configures the on-chip network and finite field arithmetic operation unit, and builds the data path and operation structure adapted to the target computing operation;
[0035] 2) The instruction controller executes the memory access instruction and loads the input data from the off-chip memory to the on-chip memory through the memory access controller;
[0036] 3) The instruction controller executes the computational instructions, dispatches the data in the on-chip memory to the configured finite field arithmetic operation unit array, completes the computational processing within the finite field on the target hardware acceleration pipeline, and writes the results back to the on-chip memory;
[0037] 4) The instruction controller executes the memory access instruction and writes the calculation result from the on-chip memory back to the off-chip memory through the memory access controller;
[0038] 5) Repeat steps 2) to 4) until the data required for the current calculation operation is processed;
[0039] 6) Repeat steps 1) to 5) until all instruction sequences are executed, completing the zero-knowledge proof generation process.
[0040] The beneficial effects achieved by the present invention are as follows:
[0041] 1. The present invention adopts a uniformly designed fine-grained finite field arithmetic operation unit to replace the traditional coarse-grained dedicated hardware module, thereby improving the flexibility and long-term availability of the system structure.
[0042] 2. The present invention connects multiple unified finite field arithmetic operation units through an on-chip interconnection device, and can flexibly configure the hardware acceleration pipeline according to the data flow mode of different computing operations, thereby improving the algorithm adaptability.
[0043] 3. The present invention dynamically configures the computing pipeline and memory access path through a customized accelerator instruction set, thereby realizing automatic scheduling and accelerated processing of the entire zero-knowledge proof process and reducing the complexity of software control.
[0044] 4. The present invention can fully exploit fine-grained parallelism at the finite field operation level, improve the utilization efficiency of hardware resources, and alleviate the module idleness problem caused by uneven task distribution in traditional solutions.
[0045] 5. This invention does not rely on a specific type of zero-knowledge proof protocol and has good versatility. It is applicable to zero-knowledge proof protocols based on finite field arithmetic operations, including elliptic curve-based protocols (such as Groth16 and Plonk) and hash function-based protocols (such as zkSTARK and Plonky2). BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of the overall architecture of the dynamically reconfigurable zero-knowledge proof accelerator.
[0047] Figure 2 It is a structural diagram of the finite field arithmetic unit.
[0048] Figure 3 It is a structural diagram of the accelerator's command controller. DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to specific embodiments and accompanying drawings.
[0050] like Figure 1 As shown, this embodiment discloses a dynamically reconfigurable zero-knowledge proof accelerator. Its overall structure includes: a finite field arithmetic operation unit array based on an on-chip network, a memory access controller, an instruction controller, and on-chip memory. Through the collaborative operation of hardware modules, the accelerator supports dynamic configuration and parallel acceleration of various computing operations, and is suitable for the proof generation process of various zero-knowledge proof protocols. The accelerator can be implemented using a field programmable logic device (FPGA) or an application-specific integrated circuit (ASIC).
[0051] The finite field arithmetic operation unit array is composed of multiple finite field arithmetic operation units, and each operation unit is interconnected through an on-chip network (NoC). The on-chip network supports flexible interconnection structure reconstruction, and its internal crossbar switch state is configured and managed by an interconnection mode register (not shown). By modifying the value of the interconnection mode register, different data transmission paths can be dynamically constructed to achieve the data flow mode required for different computing operations, thereby combining multiple finite field arithmetic operation units as needed into specific hardware acceleration pipelines, including but not limited to number theory transformation (NTT) pipelines, elliptic curve point addition pipelines, etc.
[0052] In a preferred embodiment, multiple finite field arithmetic operation units (FFUs) are organized into several computational clusters via an on-chip network router, which in turn form the entire FFU array. During task scheduling, computational operations with high data flow correlation are preferentially assigned to the same cluster to shorten communication paths and improve computational efficiency. Computational operations with more distant dependencies or resource constraints that prevent them from coexisting in the same cluster are assigned to different clusters, with data forwarding performed by the routers in the on-chip network. The number of computational units in each cluster, as well as the size and topology of the clusters in the entire array, can be flexibly configured based on performance objectives and hardware resource constraints, achieving an optimal trade-off between cost, power consumption, and computational performance.
[0053] Figure 2The structure of the finite field arithmetic operation unit is highly adaptable to the finite field arithmetic operation modes in different computing operations. It is used to accelerate the basic operations of proof generation, including finite field multiplication, addition, and subtraction. The core structure of the finite field arithmetic operation unit includes three basic units: a finite field multiplication unit, a finite field addition unit, and a finite field subtraction unit. The number of each basic unit can be flexibly adjusted to meet the cost constraints and performance requirements of the accelerator. Among them, the finite field multiplication unit uses a fast modular multiplication algorithm, including Montgomery modular multiplication or Barrett modular multiplication. The finite field addition unit and the finite field subtraction unit implement subtraction or addition of the modulus according to overflow conditions.
[0054] In addition to the basic unit, the finite field arithmetic operation unit also includes a calculation mode register, a decoder and a multiplexer. Among them, the input and output of the basic unit are connected through the multiplexer. The selection signal of the multiplexer is stored in a dedicated calculation mode register. By configuring the value of the calculation mode register, the finite field arithmetic operation unit can implement different finite field arithmetic operation modes, including but not limited to butterfly operations, multiplication and accumulation operations, etc. The decoder is used to decode the mode value stored in the calculation mode register, generate a control signal, and output the control signal to the multiplexer at the input and output ends of the finite field arithmetic operation unit to control it to select different data paths, thereby supporting different arithmetic operation modes.
[0055] The selection of the finite field depends on the specific zero-knowledge proof protocol and the security level requirements. For example, the zero-knowledge proof protocol based on elliptic curves usually selects a finite field with a bit width of 256-768 bits, and the zero-knowledge proof protocol based on hashing usually selects a finite field with a bit width of 64 bits. In order to adapt to operations with different bit widths, this embodiment sets the bit width of the finite field arithmetic operation unit to 768 bits, and adopts a bit width decomposition algorithm to implement large bit width operations. Therefore, the finite field arithmetic operation unit supports one 768-bit bit width operation, or multiple 384-bit bit width operations, 256-bit bit width operations, 128-bit bit width operations or 64-bit bit width operations in parallel.
[0056] like Figure 2 As shown, the finite field arithmetic operation unit disclosed in this embodiment has a high degree of adaptability to operation modes and can meet various basic computing requirements in the process of generating zero-knowledge proofs, including multiplication, addition, and subtraction operations within a finite field. The core structure of the operation unit consists of three types of configurable basic arithmetic units: a finite field multiplication unit, a finite field addition unit, and a finite field subtraction unit. The number of the above basic units can be flexibly set according to performance requirements and hardware resource constraints to achieve an optimal balance between computing efficiency and resource overhead.
[0057] The finite field multiplication unit uses efficient modular multiplication algorithms, including Montgomery modular multiplication and Barrett modular multiplication, to improve multiplication throughput and resource utilization. The finite field addition and subtraction units employ overflow detection logic to automatically perform modular addition / subtraction corrections on the results, ensuring that all calculations are performed within the target finite field.
[0058] In addition to the above-mentioned arithmetic units, the finite field arithmetic operation unit also includes a calculation mode register, a decoder, and multiple multiplexers for dynamically controlling the data path and adapting to different calculation structures. Specifically, the input and output of each basic unit are connected through a multiplexer, and the control signal of the selector is provided by the calculation mode register. The decoder is responsible for receiving the mode value stored in the calculation mode register, parsing it, and outputting the corresponding control signal to the multiplexer at the input and output ends, thereby realizing the construction and switching of different data paths, and supporting a variety of finite field arithmetic operation modes including but not limited to butterfly operations and multiply-accumulate operations (MAC).
[0059] In addition, the supported finite field bit width can be configured according to the adopted zero-knowledge proof protocol and its corresponding security level. For example: elliptic curve-based protocols usually use finite fields of 256 to 768 bits; hash-based protocols often use 64-bit finite fields. In order to be compatible with different protocol requirements, this embodiment sets the bit width of the finite field arithmetic operation unit to 768 bits, and supports the parallel execution of smaller bit width operations through a bit width decomposition algorithm. As a result, the unit not only supports a single 768-bit operation, but also can execute multiple sub-operations in parallel, including two 384-bit operations, three 256-bit operations, six 128-bit operations or twelve 64-bit operations, significantly improving computing throughput and hardware utilization efficiency.
[0060] like Figure 3 As shown, the instruction controller provided in this embodiment includes an instruction memory, a decoder, a data distributor, a mode control register table, and several instruction engines, including a configuration engine, a memory access engine, and a computation engine. The instruction controller is used to execute the various instructions required for zero-knowledge proof generation, dynamically configure the hardware pipeline structure, and schedule memory access and computation operations, thereby accelerating the entire proof generation process.
[0061] The instruction memory is used to store instruction sequences, including configuration instructions, memory access instructions, and computational instructions. Each instruction contains a type field that indicates the instruction type. The instruction controller reads and parses the instruction content through a decoder, identifies the instruction type, and dispatches the instruction to the corresponding instruction engine to execute the corresponding operation.
[0062] Specifically, the configuration instruction includes a mode control register table subscript field, which is used to read the corresponding computation mode value and interconnection mode value from the mode control register table. The computation mode value is used to configure the computation mode register of the finite field arithmetic operation unit, and the interconnection mode value is used to configure the interconnection mode register of the on-chip network, thereby dynamically building a hardware acceleration pipeline for specific computational operations. The mode control register table is used to store multiple preset computation and interconnection configuration combinations. The configuration engine selects the corresponding configuration from the subscript provided by the instruction and completes the writing.
[0063] The memory access instruction contains the off-chip memory address field, the on-chip memory address field and the data size field. The memory access engine controls the memory access controller to complete the data transmission between the off-chip and on-chip memories, thereby loading a certain size of data from the off-chip memory to the on-chip memory, or writing the calculation results from the on-chip memory back to the off-chip memory.
[0064] Computational instructions contain an opcode field, a target pipeline index field, an on-chip memory address field, and a data size field. These instructions are used to schedule the corresponding pipeline to read data from the on-chip memory, perform specific computations, and write the results back to the on-chip memory. The computation engine controls the coordinated operation of the various finite field arithmetic units in the pipeline based on the opcode and target pipeline information, completing basic finite field operations such as multiplication, addition, and subtraction, as well as operations such as number theoretic transformations (NTTs) and elliptic curve point additions that combine these basic operations.
[0065] In the instruction controller, the data distributor acts as a coordination hub, receiving instruction fields from the decoder and routing information such as address, size, and control parameters to the configuration engine, memory access engine, or compute engine based on type. For example, if the instruction is a configuration instruction, the data distributor transmits the mode register index to the configuration engine; if it is a memory access instruction, it transmits the address and data size information to the memory access engine; if it is a compute instruction, it distributes the operation parameters and pipeline identifier to the compute engine.
[0066] Through the collaborative work between the above modules, the instruction controller can accurately configure and schedule the data path and functional units of the entire accelerator, support the rapid switching and parallel execution of different computing tasks, and thus achieve efficient hardware acceleration of multiple stages in the zero-knowledge proof generation process.
[0067] This embodiment also discloses a zero-knowledge proof acceleration method based on a dynamically reconfigurable architecture, which relies on the aforementioned accelerator structure to accelerate the entire proof generation process. Before the calculation begins, the input data and key used for zero-knowledge proof generation are loaded into an off-chip memory, the instruction sequence is loaded into the instruction memory in the instruction controller, and the calculation mode and interconnection mode are written into the mode control register table. The specific processing steps of this method include:
[0068] 1) The instruction controller parses and executes configuration instructions: The configuration engine extracts configuration instructions from the instruction memory and reads the corresponding computation mode and interconnection mode values from the mode control register table based on the subscript field of the mode control register table. The configuration engine writes these mode values into the computation mode register of the finite field arithmetic operation unit and the interconnection mode register of the on-chip network, respectively. The configuration engine dynamically constructs an acceleration pipeline for the current computation task in the finite field arithmetic operation unit array, such as a number theoretic transform (NTT) pipeline or an elliptic curve point addition pipeline.
[0069] 2) The memory access engine parses and executes memory access instructions: The memory access instruction specifies the off-chip address, on-chip address and data size; the memory access controller reads the specified data block from the off-chip memory according to the instruction and transfers it to the on-chip memory through the data path to provide input data for subsequent computing operations.
[0070] 3) The computing engine parses and executes the computing instructions: extracting the opcode, target pipeline index, input data address and data size from the instruction; the computing engine controls the data distributor to retrieve the required data from the on-chip memory and sends it to the target hardware pipeline composed of finite field arithmetic units; after the pipeline completes the specified calculation, the result is written back to the on-chip memory again.
[0071] 4) The memory access engine executes the memory access instruction again: calling the memory access controller to write back some calculation results from the on-chip memory to the off-chip memory; achieving continuous exchange with the external storage and providing intermediate results for the next round of data loading or subsequent external processing.
[0072] 5) The above steps 2) to 4) are executed in a loop until all data involved in the current calculation operation are processed.
[0073] 6) After completing the current computation, the instruction controller continues to read and execute the next configuration instruction, repeating steps 1) to 5) until all instruction sequences are executed, completing the entire proof generation task.
[0074] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Those skilled in the art may modify or replace the technical solutions of the present invention with equivalents without departing from the spirit and scope of the present invention. The scope of protection of the present invention shall be based on the claims.
Claims
1. A dynamically reconfigurable zero-knowledge proof accelerator, characterized in that: include: a finite field arithmetic operation unit array for performing addition, subtraction, and multiplication operations within a finite field; an on-chip network, configured to interconnect a plurality of finite field arithmetic operation units in the finite field arithmetic operation unit array to construct data paths corresponding to different computing operations; The instruction controller supports customized accelerator instruction sets, is used to read configuration instructions, memory access instructions, and calculation instructions, and controls the execution process of the accelerator based on the instruction type.
2. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 1, wherein: Also includes a memory access controller and on-chip memory; The memory access controller is used to perform data loading or writing operations between the off-chip memory and the on-chip memory according to the memory access instruction; The on-chip memory is used to temporarily store intermediate data and result data required for calculation.
3. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 1, wherein: The finite field arithmetic operation unit array is divided into several operation clusters. Each operation cluster is composed of multiple finite field arithmetic operation units interconnected by an on-chip network router, which is used to perform computing operations with high correlation in data streams. Different operation clusters are used to carry computing tasks with small dependencies or limited resources that cannot coexist in the same cluster. The on-chip network router is responsible for completing cross-cluster data forwarding and communication.
4. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 1, wherein: The finite field arithmetic operation unit includes three basic operation units: a finite field addition unit, a finite field subtraction unit and a finite field multiplication unit, which are used to perform corresponding arithmetic operations in different operation modes.
5. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 4, wherein: The finite field arithmetic operation unit further includes a calculation mode register, a decoder and a multiplexer; The calculation mode register is used to store the current calculation mode value; The decoder is used to decode the calculation mode value into a control signal; The multiplexer is used to select the input path and output path of each basic operation unit in the finite field arithmetic operation unit according to the control signal to realize different operation modes.
6. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 1, wherein: The finite field arithmetic operation unit supports finite field operations of multiple bit widths, including 768 bits, 384 bits, 256 bits, 128 bits and 64 bits; When performing 768-bit operations, the finite field arithmetic operation unit performs large bit width operations in a single task mode; When executing 384-bit, 256-bit, 128-bit, and 64-bit operations, the finite field arithmetic operation unit adopts a bit width decomposition method to support parallel execution of multiple operation tasks.
7. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 1, wherein: The interconnection state of the on-chip network is controlled by an interconnection mode register, which is used to dynamically configure the connection state of each crossbar switch in the on-chip network to adapt to different data flow modes.
8. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 1, wherein: The instruction controller includes an instruction memory, a decoder, a data distributor, a mode control register table and an instruction engine; The instruction memory is used to store various instruction sequences in the zero-knowledge proof generation process, including configuration instructions, memory access instructions and calculation instructions; The decoder is used to parse the instruction content in the instruction memory, identify the instruction type and extract the corresponding operation parameters; The data distributor is used to send the operation parameters of the instruction to the corresponding instruction engine according to the instruction type; The mode control register table is used to store a plurality of preset calculation mode values and interconnection mode values; The instruction engine is used to call the corresponding hardware module to perform corresponding configuration, memory access or calculation operations according to the instruction type and parameters output by the decoder.
9. The dynamically reconfigurable zero-knowledge proof accelerator according to claim 8, wherein: The instruction engine includes a configuration engine, a memory access engine and a calculation engine; The configuration engine is used to execute configuration instructions to complete dynamic configuration of hardware; The memory access engine is used to execute memory access instructions to control the transmission of data between the off-chip memory and the on-chip memory; The computing engine is used to execute computing instructions to schedule the finite field arithmetic operation unit array to perform specific computing operations on the data in the on-chip memory.
10. A zero-knowledge proof acceleration method based on a dynamically reconfigurable architecture, executed based on the dynamically reconfigurable zero-knowledge proof accelerator according to any one of claims 1 to 9, characterized in that: The following steps are involved: 1) The instruction controller executes configuration instructions, calls the computing mode values and interconnection mode values in the mode control register table, configures the on-chip network and finite field arithmetic operation unit, and builds the data path and operation structure adapted to the target computing operation; 2) The instruction controller executes the memory access instruction and loads the input data from the off-chip memory to the on-chip memory through the memory access controller; 3) The instruction controller executes the computational instructions, dispatches the data in the on-chip memory to the configured finite field arithmetic operation unit array, completes the computational processing within the finite field on the target hardware acceleration pipeline, and writes the results back to the on-chip memory; 4) The instruction controller executes the memory access instruction and writes the calculation result from the on-chip memory back to the off-chip memory through the memory access controller; 5) Repeat steps 2) to 4) until the data required for the current calculation operation is processed; 6) Repeat steps 1) to 5) until all instruction sequences are executed, completing the zero-knowledge proof generation process.