Hardware acceleration method for elliptic curve addition group and related device

By merging and splitting the elliptic curve addition group task, multiple subtasks can be processed in parallel, solving the problem of insufficient bit width of hardware modules, improving computational efficiency and resource utilization, and making it suitable for high-performance computing scenarios such as blockchain and digital signatures.

CN118819640BActive Publication Date: 2026-01-02ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411267730.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2026-01-02
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

In existing technologies, the bit width of the elliptic curve addition group hardware module is insufficient to meet the data processing requirements of modern encryption algorithms, resulting in excessive consumption of computing resources and affecting computational efficiency.

Method used

The PADD and PDBL tasks of the elliptic curve additive group are merged into a target acceleration task, which is then broken down into multiple subtasks. These subtasks are then processed in parallel by the computing core module in the hardware acceleration system, utilizing various hardware processing units to perform different types of computational acceleration.

Benefits of technology

It improves the efficiency and processing power of elliptic curve addition group operations, optimizes hardware resource utilization, reduces computational latency, and enhances the system's flexibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118819640B_ABST
    Figure CN118819640B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of elliptic curve addition group hardware acceleration method and related device, belong to hardware acceleration field.The method comprises: the PADD task and PDBL task to be handled are merged into target acceleration task;The target acceleration task is split into multiple subtasks;The multiple subtasks are respectively used to execute the hardware acceleration of different types of task objects in the target acceleration task;Multiple subtasks are processed by the computing core module in the hardware acceleration system to realize the hardware acceleration of multiple elliptic curve addition group operations;The computing core module includes multiple hardware processing units that can be scheduled in parallel, and each hardware processing unit is used to execute the operation acceleration of the corresponding task type.Depend on the hardware acceleration of computing core module, realize the task optimization of elliptic curve addition group, improve the utilization of hardware resources, improve the efficiency and processing capacity of elliptic curve addition group operation, with the flexibility and scalability of hardware acceleration system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of hardware acceleration, and in particular to a hardware acceleration method for elliptic curve additive groups and related devices. BACKGROUND

[0002] An elliptic curve additive group is a mathematical structure used to describe the addition of points on an elliptic curve. This structure has important applications in modern cryptography, especially in the fields of public key encryption, digital signatures, etc. The mathematical foundation and structure of the elliptic curve additive group make it have advantages in security, efficiency and resource use, and it is an important encryption technology in modern cryptography.

[0003] In related technologies, the hardware module of the elliptic curve additive group is mainly for 32-bit or 64-bit prime fields. These bit widths are difficult to meet the growing demand for data encryption. For example, the elliptic curves used by modern encryption algorithms (such as BLS12-381, BLS12-377) need to handle bit widths far beyond the processing capacity of these existing technologies, limiting the practical application of such hardware modules in processing high-security encryption tasks. It can be understood that the security provided by a smaller bit width (such as 32-bit or 64-bit) is not enough to cope with modern computing power and attack techniques. Encryption strength is usually proportional to bit width, and the larger the bit width, the more computing power is required to break it. Modern encryption standards usually use larger bit widths, such as 256 bits, 384 bits or higher, to ensure higher security.

[0004] However, in related technologies, elliptic curves with large bit widths involve more complex mathematical operations, such as addition, multiplication, and modulus of large integers, and simply expanding the processing capacity of bit width will consume a large amount of computing resources, making it difficult for many devices to efficiently run these algorithms. This is because directly expanding the bit width means more storage and computing resources, and many devices do not have enough hardware resources, requiring larger memory and cache to store this data, resulting in increased burden on the storage system and computing resources.

[0005] Therefore, there is an urgent need for a technical solution to solve at least one of the above technical problems. SUMMARY

[0006] The main purpose of the embodiments of the present application is to provide a hardware acceleration method for elliptic curve additive groups and related devices, aiming to solve the problem that the bit width in related technologies is difficult to meet the growing demand for data processing, the burden on the storage system and computing resources is too heavy, and the efficiency of elliptic curve additive group operations is affected.

[0007] In a first aspect, the embodiments of the present application provide a hardware acceleration method applied to a hardware acceleration system, the method comprising:

[0008] merge a PADD task and a PDBL task to be processed into a target acceleration task;

[0009] split the target acceleration task into a plurality of sub-tasks; the plurality of sub-tasks are respectively used for performing hardware acceleration on different types of task objects in the target acceleration task;

[0010] process the plurality of sub-tasks through a computing core module in the hardware acceleration system to achieve hardware acceleration on a plurality of elliptic curve additive group operations; the computing core module includes a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is used for performing operation acceleration on a corresponding task type.

[0011] In a second aspect, an embodiment of the present application provides a hardware acceleration device applied to a hardware acceleration system, and the device includes:

[0012] a merging module configured to merge a PADD task and a PDBL task to be processed into a target acceleration task;

[0013] a splitting module configured to split the target acceleration task into a plurality of sub-tasks; the plurality of sub-tasks are respectively used for performing hardware acceleration on different types of task objects in the target acceleration task;

[0014] an acceleration module configured to process the plurality of sub-tasks through a computing core module in the hardware acceleration system to achieve hardware acceleration on a plurality of elliptic curve additive group operations; the computing core module includes a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is used for performing operation acceleration on a corresponding task type.

[0015] In a third aspect, an embodiment of the present application further provides a terminal device including a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus used for realizing connection communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of any one of the hardware acceleration methods provided in the specification of the present application are realized.

[0016] In a fourth aspect, an embodiment of the present application further provides a storage medium for computer readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the steps of any one of the hardware acceleration methods provided in the specification of the present application.

[0017] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor coupled with a transceiver, and is configured to implement the hardware acceleration method of the elliptic curve additive group provided in the first aspect of the present application. In a possible design, the chip can also be a special hardware structure for implementing the hardware acceleration method of the elliptic curve additive group provided in the first aspect above, for example, processing of a neural network model can be implemented by a special neural network processor or a graphic processor.

[0018] In a sixth aspect, an embodiment of the present application provides a chip system, which includes a processor configured to implement the functions involved in the first aspect above, for example, generating or processing the information involved in the hardware acceleration method of the elliptic curve additive group provided in the first aspect.

[0019] In a possible design, the chip system further includes a memory connected to the processor through a circuit structure. The memory is configured to store program instructions and data necessary for the terminal. The chip system can be composed of a chip, or can include the chip and other discrete devices. Further optionally, the chip further includes a communication interface connected to the processor. The communication interface is configured to receive data and / or information to be processed, and the processor obtains the data and / or information from the communication interface, processes the data and / or information, and outputs the processing result through the communication interface. The communication interface can be an input / output interface.

[0020] The embodiments of the present application provide a hardware acceleration method of an elliptic curve additive group and related apparatuses, which combines a PADD task and a PDBL task to be processed into a target acceleration task, splits the target acceleration task into a plurality of subtasks, and uses the plurality of subtasks to perform hardware acceleration on different types of task objects in the target acceleration task. The plurality of subtasks are processed by a computing core module in the hardware acceleration system to implement hardware acceleration on a plurality of elliptic curve additive group operations. The computing core module includes a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is configured to perform operation acceleration on a corresponding task type. The scheme can rely on hardware acceleration of the computing core module to optimize the elliptic curve additive group task, improve the utilization of hardware resources, improve the efficiency and processing capacity of the elliptic curve additive group operation, and have flexibility and scalability of the hardware acceleration system. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0022] Figure 1 A flow chart of a hardware acceleration method of an elliptic curve additive group provided by an embodiment of the present application is shown in FIG. 1.

[0023] Figure 2 A principle diagram of a hardware acceleration method of an elliptic curve additive group provided by an embodiment of the present application is shown in FIG. 2.

[0024] Figure 3 A structure diagram of a core computing module provided for implementing the embodiment is shown in FIG. 3.

[0025] Figure 4 A principle diagram of a core computing module provided for implementing the embodiment is shown in FIG. 4.

[0026] Figure 5 A module structure diagram of a hardware acceleration device of an elliptic curve additive group provided by an embodiment of the present application is shown in FIG. 5.

[0027] Figure 6 A structure diagram of a terminal device provided by an embodiment of the present application is shown in FIG. 6. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0029] The flow chart shown in the drawings is only an example and does not necessarily include all the contents and operations / steps, nor does it have to be executed in the described order. For example, some operations / steps can be further decomposed, combined or partially merged, so the actual execution order can be changed according to the actual situation.

[0030] It should be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms “a”, “an” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0031] An elliptic curve additive group is a mathematical structure used to describe the addition operation of points on an elliptic curve. This structure has important applications in modern cryptography, especially in the fields of public-key encryption, digital signatures, etc. Because the addition operation on an elliptic curve is relatively simple, but its inverse operation (i.e. doubling of points) is computationally complex, it is possible to design a computationally difficult discrete logarithm problem, which is very important in encryption protocols. For example, elliptic curve cryptography (ECC) is based on the mathematical properties of this additive group to implement encryption and signature algorithms. The mathematical foundation and structure of the elliptic curve additive group make it an important encryption technology in modern cryptography, with advantages in security, efficiency and resource usage.

[0032] In the prior art, the hardware modules of the elliptic curve additive group are mainly designed for 32-bit or 64-bit prime fields. These bit widths are difficult to meet the growing demand for data encryption. For example, the elliptic curves used in modern encryption algorithms (such as BLS12-381, BLS12-377) require a bit width that far exceeds the processing capacity of these existing technologies, limiting the practical application of such hardware modules in processing high-security encryption tasks. In addition, smaller bit widths (such as 32-bit or 64-bit) provide insufficient security to cope with modern computing power and attack techniques. Encryption strength is usually proportional to bit width, and the larger the bit width, the more computation is required to break it. Modern encryption standards usually use larger bit widths, such as 256 bits, 384 bits or higher, to ensure higher security.

[0033] However, in the prior art, elliptic curves with large bit widths involve more complex mathematical operations, such as addition, multiplication, and modulo of large integers, and simply expanding the processing capacity of the bit width will consume a large amount of computing resources, making it difficult for many devices to efficiently run these algorithms. This is because directly expanding the bit width means more storage and computing resources, and many devices do not have enough hardware resources, requiring larger memory and cache to store this data, resulting in increased burden on the storage system and computing resources.

[0034] Therefore, there is an urgent need for a technical solution to solve at least one of the above technical problems.

[0035] To solve the above technical problems, the embodiments of the present application provide a hardware acceleration method for an elliptic curve additive group and related devices. The hardware acceleration method can be applied in electronic components or devices, which can be terminal devices such as tablets, laptops, desktop computers, personal digital assistants, and wearable devices. The electronic components can be chips or electronic circuit units mounted in the terminal devices. The terminal device can be a server or a server cluster.

[0036] Some embodiments of the present application will be described in detail with reference to the drawings. The following embodiments and features of the embodiments described below can be combined with each other in the case of no conflict.

[0037] Please refer to Figure 1 , Figure 1 A flowchart of a hardware acceleration method provided by an embodiment of the present application.

[0038] As Figure 1 shown, the hardware acceleration method includes steps S101 to S103.

[0039] Step S101, merging a PADD task and a PDBL task to be processed into a target acceleration task.

[0040] Step S102, splitting the target acceleration task into a plurality of subtasks.

[0041] In the embodiment of the present application, the plurality of subtasks are respectively used to perform hardware acceleration on different types of task objects in the target acceleration task.

[0042] Step S103, processing the plurality of subtasks by a computing core module in the hardware acceleration system to realize hardware acceleration on a plurality of elliptic curve addition group operations.

[0043] In the embodiment of the present application, Point Addition (PADD) and Point Doubling (PDBL) are two basic operation operations contained in the elliptic curve addition group. The two tasks will be introduced in detail below. Point addition refers to the addition operation of two points on an elliptic curve, and the result is a new point on the curve. Point doubling refers to the addition of a point on an elliptic curve with itself. Doubling operation is usually taken as a special case of addition operation.

[0044] In the embodiment of the present application, hardware acceleration is used to improve the efficiency of elliptic curve addition group operation. Point Addition (PADD) task and Point Doubling (PDBL) task are two basic operations in elliptic curve addition group, which need to be repeatedly executed in encryption, decryption, key exchange and other operations. Since PADD and PDBL are essentially point operation operations, they can be processed uniformly.

[0045] In step S101, a plurality of PADD tasks (point addition) and PDBL tasks (point multiplication) to be processed are combined together and merged into a unified target acceleration task. By merging, the dispersed operation management overhead can be reduced and the subsequent parallel computing is prepared. The task scheduling is optimized, and similar types of tasks are merged into one whole to speed up the task processing speed and reduce the resource consumption of context switching. It can be seen that, in step S101, by merging the PADD tasks and the PDBL tasks, the task scheduling is optimized and the processing efficiency is improved.

[0046] In the embodiments of the present application, the acceleration of elliptic curve operation depends on the parallel processing of tasks. Exemplarily, the target acceleration task includes different types of calculations, such as point addition, point multiplication, and modulus operation. In step S102, the merged target acceleration task is split into a plurality of small subtasks. These subtasks can be split based on different calculation properties, for example, a subtask for processing point addition; a subtask for processing point multiplication; and a mathematical subtask related to modulus operation and inverse element calculation. In step S102, by splitting the complex calculation task into a plurality of independent subtasks, parallel processing can be performed to improve the operation efficiency. In addition, this decomposition strategy can more flexibly adapt to the characteristics of the hardware acceleration unit, maximizing resource utilization. In step S102, the merged target task is split into a plurality of subtasks, which can be processed in parallel in the hardware acceleration system, further improving the overall computing capability.

[0047] In the embodiments of the present application, the hardware acceleration system includes a plurality of processing cores, which can process tasks in parallel to achieve efficient calculation. In step S103, the computing core module of the hardware acceleration system is responsible for parallel processing of a plurality of subtasks. The computing core module can include a plurality of hardware processing units, for example: an operation unit dedicated to elliptic curve operation; a unit dedicated to modulus operation; and a general-purpose processing unit for other necessary calculations. Each hardware processing unit executes the corresponding subtask according to its expertise. For example, one unit can be responsible for point addition on the elliptic curve, and another unit can focus on multiplication operation. By parallel scheduling of these hardware units, a plurality of subtasks can be processed simultaneously, significantly improving the calculation speed. By fully utilizing the parallel computing capability in the hardware acceleration system, a large number of complex elliptic curve addition group operations can be completed in a short time, improving the performance of the encryption system. In addition, by reasonable task scheduling and distribution, hardware resources can be fully utilized, resource waste can be reduced, and system energy consumption can be reduced. In step S103, the parallel processing core of the hardware acceleration system efficiently executes a plurality of subtasks, fully utilizes hardware resources, reduces calculation delay, and improves encryption operation performance.

[0048] The above steps S101-S103, first, by combining the PADD (elliptic curve addition) task and the PDBL (elliptic curve multiplication) task into a target acceleration task, and then splitting it into multiple sub-tasks for parallel processing, the overall computing efficiency of elliptic curve addition group operation can be significantly improved. Each sub-task can be executed simultaneously in different hardware processing units, thereby shortening the overall computing time. Further, the various hardware processing units in the computing core module can be scheduled in parallel to adapt to the processing needs of different types of tasks. This can make more efficient use of hardware resources and avoid the phenomenon of idle or excessive load of processing units due to different types of tasks. Finally, the task is divided into multiple sub-tasks and distributed to different processing units, making the hardware acceleration system more flexible and scalable. In the future, if more types of elliptic curve operations need to be supported, additional processing units or adjustment of task allocation strategies can be used. Since the computing core module of the hardware acceleration system can process multiple sub-tasks in parallel, this parallel processing method can significantly reduce the computing delay, especially when processing large-scale encryption tasks, it can obtain faster response time. Through hardware acceleration, not only the processing speed can be improved, but also the overall processing capacity of the system when facing complex encryption operations can be improved, so that it can cope with high load and high density computing requirements. In summary, the combination of these steps helps to achieve efficient hardware acceleration in elliptic curve addition groups, especially suitable for application scenarios that require a large amount of computation, such as blockchain, digital signature, encrypted communication, etc.

[0049] In the embodiments of the present application, the target type matches the type of the task pipeline in the computing core module. Specifically, each task pipeline in the computing core module is optimized for a specific type of operation, such as modular multiplication, modular addition, modular subtraction, etc. These pipelines are usually designed according to the specific characteristics of mathematical operations and contain specialized hardware resources (such as multipliers, adders, etc.) optimized for such operations. When the target type matches the pipeline type, it can be ensured that the task runs on the most suitable hardware resources, maximizing the use of hardware resources. For example, a modular multiplication task is assigned to a modular multiplication pipeline, and the hardware structure of this pipeline can perform modular multiplication operations in the most efficient way, avoiding unnecessary resource waste or inefficiency. Each type of task has its specific computing mode and resource requirements. Assigning tasks to matching task pipelines can maximize the computing efficiency of the hardware accelerator when executing these tasks. For example, modular multiplication operations are usually more complex than modular addition and subtraction, so modular multiplication task pipelines are designed to handle high-complexity operations. By precisely matching tasks to the corresponding pipeline type, tasks can be completed in the shortest time, thereby improving the overall computing efficiency of the system. In a hardware acceleration system, different task types may cause additional computational delays if assigned to inappropriate pipelines. For example, if a modular multiplication task is assigned to a modular addition pipeline, the pipeline may need to perform additional processing steps (such as data format conversion, additional operations, etc.) to complete the task, resulting in unnecessary delays. By ensuring that the target task type matches the pipeline type, these additional processing steps can be avoided, thereby reducing overall computational delay.

[0050] In practical applications, new task types or hardware modules can be added as needed in a hardware acceleration system, as long as their task pipeline types match the new tasks. This modular design allows the system to flexibly handle different operation requirements and is easier to extend in the future. For example, if a new type of task needs to be processed, only the corresponding task pipeline module needs to be added, without the need for major modifications to the entire system architecture.

[0051] In this way, matching the target type with the type of the task pipeline in the computing core module optimizes hardware resource utilization, improves computing efficiency, reduces computing delay, and enhances the flexibility and scalability of the system. In addition, it also reduces power consumption and heat generation. This matching design has significant advantages for high-performance computing tasks, especially for encryption algorithms (such as ECC) involving complex mathematical operations.

[0052] In the embodiments of the present application, further optionally, the computing core module at least includes a Montgomery multiplication task pipeline for processing Montgomery multiplication operation objects, and the Montgomery multiplication task pipeline at least includes a scheduling module and a Montgomery multiplication computing module. Specifically, the Montgomery multiplication computing module is configured to process at least two groups of input data in the same clock unit.

[0053] It can be understood that in the computing core module, the Montgomery multiplication task pipeline is a hardware pipeline specially designed for processing Montgomery multiplication operations. In order to further improve the computing efficiency, the Montgomery multiplication computing module can process multiple groups of input data in the same clock unit. By processing two groups of input data in the same clock cycle, the Montgomery multiplication computing module effectively reduces the task processing time by half. Without such parallel processing capability, it would take two clock cycles to complete the two tasks. Through parallel processing, not only the computing speed is improved, but also the utilization of hardware is improved. By implementing the capability of processing multiple groups of data in the same clock unit in the Montgomery multiplication computing module, the hardware acceleration system can significantly improve the operation efficiency without increasing the clock frequency. Such design enables complex encryption operations to be completed more quickly, thereby improving the performance of the entire system.

[0054] Exemplarily, the Montgomery multiplication task pipeline is a special hardware structure in the hardware accelerator for efficiently performing Montgomery multiplication operations. Montgomery multiplication, i.e., modular multiplication, can be applied to cryptography algorithms (such as RSA, ECC) and other operations involving large integers. Through pipeline technology and hardware optimization design, the Montgomery multiplication task pipeline can greatly improve the speed and efficiency of Montgomery multiplication operations.

[0055] The specific principle is that pipeline technology is a method of dividing a computing process into multiple stages, each stage being processed in parallel in different hardware units, so as to achieve efficient computation. Each stage completes a part of the computing process, and all stages work simultaneously to achieve parallel processing of tasks and improve the throughput of overall operations. In the embodiments of the present application, in the Montgomery multiplication task pipeline, the Montgomery multiplication operation is divided into multiple steps, each step being processed by a special hardware unit. When the first data starts operation, the subsequent data can enter the initial stage of the pipeline, thereby achieving the effect of parallel processing. In hardware implementation, these steps can be further divided into smaller operations, such as partial product generation, partial accumulation, carry processing, modulus reduction, etc. The Montgomery multiplication task pipeline includes the following modules, for example:

[0056] Input module: receives operands that need to be subjected to Montgomery multiplication operations.

[0057] Multiplier: performs large integer multiplication operations, which is the core part of Montgomery multiplication operations.

[0058] Modulus processor: performs modulus operations to ensure that the multiplication result is within a given modulus range.

[0059] Scheduling module: responsible for assigning input data to appropriate pipeline stages and coordinating the work of each stage.

[0060] Register file: used to store intermediate results and ensure correct data transfer between pipeline stages.

[0061] In the embodiments of the present application, the key advantage of the Montgomery multiplication task pipeline lies in its parallel processing capability. By breaking down the Montgomery multiplication operation into multiple stages, each stage can handle different operands within the same clock cycle. This parallel processing capability greatly improves the throughput of the pipeline. The depth of the pipeline determines how many stages of operations can be processed simultaneously. A deeper pipeline can handle more complex Montgomery multiplication operations, but may increase the delay and hardware complexity. The Montgomery multiplication task pipeline employs various optimization techniques to further improve performance. In the present application, grouping technology is used to divide large integers into multiple smaller parts, calculate them separately and then combine them to improve hardware utilization and calculation speed. In addition, parallel design is used, with multiple pipeline modules working simultaneously to handle different Montgomery multiplication tasks, improving throughput. Furthermore, in certain conditions, unnecessary pipeline stages are turned off in the embodiments of the present application to save power consumption.

[0062] In fact, the Montgomery multiplication task pipeline is widely used in encryption algorithms, such as the Montgomery multiplication operation in the RSA algorithm and the point multiplication operation in the ECC algorithm. These algorithms require a large number of Montgomery multiplication operations, and the efficient processing of the pipeline allows these algorithms to be completed within a reasonable time, thus meeting the needs of practical applications.

[0063] In this way, the Montgomery multiplication task pipeline greatly improves the efficiency of Montgomery multiplication operations by breaking down complex Montgomery multiplication operations into multiple parallelizable stages and optimizing the design of dedicated hardware modules. It plays an important role in high-performance computing and cryptography applications, and through the use of pipeline technology, the system's calculation speed and hardware utilization have been significantly improved.

[0064] As an optional embodiment, the computing core module includes at least two task pipelines, each of which includes at least one scheduling module.

[0065] Specifically, further, each scheduling module includes at least one counter for indicating the position of the current clock cycle in the current clock unit. Each scheduling module also includes a register for storing the operation result or intermediate processing result of any of the modulo multiplication calculation module, the modulo addition calculation module, and the modulo subtraction calculation module. Each scheduling module also includes a selector for selecting the corresponding register to store the data participating in the operation processing task currently executed by any of the modulo multiplication calculation module, the modulo addition calculation module, and the modulo subtraction calculation module, or selecting the corresponding register to output the data to the corresponding data output port according to the clock cycle position indicated by the counter.

[0066] Further optionally, the calculation core module further includes at least a modulo subtraction task pipeline for processing the modulo addition operation object and / or the modulo subtraction operation object. The modulo subtraction task pipeline includes at least a scheduling module, a modulo addition calculation module, and a modulo subtraction calculation module.

[0067] In the calculation core module, the modulo addition and subtraction task pipeline is a dedicated hardware pipeline for processing modulo addition and subtraction operations. The modulo addition and subtraction task pipeline can parallelly process multiple modulo addition and subtraction operations in the same clock cycle, thereby greatly improving the calculation speed. If these tasks are sequentially executed in different pipelines, more clock cycles will be required. Through parallel processing, complex operation tasks can be completed in a shorter time. The modulo addition and subtraction task pipeline, through the cooperative work of the scheduling module, the modulo addition calculation module, and the modulo subtraction calculation module, parallelly processes multiple modulo addition and subtraction operations in the same clock cycle, greatly improving the operation efficiency. In fact, this structure is particularly suitable for high-performance computing applications, such as elliptic curve cryptography, encryption and decryption operations, and blockchain technology, etc.

[0068] For example, multiple input data can be parallelly processed by the calculation core module, and the structure of the calculation core module is as shown in Figure 3 For example, in Figure 3 , input data 1 and input data 2 can be parallelly configured to the calculation core module, and each input data is allocated to the corresponding task pipeline for processing to obtain output data 1.

[0069] In Figure 3 , the calculation core module includes at least a modulo multiplication task pipeline and a modulo subtraction task pipeline. The modulo multiplication task pipeline includes at least one scheduling module and one modulo multiplication calculation module. The modulo subtraction task pipeline includes at least one scheduling module, one modulo addition calculation module, and one modulo subtraction calculation module. The above two task pipelines are configured and scheduled by the control module.

[0070] To efficiently utilize the computing resources of the core computing module, the system time-division multiplexes the computing resources of the core computing module and introduces the concept of clock unit.

[0071] As an optional embodiment, the processing duration of the target acceleration task is determined by the clock unit length and / or the number of computing core modules that the hardware acceleration system can schedule.

[0072] It can be understood that in a hardware acceleration system, the processing duration of a target acceleration task can be affected by multiple factors, among which the clock unit length and the number of computing core modules are key factors. The clock unit length refers to the period length or clock frequency of the clock signal in the system, which determines how many operation cycles the hardware can perform per second. The shorter the clock cycle (the higher the frequency), the more operations can be performed in each clock cycle, thereby speeding up the processing. Conversely, a longer clock cycle will result in a decrease in the number of operations per second and an increase in processing duration. A higher clock frequency can improve the processing speed of the computing core, but it will also increase power consumption and heat. Optimizing the clock frequency is one of the keys to improving system performance.

[0073] For example, in a hardware acceleration system, the adjustment of the clock frequency directly affects the execution speed of the computing task. For example, when performing complex modulo addition, subtraction, or multiplication operations, the system needs to run at a higher clock frequency to reduce the total duration of the task processing.

[0074] The number of computing core modules refers to the number of computing units or cores available in the hardware acceleration system. Each core module can independently perform a computing task, process different data, or perform the same operation in parallel. Further optionally, increasing the number of computing core modules can process more computing tasks in parallel, thereby shortening the overall processing duration. For example, for a modulo addition operation task, the task can be divided into multiple cores to execute simultaneously, improving the overall operation speed. Further optionally, dividing the task into multiple cores for load balancing can reduce the amount of work each core processes, thereby improving processing efficiency and speed. For example, in a hardware acceleration system, increasing the number of computing core modules can significantly improve processing capacity. For example, if there are multiple cores in the system that can simultaneously perform modulo addition, subtraction, or multiplication operations, large-scale operation tasks can be allocated to these cores for parallel processing, significantly reducing the total processing duration of the task.

[0075] In a hardware acceleration system, the clock unit length and the number of computing core modules directly determine the processing duration of the target acceleration task. Clock frequency affects the number of operations per second, while the number of computing cores affects parallel processing capability. By optimizing these factors, the overall performance of the system can be improved, and the processing duration of the computing task can be shortened.

[0076] As an optional embodiment, the clock unit length is at least greater than the total number of clock cycles for the hardware processing unit to process a single type of target task object. Exemplarily, each clock unit contains at least 14 clock cycles.

[0077] In this optional embodiment, the concept of clock unit and how to use the clock unit to optimize the modular multiplication and modular addition / subtraction operations are introduced in detail through the execution process of a specific algorithm.

[0078] In this embodiment, the system needs to perform complex elliptic curve operations. The specific algorithm requires 14 modular multiplications and at most 14 modular additions or subtractions. The clock unit is a basic computing unit composed of 14 consecutive clock cycles. Each cycle in each clock unit corresponds to a specific modular multiplication operation step. The clock unit length is composed of 14 clock cycles, which ensures that all steps of a modular multiplication module can be completed within a clock unit. The clock unit starts at the start time of each clock unit (T1 time), the system receives a new set of data and starts calculating the first stage of modular multiplication.

[0079] Taking the modular multiplication operation as an example, the modular multiplication operation can be divided into multiple stages, and each stage is executed in different parts of the clock unit:

[0080] First stage: Calculate four modular multiplications at T1-T4 of the first clock unit.

[0081] Second stage: Calculate four modular multiplications of the second stage at T5-T8 of the K+1th clock unit.

[0082] Third stage and subsequent stages: In succession, until all 14 modular multiplication calculation steps are completed.

[0083] By dividing the computing resources of the modular multiplication module into 14 parts in time, each clock unit can perform a part of the modular multiplication operation, and the whole modular multiplication task can be completed through multiple iterations of the clock unit.

[0084] In order to further optimize the operation efficiency, the system introduces the strategy of time division multiplexing and K value adjustment. The computing resources of the modular multiplication module are divided into 14 parts in time, and these computing resources are used by each clock unit in time sharing. The system can process multiple modular multiplication tasks under the same hardware resources. K is a positive integer, which represents different stages of modular multiplication operations in different clock units. By adjusting the value of K, the system can flexibly change the total number of cycles required for modular multiplication, thereby adapting to different operation frequency requirements.

[0085] Reference Figure 4As shown, assuming that the Montgomery multiplication requires 14K clock cycles from input to output, the specific execution process can be implemented as follows: at time T1 (the start of the first clock unit), the system receives new data, and in clock unit 1, the Montgomery multiplication operation of the first stage is started (4 Montgomery multiplications of the first stage are completed at times T1-T4). In clock unit 2, at time T5 (the K+1th clock unit), the Montgomery multiplication operation of the second stage is started (4 Montgomery multiplications of the second stage are completed at times T5-T8). In clock unit 3, at time T9, the Montgomery multiplication operation of the third stage is started (4 Montgomery multiplications of the third stage are completed at times T9-T11). In clock unit 4, at time T12, the Montgomery multiplication operation of the fourth stage is started (4 Montgomery multiplications of the fourth stage are completed at times T12-T14). In a similar manner, the remaining Montgomery multiplication steps are continued to be completed in subsequent clock units until the 14th Montgomery multiplication operation is completed.

[0086] By this method, the system can simultaneously process multiple Montgomery multiplication tasks on the same calculation module, and improve the calculation efficiency through reasonable clock unit division and time division multiplexing.

[0087] In practical applications, by adjusting the value of K, the system can control the total number of cycles required for Montgomery multiplication operation. This allows the system to adjust the calculation frequency in different operation scenarios to optimize performance and resource utilization. High-frequency operation can be that a smaller K value corresponds to a smaller total cycle number, which is suitable for scenarios that require fast completion of Montgomery multiplication operation. Low-frequency operation can be that a larger K value corresponds to a larger total cycle number, which is suitable for scenarios that do not require tight calculation time but need to save power consumption.

[0088] By time division multiplexing the calculation resources of the core calculation module and introducing the concept of clock unit, the system can efficiently process complex elliptic curve operations. By adjusting the value of K, the system can flexibly change the number of cycles required for Montgomery multiplication to adapt to different operation frequencies and performance requirements. This method effectively optimizes the resource utilization of the hardware acceleration system and reduces the processing time.

[0089] Further optionally, before the PADD task and the PDBL task to be processed are merged into the target acceleration task in S101, an initial task can also be obtained; the initial task at least includes an operation task of an elliptic curve additive group. Further, the initial task is subjected to task type identification to extract the PADD task and the PDBL task to be processed.

[0090] For example, by obtaining initial tasks and performing task type identification, the system can extract and optimize PADD and PDBL tasks in an elliptic curve additive group. By merging these tasks into a target acceleration task, the calculation efficiency of the hardware accelerator can be improved, and faster elliptic curve operations can be achieved. This method can be applied in scenarios such as ECC that require a large number of operations to effectively shorten the processing time of encryption, decryption, and signature operations.

[0091] As an optional embodiment, in S101, the to-be-processed PADD task and the PDBL task are merged into a target acceleration task, as shown in Figure 2 , which can be implemented as:

[0092] S201, identifying a target task object to be processed from the to-be-processed PADD task and the PDBL task;

[0093] S202, generating configuration information of the target task object to obtain the target acceleration task.

[0094] In the embodiments of the present application, the target task object is one of a modulo addition operation object, a modulo subtraction operation object, and a modulo multiplication operation object. In an elliptic curve additive group, modulo addition, modulo subtraction, and modulo multiplication operations are key operation operations, each having specific characteristics and applications. In an elliptic curve, modulo addition operation refers to the addition operation of two points on the curve. The addition operation is the basis of elliptic curve cryptography, involving adding two points to obtain a new point. Modulo addition operation is a closed operation of an elliptic curve group, that is, the sum of two points is still within the curve group. It is a basic operation in an elliptic curve group and is also a building block for other complex operations. For example, it is used for calculating public keys, signature verification, and other operations. In elliptic curve cryptography (ECC), modulo addition operation is the core of calculating key values and performing encryption / decryption operations.

[0095] In an elliptic curve, modulo subtraction operation can be regarded as the inverse operation of addition operation, that is, finding the negative point of a point and performing addition operation. Modulo subtraction operation in an elliptic curve is equivalent to a combination operation of point addition, which is completed by using addition and negative point calculation. For example, it is used for calculating the distance between points, inverse elements, and other operations. In a signature algorithm, modulo subtraction operation is used to process the operation relationship between a private key and a public key.

[0096] In an elliptic curve, modulo multiplication operation refers to scalar multiplication of a point on an elliptic curve, that is, multiplying a point on the curve by an integer (scalar). Modulo multiplication operation is scalar multiplication on an elliptic curve, which is implemented by repeated addition (point addition). Scalar multiplication is the most important operation in elliptic curve encryption algorithms, which has high calculation complexity (for example, optimized by fast addition algorithm). The result of modulo multiplication operation is a point on an elliptic curve.

[0097] In this embodiment, the configuration information is used to indicate the target type to which the hardware processing unit of the target task object belongs. Specifically, the configuration information plays an important role in the hardware accelerator, instructing how the hardware should process specific computing tasks to ensure efficient task execution. The configuration information can optimize the use of hardware resources and adjust hardware operations according to the specific requirements of the task, thereby improving overall computing performance.

[0098] For example, in a computing task accelerator, configuration information is crucial information used to guide the hardware processing unit on how to handle a specific task. For hardware accelerators involving operations such as modular addition, modular subtraction, and modular multiplication, the configuration information contains detailed descriptions of the task object to ensure that the hardware can execute the corresponding operations correctly and efficiently.

[0099] For example, configuration information is used to instruct hardware processing units how to process parameters or instructions for a specific computational task, including task type, data format and size, operation mode, hardware resource allocation, and control signals. The task type specifies whether the task involves modular addition, modular subtraction, or modular multiplication. The data format and size define the format and size of the input data, such as the data bit width and integer length. The operation mode indicates how the operation is performed, such as whether partial result accumulation or carry handling is required. Hardware resource allocation specifies which hardware units are responsible for executing the specific task, such as adders and analog-to-digital processors. Control signals are used to manage the hardware's operational flow, including start signals, completion signals, etc., to ensure the correct and efficient execution of the task.

[0100] For example, the merged target acceleration tasks can use the same hardware resources for computation, such as... Figure 3 The computing core module shown is capable of processing the input data of two or more tasks in parallel, or it can process the input data of a single task, as is not limited in this embodiment.

[0101] Optionally, in S101, before merging the PADD and PDBL tasks to be processed into the target acceleration task, an initial task can be obtained; the initial task includes at least an elliptic curve addition group operation task. Then, the initial task is subjected to task type identification to extract the PADD and PDBL tasks to be processed.

[0102] For example, by acquiring the initial task and identifying its type, the system can extract and optimize the PADD and PDBL tasks within the elliptic curve addition group. Merging these tasks into a single target acceleration task can improve the computational efficiency of the hardware accelerator, enabling faster elliptic curve operations. This method can be applied in computationally intensive scenarios such as ECC to effectively shorten the processing time for operations like encryption, decryption, and signing.

[0103] Further, it is assumed that the computing core module comprises at least two task pipelines, each of which comprises at least a scheduling module and at least one hardware processing unit. Based on the above assumption, in S102, the target acceleration task is split into multiple subtasks, as shown in Figure 2 It can be implemented as:

[0104] In S203, the target task object is configured to the corresponding task pipeline in the computing core module based on the configuration information.

[0105] As an optional embodiment, it is assumed that the operation time of different types of subtasks is different. Based on the above assumption, before S203, the target task object is configured to the corresponding task pipeline in the computing core module based on the configuration information, the operation processing time of each target task object can be identified; the target task object is encapsulated into a task unit corresponding to different time intervals according to the operation processing time; and each task unit is transmitted to the scheduling module in the computing core module for distribution.

[0106] In this embodiment, it is assumed that the computing core module of the hardware acceleration system comprises at least two task pipelines, each of which contains a scheduling module and at least one hardware processing unit. In order to optimize the execution efficiency of the task, the target acceleration task is split into multiple subtasks, and the subtasks are distributed and scheduled according to the operation processing time. The following will introduce this process through a specific example.

[0107] In an optional example, it is assumed that the computing core module contains two task pipelines, each of which has a scheduling module and multiple hardware processing units. It is assumed that the target task object is composed of multiple subtasks of different types, and the operation time of these subtasks is different. The configuration information is used to describe the operation characteristics of each subtask, including the operation processing time. Before task scheduling, the system first needs to identify the operation processing time corresponding to each target task object (subtask). For example, subtask A is a modulo addition operation, which requires a processing time of 5 clock cycles. Subtask B is a modulo subtraction operation, which requires a processing time of 7 clock cycles. Subtask C is a modulo multiplication operation, which requires a processing time of 14 clock cycles. Then, according to the identified operation processing time, the system encapsulates these subtasks into corresponding task units. These task units are classified according to different time intervals for subsequent scheduling. For example, the short-time task unit is assumed to encapsulate subtask A (5 cycles) and subtask B (7 cycles). The long-time task unit is assumed to encapsulate subtask C (14 cycles).

[0108] Based on the above assumptions, in step S203, the system transmits the encapsulated task units to the scheduling module in the computing core module for distribution. The scheduling module dynamically allocates task units to appropriate task pipelines according to the duration of the task units and the current load of the pipeline. For example, in pipeline 1, shorter duration task units are processed, so subtask A and subtask B are preferentially allocated. In pipeline 2, longer duration task units are processed, so subtask C is allocated.

[0109] Suppose the system needs to process a complex elliptic curve operation task, which is divided into multiple subtasks such as modulo addition, modulo subtraction, and modulo multiplication. Identify the operation processing duration required by each subtask and classify it as a short duration and a long duration task unit. According to the duration interval, subtask A and subtask B are encapsulated as short duration task units, and subtask C is encapsulated as long duration task units. Short duration task units are allocated to pipeline 1 for processing, and the scheduling module of pipeline 1 processes subtask A and subtask B in order. Long duration task units are allocated to pipeline 2 for processing, and pipeline 2 focuses on processing subtask C.

[0110] The advantage of this method is that by identifying the processing duration of the task and allocating it, the system can more evenly utilize the processing capacity of different task pipelines. Different task pipelines can process different types of task units at the same time, reducing the waiting time between tasks and improving overall processing efficiency. The scheduling module can adjust task allocation according to real-time conditions, adapt to different load conditions, and optimize resource utilization. By identifying the operation processing duration of subtasks and allocating task units according to the duration, the hardware acceleration system can significantly optimize the processing flow of tasks. This method ensures that the task pipelines in the computing core module can fully utilize their computing power, reduces the total duration of task processing, and improves the overall performance of the system.

[0111] Further, in S103, multiple subtasks are processed by the computing core module in the hardware acceleration system to achieve hardware acceleration of multiple elliptic curve additive group operations, as shown in Figure 2 which can be implemented as:

[0112] S204, multiple subtasks are input into corresponding task pipelines;

[0113] S205, multiple task pipelines are used to process corresponding subtasks in parallel to obtain the operation results corresponding to multiple subtasks, to achieve time division multiplexing of multiple subtasks on the computing core module.

[0114] In this embodiment, the step S103 processes multiple sub-tasks by the computing core module in the hardware acceleration system, thereby achieving hardware acceleration of multiple elliptic curve additive group operations. To this end, the sub-tasks are input into different task pipelines, and the multiple task pipelines are used to process the sub-tasks in parallel, and the time division multiplexing mechanism is used to improve the computing efficiency. Specifically, the computing core module includes multiple task pipelines, and each pipeline can independently process different sub-tasks. The target task includes multiple sub-tasks, which involve elliptic curve additive group operations such as point addition (PADD) and point doubling (PDBL). The system assigns multiple sub-tasks to different task pipelines in the computing core module for processing.

[0115] For example, it is assumed that the target task includes the following sub-tasks:

[0116] Sub-task 1: Calculate point P + Q (PADD task)

[0117] Sub-task 2: Calculate point 2P (PDBL task)

[0118] Sub-task 3: Calculate point P + R (PADD task)

[0119] Sub-task 4: Calculate point 3P (PDBL task)

[0120] The system assigns these sub-tasks to different task pipelines to maximize the use of computing resources. For example, in task pipeline 1, sub-task 1 and sub-task 3 (PADD tasks) are processed. In task pipeline 2, sub-task 2 and sub-task 4 (PDBL tasks) are processed. In this step, the system processes the sub-tasks in parallel through multiple task pipelines, achieving efficient time division multiplexing.

[0121] Task pipeline 1: Simultaneously processes sub-task 1 and sub-task 3. In pipeline 1, two PADD tasks are executed in parallel, and multiple hardware processing units in the pipeline are used to achieve efficient computation.

[0122] Task pipeline 2: Simultaneously processes sub-task 2 and sub-task 4. In pipeline 2, two PDBL tasks are executed in parallel, and the doubling module in the pipeline is used to achieve fast computation.

[0123] From the above examples, in the pipeline mechanism, each task pipeline executes different subtask operations in different clock cycles. For example, pipeline 1 processes the first part of subtask 1 in the first clock cycle, and the first part of subtask 3 in the next clock cycle, and so on. This time-division multiplexing mechanism allows multiple subtasks to be processed in the same pipeline, thereby improving computational efficiency. Through parallel processing of multiple task pipelines, the system can obtain the operation results of multiple subtasks in a shorter time. For example, pipeline 1 outputs the results of subtask 1 and subtask 3, and pipeline 2 outputs the results of subtask 2 and subtask 4.

[0124] This multi-task pipeline parallel processing and time-division multiplexing method brings significant optimization effect. Through parallel processing and time-division multiplexing, multiple subtasks can be executed simultaneously, significantly shortening the total calculation time. Multiple task pipelines can fully utilize the hardware resources in the calculation core module, avoiding resource idling. The system can dynamically adjust the allocation and scheduling of subtasks according to the nature of the task and the current load, achieving optimal performance. By inputting multiple subtasks into different task pipelines and using parallel processing and time-division multiplexing mechanism, the hardware acceleration system can efficiently process multiple elliptic curve additive group operations. This method not only improves the calculation efficiency, but also optimizes the utilization of system resources, which is an effective way to achieve hardware acceleration in complex computing tasks.

[0125] As an optional embodiment, in S205, multiple task pipelines are used to process corresponding subtasks in parallel to obtain operation results corresponding to multiple subtasks, which can be implemented as follows:

[0126] For each task pipeline, the processing order of each target task object in each clock unit in the corresponding subtask is configured through the scheduling module in the task pipeline; each clock unit includes at least multiple clock cycles; based on the processing order, an enable signal corresponding to each target task object in each subtask is generated; the enable signal is used to indicate the clock cycle for processing each target task object in each subtask; based on the enable signal, each hardware processing unit in the task pipeline is controlled to process the corresponding target task object in the corresponding clock cycle; wherein the type of the target task object matches the type of the task corresponding to the hardware processing unit.

[0127] In this optional embodiment, multiple task pipelines are used to process corresponding subtasks in parallel, and efficient task processing is achieved by configuring the scheduling module in the task pipeline, generating the enable signal, and controlling the hardware processing unit. The following is a specific introduction to this process, and an example is given to help understand.

[0128] Assume each task pipeline includes a scheduling module and multiple hardware processing units. Each hardware processing unit is good at processing a specific type of task. Assume the target task is split into multiple sub-tasks, each of which contains several target task objects. Assume each target task object needs to be processed by a corresponding hardware processing unit within a specific clock cycle.

[0129] For each task pipeline, the processing order of each target task object in each clock unit within each sub-task is configured by the scheduling module in the task pipeline. A clock unit here contains multiple clock cycles. The scheduling module assigns a processing order to each target task object according to the type of the target task object and the required computing resources. For example, in a clock unit, task A is executed in the 1st clock cycle and task B is executed in the 2nd clock cycle. Based on the configured processing order, the scheduling module generates enable signals. These enable signals are used to indicate the clock cycle in which each target task object in each sub-task is processed. The enable signals can be seen as control instructions that tell the hardware processing unit when (specific clock cycle) to process which target task object. For example, enable signal S1 can indicate that target task object C is processed in the 3rd clock cycle. Based on the generated enable signals, the scheduling module controls each hardware processing unit in the task pipeline to process the corresponding target task object in the corresponding clock cycle. The task type of the hardware processing unit must match the type of the target task object. The enable signal activates the corresponding hardware processing unit to process the corresponding target task object in the specified clock cycle. For example, the multiplier hardware unit will process the target task object of multiplication operation in the 4th clock cycle when it receives enable signal S2.

[0130] For example, assume the system needs to process an elliptic curve additive group operation task, which is split into the following sub-tasks:

[0131] Sub-task 1: contains target task object A (addition operation) and target task object B (multiplication operation).

[0132] Sub-task 2: contains target task object C (addition operation) and target task object D (multiplication operation).

[0133] Clock unit 1 of task pipeline 1 is to process target task object A (addition operation) in the 1st clock cycle and target task object B (multiplication operation) in the 2nd clock cycle. Clock unit 1 of task pipeline 2 is to process target task object C (addition operation) in the 1st clock cycle and target task object D (multiplication operation) in the 2nd clock cycle.

[0134] The enable signal S1 is used to indicate that the adder of the task pipeline 1 processes the target task object A in the first clock cycle. The enable signal S2 is used to indicate that the multiplier of the task pipeline 1 processes the target task object B in the second clock cycle. The enable signal S3 is used to indicate that the adder of the task pipeline 2 processes the target task object C in the first clock cycle. The enable signal S4 is used to indicate that the multiplier of the task pipeline 2 processes the target task object D in the second clock cycle.

[0135] The adder of the task pipeline 1 processes the target task object A in the first clock cycle according to the enable signal S1. The multiplier of the task pipeline 1 processes the target task object B in the second clock cycle according to the enable signal S2. The adder of the task pipeline 2 processes the target task object C in the first clock cycle according to the enable signal S3. The multiplier of the task pipeline 2 processes the target task object D in the second clock cycle according to the enable signal S4.

[0136] By configuring the scheduling module in the task pipeline, generating the enable signal, and controlling the hardware processing unit, this embodiment realizes efficient operation of multiple task pipelines in parallel processing sub-tasks. By accurately controlling the processing timing of each target task object, the system can maximize the utilization of hardware resources, reduce processing delay, and improve operation efficiency.

[0137] For example, referring to Table 1 shown below, the input parameters are the parameters a and the prime number field Prime of the curve YY = XXX + aX + b, Point1(X1, Y1, ZZ1, ZZZ1), and Point2(X2, Y2, ZZ2, ZZZ2). Based on this, the output parameters are Point3(X3, Y3, ZZ3, ZZZ3).

[0138] Table 1 Elliptic Curve Additive Group Operation Algorithm

[0139]

[0140] Based on the above elliptic curve additive group operation, it is assumed that That is, The assumption of can be applied to the paired elliptic curve additive group operation, which has higher calculation efficiency. For example, the BN curve is suitable for implementing identity-based encryption and paired applications, and is widely used in blockchain technology and attack prevention. The BLS curve is not suitable for all BLS curves In general, The BLS curve of is commonly used to implement complex encryption applications such as zero-knowledge proof, pairing-based cryptographic protocol, etc. The MNT curve is widely used in cryptography applications involving pairing.

[0141] Based on the above assumptions, first, step one calculates U1, U2, S1 and S2, and then generates a selection signal by comparing the difference P and R between U2 and U1, and S2 and S1. If P and R are both zero, the PDBL operation is performed, otherwise the PADD operation is performed. In order to optimize the utilization of hardware resources, the time-consuming multiplication operation is separated from the time-consuming addition and subtraction operation, and is packaged into a modular multiplication pipeline module and a modular addition and subtraction pipeline module respectively, and the corresponding values are input into the two modules through the scheduling module according to the selection signal. In the combined algorithm, the flow of the two operations is consistent, that is, the selection signal can be used to determine whether to perform PADD or PDBL, and the selection signal can be obtained in step one. Therefore, the same set of hardware can be used to perform all the addition group operations.

[0142] Based on the operation flow shown in Table 1, the multiplication by 2, multiplication by 3, division by 2 and subtraction operations are time-consuming operations, and the remaining multiplication operations are time-consuming operations and have large resource consumption. The time-consuming modular multiplication operation can be separated from the time-consuming modular multiplication operation, and packaged into a modular multiplication pipeline module and a modular addition and subtraction pipeline module respectively, and the scheduling module is responsible for inputting the corresponding values into the two modules according to the selection signal.

[0143] In this way, the entire addition group operation can be completed using the same set of hardware, improving the calculation efficiency and reducing the resource consumption.

[0144] The embodiment of the application provides a hardware acceleration method of an elliptic curve addition group, first, the PADD task and the PDBL task to be processed are combined into a target acceleration task; then, the target acceleration task is split into a plurality of subtasks; the plurality of subtasks are respectively used for executing hardware acceleration on different types of task objects in the target acceleration task; finally, a plurality of subtasks are processed by a calculation core module in the hardware acceleration system, so as to realize hardware acceleration on a plurality of elliptic curve addition group operations; the calculation core module includes a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is used for executing operation acceleration on a corresponding task type. The method can rely on hardware acceleration of the calculation core module to realize task optimization of the elliptic curve addition group, and can effectively solve the problems in the related art that the bit width cannot meet the growing data processing demand, the storage system and the calculation resource are too heavy, and the efficiency of the elliptic curve addition group operation is affected. Further improve the utilization rate of hardware resources, improve the efficiency and processing capacity of the elliptic curve addition group operation, and have the flexibility and scalability of the hardware acceleration system.

[0145] Please refer to Figure 5 , Figure 5Provided in the embodiments of the present application is a hardware acceleration device 200 of an elliptic curve additive group, which is applied to a hardware acceleration system and includes the following modules:

[0146] a merging module configured to merge a PADD task and a PDBL task to be processed into a target acceleration task;

[0147] a splitting module configured to split the target acceleration task into a plurality of subtasks, which are respectively used to perform hardware acceleration on different types of task objects in the target acceleration task;

[0148] an acceleration module configured to process the plurality of subtasks by using a computing core module in the hardware acceleration system to achieve hardware acceleration on a plurality of elliptic curve additive group operations; the computing core module includes a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is used to perform operation acceleration on a corresponding task type.

[0149] In some embodiments, when the merging module merges the PADD task and the PDBL task to be processed into the target acceleration task, the merging module is specifically configured to:

[0150] identify a target task object to be processed from the PADD task and the PDBL task to be processed; the target task object is one of a modulo addition operation object, a modulo subtraction operation object and a modulo multiplication operation object;

[0151] generate configuration information of the target task object to obtain the target acceleration task; the configuration information is used to indicate a target type of a hardware processing unit to which the target task object belongs.

[0152] In some embodiments, the computing core module includes at least two task pipelines, and each task pipeline includes at least a scheduling module and at least one hardware processing unit; when the splitting module splits the target acceleration task into the plurality of subtasks, the splitting module is specifically configured to:

[0153] based on the configuration information, configure the target task object to a corresponding task pipeline in the computing core module;

[0154] when the acceleration module processes the plurality of subtasks by using the computing core module in the hardware acceleration system to achieve hardware acceleration on the plurality of elliptic curve additive group operations, the acceleration module is specifically configured to:

[0155] input the plurality of subtasks into the corresponding task pipelines respectively;

[0156] parallel process the corresponding subtasks by using the plurality of task pipelines to obtain operation results corresponding to the plurality of subtasks, so as to realize time division multiplexing of the plurality of subtasks on the computing core module.

[0157] In some embodiments, the target type matches the type of the task pipeline in the computing core module.

[0158] In some embodiments, the computing core module includes at least a modulo multiplication task pipeline for processing modulo multiplication operation objects, the modulo multiplication task pipeline including at least a scheduling module and a modulo multiplication computing module; the modulo multiplication computing module is configured to process at least two groups of input data in the same clock unit.

[0159] The computing core module further includes at least a modulo subtraction task pipeline for processing modulo addition operation objects and / or modulo subtraction operation objects, the modulo subtraction task pipeline including at least a scheduling module, a modulo addition computing module, and a modulo subtraction computing module.

[0160] In some embodiments, the operation time of different types of subtasks is different; before the splitting module configures the target task objects to the corresponding task pipelines in the computing core module based on the configuration information, the splitting module is further configured to:

[0161] identify the respective operation processing time of the target task objects;

[0162] encapsulate the target task objects into task units corresponding to different time intervals according to the operation processing time;

[0163] transmit each task unit to the scheduling module in the computing core module for distribution.

[0164] In some embodiments, when the acceleration module uses multiple task pipelines to process corresponding subtasks in parallel to obtain operation results of multiple subtasks, the acceleration module is specifically configured to:

[0165] for each task pipeline, configure the processing order of each target task object in each clock unit in the corresponding subtask through the scheduling module in the task pipeline; each clock unit includes at least multiple clock cycles;

[0166] generate an enable signal corresponding to each target task object in each subtask based on the processing order; the enable signal is used to indicate the clock cycle for processing each target task object in each subtask;

[0167] control each hardware processing unit in the task pipeline to process the corresponding target task object in the corresponding clock cycle based on the enable signal; wherein the type of the target task object matches the task type corresponding to the hardware processing unit.

[0168] In some embodiments, the processing duration of the target acceleration task is determined by the clock unit length and / or the number of the computing core modules that can be scheduled by the hardware acceleration system.

[0169] In some embodiments, the clock unit length is at least greater than the total number of clock cycles for the hardware processing unit to process a single type of target task object.

[0170] In some embodiments, each clock unit contains at least 14 clock cycles.

[0171] In some embodiments, at least two task pipelines are included in the computing core module, and each task pipeline includes at least one scheduling module.

[0172] In each scheduling module, at least one counter is included, which is used to indicate the position of the current clock cycle in the current clock unit.

[0173] In each scheduling module, a register is also included, which is used to store the operation result or intermediate processing result of any one of the modulo multiplication calculation module, the modulo addition calculation module and the modulo subtraction calculation module.

[0174] In each scheduling module, a selector is also included, which is used to select the corresponding register storage data to participate in the current execution of the operation processing task of any one of the modulo multiplication calculation module, the modulo addition calculation module and the modulo subtraction calculation module according to the clock cycle position indicated by the counter, or to select the corresponding register storage data to output to the corresponding data output port.

[0175] In some embodiments, before the merging module merges the PADD task and the PDBL task into the target acceleration task, the merging module is also used to obtain an initial task, and the initial task includes at least an operation task of an elliptic curve additive group.

[0176] The initial task is subjected to task type identification to extract the PADD task and the PDBL task to be processed.

[0177] In some embodiments, the hardware acceleration device 200 can be applied to a terminal device, or to a chip, or other microelectronic components.

[0178] It should be noted that, for the convenience and brevity of description, the specific working process of the hardware acceleration device 200 described above can refer to the corresponding process in the foregoing hardware acceleration method embodiments, which will not be described here.

[0179] Please refer to Figure 6 , Figure 6 A structural schematic block diagram of a terminal device provided by an embodiment of the present application is shown in FIG. 2.

[0180] As shown in Figure 6 The terminal device 300 includes a processor 301 and a memory 302, which are connected through a bus 303, such as an I2C (Inter-integrated Circuit) bus.

[0181] Specifically, the processor 301 is configured to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0182] Specifically, the memory 302 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a U disk, or a mobile hard disk, etc.

[0183] Those skilled in the art can understand that Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the embodiment of the present application, and does not constitute a limitation on the terminal device to which the embodiment of the present application is applied. The specific server can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0184] The processor is configured to run a computer program stored in the memory, and implement any one of the hardware acceleration methods provided by the embodiments of the present application when the computer program is executed.

[0185] In an embodiment, the processor is configured to run a computer program stored in the memory and implement the following steps when executing the computer program: merging the PADD task and the PDBL task to be processed into a target acceleration task; splitting the target acceleration task into a plurality of subtasks; the plurality of subtasks are respectively used to perform hardware acceleration on different types of task objects in the target acceleration task; processing the plurality of subtasks by a computing core module in the hardware acceleration system to implement hardware acceleration on a plurality of elliptic curve additive group operations; the computing core module includes a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is used to perform operation acceleration on a corresponding task type. It should be noted that, for the convenience and brevity of description, the specific working process of the terminal device described above can refer to the corresponding process in the foregoing hardware acceleration method embodiments, which will not be described here.

[0186] The embodiment of the present application further provides a storage medium for computer readable storage, the storage medium storing one or more programs, the one or more programs being executable by one or more processors to implement the steps of any one of the hardware acceleration methods provided in the specification of the embodiment of the present application.

[0187] The storage medium can be an internal storage unit of the terminal device, such as a hard disk or a memory of the terminal device. The storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0188] Those skilled in the art can understand that all or some of the steps in the methods disclosed above, the functional modules / units in the systems and devices can be implemented by software, firmware, hardware, or a combination thereof. In hardware embodiments, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer-readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and can include any information delivery media.

[0189] It should be understood that the term "and / or" as used herein refers to any combination of associated listed items, and all possible combinations, and includes these combinations. It should be noted that the terms "comprising", "including", or any other variant thereof, are intended to cover non-exclusive inclusion, so that processes, methods, articles, or systems including a series of elements not only include those elements, but also include other elements not explicitly listed, or other elements inherent to such processes, methods, articles, or systems. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or system including the element.

[0190] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of hardware acceleration of an elliptic curve additive group, the method comprising: The method is applied to a hardware acceleration system, and the method comprises: combining an elliptic curve addition PADD task and an elliptic curve multiplication PDBL task to be processed into a target acceleration task; splitting the target acceleration task into a plurality of subtasks; the plurality of subtasks are respectively used for executing hardware acceleration on different types of task objects in the target acceleration task; processing the plurality of subtasks by a computing core module in the hardware acceleration system to realize hardware acceleration on a plurality of elliptic curve addition group operations; the computing core module comprises a plurality of hardware processing units that can be scheduled in parallel, and each hardware processing unit is used for executing operation acceleration on a corresponding task type; wherein the combining the PADD task and the PDBL task to be processed into the target acceleration task comprises: identifying a target task object to be processed from the PADD task and the PDBL task to be processed; the target task object is one of a modulo addition operation object, a modulo subtraction operation object, and a modulo multiplication operation object; generating configuration information of the target task object to obtain the target acceleration task; the configuration information is used to indicate a target type of a hardware processing unit to which the target task object belongs; the target type matches a type of a task pipeline in the computing core module; wherein the computing core module comprises at least two task pipelines, and each task pipeline comprises at least a scheduling module and at least one hardware processing unit; the splitting the target acceleration task into a plurality of subtasks comprises: based on the configuration information, the target task object is configured to a corresponding task pipeline in the computing core module; the processing the plurality of subtasks by the computing core module in the hardware acceleration system to realize hardware acceleration on a plurality of elliptic curve addition group operations comprises: inputting the plurality of subtasks into corresponding task pipelines respectively; adopting a plurality of task pipelines to process corresponding subtasks in parallel to obtain operation results corresponding to the plurality of subtasks, so as to realize time division multiplexing of the plurality of subtasks on the computing core module; wherein a processing time length of the target acceleration task is determined by a clock unit length and / or a number of computing core modules that can be scheduled by the hardware acceleration system; the clock unit length is at least greater than a total number of clock periods of a hardware processing unit processing a single type of target task object; any clock unit comprises at least 14 clock periods. The calculation core module includes at least two task pipelines, each of which includes at least one scheduling module; each scheduling module includes at least one counter, which is used to indicate the position of the current clock cycle in the current clock unit; each scheduling module further includes a register, which is used to store the operation result or intermediate processing result of any one of the modulo multiplication calculation module, the modulo addition calculation module and the modulo subtraction calculation module; each scheduling module further includes a selector, which is used to select the corresponding register to store the data participating in the operation processing task currently executed by any one of the modulo multiplication calculation module, the modulo addition calculation module and the modulo subtraction calculation module, or select the corresponding register to output the data to the corresponding data output port according to the clock cycle position indicated by the counter.

2. The method of claim 1, wherein, The calculation core module includes at least a modulo multiplication task pipeline for processing modulo multiplication operation objects, which includes at least a scheduling module and a modulo multiplication calculation module; the modulo multiplication calculation module is used to process at least two groups of input data in the same clock unit. The calculation core module further includes at least a modulo subtraction task pipeline for processing modulo addition operation objects and / or modulo subtraction operation objects, which includes at least a scheduling module, a modulo addition calculation module and a modulo subtraction calculation module.

3. The method of claim 1, wherein, The operation time of different types of subtasks is different; Before the target task objects are respectively configured to the corresponding task pipelines in the calculation core module based on the configuration information, the method further includes: identifying the operation processing time corresponding to each of the target task objects; packaging the target task objects into task units corresponding to different time length intervals according to the operation processing time; transmitting each task unit to the scheduling module in the calculation core module for distribution.

4. The method of claim 1, wherein, The multiple task pipelines are used to process corresponding subtasks in parallel to obtain operation results corresponding to the multiple subtasks, including: For each task pipeline, the processing order of each target task object in each clock unit in the corresponding subtask is configured by the scheduling module in the task pipeline; each clock unit includes at least multiple clock cycles; based on the processing order, an enable signal corresponding to each target task object in each subtask is generated; the enable signal is used to indicate the clock cycle in which each target task object in each subtask is processed; based on the enable signal, each hardware processing unit in the task pipeline is controlled to process the corresponding target task object in the corresponding clock cycle; wherein the type of the target task object matches the type of the task corresponding to the hardware processing unit.

5. The method according to any one of claims 1 to 4, characterized in that, Before the PADD task and the PDBL task to be processed are merged into the target acceleration task, the method further includes: obtaining an initial task; the initial task includes at least an operation task of an elliptic curve addition group; performing task type identification on the initial task to extract the PADD task and the PDBL task to be processed.

6. A hardware acceleration device for elliptic curve addition groups, characterized in that, The device is applied to a hardware acceleration system, and the device includes: a merging module configured to merge the PADD task and the PDBL task to be processed into the target acceleration task; The splitting module is configured to split the target acceleration task into a plurality of sub-tasks, and the plurality of sub-tasks are respectively used to perform hardware acceleration on different types of task objects in the target acceleration task. The acceleration module is configured to process the plurality of sub-tasks by a computing core module in the hardware acceleration system to achieve hardware acceleration on a plurality of elliptic curve additive group operations. The merging module is configured to, when merging the PADD task and the PDBL task into the target acceleration task, specifically perform the following operations: identify the target task object to be processed from the PADD task and the PDBL task to be processed, wherein the target task object is one of a modulo addition operation object, a modulo subtraction operation object, and a modulo multiplication operation object; generate configuration information of the target task object to obtain the target acceleration task, wherein the configuration information is used to indicate a target type of a hardware processing unit to which the target task object belongs, and the target type matches a type of a task pipeline in the computing core module; The computing core module includes at least two task pipelines, and each task pipeline includes at least a scheduling module and at least one hardware processing unit. The splitting module is configured to, when splitting the target acceleration task into a plurality of sub-tasks, specifically perform the following operations: based on the configuration information, configure the target task object to a corresponding task pipeline in the computing core module; The acceleration module is configured to process the plurality of sub-tasks by the computing core module in the hardware acceleration system to achieve hardware acceleration on a plurality of elliptic curve additive group operations, including: input the plurality of sub-tasks into corresponding task pipelines; parallel processing of the plurality of sub-tasks by the plurality of task pipelines to obtain operation results corresponding to the plurality of sub-tasks to achieve time division multiplexing of the plurality of sub-tasks on the computing core module; The processing time length of the target acceleration task is determined by a clock unit length and / or a number of schedulable computing core modules of the hardware acceleration system. The computing core module includes at least two task pipelines, and each task pipeline includes at least one scheduling module. Each scheduling module includes at least one counter configured to indicate a position of a current clock cycle in a current clock unit, at least one register configured to store operation results or intermediate processing results of any one of a modulo multiplication calculation module, a modulo addition calculation module, and a modulo subtraction calculation module, and a selector configured to select corresponding register storage data to participate in operation processing tasks of any one of the modulo multiplication calculation module, the modulo addition calculation module, and the modulo subtraction calculation module, or select corresponding register storage data to output to a corresponding data output port according to the clock cycle position indicated by the counter.

7. A computing device, comprising: The terminal device comprises a processor, a memory; The memory is configured to store a computer program; The processor is configured to execute the computer program and implement the hardware acceleration method of the elliptic curve additive group according to any one of claims 1 to 5 when the computer program is executed.

8. A computer storage medium for computer storage, characterized in that The computer storage medium stores at least one program, and the at least one processor is configured to execute the at least one program to implement the steps of the hardware acceleration method of the elliptic curve additive group according to any one of claims 1 to 5.

9. A chip system, characterized by The chip system comprises: A communication interface configured to input and / or output information; A processor configured to execute a computer executable program, so that a device installed with the chip system executes the hardware acceleration method of the elliptic curve additive group according to any one of claims 1 to 5.

10. A computer program product comprising computer instructions, characterized in that, The computer instructions are executed by the processor to implement the hardware acceleration method of the elliptic curve additive group according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Acceleration method and device based on data parallelism and task parallelism and storage medium

    CN118567803A