Hardware acceleration system and method based on multi-core architecture and related equipment

By adopting a multi-core architecture in the field programmable gate array, the multi-scalar multiplication task is split into sub-tasks and processed by multiple acceleration cores, the problem of waste of hardware resources under the single-core architecture is solved, and efficient utilization of hardware resources and computational acceleration is achieved.

CN120336242APending Publication Date: 2025-07-18JIANGXI ZHENGDUZHE NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311696755.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The hardware acceleration equipment based on a single-core architecture in the prior art results in wasting hardware resources on the field programmable gate array, especially when performing elliptic curve multiplication calculations.

Method used

Using a hardware acceleration system based on a multi-core architecture, the multi-scalar multiplication task is split into multiple multi-scalar multiplication subtasks through the CPU, and these subtasks are handled separately by multiple multi-scalar computation acceleration cores in the field programmable gate array of the multi-core architecture.

Benefits of technology

The hardware resource utilization rate of field programmable gate arrays is improved, and the accelerated processing of multi-scalar multiplication tasks is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336242A_ABST
    Figure CN120336242A_ABST
Patent Text Reader

Abstract

The invention discloses a hardware acceleration system and method based on a multi-core architecture and related equipment, and relates to the technical field of computers, the hardware acceleration system based on the multi-core architecture comprises a CPU and a field programmable gate array based on the multi-core architecture, the field programmable gate array based on the multi-core architecture comprises a plurality of multi-scalar calculation acceleration cores; the CPU splits a multi-scalar multiplication task into a plurality of multi-scalar multiplication subtasks, and sends the plurality of multi-scalar multiplication subtasks to the field programmable gate array based on the multi-core architecture; the field programmable gate array based on the multi-core architecture receives the plurality of multi-scalar multiplication subtasks, and respectively processes the corresponding split multi-scalar multiplication subtasks based on the plurality of multi-scalar calculation acceleration cores; according to the invention, the hardware resource utilization rate of the field programmable gate array is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a hardware acceleration system, method, and related devices based on a multi-core architecture. Background Art

[0002] ZKP (Zero Knowledge Proofs) is a class of cryptographic protocols that enables one party to prove the correctness of its statements to another party without disclosing any information related to the statements. Currently, due to its extremely high privacy and simplicity, zero-knowledge proofs are widely used in fields such as blockchain, digital identity verification, and financial transactions.

[0003] The generation of zero-knowledge proofs requires computing a large number of elliptic curve multi-scalar multiplications of large-bitwidth data over a finite field. Existing technologies use FPGAs (Field Programmable Gate Arrays) with a pipeline architecture to accelerate elliptic curve multi-scalar multiplications. However, dozens of large integer modular multiplication hardware circuits are required inside its multi-scalar calculation acceleration core, which is a single-core large computing power architecture that requires a large amount of hardware resources and will cause a large waste of hardware resources on the field programmable gate array. Summary of the Invention

[0004] In view of this, the main purpose of this application is to provide a hardware acceleration system, method, and related devices based on a multi-core architecture, aiming to solve the problem of a large waste of hardware resources on the field programmable gate array caused by hardware acceleration devices based on a single-core architecture.

[0005] To achieve the above object, this application provides a hardware acceleration system based on a multi-core architecture. The hardware acceleration system based on a multi-core architecture includes a CPU and a field programmable gate array based on a multi-core architecture. The field programmable gate array based on a multi-core architecture includes multiple multi-scalar calculation acceleration cores;

[0006] The CPU splits the multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture;

[0007] The field programmable gate array based on a multi-core architecture receives the multiple multi-scalar multiplication subtasks and processes the corresponding split multi-scalar multiplication subtasks respectively based on the multiple multi-scalar calculation acceleration cores.

[0008] Optionally, the calculation unit of the multi-scalar calculation acceleration core only includes a modular multiplication circuit, a modular addition circuit, and a modular subtraction circuit.

[0009] Optionally, the multiple multi-scalar calculation acceleration cores have independent clock domains.

[0010] In addition, to achieve the above object, the present application also provides a hardware acceleration method based on a multi-core architecture, which is applied to a field programmable gate array based on a multi-core architecture. The hardware acceleration method based on a multi-core architecture includes the following steps:

[0011] Receive multiple multi-scalar multiplication subtasks sent by the CPU, where the CPU splits the multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture;

[0012] Based on multiple multi-scalar calculation acceleration cores, respectively process the corresponding split multi-scalar multiplication subtasks to obtain the processing results of the multiple multi-scalar multiplication subtasks, where the field programmable gate array based on a multi-core architecture includes multiple multi-scalar calculation acceleration cores.

[0013] Optionally, the multiple multi-scalar multiplication subtasks are independent of each other, and the multiple multi-scalar calculation acceleration cores process the corresponding split multi-scalar multiplication subtasks synchronously or in parallel.

[0014] Optionally, the hardware acceleration method based on a multi-core architecture further includes the following steps:

[0015] Record the state of the field programmable gate array based on a multi-core architecture, where the state includes a busy state, a pending read state, and an idle state;

[0016] When the state is the busy state, stop receiving the multiple multi-scalar multiplication subtasks sent by the CPU;

[0017] When the state is the pending read state, notify the CPU to read the processing results of the multiple multi-scalar multiplication subtasks for the CPU to merge and output the processing results;

[0018] When the state is the idle state, receive the multiple multi-scalar multiplication subtasks sent by the CPU.

[0019] Optionally, after the step of, based on multiple multi-scalar calculation acceleration cores, respectively process the corresponding split multi-scalar multiplication subtasks to obtain the processing results of the multiple multi-scalar multiplication subtasks, includes:

[0020] Cache the processing results of the multiple multi-scalar multiplication subtasks for the CPU to read the cached processing results;

[0021] When it is detected that all the cached processing results have been read by the CPU, clear the cached processing results.

[0022] In addition, to achieve the above object, the present application further provides a hardware acceleration device based on a multi-core architecture. The hardware acceleration device based on a multi-core architecture includes:

[0023] A task receiving module, configured to receive a plurality of multi-scalar multiplication subtasks sent by a CPU;

[0024] A task processing module, configured to process the corresponding split multi-scalar multiplication subtasks respectively based on a plurality of multi-scalar calculation acceleration cores to obtain processing results of the plurality of multi-scalar multiplication subtasks.

[0025] In addition, to achieve the above object, the present application further provides a hardware acceleration device based on a multi-core architecture. The device includes: a memory, a processor, and a hardware acceleration program based on a multi-core architecture stored on the memory and executable on the processor. The hardware acceleration program based on a multi-core architecture is configured to implement the steps of the hardware acceleration method based on a multi-core architecture as described in any one of the above.

[0026] In addition, to achieve the above object, the present application further provides a storage medium. A hardware acceleration program based on a multi-core architecture is stored on the storage medium. When the hardware acceleration program based on a multi-core architecture is executed by a processor, the steps of the hardware acceleration method based on a multi-core architecture as described in any one of the above are implemented.

[0027] The present application provides a hardware acceleration system, method, and related device based on a multi-core architecture. The hardware acceleration system based on a multi-core architecture includes a CPU and a field programmable gate array based on a multi-core architecture. The field programmable gate array based on a multi-core architecture includes a plurality of multi-scalar calculation acceleration cores. The CPU splits a multi-scalar multiplication task into a plurality of multi-scalar multiplication subtasks and sends the plurality of multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture. The field programmable gate array based on a multi-core architecture receives the plurality of multi-scalar multiplication subtasks and processes the corresponding split multi-scalar multiplication subtasks respectively based on the plurality of multi-scalar calculation acceleration cores. The present application splits the multi-scalar multiplication task and processes the split multi-scalar multiplication subtasks by a plurality of multi-scalar calculation acceleration cores on the field programmable gate array based on a multi-core architecture, so as to achieve accelerated processing of the multi-scalar multiplication task with fewer hardware resources and improve the utilization rate of hardware resources of the field programmable gate array. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a structural block diagram of the hardware acceleration system based on a multi-core architecture of the present application;

[0029] Figure 2 is a schematic diagram of the process of splitting a multi-scalar multiplication task in the present application;

[0030] Figure 3 The first flow diagram of the first embodiment of the hardware acceleration method based on a multi-core architecture in this application;

[0031] Figure 4 The first scenario diagram of the first embodiment of the hardware acceleration method based on a multi-core architecture in this application;

[0032] Figure 5 The second flow diagram of the second embodiment of the hardware acceleration method based on a multi-core architecture in this application;

[0033] Figure 6 The second scenario diagram of the second embodiment of the hardware acceleration method based on a multi-core architecture in this application;

[0034] Figure 7 The structural block diagram of the hardware acceleration device based on a multi-core architecture in this application;

[0035] Figure 8 The structural diagram of the hardware operating environment involved in the solution of the embodiment of this application.

[0036] The realization of the purpose of this application, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0037] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0038] Refer to Figure 1 , Figure 1 This is a hardware acceleration system based on a multi-core architecture in this application. In this embodiment, the hardware acceleration system based on a multi-core architecture includes a CPU and a field programmable gate array based on a multi-core architecture. The field programmable gate array based on a multi-core architecture includes multiple multi-scalar computing acceleration cores;

[0039] In this embodiment, the multi-core architecture is to deploy multiple multi-scalar computing acceleration cores on the field programmable gate array based on the hardware resources of the field programmable gate array. The field programmable gate array is a new type of programmable logic device, and the multi-scalar multiplication task refers to the task of simultaneously performing multiple scalar multiplication operations on different points of an elliptic curve.

[0040] Specifically, the CPU splits the multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture;

[0041] In this embodiment, the hardware acceleration system based on the multi-core architecture reads the points and scalar data of the elliptic curve to be calculated from the memory through the CPU, and then the CPU splits the multi-scalar multiplication task into multiple multi-scalar multiplication subtasks according to the splitting program pre-stored by the user, and then sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on the multi-core architecture through the PCI-e (Peripheral Component Interconnect Express, high-speed serial computer expansion bus standard) bus.

[0042] For example, with reference to Figure 2 , the multi-scalar multiplication task containing T elliptic curve points input by the user is split to obtain N multi-scalar multiplication subtasks. Each multi-scalar multiplication subtask includes a multi-scalar multiplication subtask containing T / N elliptic curve points, and there is no data dependency between the N multi-scalar multiplication subtasks.

[0043] Specifically, the field programmable gate array based on the multi-core architecture receives the multiple multi-scalar multiplication subtasks, and processes the corresponding split multi-scalar multiplication subtasks respectively based on the multiple multi-scalar calculation acceleration cores.

[0044] In this embodiment, the field programmable gate array based on the multi-core architecture receives the multiple multi-scalar multiplication subtasks through the PCI-e (Peripheral Component Interconnect Express, high-speed serial computer expansion bus standard) bus. The field programmable gate array based on the multi-core architecture preprocesses the multiple multi-scalar multiplication subtasks. The preprocessing refers to converting the multiple scalar multiplication operations of different points of the elliptic curve in the multiple multi-scalar multiplication subtasks into multiple scalar point addition operations of different points of the elliptic curve. The field programmable gate array based on the multi-core architecture processes the preprocessed multi-scalar multiplication subtasks respectively based on the multiple multi-scalar calculation acceleration cores. There is no data dependency between the multi-scalar multiplication subtasks. The field programmable gate array based on the multi-core architecture synchronously or parallelly processes the multiple multi-scalar multiplication subtasks on multiple multi-scalar calculation acceleration cores, and completes the processing of any multi-scalar multiplication subtask through thousands of elliptic curve point addition calculations.

[0045] For example, the field programmable gate array based on the multi-core architecture includes N multi-scalar calculation acceleration cores, and the field programmable gate array based on the multi-core architecture synchronously or parallelly processes the corresponding split N multi-scalar multiplication subtasks respectively based on the N multi-scalar calculation acceleration cores.

[0046] Furthermore, the calculation unit of the multi-scalar calculation acceleration core may only include a modular multiplication circuit, a modular addition circuit and a modular subtraction circuit.

[0047] In this embodiment, the computing unit in the field programmable gate array based on a multi-core architecture is used to process the corresponding split multi-scalar multiplication subtasks. Compared with PipeZK in the prior art, in which each acceleration core contains 16 modulo multiplication computing units with a relatively large circuit scale, the computing unit in the field programmable gate array based on a multi-core architecture has a smaller circuit scale.

[0048] Furthermore, the multiple multi-scalar computing acceleration cores can have independent clock domains.

[0049] In this embodiment, the clock domain refers to the area in a circuit controlled by the same clock signal. The independent clock domain means that between the multi-scalar computing acceleration cores is an asynchronous clock system and inside the multi-scalar computing acceleration cores is a synchronous clock system, that is, there is no data interaction between the multiple multi-scalar computing acceleration cores in the field programmable gate array based on a multi-core architecture, and the clock frequencies of the multiple multi-scalar computing acceleration cores can be optimized independently. The clock frequency refers to the basic frequency of the clock in a synchronous circuit.

[0050] For example, when processing a multi-scalar multiplication task based on the BLS12-381 elliptic curve, the field programmable gate array based on a multi-core architecture includes 5 multi-scalar computing acceleration cores. The 5 multi-scalar computing acceleration cores can include combination one: 4 multi-scalar computing acceleration cores with a clock frequency of 180 MHz and 1 multi-scalar computing acceleration core with a clock frequency of 100 MHz. The 5 multi-scalar computing acceleration cores can also include combination two: 5 multi-scalar computing acceleration cores with a unified clock frequency of 100 MHz. The field programmable gate array based on a multi-core architecture with independent clock domains and including combination one has a 48% higher computing power compared to the field programmable gate array based on a multi-core architecture including combination two.

[0051] The purpose of this embodiment is to use the CPU in the hardware acceleration system based on a multi-core architecture and the field programmable gate array based on a multi-core architecture to accelerate the processing of multi-scalar multiplication tasks, and to achieve the acceleration processing of multi-scalar multiplication tasks with fewer hardware resources through the CPU and the field programmable gate array based on a multi-core architecture, thereby improving the utilization rate of the hardware resources of the field programmable gate array.

[0052] Furthermore, based on the above embodiment, another embodiment of the present application is provided. In this embodiment, referring to Figure 3 , applied to a field programmable gate array based on a multi-core architecture, the hardware acceleration method based on a multi-core architecture includes the following steps:

[0053] Step S10, receive multiple multi-scalar multiplication subtasks sent by the CPU, where the CPU splits the multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture;

[0054] It should be noted that the execution subject of the method in this embodiment is a field programmable gate array based on a multi-core architecture. The field programmable gate array based on a multi-core architecture may be subordinate to a hardware acceleration device based on a multi-core architecture. The hardware acceleration device based on a multi-core architecture may be a mobile terminal or a computer. The hardware acceleration device based on a multi-core architecture may be the field programmable gate array based on a multi-core architecture on the mobile terminal or the computer. No specific limitation is made in this application.

[0055] It can be understood that the field programmable gate array based on a multi-core architecture receives the multiple multi-scalar multiplication subtasks sent by the CPU. The multiple multi-scalar multiplication subtasks include the multiple multi-scalar multiplication subtasks obtained by the CPU splitting the multi-scalar multiplication task.

[0056] In a specific implementation, referring to Figure 4 , the field programmable gate array based on a multi-core architecture includes a data distribution module, a data bus, a control bus, a data collection module, and multiple multi-scalar calculation acceleration cores. The multi-scalar calculation acceleration core includes a handshake protocol, a data cache, a status register, and a calculation unit. The data distribution module is used to preprocess the multiple multi-scalar multiplication subtasks and distribute the preprocessed multiple multi-scalar multiplication subtasks to the multiple multi-scalar calculation acceleration cores. The data bus is used to transmit the multiple multi-scalar multiplication subtasks and the processing results of the multiple multi-scalar multiplication subtasks. The control bus is used to transmit the working status of the multi-scalar calculation acceleration core and notify the CPU to read the processing results of the multiple multi-scalar multiplication subtasks. The data collection module is used to collect the processing results of the multiple multi-scalar multiplication subtasks. The multiple multi-scalar calculation acceleration cores are used to process the corresponding split multi-scalar multiplication subtasks respectively. The handshake protocol is used for the multi-scalar calculation acceleration core to receive the multi-scalar multiplication subtask based on the handshake protocol. The data cache is used to cache the multiple multi-scalar multiplication subtasks and the processing results of the multiple multi-scalar multiplication subtasks. The status register is used to record the working status of the field programmable gate array based on a multi-core architecture. The calculation unit is used to perform a series of elliptic curve point addition processes on the corresponding split multi-scalar multiplication subtasks and obtain the processing results of the multi-scalar multiplication subtasks.

[0057] Specifically, the field programmable gate array based on a multi-core architecture receives multiple multi-scalar multiplication subtasks sent by the CPU through a PCI-e (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) bus, and then preprocesses the multiple multi-scalar multiplication subtasks through a data distribution module in the field programmable gate array based on the multi-core architecture. The preprocessing includes converting multiple scalar multiplication operations of different points on an elliptic curve in the multiple multi-scalar multiplication subtasks into multiple scalar point addition operations of different points on the elliptic curve using the Pippenger algorithm. The field programmable gate array based on the multi-core architecture distributes the preprocessed multiple multi-scalar multiplication subtasks to the multiple multi-scalar calculation acceleration cores through Data Bus 1. The multi-scalar calculation acceleration cores in the field programmable gate array based on the multi-core architecture receive the multi-scalar multiplication subtasks containing T / N elliptic curve points based on a handshake protocol and cache the received multi-scalar multiplication subtasks. The handshake protocol refers to a type of network protocol used to let the client and the server confirm each other's identities, or to assist both sides in selecting the encryption algorithm, MAC algorithm (Message Authentication Codes), and related keys used when connecting.

[0058] Step S20: Process the corresponding split multi-scalar multiplication subtasks based on multiple multi-scalar calculation acceleration cores respectively to obtain the processing results of the multiple multi-scalar multiplication subtasks, where the field programmable gate array based on the multi-core architecture includes multiple multi-scalar calculation acceleration cores.

[0059] It should be noted that the field programmable gate array based on the multi-core architecture processes the corresponding split multi-scalar multiplication subtasks based on multiple multi-scalar calculation acceleration cores respectively, and the multiple multi-scalar multiplication subtasks are independent of each other.

[0060] It can be understood that with reference to Figure 4, the field programmable gate array based on a multi-core architecture performs a series of elliptic curve point addition processes on corresponding split multi-scalar multiplication subtasks by multiple multi-scalar computing acceleration cores synchronously or in parallel. Meanwhile, the status register in the field programmable gate array based on a multi-core architecture records that the field programmable gate array based on a multi-core architecture is in a busy state. Then, the field programmable gate array based on a multi-core architecture obtains the processing results of the multiple multi-scalar multiplication subtasks by multiple multi-scalar computing acceleration cores and caches the processing results. Meanwhile, the status register in the field programmable gate array based on a multi-core architecture records that the field programmable gate array based on a multi-core architecture is in a state to be read. The field programmable gate array based on a multi-core architecture notifies the CPU to read the processing results of the multiple multi-scalar multiplication subtasks through a control bus, and collects the processing results of the multiple multi-scalar multiplication subtasks to a data collection module through a data bus 2.

[0061] In this embodiment, multiple multi-scalar multiplication subtasks sent by the CPU are received. The CPU splits a multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture. Each of multiple multi-scalar computing acceleration cores processes the corresponding split multi-scalar multiplication subtasks to obtain the processing results of the multiple multi-scalar multiplication subtasks. The field programmable gate array based on a multi-core architecture includes multiple multi-scalar computing acceleration cores. The above process splits the multi-scalar multiplication task and processes the split multi-scalar multiplication subtasks by multiple multi-scalar computing acceleration cores on the field programmable gate array based on a multi-core architecture, realizing the acceleration processing of the multi-scalar multiplication task with fewer hardware resources and improving the utilization rate of the hardware resources of the field programmable gate array.

[0062] Further, based on the above embodiment, another embodiment of the present application is provided. In this embodiment, referring to Figure 5 , applied to a field programmable gate array based on a multi-core architecture, the hardware acceleration method based on a multi-core architecture further includes the following steps:

[0063] Step A1, record the status of the field programmable gate array based on a multi-core architecture. The status includes a busy state, a state to be read, and an idle state;

[0064] It should be noted that the field programmable gate array based on a multi-core architecture records the status of the field programmable gate array through a status register in the multi-scalar computing acceleration core. The status includes a busy state, a state to be read, and an idle state.

[0065] It can be understood that when the field programmable gate array based on a multi-core architecture processes the corresponding split multi-scalar multiplication subtasks respectively based on multiple multi-scalar computing acceleration cores, the field programmable gate array based on the multi-core architecture records the status as the busy state; when the field programmable gate array based on the multi-core architecture obtains the processing results of the multiple multi-scalar multiplication subtasks, the field programmable gate array based on the multi-core architecture records the status as the to-be-read state; when it is detected that all the processing results of the multiple multi-scalar multiplication subtasks are read by the CPU, the field programmable gate array based on the multi-core architecture records the status as the idle state.

[0066] In a specific implementation, referring to Figure 6 , the multi-scalar computing acceleration cores of the field programmable gate array based on the multi-core architecture include a data reception cache, a data transmission cache, a computing unit, and a status register. The data reception cache is used to cache the multi-scalar multiplication subtasks received by the multi-scalar computing acceleration core. The data transmission cache is used to cache the processing results of the multi-scalar multiplication subtasks. The computing unit only includes a modular multiplication module, a modular addition module, and a modular subtraction module. The computing unit performs a series of elliptic curve point addition processes on the corresponding split multi-scalar multiplication subtasks through the modular multiplication module, the modular addition module, and the modular subtraction module to obtain the processing results of the multi-scalar multiplication subtasks. The status register is used to record the status of the field programmable gate array based on the multi-core architecture. The field programmable gate array based on the multi-core architecture transmits the multiple multi-scalar multiplication subtasks and the processing results of the multiple multi-scalar multiplication subtasks through a data bus. The field programmable gate array based on the multi-core architecture transmits the working status of the multi-scalar computing acceleration core through a control bus and notifies the CPU to read the processing results of the multiple multi-scalar multiplication subtasks.

[0067] Specifically, the data reception cache in the field programmable gate array based on the multi-core architecture caches the multi-scalar multiplication subtasks received by the multi-scalar computing acceleration core. The field programmable gate array based on the multi-core architecture processes the multi-scalar multiplication subtasks. At the same time, the status register in the field programmable gate array based on the multi-core architecture records that the field programmable gate array based on the multi-core architecture is in the busy state. Then, the field programmable gate array based on the multi-core architecture obtains the processing results of the multiple multi-scalar multiplication subtasks and caches the processing results of the multi-scalar multiplication subtasks through the data transmission cache. At the same time, the status register in the field programmable gate array based on the multi-core architecture records that the field programmable gate array based on the multi-core architecture is in the to-be-read state. The field programmable gate array based on the multi-core architecture notifies the CPU to read the processing results of the multiple multi-scalar multiplication subtasks through the control bus.

[0068] Step A2, when the state is the busy state, stop receiving multiple multi-scalar multiplication subtasks sent by the CPU;

[0069] It should be noted that when the state is the busy state, the field programmable gate array based on the multi-core architecture stops receiving multiple multi-scalar multiplication subtasks sent by the CPU.

[0070] It can be understood that when the state is the busy state, the field programmable gate array based on the multi-core architecture sends the recorded busy state to the CPU through the control bus, so that the CPU caches the multiple multi-scalar multiplication subtasks into the to-be-sent queue in the database.

[0071] Step A3, when the state is the to-be-read state, notify the CPU to read the processing results of the multiple multi-scalar multiplication subtasks, so that the CPU can merge and output the processing results;

[0072] It should be noted that when the state is the to-be-read state, the field programmable gate array based on the multi-core architecture notifies the CPU to read the processing results of the multiple multi-scalar multiplication subtasks through the control bus, so that the CPU can merge and output the processing results.

[0073] It can be understood that when the state is the to-be-read state, the data collection module in the field programmable gate array based on the multi-core architecture collects the processing results of the multiple multi-scalar multiplication subtasks through data bus 2 and waits for the CPU to read through the PCI-e (Peripheral Component Interconnect Express) bus.

[0074] In a specific implementation, when it is detected that all the cached processing results have been read by the CPU, the field programmable gate array based on the multi-core architecture clears the processing results cached in the data transmission cache in the multi-scalar calculation acceleration core. When it is detected that not all the cached processing results have been read by the CPU, the field programmable gate array based on the multi-core architecture notifies the CPU to re-read the processing results of the multiple multi-scalar multiplication subtasks through the control bus again. The data collection module in the multi-scalar calculation acceleration core of the field programmable gate array based on the multi-core architecture re-collects the processing results of the multiple multi-scalar multiplication subtasks and waits for the CPU to re-read through the PCI-e (Peripheral Component Interconnect Express) bus.

[0075] Step A4, when the state is the idle state, receive multiple multi-scalar multiplication subtasks sent by the CPU.

[0076] It should be noted that when the state is the idle state, the field programmable gate array based on the multi-core architecture receives multiple multi-scalar multiplication subtasks sent by the CPU.

[0077] It can be understood that when the state is the idle state, the field programmable gate array based on the multi-core architecture sends the idle state recorded in the state register to the CPU, so that the CPU can sequentially extract multiple multi-scalar multiplication subtasks split from a multi-scalar multiplication task from the to-be-sent queue in the database and send the multiple multi-scalar multiplication subtasks to the field programmable gate array based on the multi-core architecture.

[0078] In this embodiment, record the state of the field programmable gate array based on the multi-core architecture. The state includes the busy state, the to-be-read state, and the idle state. When the state is the busy state, stop receiving multiple multi-scalar multiplication subtasks sent by the CPU. When the state is the to-be-read state, notify the CPU to read the processing results of the multiple multi-scalar multiplication subtasks, so that the CPU can merge and output the processing results. When the state is the idle state, receive multiple multi-scalar multiplication subtasks sent by the CPU. The above process records the state of the field programmable gate array based on the multi-core architecture and sends the recorded state to the CPU, so that the CPU can make different responses to the field programmable gate array based on the multi-core architecture in different states, thereby realizing the interaction between the CPU and the field programmable gate array based on the multi-core architecture, so as to facilitate the operation of the hardware acceleration system based on the multi-core architecture.

[0079] In addition, an embodiment of the present application also proposes a hardware acceleration device based on a multi-core architecture. Refer to Figure 7 , the hardware acceleration device based on the multi-core architecture includes:

[0080] A task receiving module 10, configured to receive multiple multi-scalar multiplication subtasks sent by the CPU;

[0081] A task processing module 20, configured to process the corresponding split multi-scalar multiplication subtasks respectively based on multiple multi-scalar calculation acceleration cores to obtain the processing results of the multiple multi-scalar multiplication subtasks

[0082] Optionally, the hardware acceleration device based on the multi-core architecture further includes:

[0083] A status recording module, configured to record the status of the field programmable gate array based on a multi-core architecture, where the status includes a busy status, a to-be-read status, and an idle status;

[0084] A busy status stopping module, configured to stop receiving multiple multi-scalar multiplication subtasks sent by the CPU when the status is the busy status;

[0085] A to-be-read status notifying module, configured to notify the CPU to read the processing results of the multiple multi-scalar multiplication subtasks when the status is the to-be-read status, for the CPU to merge and output the processing results;

[0086] An idle status receiving module, configured to receive multiple multi-scalar multiplication subtasks sent by the CPU when the status is the idle status.

[0087] Optionally, the hardware acceleration device based on a multi-core architecture further includes:

[0088] A processing result caching module, configured to cache the processing results of the multiple multi-scalar multiplication subtasks for the CPU to read the cached processing results;

[0089] A cached data clearing module, configured to clear the cached processing results when it is detected that all the cached processing results have been read by the CPU.

[0090] In this embodiment, through the above solution, multiple multi-scalar multiplication subtasks sent by the CPU are received, where the CPU splits a multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the field programmable gate array based on a multi-core architecture; the multiple multi-scalar multiplication subtasks are respectively processed by multiple multi-scalar calculation acceleration cores to obtain the processing results of the multiple multi-scalar multiplication subtasks, where the field programmable gate array based on a multi-core architecture includes multiple multi-scalar calculation acceleration cores; the above process splits the multi-scalar multiplication task and processes the split multi-scalar multiplication subtasks based on the multiple multi-scalar calculation acceleration cores on the field programmable gate array based on a multi-core architecture, so as to achieve fast processing of the multi-scalar multiplication task with less hardware resources and improve the utilization rate of the hardware resources of the field programmable gate array.

[0091] The specific implementation manners of the hardware acceleration device based on a multi-core architecture in this application are basically the same as those of the embodiments of the above-mentioned hardware acceleration method based on a multi-core architecture, and will not be elaborated here.

[0092] Refer to Figure 8 , Figure 8 which is a schematic structural diagram of a hardware acceleration device based on a multi-core architecture for the hardware operating environment involved in the solution of the embodiment of this application.

[0093] As Figure 8 shown, the hardware acceleration device based on a multi-core architecture may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to implement connection communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (WI-FI) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) memory, or may be a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0094] Those skilled in the art can understand that Figure 8 the structure shown in

[0095] does not constitute a limitation on the hardware acceleration device based on a multi-core architecture, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 8 As

[0096] shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a hardware acceleration program based on a multi-core architecture.

[0097] In Figure 8 the hardware acceleration device based on a multi-core architecture shown, the hardware acceleration device based on a multi-core architecture calls the hardware acceleration program stored in the memory 1005 through the processor 1001 to implement the steps of the hardware acceleration method based on a multi-core architecture described in any one of the above.

[0098] The specific implementation of the hardware acceleration device based on the multi-core architecture in this application is basically the same as each embodiment of the above-mentioned hardware acceleration method based on the multi-core architecture, and will not be elaborated here.

[0099] In addition, an embodiment of the present invention also proposes a storage medium. An embodiment of the present application provides a storage medium, and the storage medium stores one or more programs, and the one or more programs can also be executed by one or more processors to implement the steps of the hardware acceleration method based on the multi-core architecture described in any one of the above.

[0100] The specific implementation of the storage medium of this application is basically the same as each embodiment of the above-mentioned hardware acceleration method based on the multi-core architecture, and will not be elaborated here.

[0101] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0102] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0104] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the description of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.

Claims

1. A hardware acceleration system based on a multi-core architecture, characterized in that, The hardware acceleration system based on a multi-core architecture includes a CPU and a field programmable gate array (FPGA) based on a multi-core architecture. The FPGA based on a multi-core architecture includes multiple multi-scalar computing acceleration cores; The CPU splits a multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the FPGA based on a multi-core architecture; The FPGA based on a multi-core architecture receives the multiple multi-scalar multiplication subtasks and processes the corresponding split multi-scalar multiplication subtasks respectively based on the multiple multi-scalar computing acceleration cores.

2. The hardware acceleration system based on a multi-core architecture according to claim 1, wherein The computing unit of the multi-scalar computing acceleration core only includes a modular multiplication circuit, a modular addition circuit and a modular subtraction circuit.

3. The hardware acceleration system based on a multi-core architecture according to claim 1, characterized in that The multiple multi-scalar computing acceleration cores have independent clock domains.

4. A hardware acceleration method based on a multi-core architecture, characterized in that, Applied to an FPGA based on a multi-core architecture, the hardware acceleration method based on a multi-core architecture includes the following steps: Receiving multiple multi-scalar multiplication subtasks sent by the CPU, where the CPU splits a multi-scalar multiplication task into multiple multi-scalar multiplication subtasks and sends the multiple multi-scalar multiplication subtasks to the FPGA based on a multi-core architecture; Processing the corresponding split multi-scalar multiplication subtasks respectively based on multiple multi-scalar computing acceleration cores to obtain the processing results of the multiple multi-scalar multiplication subtasks, where the FPGA based on a multi-core architecture includes multiple multi-scalar computing acceleration cores.

5. The hardware acceleration method based on a multi-core architecture according to claim 4, wherein The multiple multi-scalar multiplication subtasks are independent of each other, and the multiple multi-scalar computing acceleration cores process the corresponding split multi-scalar multiplication subtasks synchronously or in parallel.

6. The hardware acceleration method based on a multi-core architecture according to claim 4, wherein The hardware acceleration method based on a multi-core architecture further includes the following steps: Recording the state of the FPGA based on a multi-core architecture, where the state includes a busy state, a state to be read, and an idle state; When the state is the busy state, stop receiving the multiple multi-scalar multiplication subtasks sent by the CPU; When the state is the state to be read, notify the CPU to read the processing results of the multiple multi-scalar multiplication subtasks for the CPU to merge and output the processing results; When the state is the idle state, receive the multiple multi-scalar multiplication subtasks sent by the CPU.

7. The hardware acceleration method based on a multi-core architecture according to any one of claims 4 to 6, characterized in that After the step of processing the corresponding split multi-scalar multiplication subtasks respectively based on the multiple multi-scalar computing acceleration cores to obtain the processing results of the multiple multi-scalar multiplication subtasks, it includes: Caching the processing results of the multiple multi-scalar multiplication subtasks for the CPU to read the cached processing results; When it is detected that all the cached processing results have been read by the CPU, clear the cached processing results.

8. A hardware acceleration device based on a multi-core architecture, characterized in that, The hardware acceleration device based on a multi-core architecture includes: A task receiving module, configured to receive multiple multi-scalar multiplication subtasks sent by the CPU; A task processing module, configured to process the corresponding split multi-scalar multiplication subtasks respectively based on multiple multi-scalar computing acceleration cores to obtain the processing results of the multiple multi-scalar multiplication subtasks.

9. A hardware acceleration device based on a multi-core architecture, characterized in that, The device includes: a memory, a processor, and a hardware acceleration program based on a multi-core architecture stored on the memory and executable on the processor, where the hardware acceleration program based on the multi-core architecture is configured to implement the steps of the hardware acceleration method based on the multi-core architecture according to any one of claims 4 to 7.

10. A storage medium, characterized in that, A hardware acceleration program based on a multi-core architecture is stored on the storage medium, and when the hardware acceleration program based on the multi-core architecture is executed by a processor, the steps of the hardware acceleration method based on the multi-core architecture according to any one of claims 4 to 7 are implemented.

Citation Information

Cited By

  • A heterogeneous computing signal processing method and system for high-speed communication scenarios

    CN122554288A

  • A heterogeneous computing signal processing method and system for high-speed communication scenarios

    CN122554288B