Public key processing device and method, processing system, electronic equipment and storage medium

Through the multi-threaded collaboration of the control processor and the computing acceleration module, the problem that the public key cryptography system cannot be fully uninstalled in the server is solved, and an efficient public key processing process is realized, which improves computing performance and flexibility.

CN120358035APending Publication Date: 2025-07-22XIODUAN COMPUTING (NANJING) TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410071037.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, in the large-scale deployment of public key cryptography systems in the data center server, due to the high bit width and complex modulus operation requirements of RSA and ECC algorithms, CPU computing resources cannot be fully unloaded, affecting the efficiency of other user applications, and the single-threaded processing solution cannot meet the high-performance multi-concurrency requirements.

Method used

The control processor and the operation acceleration module work together, and multiple computing engines are controlled through multiple threads to realize parallel processing of public key processing requests, including sleep state management and thread switching, and optimize the utilization rate of the computing engine.

Benefits of technology

It realizes complete offloading of the CPU public key computing load, improves the concurrency and encryption computing performance of the computing engine, and is suitable for high-performance public key algorithm processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358035A_ABST
    Figure CN120358035A_ABST
Patent Text Reader

Abstract

The invention discloses a public key processing device and method, a processing system, electronic equipment and a storage medium. The public key processing device comprises: a control processor configured to receive one or more public key processing requests of a CPU, and generate a corresponding public key protocol processing flow to control an operation acceleration module; the operation acceleration module is configured to respond to the control of the control processor, execute operation processing, obtain an operation result aiming at the public key processing request, and return the operation result to the control processor; the control processor is further configured to receive an operation result and return the operation result to the CPU; the operation acceleration module comprises a plurality of operation engines, and at least one operation engine is configured to execute operation processing of operators; the control processor is configured to multi-thread control the plurality of operation engines for one or more public key processing requests. According to the invention, unloading of the public key calculation load by the CPU and high-flexibility and high-performance encryption operation processing can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of public key computing, and in particular, to a public key processing device and method, a processing system, an electronic device, and a storage medium. Background Art

[0002] Public key cryptosystems have been widely used in today's Information and Communication Technology (ICT) systems. Their main applications include processes such as key generation, key exchange, encryption / decryption, signature / verification, etc. Currently, the mainstream algorithms for public key encryption are mainly divided into two categories: RSA (Rivest-Shamir-Adleman) and Elliptic Curve Cryptography (ECC). Due to the large bit widths introduced by these two types of algorithms (for example, the bit width of a secure RSA algorithm is at least 2Kb) and the complex modular arithmetic, the intensive computing power required has become the main bottleneck for the large-scale deployment of public key cryptosystems in data center servers.

[0003] Currently, the acceleration technology solutions for computing power operations mostly adopt the following two methods:

[0004] One is the solution of collaborative processing by a Central Processing Unit (CPU) and a Coprocessor. In this solution, the CPU undertakes tasks such as process chaining and intermediate result processing, while the Coprocessor mainly undertakes the calculation task offloading of modular arithmetic. However, in this solution, the CPU still needs to undertake tasks such as process control. The repeated interaction with the Coprocessor requires the CPU to constantly perform process switching or interrupt processing, and it cannot achieve the complete offloading of public key data processing and fully release the precious computing resources of the CPU. Therefore, it affects the efficiency of the CPU in executing other user applications.

[0005] The other is the solution of using a single-threaded Microcontroller Unit (MCU) and a Coprocessor. However, this solution can only process a series of public key operation requests one by one and serially, with low processing performance, and is only applicable to resource-constrained application scenarios (such as Internet of Things applications, etc.). For high-performance multi-concurrent public key calculations on the server side, this solution cannot meet its requirements. Summary of the Invention

[0006] The purpose of the present invention is to solve at least one of the above technical problems to a certain extent.

[0007] To achieve the above object, the present invention provides a public key processing apparatus, including a control processor and an operation acceleration module; the control processor is configured to receive one or more public key processing requests from the CPU, generate corresponding public key protocol processing flows, and control the operation acceleration module based on the generated public key protocol processing flows; the operation acceleration module is configured to perform operation processing in response to the control of the control processor, obtain an operation result for the public key processing request, and return the operation result to the controller processor; and the control processor is configured to receive and return the operation result to the CPU; wherein: the operation acceleration module includes a plurality of operation engines, and at least one operation engine is configured to perform operation processing of operators; and the control processor is further configured to perform multi-thread control on the plurality of operation engines for one or more public key processing requests.

[0008] The control processor is further configured to perform multi-thread control on the plurality of operation engines for one or more public key processing requests, including: the control processor calls a plurality of operation engines to perform operator operations based on pre-defined operators to be processed, and controls the running states of the threads where the plurality of operation engines are located, where the running states include a sleep state and a processing state.

[0009] The control processor is further configured to: obtain the operation conditions of the operation engines that perform operator operation processing in the threads in the sleep state, and if the operation conditions of the operation engines are that the operations are completed, wake up the threads in the sleep state to continue controlling the threads to perform public key processing.

[0010] The control processor is further configured to perform multi-thread control on the plurality of operation engines for one or more public key processing requests, including: determining whether there are operation engines available for performing the current operator processing, and if not, putting the current thread to sleep to switch to controlling the processing of other threads.

[0011] The control processor is further configured to: when it is obtained that there are idle and available operation engines, wake up the threads in the sleep state to control the idle and available operation engines to perform the operation processing of the operators of the threads.

[0012] The operation acceleration module further includes: a Hash calculation unit and / or a random number generation unit.

[0013] The arithmetic operation engine includes: a macro instruction processing unit configured to pre-decompose the macro instruction for executing the operator sent by the control processor into a corresponding plurality of micro instructions; a micro instruction processing unit configured to sequentially retrieve the plurality of micro instructions and send them to the arithmetic logic unit; the arithmetic logic unit configured to perform arithmetic processing of the corresponding operator based on the received micro instructions to obtain an arithmetic result of the operator; and a register unit configured to store the arithmetic result of the operator.

[0014] To achieve the above object, on the other hand, the present invention provides a public key processing method for a public key processing device. The public key processing device includes a control processor and an arithmetic acceleration module, and the arithmetic acceleration module includes a plurality of arithmetic operation engines. The public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates a corresponding public key protocol processing flow, and performs multi-threaded control on the plurality of arithmetic operation engines based on the generated public key protocol processing flow; the plurality of arithmetic operation engines, in response to the multi-threaded control, perform arithmetic processing of the operator to obtain an arithmetic result for the one or more public key processing requests, and return the arithmetic result to the control processor; and the control processor receives and returns the arithmetic result to the CPU.

[0015] The control processor performs multi-threaded control on the plurality of arithmetic operation engines, including: the control processor invokes a plurality of arithmetic operation engines to perform operator operations based on the predefined operator to be processed, and controls the running states of the threads where the plurality of arithmetic operation engines are located, where the running states include a sleep state and a processing state.

[0016] The control processor performs multi-threaded control on the plurality of arithmetic operation engines, and further includes: obtaining the execution state of the arithmetic operation engine that performs operator operation processing in the thread in the sleep state. If the execution state of the arithmetic operation engine is execution completed, the thread in the sleep state is woken up to continue controlling the thread to perform public key processing.

[0017] The control processor performs multi-threaded control on the plurality of arithmetic operation engines, and further includes: determining whether there is an arithmetic operation engine available for executing the current operator processing. If not, the corresponding thread is put to sleep to switch to controlling the processing of other threads.

[0018] The control processor performs multi-threaded control on the plurality of arithmetic operation engines, and further includes: when an idle and available arithmetic operation engine is obtained, waking up the thread in the sleep state to control the idle and available arithmetic operation engine to perform the operator operation of the thread.

[0019] For the arithmetic acceleration module, the public key processing method may further include: Hash calculation or generating a random number.

[0020] The operation engine is configured to: pre - disassemble the macro - instruction for executing the operator sent by the control processor into a corresponding plurality of micro - instructions; sequentially retrieve the plurality of micro - instructions, execute the operations of the corresponding operator to obtain the operation result of the operator; and store the operation result of the operator.

[0021] To achieve the above object, in one aspect, the present invention provides a public - key processing system, which includes: a central processing unit; and the aforementioned public - key processing device.

[0022] To achieve the above object, in another aspect, the present invention provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the public - key processing method described above in the present invention is implemented.

[0023] To achieve the above object, the present invention also provides a non - transitory computer - readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the aforementioned public - key processing method.

[0024] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The above - mentioned and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, wherein:

[0026] Figure 1 is a schematic block diagram of a public - key processing device shown according to an exemplary embodiment;

[0027] Figure 2 is a schematic block diagram of an operation engine shown according to an exemplary embodiment;

[0028] Figure 3 is a schematic block diagram of another public - key processing device shown according to an exemplary embodiment;

[0029] Figure 4 is a schematic diagram of the digital signature process of an SM2 public - key algorithm shown according to an exemplary embodiment;

[0030] Figure 5 is a schematic diagram of multi - thread control shown according to an exemplary embodiment; and

[0031] Figure 6 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0033] The present invention provides a public key processing device, as Figure 1 shown. The public key processing device includes a control processor 10 and an arithmetic acceleration module 20. The control processor 10 is configured to receive one or more public key processing requests from a central processing unit (CPU), generate a corresponding public key protocol processing flow, and control the arithmetic acceleration module 20 based on the generated public key protocol processing flow. The arithmetic acceleration module 20 is configured to perform arithmetic processing in response to the control of the control processor 10, obtain an arithmetic result for the public key processing request, and return the arithmetic result to the control processor 10. The control processor 10 is further configured to receive and return the arithmetic result to the CPU. Wherein: the arithmetic acceleration module 20 includes a plurality of arithmetic engines 201, and at least one arithmetic engine 201 is configured to perform arithmetic processing of operators; and the control processor 10 is further configured to perform multi-threaded control on the plurality of arithmetic engines 201 for one or more public key processing requests.

[0034] For example, as Figure 1 shown, the control processor 10 communicates with the CPU. It can receive one or more public key processing requests sent by the CPU and return the arithmetic result to the CPU. Specifically, when the control processor 10 receives a public key processing request from the CPU, it first parses the public key processing request to obtain a corresponding public key protocol processing flow. Then, the control processor 10 controls the arithmetic acceleration module 20 to perform arithmetic processing based on the obtained public key protocol processing flow. Specifically, the control processor 10 is used for the serial connection control of the protocol process, which supports programmability and interacts with the arithmetic acceleration module 20 to send corresponding control instructions to the arithmetic acceleration module 20 according to the arithmetic operations to be performed in each step of the protocol process, so that the arithmetic acceleration module 20 performs corresponding arithmetic processing based on different control instructions. Finally, the control processor 10 obtains the final encryption arithmetic result from the arithmetic acceleration module 20 and returns it to the CPU.

[0035] The operation engine 201 is mainly used to execute the operation processing of operators, such as various operations like modulo operation, addition, subtraction, multiplication, and division. The operation engine 201 supports programmability and can execute the processing of different operators for different algorithms. In the embodiment of the present invention, the operation acceleration module 20 includes multiple operation engines 201, and the multiple operation engines 201 can respectively interact with the control processor 10 to execute the control instructions of the control processor. For the processing of operators for the same public key processing request, it can be executed by the same operation engine 201 or jointly completed by multiple operation engines 201.

[0036] It can be seen that multiple operation engines 201 can process the operation processing of the same public key processing request or the operation processing of multiple different public key processing requests. Therefore, when multiple operation engines 201 process one or more different public key processing requests, they are in different processing threads. Therefore, in order to ensure the parallel processing of the operator operations of multiple operation engines 201 for one or more public key processing requests, the control processor 10 in the embodiment of the present invention can also perform multi-thread control on multiple operation engines 201 for one or more public key processing requests. For example, the control processor 10 can switch control among multiple processing threads to achieve efficient control of multiple operation engines 201.

[0037] Since the operation engine for executing complex operations is introduced in the embodiment of the present invention and the control processor performs multi-thread control on multiple operation engines, the public key processing device in the embodiment of the present invention can process public key algorithms such as RSA and ECC that require high-performance operations.

[0038] In the present invention, the control processor in the public key processing device is responsible for the control of the public key protocol process, and the operation acceleration module is responsible for operation processing. Therefore, when the public key processing device obtains the public key processing request of the CPU, the CPU can immediately process other computing tasks and no longer needs to be responsible for the relevant matters of the public key processing. On the one hand, the public key processing device realizes the overall offloading of the public key calculation load of the CPU. On the other hand, the control processor and the operation acceleration module cooperate effectively, perform multi-thread control on multiple operation engines, and then enable the efficient concurrent processing of multiple operation engines, effectively realizing high-performance and highly flexible encryption operation processing.

[0039] The above public key processing device will be further described in detail below with reference to the drawings and specific examples.

[0040] As mentioned in the above embodiment, the control processor of the present invention can perform multi-thread control on multiple operation engines to control the parallel processing of the operation processing of multiple public key processing requests by multiple operation engines. The specific content of the multi-thread control will be introduced in detail below with reference to the drawings and examples.

[0041] In a preferred embodiment, the control processor is further configured to perform multi-thread control on the multiple arithmetic engines for one or more public key processing requests, including: the control processor invokes the multiple arithmetic engines to perform operations on the operator based on a predefined operator to be processed, and controls the running states of the threads where the multiple arithmetic engines are located, where the running states include a sleep state and a current processing state.

[0042] For example, when the control processor controls the arithmetic engine to perform public key processing according to the public key protocol processing flow, the control processor will call multiple arithmetic engines to operate according to the predefined operator to be processed, and then control the running states of the threads where the arithmetic engines are located. The predefined here can be understood as the user's pre-analysis of the public key processing flow and extraction of complex operators with high computing power requirements, so as to call the arithmetic engines for acceleration. For example, the operator to be processed currently belongs to a complex operator, and the operation of this operator needs to be jointly completed by multiple arithmetic engines, or the processing of this complex operator takes a long time, such as a high-complexity point multiplication operation. When the arithmetic engine processes the point multiplication operation, it takes a long time to obtain the operation result (for example, for 256b SM2, it may take several operation cycles to complete). In this case, if the existing single-thread control is adopted, the control processor needs to wait for a long time, cannot process other public key processing requests, resulting in a reduction in the overall public key processing efficiency, and at the same time, it cannot effectively execute multiple different operator operations. Also, for example, when dealing with a certain public key processing request, its public key processing flow has different operation stages, and if the subsequent processing flow of this public key processing request needs to depend on the operation result of the current operator to continue processing, then at this time, it is necessary to wait for the arithmetic engine to complete the operation of the current operator before performing the next step of calculation of the next step of this processing flow. If the control processor stays in this thread control at this time, then the controller cannot process other public key processing requests to control other arithmetic engines.

[0043] Therefore, for the above several situations, the present invention adopts multi-thread control of multiple arithmetic engines to achieve switching between different threads by controlling the running states of the threads where they are located. Specifically, the present invention is described in the following two aspects:

[0044] 1. For multiple public key processing requests.

[0045] For different public key processing requests, they correspond to different threads for public key processing. When the control processor controls multiple threads, if all the operators processed by the currently controlled first thread are complex operators, which take a long time to compute, and the subsequent processing of the first thread depends on the computation result of the complex operator. Then the control processor puts the first thread to sleep. Here, the sleep does not mean that the current thread stops running, but rather it is transferred to background processing. During the process when the first thread is in the sleep state, the computing engine will continuously perform computing processing on the complex operator. In this way, the control processor does not need to stay and control the first thread, but can switch to other processing threads, such as the second thread, etc., to control other computing engines to perform operations on other operators. Thus, the control processor can switch between multiple threads to control multiple computing engines to achieve efficient processing of multiple public key processing requests.

[0046] In a second aspect, for a public key processing request.

[0047] For a public key processing request, it corresponds to a general thread for public key processing. When the control processor obtains this public key processing request, based on the public key processing protocol, it determines that multiple computing engines are required to execute the processing of this public key processing request, and the multiple computing engines perform different operator operations. The operation of each operator can be understood as each sub-thread under a general thread for public key processing. For example, the computing engine of the first sub-thread performs the operation of the current operator. If there is an operation that needs to be executed based on the computation result of the first sub-thread, then the first sub-thread can be put to sleep and transferred to continue running in the background. Further, if the operation of the second sub-thread does not depend on the previous computation result, then the control processor switches to the second sub-thread to control other computing engines to perform the corresponding operator operations. Thus, the control processor can switch between multiple sub-threads and parallelly control different computing engines to perform operator operations to achieve efficient processing of this public key processing request.

[0048] In addition, the running states of the above threads include not only the sleep state but also the processing state. For the processing state, for example, when the control processor learns that the operator to be computed and processed currently is not a complex operator but a simple operator, the controller processor can directly process this simple operator (without calling the computing engine), thereby maintaining the processing of the current thread and not putting the thread to sleep, thus avoiding unnecessary excessive switching of the thread.

[0049] More preferably, the control processor is further configured to: obtain the execution situation of the computing engine that performs operator computing processing in the thread in the sleep state. If the execution situation of the computing engine is that the execution is completed, then wake up the thread in the sleep state to continue controlling the thread to perform public key processing.

[0050] For example, after the control processor switches to another thread, it continues to obtain the operation status of the arithmetic engine in the thread that has been put into the sleep state. Taking the above example, the control processor continuously obtains the processing status of the dot product operation of the arithmetic engine. If the dot product operation is completed, the sleeping thread can be awakened, that is, the processing thread is restored and released to control the thread to continue to execute the subsequent operation steps of the corresponding public key protocol processing flow.

[0051] In a preferred embodiment, the control processor is further configured to perform multi-thread control on the multiple arithmetic engines for one or more public key processing requests, and further includes: determining whether there is an arithmetic engine available for performing the current operator operation. If not, the current thread is put to sleep to switch to controlling the processing of other threads.

[0052] For example, when the control processor controls multiple arithmetic engines, it determines which arithmetic engine will perform the operation of the operator according to the pre-defined current operator to be processed and the busy or idle status of each arithmetic engine. When the control processor obtains that all multiple arithmetic engines are in a busy state, that is, all multiple arithmetic engines are performing operator operations and there is no available arithmetic engine to perform the operation of the operator to be processed. In this case, if the existing single-thread control is adopted, the control processor needs to keep waiting in the current thread until there is an idle and available arithmetic engine. In this way, the control processor cannot control other arithmetic engines to perform the operator operation processing of other public key processing requests. Based on this, in the present invention, when it is obtained that there is no available arithmetic engine, the current thread is put to sleep, and then the control is switched to the arithmetic engines of other threads. For example, when the control processor needs to call an arithmetic engine to perform a dot product operation, at this time, the control processor needs to obtain the status of multiple arithmetic engines to determine whether there is an available arithmetic engine to execute. If there is no available arithmetic engine, the control processor puts the current processing thread to sleep at this time and then switches to other processing threads. It should be noted that the "sleep" in the embodiment of the present invention can be understood as that the current processing thread is in a waiting state for processing, that is, the entire encryption protocol processing flow of the current processing thread is in a waiting state. At this time, the current processing thread interrupts and stops the operation processing, and the control processor temporarily does not control the current processing thread, but switches to other processing threads for control according to the running state of other processing threads.

[0053] More preferably, the control processor is further configured to: when it is obtained that there is an idle and available arithmetic engine, wake up the sleeping thread to control the idle and available arithmetic engine to perform the operation processing of the operator of the thread.

[0054] For example, when the control processor switches to another processing thread, it also continues to check whether there are any idle computing engines available, that is, it continuously obtains the status of multiple computing engines. If it is found that a certain computing engine is idle after completing a certain computing step of another processing thread, at this time, the CP wakes up the sleeping processing thread to continue controlling the processing thread, and then calls the idle computing engine to continue executing the operator processing required by the processing thread.

[0055] It should be noted that the above is only an example of the multi-thread control of the present invention. The multi-thread control of the present invention is not limited to the above example. Those skilled in the art can adopt corresponding multi-thread control methods according to actual application needs for other different situations to achieve concurrent processing of multiple computing engines, and no further limitation is made here.

[0056] It can be seen from the above embodiments that the control processor can implement the switching of multiple processing threads according to the execution situation of the operator operations by the computing engines, so that the control processor continuously switches control among multiple threads, and the processing processes of multiple threads are continuously intertwined. As a result, the computing engines are at a relatively high degree of concurrency, achieving efficient utilization of the computing engines, and thus achieving high-performance processing of encryption operations.

[0057] In a preferred embodiment, the operation acceleration module further includes a Hash calculation unit and / or a random number generation unit.

[0058] For example, in the processing flow of digital signature in the SM2 public key algorithm, it also includes Hash calculation for secure encryption and / or random number generation. Therefore, the operation acceleration module also needs to be configured with corresponding Hash calculation units and / or random number generation units to perform the corresponding operation processing in the public key processing protocol flow. Similarly, the control processor can implement multi-thread control of multiple computing engines, Hash calculation units, and random number generation units.

[0059] It should be noted that this preferred embodiment only illustrates other functional operation units in the operation acceleration module by taking SM2 as an example. The embodiments of the present invention are not limited to the above two units. Those skilled in the art can configure other functional operation units for public key operations in the operation acceleration module according to the processing requirements of actual public key algorithms, and no further limitation is made here.

[0060] In a preferred embodiment, as Figure 2As shown, the operation engine 201 includes: a macro-instruction processing unit 2011 configured to pre-disintegrate the macro-instructions of the execution operators sent by the control processor into corresponding multiple micro-instructions; a micro-instruction processing unit 2012 configured to sequentially retrieve the multiple micro-instructions and send them to the arithmetic logic unit; an arithmetic logic unit 2013 configured to perform arithmetic processing of the corresponding operators based on the received micro-instructions to obtain the arithmetic results of the operators; and a register unit 2014 configured to store the processing results of the operators.

[0061] For example, when the operation engine 201 obtains the control instructions of the control processor, since the control instructions are macro-instructions, they need to be further disintegrated into fine-grained micro-instructions by the macro-instruction processing unit 2011 for arithmetic processing of complex operators. After being disintegrated into multiple micro-instructions, the multiple micro-instructions are stored in the instruction memory, and the micro-instruction processing unit 2012 further sequentially retrieves the multiple micro-instructions from the instruction memory and sequentially sends them to the arithmetic logic unit 2013 for arithmetic processing, and then the arithmetic logic unit 2013 generates the processing results and stores them in the register unit 2014. Specifically, the instruction memory of the operation engine includes multiple macro-instruction program segments, such as Figure 3 Macro_i, Macro_j, etc. in, and each macro-instruction program segment includes multiple micro-instructions. Taking Macro_i as an example, when the operation engine is called to process the complex operator corresponding to the macro-instruction, the macro-instruction processing unit 2011 parses the macro-instruction Macro_i to obtain the corresponding starting micro-instruction address. According to this address, the micro-instruction processing unit 2012 sequentially retrieves the micro-instructions corresponding to the macro-instruction Macro_i and sends them to the arithmetic logic unit 2013 for arithmetic processing. When all the micro-instructions corresponding to Macro_i are arithmetic processed, the arithmetic results are stored in the register unit 2014, waiting for the control processor to obtain the arithmetic results.

[0062] In the embodiment of the present invention, the operation engine can disintegrate the coarse-grained macro-instructions into fine-grained logical operation instructions for processing, and can disintegrate large-bit-width data into multiple small-bit-width logical operation instructions for processing, thereby effectively realizing the arithmetic processing of complex operators.

[0063] Figure 3 is a schematic block diagram of a more specific public key processing device, such as Figure 3As shown in the figure, the public key processing device includes a control processor, multiple arithmetic engines, a random number generator, and a cryptographic HASH calculation module. Among them, the control processor interacts with the CPU. The CPU sends a public key processing request to the control processor, and the control processor responds to the public key processing request of the CPU. The cryptographic HASH calculation module performs Hash calculation on the public key processing data sent by the CPU. The random number generator can generate different random numbers to ensure the security of the entire public key processing process. The thread context in the control processor includes multiple processing threads for multiple public key processing requests, and interacts with multiple arithmetic engines respectively based on the multiple processing threads to control the multiple arithmetic engines to execute the operations of the operators in different processing threads. The arithmetic engine receives the macro instruction from the control processor, and the macro instruction includes the operator to perform the operation. The arithmetic engine first splits the macro instruction into multiple micro instruction segments in advance through macro instruction processing and stores the micro instruction segments in the instruction memory. The micro instruction processing unit then calls each micro instruction from the instruction memory one by one and sends it to the arithmetic logic unit for operation. When the arithmetic logic unit finishes the operation of the current micro instruction, it stores the current operation result in the register bank. Furthermore, when the arithmetic logic unit continues to execute the next micro instruction, it can continue the operation based on the previous operation result. After finishing all the micro instruction operations, the arithmetic logic unit stores the final operation result in the register bank and records the storage address. Then it waits for the control processor to obtain the final operation result.

[0064] To more clearly illustrate the multi-thread control, the following will be combined with Figure 4 and Figure 5 to illustrate more specifically.

[0065] Figure 4 is a flowchart of the generation process of the digital signature of the commercial cipher SM2 algorithm. As Figure 4 shown, the control processor receives the SM2_Sign protocol request sent by the CPU, and the CP starts to execute the processing of the SM2_Sign protocol. Specifically, the received original data includes elliptic curve parameters, the information to be signed, keys, etc., where the order of the elliptic curve is n.

[0066] Furthermore, the control processor is used to control Figure 4 the process connection between each operation step in Figure 4 . In the process of Figure 2 , the second step is executed by the cryptographic Hash calculation module in Figure 3 , the third step is completed by the random number generator in Figure 3 , and the dot multiplication, modular addition, and inverse and modular multiplication operations in the fourth to sixth steps are completed by the arithmetic engine in Figure 3In steps 4-6, the control processor can control the processing by the same arithmetic engine or by multiple arithmetic engines.

[0067] If the dot product operation in step 4 requires the arithmetic engine to execute for a long time, such as SM2 with 256b, it may take several cycles to complete. At this time, the thread of this SM2 goes to sleep, and the control processor turns to process other new thread requests. After the operation in step 4 is completed, the control processor wakes up the processing thread of this SM2 and continues to send arithmetic instructions to the corresponding arithmetic engine to execute the operation in step 5.

[0068] Finally, through the control of the arithmetic engine by the control processor, after each call to the arithmetic engine for operation, the operation result is stored in the context memory of this thread. Finally, the calculation result of the digital signature is determined and the determined digital signature is sent to the CPU.

[0069] Figure 5 is a schematic diagram of multi-thread control shown according to an exemplary embodiment, which is for a single-issue, coarse-grained multi-thread switching example, such as Figure 5 shown, and its thread switching process is as follows:

[0070] Take Figure 4 the SM2 signature process as an example, Figure 5 In it, T1 to TN represent N processing threads, which execute different SM2 signatures in parallel, and each processing thread is in a different running state.

[0071] Furthermore, assume that the processing thread T1 enters the main process of the control processor and continues to execute after being scheduled by the scheduler (ellipse) in this cycle. The program counter pointer (PC pointer) and the register file (Reg File) status value need to be restored from Ctx Mem.1. The previous state of the processing thread T1 was Figure 3 the calculation in step 3 in it was completed, the PC pointed to the instruction corresponding to step 4. After extracting the instruction and decoding, through the load-store unit (LSU), Figure 4 the application programming interface (API) of the dot product operation ((x_1, y_1) = [k]G) in step 4 in it (including the elliptic curve system parameters, the multiple k, etc., all exist in CtxMem.1) is sent to one of the available arithmetic engines through interaction and the arithmetic engine is started to perform the operator operation.

[0072] Furthermore, since the dot product operation takes a long time, the control processor needs to perform a thread switch at this time, storing the PC pointer and the register bank (RegFile) status value of the processing thread T1 into Ctx Mem.1. It should be noted that when the control processor controls other processing threads to execute after the switch, the dot product operation of the processing thread T1 is still running and calculating. At this time, the arithmetic engine continues to continuously read and update Ctx Mem.1.

[0073] Suppose the processing thread T2 enters the main process of the control processor and continues to execute after being scheduled by the scheduler in this cycle. Its PC pointer and the register Reg File status value need to be restored from Ctx Mem.2. The previous state of the processing thread T2 was Figure 4 The calculation in step 6 in was completed. At this time, the PC points to the instruction corresponding to after step 6 (that is, judging s = 0?, and the operation result after the previous processing thread T2 based on the calculation of step 6 exists in Ctx Mem.2); if s!= 0, the processing thread T2 executes to step 7, and the signature calculation ends. The operation result (r, s) exists in Ctx Mem.2. After waiting for the caller of this thread to retrieve it, the thread T2 can be released.

[0074] And so on. When the arithmetic engine executes a long public key operation, the control processor switches to process other executable threads, and the thread execution processes are constantly intertwined, achieving a very high degree of concurrency. Especially at the highest performance, all arithmetic engines can work simultaneously.

[0075] In summary, the public key processing device disclosed in the present invention has the following advantages:

[0076] 1) The public key processing flow is completed by the cooperation of the control processor and the arithmetic acceleration module. After the CPU sends a public key processing request, it does not need to perform subsequent encryption transaction processing, effectively realizing the complete offloading of the CPU's public key calculation load;

[0077] 2) The control processor uses multi-thread control for multiple arithmetic engines, improving the concurrency of the arithmetic engines in processing public key calculations, being able to fully utilize the computing power of the arithmetic engines, and realizing highly flexible and high-performance public key operation processing for various public key algorithms.

[0078] Based on the same inventive concept as the above public key processing device, the present invention provides a public key processing method for a public key processing device. The public key processing device includes a control processor and an arithmetic acceleration module. The arithmetic acceleration module includes a plurality of arithmetic engines. The public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing flows, and performs multi-threaded control on the plurality of arithmetic engines based on the generated public key protocol processing flows; the plurality of arithmetic engines, in response to the multi-threaded control, perform arithmetic processing of operators, obtain arithmetic results for the one or more public key processing requests, and return the arithmetic results to the control processor; and the control processor receives and returns the encryption arithmetic results to the CPU.

[0079] The control processor performs multi-threaded control on the plurality of arithmetic engines, including: the control processor calls a plurality of arithmetic engines to perform operator arithmetic based on the predefined operators to be processed, and controls the running states of the threads where the plurality of arithmetic engines are located, where the running states include a sleep state and a processing state

[0080] The control processor performs multi-threaded control on the plurality of arithmetic engines, and further includes: obtaining the execution state of the arithmetic engine that performs operator arithmetic processing in the thread in the sleep state. If the execution state of the arithmetic engine is execution completed, wake up the thread in the sleep state to continue controlling the thread to perform public key processing.

[0081] The control processor performs multi-threaded control on the plurality of arithmetic engines, and further includes: determining whether there is an arithmetic engine available for executing the current operator processing. If not, put the corresponding thread to sleep to switch to controlling the processing of other threads.

[0082] The control processor performs multi-threaded control on the plurality of arithmetic engines, and further includes: when it is obtained that there is an idle and available arithmetic engine, wake up the thread in the sleep state to control the idle and available arithmetic engine to perform the operator arithmetic of the thread.

[0083] For the arithmetic acceleration module, the public key processing method further includes: Hash calculation or generating a random number.

[0084] The arithmetic engine is used to: disassemble the macro instruction for executing the operator sent by the control processor into corresponding multiple micro instructions in advance; sequentially retrieve the multiple micro instructions, perform arithmetic of the corresponding operator to obtain the arithmetic result of the operator; and store the arithmetic result of the operator.

[0085] Correspondingly, the present invention further provides a public key processing system, which includes a central processing unit and the public key processing device described in the above embodiment.

[0086] Correspondingly, the present invention further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the public key processing method described in the above embodiments.

[0087] Correspondingly, the present invention further provides an electronic device, as Figure 6 shown, which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present invention. The electronic device in the embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0088] As Figure 6 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0089] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 shown is an electronic device having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0090] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present invention are executed.

[0091] It should be noted that the above computer-readable medium of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0092] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0093] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0094] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to execute a public key processing method. The public key processing method is for a public key processing device, the public key processing device includes a control processor and an arithmetic acceleration module, the arithmetic acceleration module includes a plurality of arithmetic engines, and the public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing flows, and performs multi-threaded control over the plurality of arithmetic engines based on the generated public key protocol processing flows; the plurality of arithmetic acceleration engines, in response to the multi-threaded control, perform arithmetic processing of operators, obtain arithmetic results for one or more of the public key processing requests, and return the arithmetic results to the control processor; and the control processor receives and returns the arithmetic results to the CPU.

[0095] Alternatively, the above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to execute a public key processing method. The public key processing method is for a public key processing device, the public key processing device includes a control processor and an arithmetic acceleration module, the arithmetic acceleration module includes a plurality of arithmetic engines, and the public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing flows, and performs multi-threaded control over the plurality of arithmetic engines based on the generated public key protocol processing flows; the plurality of arithmetic acceleration engines, in response to the multi-threaded control, perform arithmetic processing of operators, obtain arithmetic results for one or more of the public key processing requests, and return the arithmetic results to the control processor; and the control processor receives and returns the arithmetic results to the CPU.

[0096] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0098] The units involved in the embodiments of the present invention may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0099] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0100] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0101] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.

[0102] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present invention. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0103] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A public key processing device, comprising a control processor and an operation acceleration module, characterized in that: The control processor is configured to receive one or more public key processing requests from a central processing unit (CPU), generate a corresponding public key protocol processing flow, and control the operation acceleration module based on the generated public key protocol processing flow; The operation acceleration module is configured to perform arithmetic processing in response to the control of the control processor, obtain an arithmetic result for the one or more public key processing requests, and return the arithmetic result to the control processor; The control processor is further configured to receive and return the encryption arithmetic result to the CPU; Wherein: The operation acceleration module includes a plurality of operation engines, and at least one operation engine is configured to perform arithmetic processing of operators; and The control processor is further configured to perform multi-thread control on the plurality of operation engines for the one or more public key processing requests.

2. The public key processing apparatus according to claim 1, wherein The control processor is further configured to perform multi-thread control on the plurality of operation engines for one or more public key processing requests, including: The control processor calls a plurality of operation engines to perform operator operations based on predefined operators to be processed, and controls the running states of the threads where the plurality of operation engines are located, where the running states include a sleep state and a processing state.

3. The public key processing device according to claim 2, characterized in that The control processor is further configured to: Obtain the arithmetic condition of the operation engine that performs operator arithmetic processing in the thread in the sleep state. If the arithmetic condition of the operation engine is that the arithmetic is completed, wake up the thread in the sleep state to continue controlling the thread to perform public key processing.

4. The public key processing device according to claim 1, characterized in that, The control processor is further configured to perform multi-thread control on the plurality of operation engines for one or more public key processing requests, and further includes: Judge whether there is an operation engine available for executing the current operator processing. If not, put the current thread to sleep to switch to controlling the processing of other threads.

5. The public key processing device according to claim 4, wherein The control processor is further configured to: When an idle and available operation engine is obtained, wake up the sleeping thread or process to control the idle and available operation engine to perform the arithmetic processing of the operator of the thread.

6. The public key processing device according to claim 1, wherein The operation acceleration module further includes: A Hash calculation unit and / or a random number generation unit.

7. The public key processing apparatus according to claim 1, wherein The operation engine includes: A macro instruction processing unit configured to pre-dissemble the macro instruction for executing the operator sent by the control processor into a corresponding plurality of micro instructions; A micro instruction processing unit configured to sequentially retrieve the plurality of micro instructions and send them to the arithmetic logic unit; The arithmetic logic unit is configured to perform arithmetic processing of the corresponding operator based on the received micro instructions to obtain an arithmetic result of the operator; and A register unit configured to store the arithmetic result of the operator.

8. A public key processing method for a public key processing device, the public key processing device comprising a control processor and an arithmetic acceleration module, characterized in that, The operation acceleration module includes a plurality of operation engines, and the public key processing method includes: The control processor receives one or more public key processing requests from the CPU, generates a corresponding public key protocol processing flow, and performs multi-thread control on the plurality of operation engines based on the generated public key protocol processing flow; The multiple operation acceleration engines, in response to the multi-thread control, perform the operation processing of the operator, obtain the operation results for one or more of the public key processing requests, and return the operation results to the control processor; and The control processor receives and returns the operation results to the CPU.

9. The public key processing method according to claim 8, wherein The control processor performs multi-thread control on the multiple operation engines, including: Based on the predefined operators to be processed, the control processor calls multiple operation engines to perform operator operations and controls the running states of the threads where the multiple operation engines are located, where the running states include a sleep state and a processing state.

10. The public key processing method according to claim 9, wherein The control processor performs multi-thread control on the multiple operation engines, and further includes: Obtain the execution state of the operation engine that performs the operator operation processing in the thread in the sleep state. If the execution state of the operation engine is completed, wake up the thread in the sleep state to continue controlling the thread to perform public key processing.

11. The public key processing method according to claim 8, wherein The control processor performs multi-thread control on the multiple operation engines, and further includes: Determine whether there is an operation engine available for executing the current operator processing. If not, put the corresponding thread to sleep to switch to controlling the processing of other threads.

12. The public key processing method according to claim 11, wherein The control processor performs multi-thread control on the multiple operation engines, and further includes: When it is obtained that there is an idle and available operation engine, wake up the thread in the sleep state to control the idle and available operation engine to perform the operator operation of the thread.

13. The public key processing method according to claim 8, wherein, For the operation acceleration module, the public key processing method further includes: Hash calculation or generation of a random number.

14. The public key processing method according to claim 8, wherein The operation engine is used for: Pre-disintegrate the macro instruction for executing the operator sent by the control processor into corresponding multiple micro instructions; Sequentially retrieve the multiple micro instructions, perform the operations of the corresponding operator, and obtain the operation results of the operator; and Store the operation results of the operator.

15. A public key processing system, characterized in that, The public key processing system includes: A central processing unit; and The public key processing device according to any one of claims 1-7.

16. An electronic device, characterized in that, Including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the public key processing method according to any one of claims 8-14.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the public key processing method according to any one of claims 8-14.

Citation Information

Patent Citations

  • Checksum calculation method and network processor

    CN106484503A

  • Scheduling method, device and system in data processing process

    CN108287759A

  • Encryption algorithm module accelerator and data high-speed encryption method

    CN112713993A

  • Data processor proceeding of accelerated synchronization between central processing unit and graphics processing unit

    KR1020180099420A