Public key processing apparatus and method, processing system, electronic device, and storage medium
Through the multi-threaded collaborative control of the control processor and the computing acceleration module, the problem of the CPU load not being offloaded in the data center server is solved, and efficient parallel public key processing is achieved, which improves processing performance and flexibility.
Patent Information
- Application Number
- PCT/CN2024/141031
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-24
AI Technical Summary
In the prior art, in the large-scale deployment of the public key cryptographic system in the data center server, the CPU load cannot be completely offloaded, which affects the efficiency of other user applications. The single-threaded processing solution has low processing performance and cannot meet the high-performance multi-concurrency public key computing needs.
The control processor and operation acceleration module work together, and multiple computing engines are controlled through multi-threading to realize efficient parallel processing of the public key processing process, including sleep state management and thread switching, and supports high-performance computing such as RSA and ECC.
It realizes complete offloading of the CPU public key computing load, improves the concurrency and processing efficiency of the computing engine, and supports high-flexibility and high-performance encryption operations.
Smart Images

Figure CN2024141031_24072025_PF_FP_ABST
Abstract
Description
Public key processing device and method, processing system, electronic device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on January 17, 2024, with application number 202410071037.7 and application name “Public key processing device and method, processing system, electronic device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of public key computing technology, and in particular to a public key processing device and method, a processing system, an electronic device, and a storage medium. Background Art
[0004] Public-key cryptography is widely used in today's information and communication technology (ICT) systems. Its primary applications include key generation, key exchange, encryption and decryption, and signing / verification. Currently, the mainstream public-key cryptography algorithms are divided into two categories: Rivest-Shamir-Adleman (RSA) and elliptic curve cryptography (ECC). Due to the large bit widths (for example, secure RSA requires at least 2 KB of arithmetic bits) and complex modular arithmetic operations introduced by these algorithms, the intensive computing power required has become a major bottleneck for the large-scale deployment of public-key cryptography in data center servers.
[0005] Currently, acceleration technologies for computing power mainly adopt the following two methods:
[0006] One approach involves collaborative processing between the central processing unit (CPU) and a coprocessor (coprocessor). In this approach, the CPU handles tasks such as process concatenation and intermediate result processing, while the coprocessor primarily offloads the computational tasks of modular operations. However, this approach still requires the CPU to handle tasks such as process control, and repeated interaction with the coprocessor requires the CPU to constantly switch processes or handle interrupts. This approach does not completely offload public key data processing and fully free up the CPU's valuable computing resources, thus affecting the CPU's efficiency in executing other user applications.
[0007] Another approach is to use a single-threaded microcontroller unit (MCU) and coprocessor. However, this approach can only process a series of public key calculation requests one by one and serially, resulting in low processing performance. It is only suitable for resource-constrained application scenarios (such as the Internet of Things). However, this approach cannot meet the needs of high-performance, multi-concurrent public key calculations on the server side. Summary of the Invention
[0008] The purpose of this application is to solve one of the above technical problems at least to a certain extent.
[0009] To achieve the above-mentioned purpose, the present application proposes a public key processing device, comprising a control processor and an operation acceleration module; the control processor is configured to receive one or more public key processing requests from the CPU, generate a corresponding public key protocol processing flow, and control the operation acceleration module based on the generated public key protocol processing flow; the operation acceleration module is configured to perform operation processing in response to the control of the control processor, obtain the operation result for the public key processing request, and return the operation result to the controller processor; and the control processor is configured to receive and return the operation result to the CPU; wherein: the operation acceleration module includes multiple operation engines, at least one operation engine is configured to perform operation processing of an operator; and the control processor is further configured to perform multi-threaded control of the multiple operation engines for one or more public key processing requests.
[0010] The control processor is further configured to perform multi-threaded control on the multiple computing engines in response to one or more public key processing requests, including: the control processor calls multiple computing engines to perform operator operations based on pre-defined operators to be processed, and controls the running status of the threads where the multiple computing engines are located, wherein the running status includes a sleep state and a processing state.
[0011] The control processor is further configured to obtain the operation status of the operation engine that executes operator operation processing in the thread in the dormant state, and if the operation status of the operation engine is completed, wake up the thread in the dormant state to continue controlling the thread to perform public key processing.
[0012] The control processor is further configured to perform multi-threaded control on the multiple computing engines in response to one or more public key processing requests, including: determining whether there is a computing engine available to execute the current operator processing; if not, putting the current thread to sleep to switch control of the processing of other threads.
[0013] The control processor is further configured to: upon obtaining that there is an idle and available computing engine, wake up the dormant thread to control the idle and available computing engine to execute computing processing of the operator of the thread.
[0014] The operation acceleration module also includes: a Hash calculation unit and / or a random number generation unit.
[0015] The operation engine includes: a macroinstruction processing unit, configured to pre-disassemble the macroinstructions for executing the operator sent by the control processor into corresponding multiple microinstructions; a microinstruction processing unit, configured to sequentially call the multiple microinstructions and send them to the arithmetic logic unit; the arithmetic logic unit, configured to perform the operation processing of the corresponding operator based on the received microinstructions to obtain the operation result of the operator; and a storage unit, configured to store the operation result of the operator.
[0016] To achieve the above-mentioned purpose, the present application proposes, on the other hand, a public key processing method for a public key processing device, wherein the public key processing device includes a control processor and an operation acceleration module, and the operation acceleration module includes multiple operation engines. The public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing processes, and performs multi-threaded control on the multiple operation engines based on the generated public key protocol processing processes; the multiple operation engines execute the operation processing of the operator in response to the multi-threaded control, obtain the operation results for the one or more public key processing requests, and return the operation results to the control processor; and the control processor receives and returns the operation results to the CPU.
[0017] The control processor performs multi-threaded control on the multiple computing engines, including: the control processor calls the multiple computing engines to perform operator operations based on pre-defined operators to be processed, and controls the running status of the threads where the multiple computing engines are located, wherein the running status includes a sleep state and a processing state.
[0018] The control processor performs multi-threaded control on the multiple computing engines, and also includes: obtaining the execution status of the computing engine that executes operator operation processing in the thread in the dormant state. If the execution status of the computing engine is completed, the dormant thread is awakened to continue controlling the thread to perform public key processing.
[0019] The control processor performs multi-threaded control on the multiple operation engines, and further includes: determining whether there is an operation engine that can be used to execute the current operator processing. If not, the corresponding thread is put to sleep to switch to control the processing of other threads.
[0020] The control processor performs multi-thread control on the multiple computing engines, and further includes: when obtaining an idle and available computing engine, waking up a dormant thread to control the idle and available computing engine to execute the operator operation of the thread.
[0021] For the operation acceleration module, the public key processing method may further include: Hash calculation or random number generation.
[0022] The operation engine is used to: pre-disassemble the macro instruction processing sent by the control processor to execute the operator into corresponding multiple micro instructions; sequentially call the multiple micro instructions, execute the corresponding operator operation, and obtain the operator operation result; and store the operator operation result.
[0023] To achieve the above objectives, the present application proposes, on one hand, a public key processing system, which includes: a central processing unit; and the aforementioned public key processing device.
[0024] To achieve the above-mentioned purpose, the present application proposes an electronic device on the other hand, which includes: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the public key processing method mentioned above in the present application is implemented.
[0025] To achieve the above-mentioned purpose, the present application also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the aforementioned public key processing method.
[0026] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0028] FIG1 is a schematic block diagram of a public key processing apparatus according to an exemplary embodiment;
[0029] FIG2 is a schematic block diagram of a computing engine according to an exemplary embodiment;
[0030] FIG3 is a schematic block diagram of another public key processing apparatus according to an exemplary embodiment;
[0031] FIG4 is a schematic diagram of a digital signature process of an SM2 public key algorithm according to an exemplary embodiment;
[0032] FIG5 is a schematic diagram showing a multi-thread control according to an exemplary embodiment; and
[0033] Fig. 6 is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0034] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0035] The present application proposes a public key processing device, as shown in Figure 1, the public key processing device includes a control processor 10 and an operation acceleration module 20; the control processor 10 is configured to receive one or more public key processing requests from a central processing unit (CPU), generate a corresponding public key protocol processing flow, and control the operation acceleration module 20 based on the generated public key protocol processing flow; the operation acceleration module 20 is configured to respond to the control of the control processor 10, perform operation processing, obtain the operation result for the public key processing request, and return the operation result to the control processor 10; the control processor 10 is also configured to receive and return the operation result to the CPU; wherein: the operation acceleration module 20 includes multiple operation engines 201, at least one operation engine 201 is configured to perform operation processing of an operator; and the control processor 10 is further configured to perform multi-threaded control of the multiple operation engines 201 for one or more public key processing requests.
[0036] For example, as shown in Figure 1, the control processor 10 communicates and interacts with the CPU, and is capable of receiving one or more public key processing requests issued by the CPU and returning the calculation results to the CPU. Specifically, when the control processor 10 receives the public key processing request from the CPU, it first parses the public key processing request to obtain the corresponding public key protocol processing flow. Then, the control processor 10 controls the operation acceleration module 20 to perform operation processing based on the obtained public key protocol processing flow. Specifically, the control processor 10 is used for serial control of the protocol flow, which supports programmability and interacts with the operation acceleration module 20 to issue corresponding control instructions to the operation acceleration module 20 according to the operation operations to be performed in each step of the protocol flow, so that the operation acceleration module 20 performs corresponding operation processing based on different control instructions. Finally, the control processor 10 obtains the final encryption operation result from the operation acceleration module 20 and returns it to the CPU.
[0037] The operation engine 201 is mainly used to perform operator operation processing, such as modular operations, addition, subtraction, multiplication and division and other operations. The operation engine 201 supports programmability and can perform different operator processing for different algorithms. In the embodiment of the present application, the operation acceleration module 20 includes multiple operation engines 201, and the multiple operation engines 201 can interact with the control processor 10 respectively to execute the control instructions of the control processor. Among them, the processing of operators for the same public key processing request can be performed by the same operation engine 201, or it can be completed by multiple operation engines 201 together.
[0038] It can be seen that multiple computing engines 201 can process the computing of the same public key processing request, or can process the computing of multiple different public key processing requests. Therefore, when multiple computing engines 201 process one or more different public key processing requests, they are in different processing threads. Therefore, in order to ensure that multiple computing engines 201 process the operator operations of one or more public key processing requests in parallel, the control processor 10 of the embodiment of the present application can also perform multi-threaded control on multiple computing engines 201 for one or more public key processing requests. For example, the control processor 10 can switch control between multiple processing threads to achieve efficient control of multiple computing engines 201.
[0039] Since an operation engine for performing complex operations is introduced in the embodiment of the present application, and the control processor controls multiple operation engines in multiple threads, the public key processing device in the embodiment of the present application can process public key algorithms such as RSA and ECC that require high-performance operations.
[0040] In this application, the control processor in the public key processing device is responsible for controlling the public key protocol process, and the operation acceleration module is responsible for operation processing. Therefore, when the public key processing device obtains the public key processing request from the CPU, the CPU can immediately process other computing tasks without being responsible for the related tasks of the public key processing. On the one hand, the public key processing device realizes the overall unloading of the public key computing load of the CPU. On the other hand, the control processor and the operation acceleration module effectively cooperate to perform multi-threaded control on multiple computing engines, thereby enabling efficient concurrent processing of multiple computing engines, effectively realizing high-performance and highly flexible encryption computing processing.
[0041] The public key processing apparatus is described in further detail below with reference to the accompanying drawings and specific examples.
[0042] As mentioned in the above embodiments, the control processor of the present application can perform multi-threaded control on multiple computing engines to control the multiple computing engines to process multiple public key processing requests in parallel. The following details the specific content of multi-threaded control with reference to the accompanying drawings and examples.
[0043] In a preferred embodiment, the control processor is further configured to perform multi-threaded control of the multiple computing engines in response to one or more public key processing requests, including: the control processor calls multiple computing engines to execute the operations of the operators based on pre-defined operators to be processed, and controls the running status of the threads where the multiple computing engines are located, wherein the running status includes a sleep state and a current processing state.
[0044] For example, when the control processor controls the computing engine to perform public key processing according to the public key protocol processing flow, the control processor will call multiple computing engines based on the pre-defined operators to be processed, thereby controlling the running state of the threads in which the computing engines are located. The pre-defined here can be understood as the user's pre-analysis of the public key processing flow and the extraction of complex operators with high computing power requirements so that the computing engines can be called for acceleration. For example, the operator currently to be processed is a complex operator, the operation of which requires multiple computing engines to complete, or the processing of the complex operator takes a long time, such as a highly complex point multiplication operation. When the computing engine processes the point multiplication operation, it takes a long time to obtain the operation result (for example, a 256b SM2 may require several computing cycles to complete). In this case, if the existing single-thread control is used, the control processor will need to wait for a long time and cannot process other public key processing requests, resulting in a decrease in the overall public key processing efficiency and the inability to effectively execute multiple different operator operations. For example, when processing a public key request, the public key processing flow has different operation stages. If the subsequent processing flow of the public key processing request depends on the result of the current operator, then the next step of the processing flow must wait until the operation engine completes the operation of the current operator before executing the next step. If the control processor remains in control of this thread at this time, the controller will be unable to process other public key processing requests and control other operation engines.
[0045] Therefore, in response to the above situations, this application adopts multi-threaded control of multiple computing engines, and switches between different threads by controlling the running status of the threads in which they are located. Specifically, this application is described in the following two aspects:
[0046] 1. Processing requests for multiple public keys.
[0047] For different public key processing requests, they correspond to different public key processing threads. When the control processor controls multiple threads, if the operators processed by the currently controlled first thread are complex operators, the calculation takes a long time, and the subsequent processing of the first thread needs to rely on the calculation results of the complex operator. Then the control processor will put the first thread to sleep. The sleep here does not mean that the current thread stops running, but switches to background processing. While the first thread is in the sleep state, the calculation engine will continue to calculate and process the complex operator. In this way, the control processor does not need to stay in the first thread for control, but can switch to other processing threads, such as the second thread, etc., to control other calculation engines to calculate other operators. In this way, the control processor can switch between multiple threads to control multiple calculation engines to achieve efficient processing of multiple public key processing requests.
[0048] Second, a request is processed for a public key.
[0049] For a public key processing request, it corresponds to a general public key processing thread. After the control processor obtains the public key processing request, based on the public key processing protocol, it determines that processing the public key processing request requires multiple computing engines to execute, and the multiple computing engines perform different operator operations. The operation of each operator can be understood as each sub-thread under a general public key processing thread. For example, the computing engine of the first sub-thread performs the operation of the current operator. If there is a need to use the calculation result of the first sub-thread to perform the next calculation operation, then the first sub-thread can be put to sleep and turned to the background to continue running. Furthermore, the calculation of the second sub-thread does not need to rely on the previous calculation result, so the control processor switches to the second sub-thread to control other computing engines to perform the corresponding operator operation. In this way, the control processor can switch between multiple sub-threads and control different computing engines in parallel to perform operator operations, thereby achieving efficient processing of the public key processing request.
[0050] In addition, the above thread running states include not only sleep states but also processing states. Regarding the processing state, for example, when the control processor knows that the operator currently requiring processing is not a complex operator but a simple operator, the controller processor can directly process the simple operator (without calling the calculation engine), thereby maintaining the processing of the current thread without putting the thread to sleep, thereby avoiding unnecessary excessive thread switching.
[0051] More preferably, the control processor is further configured to: obtain the execution status of the calculation engine that performs operator operation processing in the thread in the dormant state; if the execution status of the calculation engine is completed, wake up the thread in the dormant state to continue controlling the thread to perform public key processing.
[0052] For example, after the control processor switches to other threads, it continues to obtain the calculation status of the calculation engine in the thread that has turned to sleep state. Taking the above example, the control processor continues to obtain the processing status of the point multiplication operation of the calculation engine. If the point multiplication operation is completed, the dormant thread can be awakened, that is, the processing thread can be restored and released to control the thread to continue to execute the subsequent calculation steps of the corresponding public key protocol processing flow.
[0053] In a preferred embodiment, the control processor is further configured to perform multi-threaded control on the multiple computing engines in response to one or more public key processing requests, and also includes: determining whether there is a computing engine that can be used to execute the current operator operation. If not, the current thread is put to sleep to switch control of the processing of other threads.
[0054] For example, when the control processor controls multiple computing engines, it will determine which computing engine will perform the operator operation based on the pre-defined operator to be processed and the busy and idle status of each computing engine. When the control processor obtains that multiple computing engines are all busy, that is, multiple computing engines are all performing operator operations, there is currently no available computing engine to perform the operation of the operator to be processed. In this case, if the existing single-threaded control is adopted, the control processor needs to wait in the current thread until there is an idle computing engine available. In this way, the control processor cannot control other computing engines to perform operator operations for other public key processing requests. Based on this, in this application, when it is obtained that there is no available computing engine, the current thread is put to sleep, and then the computing engine that controls other threads is switched. For example, the control processor needs to call a computing engine to perform a point multiplication operation. At this time, the control processor needs to obtain the status of multiple computing engines to determine whether there is an available computing engine to execute. If there is no available computing engine, the control processor will put the current processing thread to sleep and then switch to other processing threads. It should be noted that the "sleep" in the embodiment of the present application can be understood as the current processing thread being in a waiting state, that is, the entire encryption protocol processing flow of the current processing thread is in a waiting state. At this time, the current processing thread interrupts and stops the calculation processing, and the control processor temporarily does not control the current processing thread, but switches to other processing threads for control according to the running status of other processing threads.
[0055] More preferably, the control processor is further configured to: upon detecting the existence of an idle and available computing engine, wake up the dormant thread to control the idle and available computing engine to execute computing processing of the operator of the thread.
[0056] For example, when the control processor switches to another processing thread, it continues to check whether there are idle computing engines, that is, it continues to obtain the status of multiple computing engines. If it is found that a computing engine has completed a computing step of another processing thread and is idle, the CP will wake up the dormant processing thread to resume control of the processing thread, and then call the idle computing engine to continue executing the operator processing required by the processing thread.
[0057] It should be noted that the above is only an example of multi-threaded control in this application, and the multi-threaded control in this application is not limited to the above example. Those skilled in the art can adopt corresponding multi-threaded control methods for other different situations according to actual application needs to achieve concurrent processing of multiple computing engines, and no excessive restrictions are made here.
[0058] It can be seen from the above embodiments that the control processor can realize the switching of multiple processing threads according to the execution status of the operator operation by the operation engine, so that the control processor continuously switches control between multiple threads, and the processing processes of multiple threads are continuously intertwined, thereby making the operation engine at a higher degree of concurrency, realizing efficient utilization of the operation engine, and thus achieving high-performance processing of encryption operations.
[0059] In a preferred embodiment, the operation acceleration module further includes a Hash calculation unit and / or a random number generation unit.
[0060] For example, the digital signature processing flow for the SM2 public key algorithm also includes hash calculation and / or random number generation for secure encryption. Therefore, the computation acceleration module also needs to be configured with a corresponding hash calculation unit and / or random number generation unit to perform the corresponding computations in the public key processing protocol flow. Similarly, the control processor can implement multi-threaded control of multiple computation engines, hash calculation units, and random number generation units.
[0061] It should be noted that this preferred embodiment only uses SM2 as an example to illustrate other functional operation units in the operation acceleration module. The embodiments of this application are not limited to the above two units. Those skilled in the art can configure other functional operation units for public key operations in the operation acceleration module according to the actual processing requirements of the public key algorithm. No excessive restrictions are made here.
[0062] In a preferred embodiment, as shown in Figure 2, the operation engine 201 includes: a macroinstruction processing unit 2011, configured to pre-disassemble the macroinstructions for executing the operator sent by the control processor into corresponding multiple microinstructions; a microinstruction processing unit 2012, configured to sequentially call the multiple microinstructions and send them to the arithmetic logic unit; an arithmetic logic unit 2013, configured to perform the operation processing of the corresponding operator based on the received microinstructions to obtain the operation result of the operator; and a storage unit 2014, configured to store the processing result of the operator.
[0063] For example, after the computing engine 201 obtains the control instruction of the control processor, since the control instruction is a macroinstruction, it needs to be further disassembled into fine-grained microinstructions by the macroinstruction processing unit 2011 to perform calculation processing on complex operators. After being disassembled into multiple microinstructions, the multiple microinstructions are stored in the instruction memory, and the microinstruction processing unit 2012 further calls multiple microinstructions from the instruction memory in sequence and sends them to the arithmetic logic unit 2013 in sequence for calculation processing, and then the arithmetic logic unit 2013 generates the processing results and stores them in the register unit 2014. Specifically, the instruction memory of the computing engine includes multiple macroinstruction program segments, such as Macro_i, Macro_j, etc. in Figure 3, and each macroinstruction program segment includes multiple microinstructions. Taking Macro_i as an example, when the calculation engine is called to process the complex operator corresponding to this macroinstruction, the macroinstruction processing unit 2011 parses the macroinstruction Macro_i and obtains the corresponding starting microinstruction address. Based on this address, the microinstruction processing unit 2012 extracts the microinstructions corresponding to Macro_i one by one and sends them to the arithmetic logic unit 2013 for calculation. When all the microinstructions corresponding to Macro_i are calculated, the calculation results are stored in the register unit 2014, waiting for the control processor to obtain them.
[0064] In an embodiment of the present application, the computing engine can decompose coarse-grained macroinstructions into fine-grained logical operation instruction processing, and can split large-bit-width data into multiple small-bit-width logical operation instruction processing, thereby effectively realizing the operation of complex operators.
[0065] Figure 3 is a schematic block diagram of a more specific public key processing device. As shown in Figure 3, the public key processing device includes a control processor, multiple computing engines, a random number generator, and a cryptographic hash calculation module. The control processor interacts with the CPU, with the CPU sending public key processing requests to the control processor, and the control processor responding to the CPU's public key processing requests. The cryptographic hash calculation module performs hash calculations on the public key processing data sent by the CPU. The random number generator can generate different random numbers to ensure the security of the entire public key processing process. The thread context in the control processor includes multiple processing threads for multiple public key processing requests. Based on these multiple processing threads, the control processor interacts with multiple computing engines, controlling the multiple computing engines to execute the operations of operators in different processing threads. The computing engines receive macroinstructions from the control processor, which include the operators to be performed. The computing engines first pre-split the macroinstructions into multiple microinstruction segments through macroinstruction processing and store the microinstruction segments in the instruction memory. The microinstruction processing unit then calls each microinstruction from the instruction memory one by one and sends them to the arithmetic logic unit for calculation. After the ALU completes the operation of the current microinstruction, it stores the current operation result in the register bank. Then, when the ALU continues to execute the next microinstruction, it can continue the operation based on the previous operation result. After executing all microinstruction operations, the ALU stores the final operation result in the register bank and records the storage address. Then, it waits for the control processor to obtain the final operation result.
[0066] In order to more clearly illustrate the multi-thread control, a more detailed description is given below with reference to FIG. 4 and FIG. 5 .
[0067] Figure 4 is a flow chart of the digital signature generation process using the commercial secret SM2 algorithm. As shown in Figure 4, the control processor receives an SM2_Sign protocol request from the CPU, and the CP begins executing the SM2_Sign protocol. Specifically, the received raw data includes elliptic curve parameters, the information to be signed, and the key, where the order of the elliptic curve is n.
[0068] Furthermore, the control processor is used to control the flow connection between the various computation steps in Figure 4. In the process of Figure 4, step 2 is performed by the cryptographic hash calculation module in Figure 2, step 3 is performed by the random number generator in Figure 3, and steps 4-6 (point multiplication, modular addition, inversion, and modular multiplication) are performed by the computation engine in Figure 3. Depending on the busyness of each computation engine, the control processor can control steps 4-6 in Figure 3 to be processed by the same computation engine or by multiple computation engines.
[0069] If the dot product in step 4 requires a long execution time, such as in a 256-bit SM2, which may take several cycles to complete, the thread in that SM2 goes dormant, and the control processor handles new thread requests. Once step 4 is complete, the control processor wakes up the SM2 thread and continues sending instructions to the corresponding engine to execute step 5.
[0070] Finally, by controlling the processor's control over the calculation engine, the control processor stores the calculation results in the context memory of this thread after each call to the calculation engine, and finally determines the calculation result of the digital signature and sends the determined digital signature to the CPU.
[0071] FIG5 is a schematic diagram of multi-thread control according to an exemplary embodiment, which is directed to a single-issue, coarse-grained multi-thread switching example. As shown in FIG5 , the thread switching process is as follows:
[0072] Taking the SM2 signature process of FIG4 as an example, T1 to TN in FIG5 represent N processing threads, which execute different SM2 signatures in parallel, and each processing thread is in a different running state.
[0073] Furthermore, assuming that processing thread T1 enters the main process of the control processor after passing through the scheduler (ellipse) in this cycle and continues to execute, its program counter pointer (PC pointer) and register group (Reg File) status values need to be restored from Ctx Mem.1. Previously, the status of processing thread T1 is that the calculation of step 3 in Figure 3 is completed, and the PC points to the instruction corresponding to step 4. After extracting and decoding the instruction, the application programming interface (API) of the point multiplication operation ((x_1, y_1) = [k]G) in step 4 of Figure 4 (including elliptic curve system parameters, multiple k, etc., all of which are present in Ctx Mem.1) is sent to one of the available operation engines through interaction through the load-store unit (LSU), and the operation engine is started to perform the operator operation.
[0074] Furthermore, because the dot product operation takes a long time, the control processor needs to perform a thread switch and store the PC pointer and register file status of processing thread T1 in Ctx Mem.1. It should be noted that after the switch, when the control processor controls other processing threads, the dot product operation of processing thread T1 continues to execute. At this time, the calculation engine continues to read and update Ctx Mem.1.
[0075] Assume that processing thread T2 enters the main process of the control processor after passing through the scheduler in this cycle and continues to execute. Its PC pointer and register Reg File status value need to be restored from Ctx Mem.2. The previous state of processing thread T2 is that the calculation of step 6 in Figure 4 is completed. At this time, PC points to the instruction corresponding to step 6 (that is, does it determine whether s = 0?, and the calculation result of processing thread T2 after step 6 is stored in Ctx Mem.2); if s! = 0, processing thread T2 executes to step 7, the signature calculation is completed, and the calculation result (r, s) is stored in Ctx Mem.2. After waiting for the caller of this thread to retrieve it, thread T2 can be released.
[0076] Similarly, when the computing engine performs a long public key calculation, the control processor switches to process other executable threads, and the thread execution process is continuously intertwined to achieve a high degree of concurrency. Especially at the highest performance, all computing engines can work simultaneously.
[0077] In summary, the public key processing device disclosed in this application has the following advantages:
[0078] 1) The public key processing process is completed by the control processor and the computing acceleration module in collaboration. After the CPU sends a public key processing request, it does not need to perform subsequent encryption transactions, effectively achieving complete offloading of the CPU's public key calculation load;
[0079] 2) The control processor adopts multi-threaded control for multiple computing engines, improving the concurrency of the computing engines in processing public key calculations, which can fully utilize the computing power of the computing engines and achieve high flexibility and high performance public key computing processing for various public key algorithms.
[0080] Based on the same inventive concept as the above-mentioned public key processing device, the present application provides a public key processing method for a public key processing device, wherein the public key processing device includes a control processor and an operation acceleration module, and the operation acceleration module includes multiple operation engines. The public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing flows, and performs multi-threaded control on the multiple operation engines based on the generated public key protocol processing flows; the multiple operation engines execute the operation processing of the operator in response to the multi-threaded control, obtain the operation results for the one or more public key processing requests, and return the operation results to the control processor; and the control processor receives and returns the encryption operation results to the CPU.
[0081] The control processor performs multi-threaded control on the multiple computing engines, including: the control processor calls the multiple computing engines to perform operator operations based on pre-defined operators to be processed, and controls the running states of the threads where the multiple computing engines are located, wherein the running states include a dormant state and a processing state.
[0082] The control processor performs multi-threaded control on the multiple computing engines, and also includes: obtaining the execution status of the computing engine that executes operator operation processing in the thread in the dormant state. If the execution status of the computing engine is completed, the dormant thread is awakened to continue controlling the thread to perform public key processing.
[0083] The control processor performs multi-threaded control on the multiple operation engines, and further includes: determining whether there is an operation engine that can be used to execute the current operator processing. If not, the corresponding thread is put to sleep to switch to control the processing of other threads.
[0084] The control processor performs multi-thread control on the multiple computing engines, and further includes: when obtaining an idle and available computing engine, waking up a dormant thread to control the idle and available computing engine to execute the operator operation of the thread.
[0085] For the operation acceleration module, the public key processing method further includes: Hash calculation or generation of random numbers.
[0086] The operation engine is used to: pre-disassemble the macro instructions for executing the operator sent by the control processor into corresponding multiple micro instructions; sequentially call the multiple micro instructions, execute the corresponding operator operations, and obtain the operator operation results; and store the operator operation results.
[0087] Correspondingly, the present application also provides a public key processing system, which includes a central processing unit and the public key processing device described in the above embodiment.
[0088] Accordingly, the present application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the public key processing method described in the above embodiment.
[0089] Accordingly, the present application also provides an electronic device, as shown in FIG6 , which shows a schematic diagram of the structure of an electronic device suitable for implementing an embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG6 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0090] As shown in Figure 6, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device are also stored in RAM 603. The processing device 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0091] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 illustrates an electronic device having various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0092] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present application are performed.
[0093] It should be noted that the computer-readable medium mentioned above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0094] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0095] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0096] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device executes a public key processing method. The public key processing method is used for a public key processing device. The public key processing device includes a control processor and an operation acceleration module. The operation acceleration module includes multiple operation engines. The public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing processes, and performs multi-threaded control on the multiple operation engines based on the generated public key protocol processing processes; the multiple operation acceleration engines, in response to the multi-threaded control, execute the operation processing of the operator, obtain the operation results for one or more of the public key processing requests, and return the operation results to the control processor; and the control processor receives and returns the operation results to the CPU.
[0097] Alternatively, the computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes a public key processing method. The public key processing method is used for a public key processing device. The public key processing device includes a control processor and an operation acceleration module. The operation acceleration module includes multiple operation engines. The public key processing method includes: the control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing flows, and performs multi-threaded control on the multiple operation engines based on the generated public key protocol processing flows; the multiple operation acceleration engines execute operator operation processing in response to the multi-threaded control, obtain operation results for one or more public key processing requests, and return the operation results to the control processor; and the control processor receives and returns the operation results to the CPU.
[0098] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0099] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] The units involved in the embodiments described in this application may be implemented in software or hardware. The name of a unit does not necessarily limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0101] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0102] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
[0104] In addition, although adopting specific order to describe each operation, this should not be interpreted as requiring these operations to be executed in the specific order shown or in sequential order.Under certain environment, multitasking and parallel processing may be advantageous.Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the application.Some features described in the context of separate embodiment can also be implemented in a single embodiment in combination.On the contrary, the various features described in the context of a single embodiment also can be implemented in multiple embodiments individually or in the mode of any suitable subcombination.
[0105] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A public key processing device, comprising a control processor and an operation acceleration module, characterized in that: The control processor is configured to receive one or more public key processing requests from a central processing unit (CPU), generate corresponding public key protocol processing flows, and control the operation acceleration module based on the generated public key protocol processing flows; The operation acceleration module is configured to perform operation processing in response to the control of the control processor, obtain operation results for the one or more public key processing requests, and return the operation results to the control processor; The control processor is further configured to receive and return the encryption operation results to the CPU; Wherein: The operation acceleration module includes a plurality of operation engines, and at least one operation engine is configured to perform operation processing of operators; and The control processor is further configured to perform multi-threaded control on the plurality of operation engines for the one or more public key processing requests.
2. The public key processing apparatus according to claim 1, wherein The control processor is further configured to perform multi-threaded control on the plurality of operation engines for one or more public key processing requests, including: The control processor invokes a plurality of operation engines to perform operator operations based on predefined operators to be processed, and controls the running states of the threads where the plurality of operation engines are located, where the running states include a sleep state and a processing state.
3. The public key processing apparatus according to claim 2, wherein, The control processor is further configured to: Obtain the operation conditions of the operation engines that perform operator operation processing in the threads in the sleep state. If the operation conditions of the operation engines are that the operations are completed, wake up the threads in the sleep state to continue controlling the threads to perform public key processing.
4. The public key processing apparatus according to claim 1, wherein, The control processor is further configured to perform multi-threaded control on the plurality of operation engines for one or more public key processing requests, further including: Determine whether there are operation engines available for performing the current operator processing. If not, put the current thread to sleep to switch to controlling the processing of other threads.
5. The public key processing device according to claim 4, wherein, The control processor is further configured to: When it is obtained that there are idle and available operation engines, wake up the sleeping threads or processes to control the idle and available operation engines to perform the operation processing of the operators of the threads.
6. The public key processing device according to claim 1, wherein, The operation acceleration module further includes: A Hash calculation unit and / or a random number generation unit.
7. The public key processing device according to claim 1, wherein, The operation engine includes: A macro instruction processing unit, configured to pre-dissemble the macro instructions for executing the operators sent by the control processor into corresponding multiple micro instructions; A micro instruction processing unit, configured to sequentially retrieve the multiple micro instructions and send them to the arithmetic logic unit; The arithmetic logic unit, configured to perform operation processing of corresponding operators based on the received micro instructions to obtain operation results of the operators; and A register unit, configured to store the operation results of the operators.
8. A public key processing method for a public key processing device, the public key processing device including a control processor and an arithmetic acceleration module, characterized in that, The operation acceleration module includes a plurality of operation engines, and the public key processing method includes: The control processor receives one or more public key processing requests from the CPU, generates corresponding public key protocol processing flows, and performs multi-threaded control on the plurality of operation engines based on the generated public key protocol processing flows; The multiple operation acceleration engines, in response to the multi-thread control, perform operation processing of operators, obtain operation results for one or more of the public key processing requests, and return the operation results to the control processor; and The control processor receives and returns the operation results to the CPU.
9. The public key processing method according to claim 8, wherein, The control processor performs multi-thread control on the multiple operation engines, including: Based on the pre-defined operators to be processed, the control processor calls multiple operation engines to perform operator operations and controls the running states of the threads where the multiple operation engines are located, where the running states include a sleeping state and a processing state.
10. The public key processing method according to claim 9, wherein, The control processor performs multi-thread control on the multiple operation engines, and further includes: Obtain the execution state of the operation engine that performs operator operation processing in the thread in the sleeping state. If the execution state of the operation engine is completed, wake up the thread in the sleeping state to continue controlling the thread to perform public key processing.
11. The public key processing method according to claim 8, wherein, The control processor performs multi-thread control on the multiple operation engines, and further includes: Determine whether there is an operation engine available for performing the current operator processing. If not, put the corresponding thread to sleep to switch to controlling the processing of other threads.
12. The public key processing method according to claim 11, wherein, The control processor performs multi-thread control on the multiple operation engines, and further includes: When it is obtained that there is an idle and available operation engine, wake up the sleeping thread to control the idle and available operation engine to perform the operator operation of the thread.
13. The public key processing method according to claim 8, wherein, For the operation acceleration module, the public key processing method further includes: Hash calculation or generation of a random number.
14. The public key processing method according to claim 8, wherein, The operation engine is used for: Pre-dismantle the macro instruction for executing the operator sent by the control processor into corresponding multiple micro instructions; Sequentially retrieve the multiple micro instructions, perform the operations of the corresponding operator, and obtain the operation result of the operator; And Store the operation result of the operator.
15. A public key processing system, characterized in that, The public key processing system includes: A central processing unit; and The public key processing device according to any one of claims 1-7.
16. An electronic device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the public key processing method according to any one of claims 8-14.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the public key processing method according to any one of claims 8-14.
Citation Information
Patent Citations
Scheduling method, device and system in data processing process
CN108287759A
Cryptographic operation method, system-on-chip and computer equipment
CN113935018A
Computing engine, data processing method and device and storage medium
CN115934031A
Scheduling system, scheduling method, and recording medium
US20160077882A1