Adaptive Scheduling and Communication Method for Heterogeneous Computing of Hardware Cryptographic Devices Based on VPP

By deploying schedulers and multiple hardware cipher devices in VPP applications, dynamically selecting the target hardware cipher devices for encryption or decryption, the problem of poor cipher operation performance in the prior art is solved, and efficient encryption and decryption operations are achieved.

CN119814480BActive Publication Date: 2025-05-27HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510310937.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In the prior art, during the encryption or decryption process, the overall password computing performance of the system is poor and it is impossible to complete the encryption or decryption operation efficiently.

Method used

Adaptive scheduling and communication methods based on VPP-based hardware cryptographic devices are adopted, and the scheduler receives the to-process data and encryption task types, obtains the scheduling parameter values ​​of each hardware cryptographic device, selects the target hardware cryptographic device, and sends the to-process data to the target hardware cryptographic device for encryption or decryption.

Benefits of technology

It improves the overall cryptographic computing performance of the system, realizes high-performance encryption or decryption, and dynamically and efficiently allocates data packet processing tasks, and takes advantage of the advantages of the heterogeneous computing framework to improve overall computing performance and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814480B_ABST
    Figure CN119814480B_ABST
Patent Text Reader

Abstract

The present application provides an adaptive scheduling and communication method for heterogeneous computing of a hardware cryptographic device based on VPP. The method includes: receiving, by a scheduler, data to be processed sent by a VPP application, and obtaining, by the scheduler, the type of encryption task corresponding to the data to be processed; obtaining, by the scheduler, the scheduling parameter values of each hardware cryptographic device, and selecting a target hardware cryptographic device from multiple hardware cryptographic devices based on the scheduling parameter values of each hardware cryptographic device; sending, by the scheduler, the data to be processed to the target hardware cryptographic device; if the data to be processed is data to be encrypted, encrypting, by the target hardware cryptographic device, the data to be encrypted; and if the data to be processed is data to be decrypted, decrypting, by the target hardware cryptographic device, the data to be decrypted. Through the solution of the present application, the overall cryptographic operation performance of the system is improved, and high-performance encryption or decryption is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technologies, and in particular, to an adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on VPP (Vector Packet Processing). Background Art

[0002] DPDK (Data Plane Development Kit) runs based on an operating system (such as the Linux operating system), which is a set of function libraries and drivers for fast packet processing, improving data processing performance and throughput, and enhancing the working efficiency of data plane applications. DPDK is a high-performance network driver component that provides a convenient, complete, and fast packet processing solution for data plane applications.

[0003] Based on the functions of DPDK, encryption or decryption of data to be processed can be implemented based on DPDK. Since DPDK bypasses the kernel protocol stack of the operating system and can implement a high-performance network packet processing framework, when implementing encryption or decryption based on DPDK, the kernel protocol stack can also be bypassed. However, although DPDK bypasses the kernel protocol stack, DPDK is only a network packet processing framework and does not have the ability to process complex network protocols. Therefore, during the encryption or decryption process, the overall cryptographic operation performance of the system is relatively poor. Summary of the Invention

[0004] This application provides an adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on vector packet processing VPP, which is used to adaptively schedule multiple hardware cryptographic devices. The method includes:

[0005] Receiving, by a scheduler, data to be processed sent by a VPP application, and obtaining, by the scheduler, an encryption task type corresponding to the data to be processed, where the encryption task type includes a cryptographic algorithm type and a data volume level;

[0006] Obtain the scheduling parameter values of each hardware cryptographic device through the scheduler, and select the target hardware cryptographic device from the multiple hardware cryptographic devices based on the scheduling parameter values of each hardware cryptographic device; wherein, the multiple hardware cryptographic devices include GPU hardware cryptographic devices, USB hardware cryptographic devices, and PCIE hardware cryptographic devices; the VPP application communicates with the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device through the OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through the DPDK link; the scheduling parameter value of the GPU hardware cryptographic device is determined based on the encryption task type, the performance parameters of the GPU hardware cryptographic device, the hardware load parameters of the GPU hardware cryptographic device, and the link load parameters of the OpenSSL link; the scheduling parameter value of the USB hardware cryptographic device is determined based on the encryption task type, the performance parameters of the USB hardware cryptographic device, the hardware load parameters of the USB hardware cryptographic device, and the link load parameters of the OpenSSL link; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the encryption task type, the performance parameters of the PCIE hardware cryptographic device, the hardware load parameters of the PCIE hardware cryptographic device, the link load parameters of the OpenSSL link, and the link load parameters of the DPDK link;

[0007] Send the data to be processed to the target hardware cryptographic device through the scheduler;

[0008] Wherein, if the data to be processed is data to be encrypted, encrypt the data to be encrypted through the target hardware cryptographic device; or, if the data to be processed is data to be decrypted, decrypt the data to be decrypted through the target hardware cryptographic device.

[0009] The present application provides an adaptive scheduling and communication device for heterogeneous computing of hardware cryptographic devices based on Vector Packet Processing (VPP), which is used for adaptively scheduling multiple hardware cryptographic devices, including:

[0010] A VPP application module, configured to send the data to be processed and the encryption task type corresponding to the data to be processed to the scheduler; wherein, the encryption task type includes a cryptographic algorithm type and a data volume level;

[0011] A scheduler, configured to receive the data to be processed and the type of encryption task, obtain the scheduling parameter values of each hardware cryptographic device, and select a target hardware cryptographic device from the multiple hardware cryptographic devices based on the scheduling parameter values of each hardware cryptographic device; wherein, the multiple hardware cryptographic devices include a GPU hardware cryptographic device, a USB hardware cryptographic device, and a PCIE hardware cryptographic device; the VPP application communicates with the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device through an OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through a DPDK link; wherein, the scheduling parameter value of the GPU hardware cryptographic device is determined based on the type of encryption task, the performance parameters of the GPU hardware cryptographic device, the hardware load parameters of the GPU hardware cryptographic device, and the link load parameters of the OpenSSL link; the scheduling parameter value of the USB hardware cryptographic device is determined based on the type of encryption task, the performance parameters of the USB hardware cryptographic device, the hardware load parameters of the USB hardware cryptographic device, and the link load parameters of the OpenSSL link; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the type of encryption task, the performance parameters of the PCIE hardware cryptographic device, the hardware load parameters of the PCIE hardware cryptographic device, the link load parameters of the OpenSSL link, and the link load parameters of the DPDK link;

[0012] The scheduler is further configured to send the data to be processed to the target hardware cryptographic device;

[0013] Wherein, if the data to be processed is data to be encrypted, the data to be encrypted is encrypted by the target hardware cryptographic device; or, if the data to be processed is data to be decrypted, the data to be decrypted is decrypted by the target hardware cryptographic device.

[0014] This application provides an electronic device, including a processor and a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement an adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on VPP.

[0015] This application provides a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on VPP.

[0016] The present application provides a machine-readable storage medium storing machine-executable instructions executable by a processor; wherein, the processor is configured to execute the machine-executable instructions to implement the above-mentioned adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on VPP.

[0017] As can be seen from the above technical solutions, in the embodiments of the present application, under the VPP application, by deploying DPDK links and OpenSSL links, it is possible to support GPU hardware cryptographic devices, USB hardware cryptographic devices, and PCIE hardware cryptographic devices simultaneously. The VPP application communicates with the GPU hardware cryptographic device, USB hardware cryptographic device, and PCIE hardware cryptographic device through the OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through the DPDK link. On this basis, by additionally deploying a scheduler, it is possible to select a target hardware cryptographic device from the GPU hardware cryptographic device, USB hardware cryptographic device, and PCIE hardware cryptographic device, and send the data to be processed to the target hardware cryptographic device for encryption or decryption, thereby improving the overall cryptographic operation performance of the system and achieving high-performance encryption or decryption. It is possible to implement heterogeneous computing of hardware cryptographic devices based on VPP, incorporate an adaptive mechanism into the scheduling algorithm, perform task scheduling based on performance parameters and load parameters, dynamically and efficiently allocate packet processing tasks to appropriate hardware cryptographic devices (i.e., target hardware cryptographic devices), and utilize the advantages of the heterogeneous computing framework to improve the overall computing performance and computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on VPP;

[0019] Figure 2 is a schematic diagram of encryption and decryption based on the kernel protocol stack in the present application;

[0020] Figure 3 is a schematic diagram of the association between VPP, DPDK, and the operating system in the present application;

[0021] Figure 4 is a schematic diagram of the encryption and decryption framework based on VPP in the present application;

[0022] Figure 5A is a schematic diagram of the registration of the USB hardware cryptographic device in the present application;

[0023] Figure 5B is a schematic diagram of the registration of the GPU hardware cryptographic device in the present application;

[0024] Figure 5C is a schematic diagram of the registration of the PCIE hardware cryptographic device in the present application;

[0025] Figure 5D It is a schematic diagram of the implementation of the PCIE hardware password device in this application;

[0026] Figure 5E It is a schematic diagram of the interface implementation of VFIO or UIO in this application;

[0027] Figure 6 It is a flowchart of the adaptive scheduling and communication method for heterogeneous computing of the hardware password device based on VPP;

[0028] Figure 7 It is the hardware structure diagram of the electronic device in an implementation manner of this application. Specific implementation manner

[0029] In the embodiments of this application, an adaptive scheduling and communication method for heterogeneous computing of a hardware password device based on VPP (such as an adaptive scheduling and communication implementation method) is proposed, which is used for adaptively scheduling multiple hardware password devices. Refer to Figure 1 As shown, it is a schematic flowchart of this method, and this method may include:

[0030] Step 101: Receive the data to be processed sent by the VPP application through the scheduler, and obtain the encryption task type corresponding to the data to be processed through the scheduler. The encryption task type may include a password algorithm type (used to represent the type of password algorithm) and a data volume level (used to represent the size of the data volume).

[0031] Step 102: Obtain the scheduling parameter values of each hardware password device through the scheduler, and select the target hardware password device from multiple hardware password devices based on the scheduling parameter values of each hardware password device.

[0032] Exemplarily, multiple hardware password devices may include but are not limited to GPU hardware password devices, USB hardware password devices, and PCIE hardware password devices. Based on this, the VPP application can communicate with the GPU hardware password device, USB hardware password device, and PCIE hardware password device respectively through the OpenSSL link, and the VPP application can communicate with the PCIE hardware password device through the DPDK link.

[0033] Exemplarily, the scheduling parameter values of the GPU hardware cryptographic device can be determined based on the encryption task type, the performance parameters of the GPU hardware cryptographic device, the hardware load parameters of the GPU hardware cryptographic device, and the link load parameters of the OpenSSL link. The scheduling parameter values of the USB hardware cryptographic device can be determined based on the encryption task type, the performance parameters of the USB hardware cryptographic device, the hardware load parameters of the USB hardware cryptographic device, and the link load parameters of the OpenSSL link. The scheduling parameter values of the PCIE hardware cryptographic device can be determined based on the encryption task type, the performance parameters of the PCIE hardware cryptographic device, the hardware load parameters of the PCIE hardware cryptographic device, the link load parameters of the OpenSSL link, and the link load parameters of the DPDK link.

[0034] Step 103: Send the data to be processed to the target hardware cryptographic device through the scheduler.

[0035] Exemplarily, if the data to be processed is data to be encrypted, the data to be encrypted can be encrypted by the target hardware cryptographic device; or, if the data to be processed is data to be decrypted, the data to be decrypted can be decrypted by the target hardware cryptographic device.

[0036] Exemplarily, the process of obtaining the scheduling parameter values of the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device can include, but is not limited to:

[0037] Obtain the first initial weight, the second initial weight, and the third initial weight corresponding to the encryption task type from the configured weight table; wherein, the weight table can include the correspondence between the encryption task type and the initial weight; the first initial weight can represent the probability that the GPU hardware cryptographic device is assigned the data to be processed, the second initial weight can represent the probability that the USB hardware cryptographic device is assigned the data to be processed, and the third initial weight can represent the probability that the PCIE hardware cryptographic device is assigned the data to be processed. Determine the first target weight based on the current performance parameters of the GPU hardware cryptographic device for the encryption task type and the first initial weight, and determine the first performance score value based on the first target weight; determine the second target weight based on the current performance parameters of the USB hardware cryptographic device for the encryption task type and the second initial weight, and determine the second performance score value based on the second target weight; determine the third target weight based on the current performance parameters of the PCIE hardware cryptographic device for the encryption task type and the third initial weight, and determine the third performance score value based on the third target weight.

[0038] Determine a first hardware load score value based on the current hardware load parameters of the GPU hardware password device, determine a second hardware load score value based on the current hardware load parameters of the USB hardware password device, and determine a third hardware load score value based on the current hardware load parameters of the PCIE hardware password device.

[0039] Determine an OpenSSL link load score value based on the current link load parameters of the OpenSSL link, and determine a first link load score value assigned to the GPU hardware password device in the OpenSSL link load score value, a second link load score value assigned to the USB hardware password device in the OpenSSL link load score value, and a third link load score value assigned to the PCIE hardware password device in the OpenSSL link load score value. Determine a DPDK link load score value based on the current link load parameters of the DPDK link.

[0040] The scheduling parameter value of the GPU hardware password device can be determined based on the first performance score value, the first hardware load score value, and the first link load score value; the scheduling parameter value of the USB hardware password device can be determined based on the second performance score value, the second hardware load score value, and the second link load score value; the scheduling parameter value of the PCIE hardware password device can be determined based on the third performance score value, the third hardware load score value, the third link load score value, and the DPDK link load score value.

[0041] In a possible implementation, if the password algorithm type in the encryption task type indicates parallel encryption or decryption of the data to be processed, and the data volume level in the encryption task type indicates that the data volume of the data to be processed is not less than the first threshold, then the first initial weight can be greater than the second initial weight, and the first initial weight can be greater than the third initial weight. If the password algorithm type in the encryption task type indicates serial encryption or decryption of the data to be processed, and the data volume level in the encryption task type indicates that the data volume of the data to be processed is not greater than the second threshold, where the second threshold is less than the first threshold, then the second initial weight can be greater than the first initial weight, and the third initial weight can be greater than the first initial weight.

[0042] Exemplarily, determining the first target weight based on the current performance parameters and the first initial weight of the GPU hardware cryptographic device for the encryption task type may include: sending a first request message to the GPU hardware cryptographic device through a scheduler, where the first request message includes the encryption task type, and receiving, through the scheduler, a first response message returned by the GPU hardware cryptographic device, where the first response message includes the current performance parameters; the current performance parameters include the target number of processes, and the target number of processes represents the maximum amount of data that the GPU hardware cryptographic device can process at most within a unit cycle when encrypting or decrypting the data to be processed of the encryption task type; wherein, if there is one GPU hardware cryptographic device, the target number of processes is the maximum number of processes supported by one GPU hardware cryptographic device, and if there are multiple GPU hardware cryptographic devices, the target number of processes is the sum of the maximum number of processes supported by multiple GPU hardware cryptographic devices; if the target number of processes is greater than a third threshold, then increase the first initial weight to obtain the first target weight; if the target number of processes is less than a fourth threshold, and the fourth threshold is less than the third threshold, then decrease the first initial weight to obtain the first target weight; if the target number of processes is not greater than the third threshold and not less than the fourth threshold, then determine the first initial weight as the first target weight.

[0043] Alternatively, input the historical performance data of the GPU hardware cryptographic device into a trained machine learning model to obtain a weight adjustment amount, and adjust the first initial weight based on the weight adjustment amount to obtain an adjusted weight; if the target number of processes is greater than the third threshold, then the adjusted weight can be increased to obtain the first target weight; if the target number of processes is less than the fourth threshold, then the adjusted weight can be decreased to obtain the first target weight; if the target number of processes is not greater than the third threshold and not less than the fourth threshold, then the adjusted weight can be determined as the first target weight; wherein, the historical performance data may at least include the historical performance parameters and the actual amount of processed data of the GPU hardware cryptographic device for the encryption task type in a historical time period.

[0044] Exemplarily, determining the first hardware load score value based on the current hardware load parameters of the GPU hardware cryptographic device may include, but is not limited to: sending a second request message to the GPU hardware cryptographic device through a scheduler, and receiving, through the scheduler, a second response message returned by the GPU hardware cryptographic device, where the second response message may include the current hardware load parameters, and the current hardware load parameters may include, but are not limited to, device utilization and processing latency; wherein, the device utilization may represent the ratio of the actual amount of data processed by the GPU hardware cryptographic device within a unit cycle to the maximum amount of data that the GPU hardware cryptographic device can process within a unit cycle, and the processing latency may represent the duration from the start of data reception by the GPU hardware cryptographic device to the completion of data processing. Then, the first hardware load score value can be determined based on the device utilization and the processing latency.

[0045] Determining the OpenSSL link load score value based on the current link load parameters of the OpenSSL link includes: statistically obtaining the current link load parameters of the OpenSSL link through a scheduler, where the current link load parameters of the OpenSSL link include a first data processing rate, a first queue length, and a first memory occupancy size; the first data processing rate represents the current rate of data transmitted through the OpenSSL link, the first queue length represents the total number of data waiting to be transmitted on the OpenSSL link, and the first memory occupancy size represents the total memory occupancy of the data waiting to be transmitted on the OpenSSL link. Determining the OpenSSL link load score value based on the first data processing rate, the first queue length, and the first memory occupancy size, where the OpenSSL link load score value is directly proportional to the first data processing rate and inversely proportional to the first queue length and the first memory occupancy size.

[0046] Determining the DPDK link load score value based on the current link load parameters of the DPDK link may include, but is not limited to: statistically obtaining the current link load parameters of the DPDK link through a scheduler, where the current link load parameters of the DPDK link include a second data processing rate, a second queue length, and a second memory occupancy size; the second data processing rate represents the current rate of data transmitted through the DPDK link, the second queue length represents the total number of data waiting to be transmitted on the DPDK link, and the second memory occupancy size represents the total memory occupancy of the data waiting to be transmitted on the DPDK link. Determining the DPDK link load score value based on the second data processing rate, the second queue length, and the second memory occupancy size, where the DPDK link load score value is directly proportional to the second data processing rate and inversely proportional to the second queue length and the second memory occupancy size.

[0047] Exemplarily, sending the data to be processed to the target hardware cryptographic device through the scheduler may include, but is not limited to: if the target hardware cryptographic device is a GPU hardware cryptographic device, the data to be processed may be sent to the GPU intermediate firmware through the scheduler based on the OpenSSL link, and the data to be processed may be sent to the GPU hardware cryptographic device through the GPU intermediate firmware. Or, if the target hardware cryptographic device is a USB hardware cryptographic device, the data to be processed may be sent to the USB intermediate firmware through the scheduler based on the OpenSSL link, and the data to be processed may be sent to the USB hardware cryptographic device through the USB intermediate firmware. Or, if the target hardware cryptographic device is a PCIE hardware cryptographic device and the link load parameter of the OpenSSL link is better than that of the DPDK link, the data to be processed may be sent to the PCIE intermediate firmware through the scheduler based on the OpenSSL link, and the data to be processed may be sent to the PCIE hardware cryptographic device through the PCIE intermediate firmware; or, if the link load parameter of the DPDK link is better than that of the OpenSSL link, the data to be processed may be sent to the PCIE intermediate firmware through the scheduler based on the DPDK link, and the data to be processed may be sent to the PCIE hardware cryptographic device through the PCIE intermediate firmware.

[0048] Exemplarily, the OpenSSL link sequentially includes the OpenSSL Crypto API (OpenSSL cryptographic application programming interface), the OpenSSL Engine (OpenSSL engine), and the Kernel Driver (kernel driver of the operating system); the DPDK link sequentially includes the DPDKCryptoDev API (DPDK cryptographic library application programming interface), the CryptoDev PMD (cryptographic library polling mode driver), the driver of the virtual function input / output VFIO or the user space input / output UIO, and the data buffer.

[0049] As can be seen from the above technical solutions, in the embodiments of the present application, under the VPP application, by deploying the DPDK link and the OpenSSL link, it is possible to support GPU hardware cryptographic devices, USB hardware cryptographic devices, and PCIE hardware cryptographic devices at the same time. The VPP application communicates with the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device through the OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through the DPDK link. On this basis, by additionally deploying a scheduler, it is possible to select a target hardware cryptographic device from the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device, and send the data to be processed to the target hardware cryptographic device for encryption or decryption, thereby improving the overall cryptographic operation performance of the system and achieving high-performance encryption or decryption. It is possible to implement heterogeneous computing of hardware cryptographic devices based on VPP, incorporate an adaptive mechanism into the scheduling algorithm, perform task scheduling according to performance parameters and load parameters, and dynamically and efficiently allocate packet processing tasks to appropriate hardware cryptographic devices (i.e., target hardware cryptographic devices), which can utilize the advantages of the heterogeneous computing framework to improve the overall computing performance and computing efficiency.

[0050] The above technical solutions of the embodiments of the present application will be described below in conjunction with specific application scenarios.

[0051] Exemplarily, refer to Figure 2 As shown, it is a schematic diagram of encryption and decryption based on the kernel protocol stack. An application (such as APP1 or APP2, etc.) can call the USB hardware cryptographic device for encryption or decryption through the interface supported by the kernel protocol stack, or can call the PCIE hardware cryptographic device for encryption or decryption.

[0052] For example, the kernel protocol stack (kernel module) may include Driver (driver), Crypto API (cryptographic service provider), Kernel (kernel), and other modules. APP1 calls the USB interface through the Driver in the kernel protocol stack, and then calls the USB hardware cryptographic device for encryption or decryption. Or, APP1 calls the PCIE interface through the Driver in the kernel protocol stack, and then calls the PCIE hardware cryptographic device for encryption or decryption. APP2 calls the USB interface through the Crypto API and Driver in the kernel protocol stack, and then calls the USB hardware cryptographic device for encryption or decryption. Or, APP2 calls the PCIE interface through the Crypto API and Driver in the kernel protocol stack, and then calls the PCIE hardware cryptographic device for encryption or decryption. By calling the hardware cryptographic device, complex cryptographic operations can be realized, thereby reducing the CPU load.

[0053] In the encryption and decryption process based on the kernel protocol stack, it is necessary to implement encryption and decryption based on the kernel protocol stack of the operating system, which has performance issues based on the kernel protocol stack, and the overall cryptographic operation performance of the system is relatively poor.

[0054] Exemplarily, DPDK runs based on the operating system. To solve the performance problem based on the kernel protocol stack, DPDK bypasses the kernel protocol stack of the operating system and implements a high-performance network packet processing framework. However, although DPDK bypasses the kernel protocol stack, DPDK is only a network packet processing framework and does not have the ability to process complex network protocols. In the encryption or decryption process based on DPDK, there is still a problem that the overall cryptographic operation performance of the system is relatively poor, and the encryption or decryption operation cannot be efficiently completed.

[0055] Based on DPDK, VPP realizes a complete high-performance network framework by separating the control plane and the data plane. For example, VPP is an implementation of a network protocol stack. VPP can run in the user space, support multiple packet receiving methods, and can realize framework extensible functions and switching / routing functions. See Figure 3 As shown, it is a schematic diagram of the association between VPP, DPDK, and the operating system (such as the Linux operating system).

[0056] Since the combination of VPP and DPDK can bypass the kernel protocol stack, the network forwarding and processing performance have been greatly improved. However, VPP only implements the interfaces of some hardware devices and cannot be compatible with hardware cryptographic devices (such as GPU hardware cryptographic devices, USB hardware cryptographic devices, PCIE hardware cryptographic devices).

[0057] In view of the above findings, an adaptive scheduling and communication implementation method for heterogeneous computing of hardware cryptographic devices based on VPP is proposed in the embodiments of the present application, which can be applied to any electronic device that supports VPP applications, and the VPP application has an encryption requirement or a decryption requirement for the data to be processed.

[0058] In this embodiment, heterogeneous computing of the hardware cryptographic device based on VPP can be achieved, and an adaptive mechanism is incorporated into the scheduling algorithm to provide optimization strategies for various types of cryptographic operations on VPP, so as to give full play to the computing advantages of various hardware cryptographic devices and improve the overall computing performance of the system. Among them, "adaptive" means that in the process of processing and analysis, the processing method, processing order, processing parameters, boundary conditions or constraint conditions are automatically adjusted according to the data characteristics of the data to be processed, so as to adapt to the statistical distribution characteristics and structural characteristics of the data to be processed, in order to achieve the best processing effect. In addition, heterogeneous computing is a parallel and distributed computing that can coordinately use hardware cryptographic devices with different performances and structures to meet different computing requirements, and can execute in a way to obtain the maximum overall performance. In addition, the hardware cryptographic device is hardware that supports cryptographic algorithms, encapsulates the cryptographic operation process in the hardware, and is called through a software interface in the application. The hardware cryptographic device can reduce the burden on the CPU caused by cryptographic operations and improve the cryptographic operation rate.

[0059] See Figure 4 As shown, it is a schematic diagram of the encryption and decryption framework based on VPP. In the encryption and decryption process based on VPP, multiple hardware cryptographic devices can be adaptively scheduled to complete encryption and decryption, that is, the target hardware cryptographic device is adaptively selected from multiple hardware cryptographic devices, and the target hardware cryptographic device is called to complete encryption and decryption.

[0060] Exemplarily, the multiple hardware cryptographic devices may include, but are not limited to, GPU (Graphic Processing Unit) hardware cryptographic devices, USB (Universal Serial Bus) hardware cryptographic devices, and PCIE (Peripheral Component Interconnect Express) hardware cryptographic devices. There is no restriction on the type of such hardware cryptographic devices. Among them, the GPU hardware cryptographic device may be a hardware cryptographic device implemented based on the GPU. The GPU is a dedicated graphics core processor that can perform parallel computing. The USB hardware cryptographic device may be a hardware cryptographic device implemented based on the USB. The USB interface is a serial bus standard. The PCIE hardware cryptographic device may be a hardware cryptographic device implemented based on the PCIE. The PCIE interface is a high-speed serial computer expansion bus standard.

[0061] Exemplarily, VPP applications have encryption requirements or decryption requirements for the data to be processed. For example, if the data to be processed is data to be encrypted, the VPP application has an encryption requirement for the data to be encrypted, and the VPP application needs to call the hardware cryptographic device to complete the encryption operation on the data to be encrypted. Or, if the data to be processed is data to be decrypted, the VPP application has a decryption requirement for the data to be decrypted, and the VPP application needs to call the hardware cryptographic device to complete the decryption operation on the data to be decrypted.

[0062] See Figure 4 As shown, the VPP application can communicate with the GPU hardware cryptographic device through an OpenSSL (Open Secure Socket Layer) link (the OpenSSL link is a transport link that supports OpenSSL), that is, the VPP application connects to the GPU intermediate firmware (such as GPU Firmware) through the OpenSSL link, and the GPU intermediate firmware is connected to the GPU hardware cryptographic device (such as GPU Crypto Device). The VPP application can communicate with the USB hardware cryptographic device through the OpenSSL link, that is, the VPP application connects to the USB intermediate firmware (such as USB Firmware) through the OpenSSL link, and the USB intermediate firmware is connected to the USB hardware cryptographic device (such as USB Crypto Device). The VPP application can communicate with the PCIE hardware cryptographic device through the OpenSSL link, that is, the VPP application connects to the PCIE intermediate firmware (such as PCIE Firmware) through the OpenSSL link, and the PCIE intermediate firmware is connected to the PCIE hardware cryptographic device (such as PCIE Crypto Device). In addition, the VPP application can communicate with the PCIE hardware cryptographic device through a DPDK link (the DPDK link is a transport link that supports DPDK), that is, the VPP application connects to the PCIE intermediate firmware through the DPDK link, and the PCIE intermediate firmware is connected to the PCIE hardware cryptographic device.

[0063] For example, OpenSSL is a software package for implementing secure communication, consisting of a set of cryptographic function libraries, which protects the confidentiality, integrity, and authentication of data by using cryptographic algorithms, and supports functions such as symmetric encryption, asymmetric encryption, digital signature, and certificate management. DPDK is a high-performance network driver component that provides a simple, convenient, complete, and fast packet processing solution for data plane applications.

[0064] In Figure 4Among them, the OpenSSL link sequentially includes the OpenSSL Crypto API (Open Secure Sockets Layer Cryptography Application Programming Interface, where API stands for Application Programming Interface), the OpenSSL Engine (Open Secure Sockets Layer Engine), and the Kernel Driver of the operating system (such as the Linux operating system). In this way, the VPP application can be connected to the GPU relay firmware, the USB relay firmware, and the PCIE relay firmware through the OpenSSL Crypto API, the OpenSSL Engine, and the Kernel Driver respectively.

[0065] For example, the OpenSSL Crypto API is the Crypto API that supports OpenSSL. The function of the Crypto API is that it is a built-in security suite used to implement functions such as advanced encryption, decryption, certificate management, and digital signature, providing developers with basic tools for handling public key infrastructure, such as operations like certificate management, digital signature, data encryption, and hashing. In this embodiment, during the process of the VPP application sending the data to be processed to the hardware cryptographic device through the OpenSSL link, the data to be processed will pass through the OpenSSL Crypto API, and the OpenSSL Crypto API will perform relevant operations, and no restrictions are imposed on this operation method.

[0066] The OpenSSL Engine is the Engine that supports OpenSSL. The Engine is an extension mechanism of the OpenSSL library that allows users or developers to extend the functions of OpenSSL through custom cryptographic algorithms, encryption algorithms, random number generators, etc. The Engine exists in the form of a shared library (dynamic link library) and can be loaded and used during the runtime of OpenSSL to provide a more flexible and customizable encryption solution. In this embodiment, during the process of the VPP application sending the data to be processed to the hardware cryptographic device through the OpenSSL link, the data to be processed will pass through the OpenSSL Engine, and the OpenSSL Engine will perform relevant operations.

[0067] The Kernel Driver is the kernel driver of the operating system, also known as the kernel driver program. The Kernel Driver is responsible for controlling hardware devices and managing system resources, enabling the operating system to communicate with the hardware. In this embodiment, during the process of the VPP application sending the data to be processed to the hardware cryptographic device through the OpenSSL link, the data to be processed will pass through the Kernel Driver, and the Kernel Driver will perform relevant operations.

[0068] In Figure 4 it, the DPDK link sequentially includes the DPDK CryptoDev API (Data Plane Development Kit Cryptographic Library Application Programming Interface), CryptoDev PMD (Cryptographic Library Poll Mode Driver), VFIO Driver (Virtual Function Input / Output Driver) or UIO Driver (User Space Input / Output Driver), and a data buffer. The data buffer includes a receive data buffer (RX Ring) and a transmit data buffer (TX Ring). In this way, the VPP application is connected to the PCIe relay firmware through the DPDK CryptoDev API, CryptoDev PMD, VFIO Driver or UIO Driver, and the data buffer (RX Ring and TX Ring).

[0069] The DPDK CryptoDev API is the CryptoDev API that supports DPDK. The DPDK CryptoDev API is a software library of DPDK, which provides a unified API for various software and hardware CryptoDev engines with different algorithms, hiding the highly optimized implementation details of various Crypto from users. The DPDK CryptoDev has a unified implementation of asymmetric enqueue and dequeue to ensure the optimization of the efficiency of hardware Crypto operations. In this embodiment, during the process of the VPP application sending the data to be processed to the hardware cryptographic device through the DPDK link, the data to be processed will pass through the DPDK CryptoDev API, and the DPDK CryptoDev API will perform related operations.

[0070] CryptoDev PMD (Poll Mode Driver, also known as the poll mode driver program) is the poll mode driver of the cryptographic library. CryptoDev PMD consists of APIs provided by the driver running in the user space and can configure the device and its respective queues. CryptoDev PMD directly accesses the RX descriptor (RX Ring) and the TX descriptor (TX Ring) without generating interrupts (except for link state change interrupts), which can ensure the fast reception, processing, and transmission of data packets in the user space application. In this embodiment, during the process of the VPP application sending the data to be processed to the hardware cryptographic device through the DPDK link, the data to be processed will pass through CryptoDev PMD, and CryptoDev PMD will perform related operations.

[0071] The VFIO Driver (Virtual Function Input Output Driver) or UIO Driver (Userspace Input Output Driver) is used to complete the operation and configuration of devices in the user space and is the basis for DPDK to implement PMD. The UIO Driver is a user-mode driver framework provided by the Linux kernel that can run device drivers in the user space without recompiling the kernel. Only a small kernel module needs to be maintained in the kernel. The VFIO Driver is a user-mode driver framework that can securely expose device I / O, interrupts, etc. to the user space. User-space processes can directly use the VFIO Driver to access hardware, thus completing the device driver framework in the user space. In this embodiment, during the process of the VPP application sending the data to be processed to the hardware cryptographic device through the DPDK link, the data to be processed will pass through the VFIO Driver or UIO Driver, and the VFIO Driver or UIO Driver will perform related operations.

[0072] Regarding the RX Ring and TX Ring of the data buffer, the TX Ring is the transmit data buffer. During the process of the VPP application sending the data to be processed to the hardware cryptographic device through the DPDK link, the data to be processed can be cached in the TX Ring. In addition, the RX Ring is the receive data buffer. During the process of the hardware cryptographic device sending the processed data (i.e., the data after encryption or decryption of the data to be processed) to the VPP application through the DPDK link, the processed data can be cached in the RX Ring.

[0073] In Figure 4 considering that CryptoDev PMD, VFIO Driver or UIO Driver are applicable to PCIE but not to GPU and USB, therefore, the VPP application can communicate with the PCIE hardware cryptographic device through the DPDK link. However, the VPP application does not communicate with the GPU hardware cryptographic device through the DPDK link, and the VPP application does not communicate with the USB hardware cryptographic device through the DPDK link.

[0074] In a possible implementation, to support Figure 4 the framework shown, enabling the VPP application to communicate with the hardware cryptographic device through the OpenSSL link or DPDK link, the following method can be adopted:

[0075] 1. OpenSSL Crypto API for the OpenSSL link. To implement the OpenSSL Crypto API for the OpenSSL link so that the VPP application can communicate with the hardware cryptographic device through the OpenSSL link, the hardware cryptographic device can be registered in the OpenSSL Engine to achieve the conversion between the standard interface and the hardware interface.

[0076] For the registration process of the USB hardware cryptographic device, see Figure 5A As shown, it is a schematic diagram of the registration of the USB hardware cryptographic device. Figure 5A The hardware is a USB hardware cryptographic device, which includes a USB host controller and a USB device. The USB device includes a hardware cryptographic module and a USB device controller.

[0077] The kernel space can include a USB controller driver, a USB device driver, USBFS (USB File System), a USB core, and a USB node ( / dev / usb). For example, USBFS is a virtual file system used for communication between the user space and the kernel space. It can abstract the USB device as a file in the file system, enabling users to access the USB device by reading and writing files. The USB node is a device node created for the USB device, and this device node is a special file located in the / dev directory. This device node allows user space programs to communicate with the USB device. The user space can include a user-mode driver, an OpenSSL Engine (Open Secure Sockets Layer Engine), and an application.

[0078] Through the cooperation among the USB hardware cryptographic device, the kernel space, and the user space, the USB hardware cryptographic device can be registered in the OpenSSL Engine to achieve the conversion between the standard interface and the hardware interface. Subsequently, the VPP application can communicate with the USB hardware cryptographic device through the OpenSSL link. There is no restriction on this implementation process. It is sufficient that the VPP application can communicate with the USB hardware cryptographic device through the OpenSSL link.

[0079] In Figure 5A the USB device controller and the hardware cryptographic module included in the USB device can complete the functions of the hardware part, receive the encryption and decryption instructions sent from the kernel space, complete the USB protocol decoding and encryption and decryption operations, and then send the calculation results through the USB protocol. The kernel space can complete the registration of the USB hardware cryptographic device, and the OpenSSL Engine is responsible for completing the protocol conversion of data writing. Finally, in the VPP application, this interface is called to communicate with the USB hardware cryptographic device through the OpenSSL link.

[0080] For the registration process of the GPU hardware cryptographic device, see Figure 5B As shown, it is a schematic diagram of the registration of the GPU hardware cryptographic device. Figure 5B The hardware is a GPU hardware cryptographic device. The GPU hardware cryptographic device includes a PCIE bridge and a GPU device, and the GPU device is connected to the PCIE bridge through a PCIE bus.

[0081] The kernel space may include a Root Complex, a PCIE device driver, a GPU kernel-mode driver, and a CPU. For example, the Root Complex is one of the core components of the PCIE bus architecture, responsible for managing external I / O devices. The CPU can be connected to the PCIE bus through the Root Complex and ultimately connected to all PCIE devices. For example, the Root Complex is used to manage external I / O devices, the Root Complex is used to initialize devices, and the Root Complex is used for resource allocation. The user space may include a GPU user-mode driver, an OpenSSL Engine, and an application.

[0082] Through the cooperation among the GPU hardware cryptographic device, the kernel space, and the user space, the GPU hardware cryptographic device can be registered in the OpenSSL Engine, realizing the conversion between the standard interface and the hardware interface. Subsequently, the VPP application can communicate with the GPU hardware cryptographic device through the OpenSSL link. There is no limitation on this implementation process, as long as the VPP application can communicate with the GPU hardware cryptographic device through the OpenSSL link.

[0083] In Figure 5B the GPU hardware cryptographic device contains a GPU arithmetic unit (inside the GPU device, not shown in the figure). By configuring it in the GPU user-mode driver, it can be configured with relevant functions for cryptographic operations and follows and implements the same Engine interface as the USB hardware cryptographic device.

[0084] For the registration process of the PCIE hardware cryptographic device, see Figure 5C As shown, it is a schematic diagram of the registration of the PCIE hardware cryptographic device. Figure 5CThe hardware is a PCIE hardware cryptographic device, which includes a PCIE bridge and a PCIE device, and the PCIE device is connected to the PCIE bridge through a PCIE bus. The kernel space may include a Root Complex, a PCIE device driver, a Crypto driver (i.e., an encryption driver, such as a Linux-based Crypto driver), and a CPU. The user space may include a Crypto user-mode driver, an OpenSSL Engine, and an application.

[0085] Through the cooperation among the PCIE hardware cryptographic device, the kernel space, and the user space, the PCIE hardware cryptographic device can be registered in the OpenSSL Engine to implement the conversion between the standard interface and the hardware interface, and then enable the VPP application to communicate with the PCIE hardware cryptographic device through the OpenSSL link. There is no limitation on this implementation process, as long as the VPP application can communicate with the PCIE hardware cryptographic device through the OpenSSL link.

[0086] 2. The DPDK CryptoDev API for the DPDK link.

[0087] For the PCIE hardware cryptographic device, in addition to the implementation method using the OpenSSL Engine (see Figure 5C shown), the kernel-mode driver of the PCIE hardware cryptographic device can also implement the VFIO or UIO interface and implement the custom data interface of the DPDK CryptoDev API for calling in the polling program of the CryptoDev PMD. On this basis, the DPDK CryptoDev API for the DPDK link can be implemented to enable the VPP application to communicate with the PCIE hardware cryptographic device through the DPDK link.

[0088] For the implementation process of the PCIE hardware cryptographic device, see Figure 5D shown, which is a schematic diagram of the implementation of the PCIE hardware cryptographic device. Figure 5DThe hardware is a PCIE hardware cryptographic device, and the PCIE hardware cryptographic device may include a PCIE bridge, a PCIE device, a UIO kernel-mode driver, and DMA (Direct Memory Access). The PCIE device and the PCIE bridge may be connected through a PCIE bus. The DMA and the PCIE bridge may be connected through a PCIE bus, and the DMA and the PCIE device may be directly connected. The kernel space may include a UIO kernel-mode driver and a VFIO kernel-mode driver. The user space may include a UIO user-mode driver, a VFIO user-mode driver, VPP Crypto (i.e., the VPP encryption component), DPDK CryptoDev (i.e., the DPDK CryptoDev API, i.e., the DPDK encryption library application programming interface), and a VPP application.

[0089] Through the cooperation among the PCIE hardware cryptographic device, the kernel space, and the user space, the DPDK CryptoDev API of the DPDK link can be implemented, and then the VPP application can communicate with the PCIE hardware cryptographic device through the DPDK link (i.e., the communication process is implemented using Figure 5D the framework). There is no limitation on this implementation process, as long as the VPP application can communicate with the PCIE hardware cryptographic device through the DPDK link.

[0090] For the PCIE hardware cryptographic device, the kernel-mode driver of the PCIE hardware cryptographic device can also implement the interfaces of VFIO or UIO. See Figure 5E as shown, which is a schematic diagram of the interface implementation of VFIO or UIO.

[0091] Create a Crypto Dev (encryption library) device in the VPP application through the rte_cryptodev_pmd_create (runtime environment encryption library polling mode driver creation) function. The Crypto Dev device needs to implement the relevant interfaces of enqueue_burst (batch enqueue) and dequeue_burst (batch dequeue).

[0092] The Crypto Dev device receives data from the RX Ring (receive data buffer) and sends data through the TX Ring (send data buffer), thus completing the interaction with the PCIE hardware cryptographic device. In the RX Ring, the first response status message can be called the head response status message (RSP Head), and so on. The last response status message can be called the tail response status message (RSP Tail). In the TX Ring, the first data request message can be called the head data request message (REQ Head), and so on. The last data request message can be called the tail data request message (REQ Tail).

[0093] In the PCIE hardware cryptographic device, it includes a PCIE Core, hardware cores or FPGA software cores of several cryptographic algorithms, such as an Asymmetric Cipher Core, a Symmetric Cipher Core, a Hasher Cipher Core, etc. In addition, through the Encap (encapsulation protocol) core and the Decap (decapsulation protocol) core, the encapsulation and decapsulation of the data format are completed to realize the data interaction between the cryptographic algorithm core and the VPP application. At the same time, the Encap / Decap core can complete the data format encapsulation / decapsulation that conforms to the OpenSSL Engine interface.

[0094] Based on the above encryption and decryption framework based on VPP, in the embodiments of this application, as shown in Figure 4 it can also deploy a Scheduler additionally. The Scheduler is located between the VPP application and the OpenSSL link, and the Scheduler is located between the VPP application and the DPDK link. The Scheduler can be a hardware-implemented Scheduler or a software-implemented Scheduler, and there is no limitation on this implementation method. In this way, when the VPP application communicates with hardware cryptographic devices such as GPUs, USBs, and PCIEs through the OpenSSL link, it needs to pass through the Scheduler first. When the VPP application communicates with the PCIE hardware cryptographic device through the DPDK link, it needs to pass through the Scheduler first.

[0095] By implementing an additional Scheduler (i.e., an adaptive heterogeneous task scheduler) in VPP, it is possible to achieve task adaptive scheduling and load balancing of hardware cryptographic devices, and allocate tasks to the most suitable hardware cryptographic device (subsequently called the target hardware cryptographic device) according to the type of encryption task, thereby improving the computing efficiency.

[0096] For example, based on factors such as the encryption task type (such as the cryptographic algorithm type and the data volume level, and the data volume level can also be referred to as the cryptographic operation scale), the computing latency requirement, and the load condition, a target hardware cryptographic device is selected from all hardware cryptographic devices, and the task is assigned to the target hardware cryptographic device.

[0097] By implementing an additional scheduler in VPP and using a scheduling algorithm to allocate tasks, the scheduling algorithm can perform task scheduling from perspectives such as performance and load condition, so as to balance performance and load and achieve adaptive scheduling of the heterogeneous computing framework, that is, an adaptive mechanism can be incorporated into the scheduling algorithm.

[0098] In an embodiment of the present application, an adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on VPP is proposed. Refer to Figure 6 As shown, it is a schematic flowchart of the method. The method may include:

[0099] Step 601: The VPP application sends the data to be processed and the corresponding encryption task type of the data to be processed to the scheduler. The data to be processed is the data to be encrypted or the data to be decrypted. The encryption task type may include the cryptographic algorithm type (used to represent the type of cryptographic algorithm) and the data volume level (used to represent the size of the data volume).

[0100] For example, the cryptographic algorithm type may be a parallel cryptographic algorithm or a serial cryptographic algorithm. The parallel cryptographic algorithm means that the hardware cryptographic device can perform parallel encryption or parallel decryption on multiple data to be processed, and the serial cryptographic algorithm means that the hardware cryptographic device performs serial encryption or serial decryption on multiple data to be processed.

[0101] For example, the data volume level may be the first level, the second level, or the third level (of course, there may be more data volume levels). The first level means that the data volume of the data to be processed is not less than the first threshold, that is, the data volume of the data to be processed is a large data volume and there is a large amount of data to be processed. The third level means that the data volume of the data to be processed is not greater than the second threshold, and the second threshold is less than the first threshold, that is, the data volume of the data to be processed is a small data volume and there is a small amount of data to be processed. The second level means that the data volume of the data to be processed is between the second threshold and the first threshold, that is, the data volume of the data to be processed is a medium data volume.

[0102] Step 602: The scheduler receives the data to be processed and the encryption task type sent by the VPP application.

[0103] Step 603: The scheduler obtains the current performance parameters of the GPU hardware cryptographic device for this encryption task type, the scheduler obtains the current performance parameters of the USB hardware cryptographic device for this encryption task type, and the scheduler obtains the current performance parameters of the PCIE hardware cryptographic device for this encryption task type.

[0104] Exemplarily, the scheduler sends a first request message to the GPU hardware cryptographic device, and the first request message includes the type of the encryption task. After receiving the first request message, the GPU hardware cryptographic device obtains the current performance parameter of this GPU hardware cryptographic device for the type of the encryption task, and returns a first response message to the scheduler, and the first response message includes the current performance parameter. The scheduler receives the first response message returned by the GPU hardware cryptographic device, and obtains the current performance parameter from the first response message.

[0105] For example, the current performance parameter may include, but is not limited to, the target number of processing times, and the target number of processing times represents the maximum amount of data that the GPU hardware cryptographic device can process per unit cycle when encrypting or decrypting the data to be processed of the type of the encryption task, such as the number of processing times per second (the unit cycle is 1 second) of the cryptographic operation. The target number of processing times is a theoretical performance index of the GPU hardware cryptographic device, representing the total processing capacity of the GPU hardware cryptographic device. For example, assuming that the target number of processing times of the type of the encryption task is 100, it means that the GPU hardware cryptographic device can process at most 100 times of data per second when encrypting or decrypting the data to be processed of the type of the encryption task, and the target number of processing times is related to the capacity of the GPU hardware cryptographic device.

[0106] When the GPU hardware cryptographic device encrypts or decrypts the data to be processed of different types of encryption tasks, the target number of processing times may be different. Therefore, the GPU hardware cryptographic device can obtain the type of the encryption task from the first request message, and then send the target number of processing times corresponding to the type of the encryption task to the scheduler.

[0107] For example, if there is one GPU hardware cryptographic device, the GPU hardware cryptographic device is the smallest hardware entity for implementing encryption and decryption, has encryption and decryption resources, and can complete encryption and decryption. Then the target number of processing times is the maximum number of processing times supported by one GPU hardware cryptographic device. The maximum number of processing times supported by the GPU hardware cryptographic device is a theoretical performance index, related to the capacity of the GPU hardware cryptographic device, and is a fixed value. Thus, the maximum number of processing times supported by the GPU hardware cryptographic device can be obtained, and the target number of processing times can be known.

[0108] If there are multiple GPU hardware cryptographic devices, the target number of processing times is the sum of the maximum number of processing times supported by all GPU hardware cryptographic devices. That is, the maximum number of processing times supported by each GPU hardware cryptographic device can be obtained, and then the sum of these maximum number of processing times is calculated to obtain the target number of processing times.

[0109] Exemplarily, the scheduler sends a first request message to the USB hardware cryptographic device, and the first request message includes the encryption task type. After receiving the first request message, the USB hardware cryptographic device obtains the current performance parameters of this USB hardware cryptographic device for the encryption task type, and returns a first response message to the scheduler, and the first response message includes the current performance parameters. The scheduler receives the first response message returned by the USB hardware cryptographic device, and obtains the current performance parameters from the first response message.

[0110] Exemplarily, the scheduler sends a first request message to the PCIE hardware cryptographic device, and the first request message includes the encryption task type. After receiving the first request message, the PCIE hardware cryptographic device obtains the current performance parameters of this PCIE hardware cryptographic device for the encryption task type, and returns a first response message to the scheduler, and the first response message includes the current performance parameters. The scheduler receives the first response message returned by the PCIE hardware cryptographic device, and obtains the current performance parameters from the first response message.

[0111] Step 604, the scheduler obtains the current hardware load parameters of the GPU hardware cryptographic device, the current hardware load parameters of the USB hardware cryptographic device, and the current hardware load parameters of the PCIE hardware cryptographic device.

[0112] Exemplarily, the scheduler sends a second request message to the GPU hardware cryptographic device. After receiving the second request message, the GPU hardware cryptographic device obtains the current hardware load parameters of this GPU hardware cryptographic device, and returns a second response message to the scheduler, and the second response message includes the current hardware load parameters of the GPU hardware cryptographic device. The scheduler receives the second response message returned by the GPU hardware cryptographic device, and obtains the current hardware load parameters of the GPU hardware cryptographic device from the second response message.

[0113] For example, the current hardware load parameters may include device utilization and processing latency. In addition to device utilization and processing latency, the current hardware load parameters may also include data such as throughput and power consumption.

[0114] The device utilization rate can represent the ratio of the actual amount of data processed by the GPU hardware password device in a unit cycle to the maximum amount of data that the GPU hardware password device can process in a unit cycle. Taking one second as an example of the unit cycle, if the actual amount of data processed by the GPU hardware password device in one second is 200, it means that the GPU hardware password device processes 200 data to be processed in one second (which can be the average value over a period of time). The maximum amount of data that can be processed is the total processing capacity of the GPU hardware password device. Assuming that the maximum amount of data that the GPU hardware password device can process in one second is 500, it means that the GPU hardware password device can process at most 500 data to be processed in one second. On this basis, the device utilization rate of the GPU hardware password device can be 40%.

[0115] The processing delay can represent the duration from the start of data reception to the completion of data processing by the GPU hardware password device. For example, the GPU hardware password device receives data to be processed at time t1, and this data to be processed is completed at time t2. Then the start time of data reception is time t1, the completion time of data processing is time t2, and the processing delay of the GPU hardware password device is t2 - t1. The average value of the processing delay over a period of time can be statistically calculated, and this average value of the processing delay is used as the processing delay of the GPU hardware password device.

[0116] Exemplarily, the scheduler sends a second request message to the USB hardware password device. After receiving the second request message, the USB hardware password device obtains the current hardware load parameters of this USB hardware password device and returns a second response message to the scheduler, and this second response message includes the current hardware load parameters of the USB hardware password device. The scheduler receives this second response message returned by the USB hardware password device and obtains the current hardware load parameters of the USB hardware password device from this second response message.

[0117] Exemplarily, the scheduler sends a second request message to the PCIE hardware password device. After receiving the second request message, the PCIE hardware password device obtains the current hardware load parameters of this PCIE hardware password device and returns a second response message to the scheduler, and this second response message includes the current hardware load parameters of the PCIE hardware password device. The scheduler receives this second response message returned by the PCIE hardware password device and obtains the current hardware load parameters of the PCIE hardware password device from this second response message.

[0118] Step 605, the scheduler statistically calculates the current link load parameters of the OpenSSL link.

[0119] Exemplarily, when the VPP application sends the data to be processed to the hardware cryptographic device through the OpenSSL link, the data to be processed needs to pass through the scheduler. Therefore, the scheduler can count the current link load parameters of the OpenSSL link. For example, the current link load parameters of the OpenSSL link can include, but are not limited to, the first data processing rate, the first queue length, and the first memory occupancy size, and there is no limitation thereto.

[0120] For example, the first data processing rate (the rate at which the OpenSSL link processes data packets) represents the current rate of the data that has been transmitted through the OpenSSL link, that is, the current data transmission rate of the OpenSSL link, indicating how much data is transmitted per second through the OpenSSL link currently. The average value of the data transmission rate over a period of time can be counted, and the average value of the data transmission rate can be used as the first data processing rate.

[0121] For example, the first queue length (that is, the length of the data packet queue waiting to be processed) represents the total number of data waiting to be transmitted on the OpenSSL link. For example, when transmitting data through the OpenSSL link, the data is first stored in the queue, and each data in the queue is read in sequence for transmission on the OpenSSL link. In this way, the total number of remaining data in the queue can be counted, and this total number of data is used as the first queue length.

[0122] For example, the first memory occupancy size (that is, the memory occupancy situation) represents the total memory occupancy size of the data waiting to be transmitted on the OpenSSL link. For example, when transmitting data through the OpenSSL link, it is necessary to cache the data in a specified storage medium (such as memory), and each data in the memory is read in sequence for transmission on the OpenSSL link. The memory size occupied by the data can be counted, and this memory size is used as the first memory occupancy size.

[0123] Step 606, the scheduler counts the current link load parameters of the DPDK link.

[0124] When the VPP application sends the data to be processed to the hardware cryptographic device through the DPDK link, the data to be processed needs to pass through the scheduler. Therefore, the scheduler can count the current link load parameters of the DPDK link. For example, the current link load parameters of the DPDK link can include, but are not limited to, the second data processing rate, the second queue length, and the second memory occupancy size. The second data processing rate represents the current rate of the data that has been transmitted through the DPDK link, the second queue length represents the total number of data waiting to be transmitted on the DPDK link, and the second memory occupancy size represents the total memory occupancy size of the data waiting to be transmitted on the DPDK link.

[0125] Step 607: The scheduler obtains the first initial weight, the second initial weight, and the third initial weight corresponding to the encryption task type from the configured weight table. The first initial weight represents the probability that the GPU hardware password device is assigned the data to be processed. The second initial weight represents the probability that the USB hardware password device is assigned the data to be processed. The third initial weight represents the probability that the PCIE hardware password device is assigned the data to be processed.

[0126] Exemplarily, the scheduler may pre-configure a weight table, and the weight table may include the correspondence between the encryption task type and the initial weight. For example, as shown in Table 1, it is an example of the weight table.

[0127] Table 1

[0128]

[0129] For each encryption task type in the weight table, the initial weight corresponding to the encryption task type can be configured according to experience, and there is no limitation on this. For example, the sum of all initial weights can be a fixed value (such as 1). For example, initial weight a1 + initial weight a2 + initial weight a3 = 1.

[0130] In a possible implementation, for encryption tasks of the large-scale parallel cryptographic algorithm type (such as CTR), the data to be processed is preferentially assigned to the GPU hardware password device. For encryption tasks of the small-scale serial cryptographic algorithm type, the data to be processed is preferentially assigned to the USB hardware password device or the PCIE hardware password device. Based on the above idea, if the cryptographic algorithm type in encryption task type a indicates that the data to be processed is encrypted or decrypted in parallel, that is, the cryptographic algorithm type is a parallel cryptographic algorithm, and the data volume level in encryption task type a indicates that the data volume of the data to be processed is not less than the first threshold, that is, the data volume level is the first level (indicating a large data volume), then the initial weight a1 can be greater than the initial weight a2, and the initial weight a1 can be greater than the initial weight a3. If the cryptographic algorithm type in encryption task type b indicates that the data to be processed is encrypted or decrypted serially, that is, the cryptographic algorithm type is a serial cryptographic algorithm, and the data volume level in encryption task type b indicates that the data volume of the data to be processed is not greater than the second threshold, that is, the data volume level is the third level (indicating a small data volume), then the initial weight b2 can be greater than the initial weight b1, and the initial weight b3 can be greater than the initial weight b1. For encryption task types in other situations, such as the cryptographic algorithm type being a parallel cryptographic algorithm, the data volume level being the second level or the third level, the cryptographic algorithm type being a serial cryptographic algorithm, and the data volume level being the first level or the second level, the initial weights can be configured arbitrarily.

[0131] Exemplarily, after obtaining the encryption task type corresponding to the data to be processed, the scheduler can query the weight table shown in Table 1 through the encryption task type to obtain the first initial weight, the second initial weight, and the third initial weight corresponding to the encryption task type. Obviously, if the cipher algorithm type in the encryption task type indicates that the data to be processed is encrypted or decrypted in parallel, and the data volume level in the encryption task type indicates that the data volume of the data to be processed is not less than the first threshold, then the first initial weight is greater than the second initial weight, and the first initial weight is greater than the third initial weight, so that the data to be processed is preferentially allocated to the GPU hardware cipher device to make full use of the parallel encryption task processing advantage of the GPU hardware cipher device. Or, if the cipher algorithm type in the encryption task type indicates that the data to be processed is encrypted or decrypted serially, and the data volume level in the encryption task type indicates that the data volume of the data to be processed is not greater than the second threshold, then the second initial weight is greater than the first initial weight, and the third initial weight is greater than the first initial weight, so that the data to be processed is preferentially allocated to the USB hardware cipher device or the PCIE hardware cipher device.

[0132] Step 608: The scheduler determines a first target weight based on the current performance parameter of the GPU hardware cipher device for the encryption task type and the first initial weight, and determines a first performance score value based on the first target weight. The scheduler determines a second target weight based on the current performance parameter of the USB hardware cipher device for the encryption task type and the second initial weight, and determines a second performance score value based on the second target weight. The scheduler determines a third target weight based on the current performance parameter of the PCIE hardware cipher device for the encryption task type and the third initial weight, and determines a third performance score value based on the third target weight.

[0133] In a possible implementation, the current performance parameters of the GPU hardware password device for this type of encryption task may include the target number of processing times. If the target number of processing times is greater than the third threshold (which can be configured according to experience, such as the sum of the maximum number of processing times supported by three GPU hardware password devices), it indicates that the performance of the GPU hardware password device has improved, the processing efficiency (computing volume / computing time) of the GPU hardware password device in the current period has been significantly improved, and there may be situations such as expansion of the GPU hardware password device (i.e., an increase in the number of GPU hardware password devices). Therefore, the first initial weight is increased to obtain the first target weight, and the first target weight represents the probability that the GPU hardware password device is assigned data to be processed, that is, the probability that the GPU hardware password device is assigned data to be processed increases. In addition, if the target number of processing times is less than the fourth threshold (configured according to experience, such as the sum of the maximum number of processing times supported by two GPU hardware password devices), and the fourth threshold is less than the third threshold, it indicates that the performance of the GPU hardware password device has deteriorated, the processing efficiency (computing volume / computing time) of the GPU hardware password device in the current period has decreased, and there may be situations such as contraction of the GPU hardware password device (i.e., a decrease in the number of GPU hardware password devices). Therefore, the first initial weight is decreased to obtain the first target weight, that is, the probability that the GPU hardware password device is assigned data to be processed decreases.

[0134] If the target number of processing times is not greater than the third threshold and not less than the fourth threshold, it indicates that the performance of the GPU hardware password device remains unchanged, and the processing efficiency in the current period remains unchanged. Therefore, the first initial weight is determined as the first target weight, that is, the probability that the GPU hardware password device is assigned data to be processed remains unchanged.

[0135] After obtaining the first target weight, the first target weight can be a value between 0 and 1. The scheduler can use the first target weight as the first performance score value, or can use a certain algorithm to process the first target weight to obtain the first performance score value, and the first performance score value is a value between 0 and 1.

[0136] In a possible implementation, the scheduler can pre-train a machine learning model (such as a deep learning model or a neural network model, etc.). The input data of the machine learning model is the historical performance data of the GPU hardware password device, and the output data of the machine learning model is the weight adjustment amount. There is no restriction on the network structure of this machine learning model, and there is no restriction on the training process of this machine learning model. In this way, the first initial weight can be dynamically adjusted by analyzing the historical performance data through the machine learning model.

[0137] The scheduler can obtain the historical performance data of the GPU hardware password device, such as the historical performance data of the GPU hardware password device over a past period of time. This historical performance data may include the historical performance parameters and the actual processed data volume of the GPU hardware password device for this encryption task type during a historical time period.

[0138] For example, the historical performance parameters may include the historical processing times, and the historical processing times indicate the maximum amount of data that the GPU hardware password device can process within a unit cycle when encrypting or decrypting the data to be processed of this encryption task type, such as the number of times processed per second. The historical processing times are similar to the target processing times. The difference between the two is that the historical processing times can represent the processing times during a historical time period, while the target processing times can represent the processing times during the current time period. Obviously, when the number of GPU hardware password devices increases or decreases, the historical processing times and the target processing times may be different.

[0139] The actual processed data volume represents the processed data volume of the GPU hardware password device for this encryption task type during a historical time period, such as how much data is processed per second, which can be the average value during the historical time period.

[0140] On this basis, the historical performance data of the GPU hardware password device can be input into a machine learning model to obtain a weight adjustment amount. There is no limit to the processing process of this machine learning model, as long as the machine learning model can output a weight adjustment amount. The weight adjustment amount can be a positive value. For example, when the historical performance data indicates better performance and the actual processed data volume indicates less data, the GPU hardware password device can process more data. Therefore, the weight adjustment amount can be a positive value. Or, the weight adjustment amount can be a negative value. For example, when the historical performance data indicates poor performance and the actual processed data volume indicates more data, the GPU hardware password device cannot meet the data processing requirements. Therefore, the weight adjustment amount can be a negative value. Or, the weight adjustment amount can be 0. For example, when the historical performance data indicates moderate performance and the actual processed data volume indicates moderate data, the GPU hardware password device can just process the current data volume. Therefore, the weight adjustment amount can be 0.

[0141] After obtaining the weight adjustment amount, the first initial weight can be adjusted based on the weight adjustment amount to obtain an adjusted weight. For example, the adjusted weight can be the sum of the first initial weight and the weight adjustment amount.

[0142] Since the current performance parameter includes the target number of processes, if the target number of processes is greater than the third threshold, the adjusted weight is increased to obtain the first target weight. Alternatively, if the target number of processes is less than the fourth threshold, the adjusted weight is decreased to obtain the first target weight. Alternatively, if the target number of processes is not greater than the third threshold and not less than the fourth threshold, the adjusted weight is determined as the first target weight. After obtaining the first target weight, the scheduler can use the first target weight as the first performance score value, or can process the first target weight using a certain algorithm to obtain the first performance score value, and there is no limitation on this.

[0143] In summary, the scheduler can obtain the first performance score value corresponding to the GPU hardware password device. Similarly, the scheduler can also obtain the second performance score value corresponding to the USB hardware password device and the third performance score value corresponding to the PCIE hardware password device, which will not be repeated here. Obviously, the scheduler can dynamically adjust the target weights of different hardware password devices according to real-time performance data, and will automatically increase the target weight of a certain hardware password device in task allocation, and allocate more appropriate tasks to this hardware password device.

[0144] Step 609: The scheduler determines the first hardware load score value based on the current hardware load parameters of the GPU hardware password device, determines the second hardware load score value based on the current hardware load parameters of the USB hardware password device, and determines the third hardware load score value based on the current hardware load parameters of the PCIE hardware password device.

[0145] For example, the current hardware load parameters of the GPU hardware password device include device utilization and processing latency, and the first hardware load score value can be determined based on the device utilization and the processing latency.

[0146] For example, the first hardware load score value is inversely proportional to the device utilization. The smaller the device utilization, the smaller the load of the GPU hardware password device, the larger the first hardware load score value, and the greater the possibility of the GPU hardware password device being the target hardware password device. The first hardware load score value is inversely proportional to the processing latency. The smaller the processing latency, the smaller the load of the GPU hardware password device, the larger the first hardware load score value, and the greater the possibility of the GPU hardware password device being the target hardware password device.

[0147] There is no limitation in this embodiment on how to determine the first hardware load score value based on the device utilization and the processing latency. For example, a weighted operation can be performed on the device utilization and the processing latency to obtain the first hardware load score value, as long as the first hardware load score value is inversely proportional to the device utilization and the processing latency.

[0148] In summary, the scheduler can obtain the first hardware load score value corresponding to the GPU hardware password device. Similarly, the scheduler can also obtain the second hardware load score value corresponding to the USB hardware password device and the third hardware load score value corresponding to the PCIE hardware password device, which will not be repeated here.

[0149] Step 610: The scheduler determines the OpenSSL link load score value based on the current link load parameters of the OpenSSL link. The scheduler determines the first link load score value assigned to the GPU hardware password device in the OpenSSL link load score value. The scheduler determines the second link load score value assigned to the USB hardware password device in the OpenSSL link load score value. And the scheduler determines the third link load score value assigned to the PCIE hardware password device in the OpenSSL link load score value.

[0150] For example, the current link load parameters of the OpenSSL link may include the first data processing rate, the first queue length, and the first memory occupancy size. In this way, the scheduler can determine the OpenSSL link load score value based on the first data processing rate, the first queue length, and the first memory occupancy size.

[0151] For example, the OpenSSL link load score value is proportional to the first data processing rate. The larger the first data processing rate, the smaller the load of the OpenSSL link, the larger the OpenSSL link load score value, and the greater the possibility of transmitting the data to be processed through the OpenSSL link. The OpenSSL link load score value is inversely proportional to the first queue length. The larger the first queue length, the greater the load of the OpenSSL link, the smaller the OpenSSL link load score value, and the smaller the possibility of transmitting the data to be processed through the OpenSSL link. The OpenSSL link load score value is inversely proportional to the first memory occupancy size. The larger the first memory occupancy size, the greater the load of the OpenSSL link, the smaller the OpenSSL link load score value, and the smaller the possibility of transmitting the data to be processed through the OpenSSL link.

[0152] Exemplarily, the first allocation weight of the GPU hardware password device, the second allocation weight of the USB hardware password device, and the third allocation weight of the PCIE hardware password device can be pre-configured. There is no limit to this allocation weight, as long as the sum of the first allocation weight, the second allocation weight, and the third allocation weight is 1.

[0153] For example, the first allocation weight, the second allocation weight, and the third allocation weight are all 1 / 3. Another example is that the first allocation weight is 0.5, and the second allocation weight and the third allocation weight are both 0.25.

[0154] Exemplarily, the product value of the OpenSSL link load fraction value and the first allocation weight is used as the first link load fraction value of the GPU hardware cryptographic device. The product value of the OpenSSL link load fraction value and the second allocation weight is used as the second link load fraction value of the USB hardware cryptographic device. The product value of the OpenSSL link load fraction value and the third allocation weight is used as the third link load fraction value of the PCIE hardware cryptographic device.

[0155] Step 611: The scheduler determines the DPDK link load fraction value based on the current link load parameters of the DPDK link. For example, the current link load parameters of the DPDK link may include the second data processing rate, the second queue length, and the second memory occupancy size. In this way, the scheduler can determine the DPDK link load fraction value based on the second data processing rate, the second queue length, and the second memory occupancy size.

[0156] For example, the DPDK link load fraction value may be directly proportional to the second data processing rate, that is, the greater the second data processing rate, the greater the DPDK link load fraction value. The DPDK link load fraction value may be inversely proportional to the second queue length, that is, the greater the second queue length, the smaller the DPDK link load fraction value. The DPDK link load fraction value may be inversely proportional to the second memory occupancy size, that is, the greater the second memory occupancy size, the smaller the DPDK link load fraction value.

[0157] Step 612: The scheduler obtains the scheduling parameter values of each hardware cryptographic device.

[0158] Exemplarily, the scheduler can determine the scheduling parameter value of the GPU hardware cryptographic device based on the first performance score value of the GPU hardware cryptographic device, the first hardware load fraction value of the GPU hardware cryptographic device, and the first link load fraction value of the GPU hardware cryptographic device. Obviously, the scheduling parameter value is directly proportional to the first performance score value, the scheduling parameter value is directly proportional to the first hardware load fraction value, and the scheduling parameter value is directly proportional to the first link load fraction value. For example, a weighted operation is performed on the first performance score value, the first hardware load fraction value, and the first link load fraction value to obtain the scheduling parameter value of the GPU hardware cryptographic device.

[0159] The scheduler can determine the scheduling parameter value of the USB hardware password device based on the second performance score value of the USB hardware password device, the second hardware load score value of the USB hardware password device, and the second link load score value of the USB hardware password device. Obviously, the scheduling parameter value is directly proportional to the second performance score value, the scheduling parameter value is directly proportional to the second hardware load score value, and the scheduling parameter value is directly proportional to the second link load score value. For example, a weighted operation is performed on the second performance score value, the second hardware load score value, and the second link load score value to obtain the scheduling parameter value of the USB hardware password device.

[0160] The scheduler can determine the scheduling parameter value of the PCIE hardware password device based on the third performance score value of the PCIE hardware password device, the third hardware load score value of the PCIE hardware password device, the third link load score value of the PCIE hardware password device, and the DPDK link load score value. Obviously, the scheduling parameter value is directly proportional to the third performance score value, the scheduling parameter value is directly proportional to the third hardware load score value, the scheduling parameter value is directly proportional to the third link load score value, and the scheduling parameter value is directly proportional to the DPDK link load score value. For example, a weighted operation is performed on the third performance score value, the third hardware load score value, the third link load score value, and the DPDK link load score value to obtain the scheduling parameter value of the PCIE hardware password device.

[0161] Step 613: Based on the scheduling parameter values of each hardware password device, the scheduler selects a target hardware password device from multiple hardware password devices. For example, the scheduler determines the hardware password device with the largest scheduling parameter value and uses the hardware password device with the largest scheduling parameter value as the target hardware password device.

[0162] Step 614: The scheduler sends the data to be processed to the target hardware password device.

[0163] Exemplarily, if the target hardware password device is a GPU hardware password device, the scheduler can send the data to be processed to the GPU transfer firmware based on the OpenSSL link. There is no restriction on the transmission process of this OpenSSL link, and the data to be processed is sent to the GPU hardware password device through the GPU transfer firmware.

[0164] Exemplarily, if the target hardware password device is a USB hardware password device, the scheduler can send the data to be processed to the USB transfer firmware based on the OpenSSL link. There is no restriction on the transmission process of this OpenSSL link, and the data to be processed is sent to the USB hardware password device through the USB transfer firmware.

[0165] Exemplarily, if the target hardware cryptographic device is a PCIE hardware cryptographic device and the link load parameter of the OpenSSL link is better than that of the DPDK link (i.e., the link load parameter of the OpenSSL link is greater than that of the DPDK link), the scheduler can send the data to be processed to the PCIE relay firmware based on the OpenSSL link, without restricting the transmission process of the OpenSSL link, and send the data to be processed to the PCIE hardware cryptographic device through the PCIE relay firmware. Or,

[0166] if the link load parameter of the DPDK link is better than that of the OpenSSL link (i.e., the link load parameter of the DPDK link is greater than that of the OpenSSL link), the scheduler can send the data to be processed to the PCIE relay firmware based on the DPDK link, without restricting the transmission process of the DPDK link, and send the data to be processed to the PCIE hardware cryptographic device through the PCIE relay firmware.

[0167] Exemplarily, after the target hardware cryptographic device receives the data to be processed, if the data to be processed is data to be encrypted, the target hardware cryptographic device encrypts the data to be encrypted and returns the encryption result (i.e., the encrypted data) to the VPP application, without restricting the transmission process.

[0168] Or, after the target hardware cryptographic device receives the data to be processed, if the data to be processed is data to be decrypted, the target hardware cryptographic device decrypts the data to be decrypted and returns the decoding result (i.e., the decrypted data) to the VPP application, without restricting the transmission process.

[0169] As can be seen from the above technical solutions, in the embodiments of the present application, under the VPP application, by deploying DPDK links and OpenSSL links, it is possible to support GPU hardware cryptographic devices, USB hardware cryptographic devices, and PCIE hardware cryptographic devices simultaneously. By additionally deploying a scheduler, the overall cryptographic operation performance of the system can be improved, and high-performance encryption or decryption can be achieved. It is possible to implement heterogeneous computing of hardware cryptographic devices based on VPP, incorporate an adaptive mechanism into the scheduling algorithm, and perform task scheduling according to performance parameters and load parameters, dynamically and efficiently distributing packet processing tasks to the target hardware cryptographic devices. The advantages of the heterogeneous computing framework can be utilized to improve the overall computing performance and computing efficiency. It is possible to achieve intelligent load balancing based on heterogeneous computing, implement an adaptive heterogeneous computing mechanism in VPP, and dynamically and efficiently distribute packet processing tasks to the target hardware cryptographic devices according to the type of encryption task, data characteristics, and load conditions, improving the computing efficiency. Optimize resource allocation according to the computing advantages of each hardware cryptographic device and the real-time load conditions of each hardware cryptographic device and core, thereby improving the overall system performance. It is possible to support multiple hardware cryptographic devices, such as GPU hardware cryptographic devices, USB hardware cryptographic devices, and PCIE hardware cryptographic devices, and support allowing users to select the most suitable hardware cryptographic device according to their own needs and hardware resources, improving flexibility. By using different types of hardware cryptographic devices, the performance and cost can be better balanced. For example, GPU hardware cryptographic devices are suitable for processing a large number of parallel encryption tasks, and USB hardware cryptographic devices provide convenient plug-and-play functions. When multiple hardware cryptographic devices are available, it is possible to ensure that encryption tasks are evenly distributed and avoid overloading a single node. Perform task scheduling according to performance and load conditions, and monitor changes in the system operation status in real time. For example, when the type of encryption task changes, from mainly serial tasks to mainly large-scale parallel tasks, the scheduling algorithm can automatically sense and adjust the task allocation strategy, complete device communication according to different types of hardware cryptographic devices, and utilize the advantages of the heterogeneous computing framework to improve the overall computing performance of the system.

[0170] Based on the same application concept as the above method, an adaptive scheduling and communication device for heterogeneous computing of hardware cryptographic devices based on VPP is proposed in an embodiment of the present application, which is used for adaptively scheduling multiple hardware cryptographic devices. The device includes: a VPP application module, configured to send data to be processed and the encryption task type corresponding to the data to be processed to a scheduler; wherein, the encryption task type includes a cryptographic algorithm type and a data volume level; the scheduler is configured to receive the data to be processed and the encryption task type, obtain the scheduling parameter values of each hardware cryptographic device, and select a target hardware cryptographic device from the multiple hardware cryptographic devices based on the scheduling parameter values of each hardware cryptographic device; wherein, the multiple hardware cryptographic devices include a GPU hardware cryptographic device, a USB hardware cryptographic device, and a PCIE hardware cryptographic device; the VPP application communicates with the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device through an OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through a DPDK link; wherein, the scheduling parameter value of the GPU hardware cryptographic device is determined based on the encryption task type, the performance parameter of the GPU hardware cryptographic device, the hardware load parameter of the GPU hardware cryptographic device, and the link load parameter of the OpenSSL link; the scheduling parameter value of the USB hardware cryptographic device is determined based on the encryption task type, the performance parameter of the USB hardware cryptographic device, the hardware load parameter of the USB hardware cryptographic device, and the link load parameter of the OpenSSL link; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the encryption task type, the performance parameter of the PCIE hardware cryptographic device, the hardware load parameter of the PCIE hardware cryptographic device, the link load parameter of the OpenSSL link, and the link load parameter of the DPDK link;

[0171] The scheduler is further configured to send the data to be processed to the target hardware cryptographic device;

[0172] Wherein, if the data to be processed is data to be encrypted, the data to be encrypted is encrypted by the target hardware cryptographic device; or, if the data to be processed is data to be decrypted, the data to be decrypted is decrypted by the target hardware cryptographic device.

[0173] Exemplarily, when the scheduler obtains the scheduling parameter values of the GPU hardware password device, the USB hardware password device, and the PCIE hardware password device, it specifically is used for: obtaining the first initial weight, the second initial weight, and the third initial weight corresponding to the encryption task type from the configured weight table; wherein, the weight table includes the correspondence between the encryption task type and the initial weight; the first initial weight represents the probability that the GPU hardware password device is allocated with data to be processed, the second initial weight represents the probability that the USB hardware password device is allocated with data to be processed, and the third initial weight represents the probability that the PCIE hardware password device is allocated with data to be processed;

[0174] Determine a first target weight based on the current performance parameter of the GPU hardware password device for the encryption task type and the first initial weight, and determine a first performance score value based on the first target weight; determine a second target weight based on the current performance parameter of the USB hardware password device for the encryption task type and the second initial weight, and determine a second performance score value based on the second target weight; determine a third target weight based on the current performance parameter of the PCIE hardware password device for the encryption task type and the third initial weight, and determine a third performance score value based on the third target weight;

[0175] Determine a first hardware load score value based on the current hardware load parameter of the GPU hardware password device, determine a second hardware load score value based on the current hardware load parameter of the USB hardware password device, and determine a third hardware load score value based on the current hardware load parameter of the PCIE hardware password device;

[0176] Determine an OpenSSL link load score value based on the current link load parameter of the OpenSSL link, and determine a first link load score value allocated to the GPU hardware password device, a second link load score value allocated to the USB hardware password device, and a third link load score value allocated to the PCIE hardware password device in the OpenSSL link load score value;

[0177] Determine a DPDK link load score value based on the current link load parameter of the DPDK link;

[0178] Determine the scheduling parameter value of the GPU hardware password device based on the first performance score value, the first hardware load score value, and the first link load score value; determine the scheduling parameter value of the USB hardware password device based on the second performance score value, the second hardware load score value, and the second link load score value; determine the scheduling parameter value of the PCIE hardware password device based on the third performance score value, the third hardware load score value, the third link load score value, and the DPDK link load score value.

[0179] Based on the same application concept as the above method, an embodiment of the present application proposes an electronic device, as shown in Figure 7 shown, including: a processor 71 and a machine-readable storage medium 72, the machine-readable storage medium 72 stores machine-executable instructions that can be executed by the processor 71; the processor 71 is used to execute the machine-executable instructions to implement an adaptive scheduling and communication method for heterogeneous computing of hardware password devices based on VPP.

[0180] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which several computer instructions are stored. When the computer instructions are executed by a processor, an adaptive scheduling and communication method for heterogeneous computing of hardware password devices based on VPP can be implemented.

[0181] Among them, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device, and can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Radom Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drive (such as a hard disk drive), solid-state drive, any type of storage disk (such as an optical disk, DVD, etc.), or a similar storage medium, or a combination thereof.

[0182] Based on the same application concept as the above method, an embodiment of the present application also provides a computer program product, the computer program product may include a computer program, and when the computer program is executed by a processor, an adaptive scheduling and communication method for heterogeneous computing of hardware password devices based on VPP is implemented.

[0183] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0184] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An adaptive scheduling and communication method for heterogeneous computing of hardware cryptographic devices based on vector packet processing (VPP), which is used to adaptively schedule multiple hardware cryptographic devices, characterized in that: The method comprises: Receiving the data to be processed sent by the VPP application through the scheduler, and obtaining the encryption task type corresponding to the data to be processed through the scheduler, wherein the encryption task type includes the cryptographic algorithm type and the data volume level; Obtaining the scheduling parameter value of each hardware cryptographic device through the scheduler, and selecting a target hardware cryptographic device from the multiple hardware cryptographic devices based on the scheduling parameter value of each hardware cryptographic device; wherein the multiple hardware cryptographic devices include a GPU hardware cryptographic device, a USB hardware cryptographic device and a PCIE hardware cryptographic device; the VPP application communicates with the GPU hardware cryptographic device, the USB hardware cryptographic device and the PCIE hardware cryptographic device through an OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through a DPDK link; the scheduling parameter value of the GPU hardware cryptographic device is determined based on the encryption task type, the performance parameter of the GPU hardware cryptographic device, the hardware load parameter of the GPU hardware cryptographic device and the link load parameter of the OpenSSL link; the scheduling parameter value of the USB hardware cryptographic device is determined based on the encryption task type, the performance parameter of the USB hardware cryptographic device, the hardware load parameter of the USB hardware cryptographic device and the link load parameter of the OpenSSL link; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the encryption task type, the performance parameter of the PCIE hardware cryptographic device, the hardware load parameter of the PCIE hardware cryptographic device, the link load parameter of the OpenSSL link and the link load parameter of the DPDK link; Sending the data to be processed to the target hardware cryptographic device through the scheduler; Among them, if the data to be processed is data to be encrypted, the data to be encrypted is encrypted by the target hardware cryptographic device; or, if the data to be processed is data to be decrypted, the data to be decrypted is decrypted by the target hardware cryptographic device.

2. The method according to claim 1, characterized in that The process of obtaining the scheduling parameter value of the GPU hardware cryptographic device, the scheduling parameter value of the USB hardware cryptographic device, and the scheduling parameter value of the PCIE hardware cryptographic device specifically includes: Obtaining a first initial weight, a second initial weight, and a third initial weight corresponding to the encryption task type from a configured weight table; wherein the weight table includes a correspondence between the encryption task type and the initial weight; the first initial weight represents a probability that the GPU hardware cryptographic device is assigned data to be processed, the second initial weight represents a probability that the USB hardware cryptographic device is assigned data to be processed, and the third initial weight represents a probability that the PCIE hardware cryptographic device is assigned data to be processed; Determine a first target weight based on the current performance parameters of the GPU hardware cryptographic device for the encryption task type and the first initial weight, and determine a first performance score value based on the first target weight; determine a second target weight based on the current performance parameters of the USB hardware cryptographic device for the encryption task type and the second initial weight, and determine a second performance score value based on the second target weight; determine a third target weight based on the current performance parameters of the PCIE hardware cryptographic device for the encryption task type and the third initial weight, and determine a third performance score value based on the third target weight; Determine a first hardware load score value based on the current hardware load parameters of the GPU hardware cryptographic device, determine a second hardware load score value based on the current hardware load parameters of the USB hardware cryptographic device, and determine a third hardware load score value based on the current hardware load parameters of the PCIE hardware cryptographic device; Determine an OpenSSL link load score value based on a current link load parameter of the OpenSSL link, and determine a first link load score value allocated to the GPU hardware cryptographic device, a second link load score value allocated to the USB hardware cryptographic device, and a third link load score value allocated to the PCIE hardware cryptographic device among the OpenSSL link load score values; Determining a DPDK link load score value based on a current link load parameter of the DPDK link; The scheduling parameter value of the GPU hardware cryptographic device is determined based on the first performance score value, the first hardware load score value and the first link load score value; the scheduling parameter value of the USB hardware cryptographic device is determined based on the second performance score value, the second hardware load score value and the second link load score value; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the third performance score value, the third hardware load score value, the third link load score value and the DPDK link load score value.

3. The method according to claim 2, characterized in that If the cryptographic algorithm type in the encryption task type indicates that the data to be processed is encrypted or decrypted in parallel, and the data amount level in the encryption task type indicates that the amount of the data to be processed is not less than a first threshold, then the first initial weight is greater than the second initial weight, and the first initial weight is greater than the third initial weight; If the cryptographic algorithm type in the encryption task type indicates that the data to be processed is encrypted or decrypted serially, and the data amount level in the encryption task type indicates that the data amount of the data to be processed is not greater than the second threshold, and the second threshold is less than the first threshold, then the second initial weight is greater than the first initial weight, and the third initial weight is greater than the first initial weight.

4. The method according to claim 2, characterized in that: The determining of the first target weight based on the current performance parameter of the GPU hardware cryptographic device for the encryption task type and the first initial weight specifically includes: Sending a first request message to the GPU hardware cryptographic device through a scheduler, wherein the first request message includes the encryption task type; receiving a first response message returned by the GPU hardware cryptographic device through the scheduler, wherein the first response message includes the current performance parameter; the current performance parameter includes a target processing number, wherein the target processing number indicates the maximum amount of data processed within a unit cycle when the GPU hardware cryptographic device encrypts or decrypts the to-be-processed data of the encryption task type; wherein, if there is one GPU hardware cryptographic device, the target processing number is the maximum processing number supported by one GPU hardware cryptographic device; and if there are multiple GPU hardware cryptographic devices, the target processing number is the sum of the maximum processing numbers supported by multiple GPU hardware cryptographic devices; If the target processing times is greater than a third threshold, the first initial weight is increased to obtain a first target weight; if the target processing times is less than a fourth threshold, and the fourth threshold is less than the third threshold, the first initial weight is reduced to obtain a first target weight; if the target processing times is not greater than the third threshold and not less than the fourth threshold, the first initial weight is determined as the first target weight; Alternatively, the historical performance data of the GPU hardware cryptographic device is input into a trained machine learning model to obtain a weight adjustment amount, and the first initial weight is adjusted based on the weight adjustment amount to obtain an adjusted weight; if the target processing number is greater than a third threshold, the adjusted weight is increased to obtain the first target weight; if the target processing number is less than a fourth threshold, the adjusted weight is reduced to obtain the first target weight; if the target processing number is not greater than the third threshold and not less than the fourth threshold, the adjusted weight is determined as the first target weight; The historical performance data at least includes historical performance parameters and actual processed data volume of the GPU hardware cryptographic device for the encryption task type in a historical time period.

5. The method according to claim 2, characterized in that: The determining of the first hardware load score value based on the current hardware load parameter of the GPU hardware cryptographic device includes: Sending a second request message to the GPU hardware cryptographic device through a scheduler, and receiving a second response message returned by the GPU hardware cryptographic device through the scheduler, wherein the second response message includes the current hardware load parameter, and the current hardware load parameter includes device utilization and processing delay; wherein the device utilization represents the ratio of the actual amount of data processed by the GPU hardware cryptographic device in a unit cycle to the maximum amount of data processed by the GPU hardware cryptographic device in a unit cycle, and the processing delay represents the time from the start of data reception to the completion of data processing by the GPU hardware cryptographic device; The first hardware load score value is determined based on the device utilization and the processing delay.

6. The method according to claim 2, characterized in that The determining of the OpenSSL link load score value based on the current link load parameter of the OpenSSL link comprises: The current link load parameters of the OpenSSL link are counted by the scheduler, and the current link load parameters of the OpenSSL link include a first data processing rate, a first queue length, and a first memory occupancy size; the first data processing rate represents the current rate of data transmitted through the OpenSSL link, the first queue length represents the total amount of data waiting to be transmitted on the OpenSSL link, and the first memory occupancy size represents the total memory occupancy size of data waiting to be transmitted on the OpenSSL link; Determine the OpenSSL link load score value based on the first data processing rate, the first queue length, and the first memory occupancy size, wherein the OpenSSL link load score value is proportional to the first data processing rate and inversely proportional to the first queue length and the first memory occupancy size; The determining of the DPDK link load score value based on the current link load parameter of the DPDK link includes: counting the current link load parameter of the DPDK link through a scheduler, wherein the current link load parameter of the DPDK link includes a second data processing rate, a second queue length, and a second memory occupancy size; the second data processing rate represents the current rate of data transmitted through the DPDK link, the second queue length represents the total amount of data waiting to be transmitted on the DPDK link, and the second memory occupancy size represents the total memory occupancy size of data waiting to be transmitted on the DPDK link; The DPDK link load score value is determined based on the second data processing rate, the second queue length, and the second memory occupancy size, wherein the DPDK link load score value is proportional to the second data processing rate and inversely proportional to the second queue length and the second memory occupancy size.

7. The method according to claim 1, characterized in that The sending the to-be-processed data to the target hardware cryptographic device through the scheduler includes: If the target hardware cryptographic device is the GPU hardware cryptographic device, sending the data to be processed to the GPU transit firmware based on the OpenSSL link through the scheduler, and sending the data to be processed to the GPU hardware cryptographic device through the GPU transit firmware; If the target hardware cryptographic device is the USB hardware cryptographic device, the data to be processed is sent to the USB transit firmware based on the OpenSSL link by the scheduler, and the data to be processed is sent to the USB hardware cryptographic device by the USB transit firmware; If the target hardware cryptographic device is the PCIE hardware cryptographic device, and the link load parameter of the OpenSSL link is better than the link load parameter of the DPDK link, the data to be processed is sent to the PCIE transit firmware based on the OpenSSL link through the scheduler, and the data to be processed is sent to the PCIE hardware cryptographic device through the PCIE transit firmware; or, if the link load parameter of the DPDK link is better than the link load parameter of the OpenSSL link, the data to be processed is sent to the PCIE transit firmware based on the DPDK link through the scheduler, and the data to be processed is sent to the PCIE hardware cryptographic device through the PCIE transit firmware.

8. The method according to any one of claims 1 to 7, characterized in that: The OpenSSL link includes an open secure socket layer encryption application programming interface OpenSSL Crypto API, an open secure socket layer engine OpenSSL Engine, and a kernel driver Kernel Driver of an operating system in sequence; The DPDK link includes in sequence the data plane development kit encryption library application programming interface DPDK CryptoDevAPI, the encryption library polling mode driver CryptoDev PMD, the virtual function input and output VFIO or the user space input and output UIO driver, and the data buffer.

9. An adaptive scheduling and communication device for heterogeneous computing of hardware cryptographic devices based on vector packet processing (VPP), used for adaptive scheduling of multiple hardware cryptographic devices, characterized in that: include: A VPP application module, used to send the data to be processed and the encryption task type corresponding to the data to be processed to the scheduler; wherein the encryption task type includes the cryptographic algorithm type and the data volume level; A scheduler is used to receive the data to be processed and the encryption task type, obtain the scheduling parameter value of each hardware cryptographic device, and select a target hardware cryptographic device from the multiple hardware cryptographic devices based on the scheduling parameter value of each hardware cryptographic device; wherein the multiple hardware cryptographic devices include a GPU hardware cryptographic device, a USB hardware cryptographic device, and a PCIE hardware cryptographic device; the VPP application communicates with the GPU hardware cryptographic device, the USB hardware cryptographic device, and the PCIE hardware cryptographic device through an OpenSSL link, and the VPP application communicates with the PCIE hardware cryptographic device through a DPDK link; wherein the scheduling parameter value of the GPU hardware cryptographic device is based on the encryption task type, The performance parameters of the GPU hardware cryptographic device, the hardware load parameters of the GPU hardware cryptographic device, and the link load parameters of the OpenSSL link are determined; the scheduling parameter value of the USB hardware cryptographic device is determined based on the encryption task type, the performance parameters of the USB hardware cryptographic device, the hardware load parameters of the USB hardware cryptographic device, and the link load parameters of the OpenSSL link; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the encryption task type, the performance parameters of the PCIE hardware cryptographic device, the hardware load parameters of the PCIE hardware cryptographic device, the link load parameters of the OpenSSL link, and the link load parameters of the DPDK link; The scheduler is further used to send the data to be processed to the target hardware cryptographic device; Among them, if the data to be processed is data to be encrypted, the data to be encrypted is encrypted by the target hardware cryptographic device; or, if the data to be processed is data to be decrypted, the data to be decrypted is decrypted by the target hardware cryptographic device.

10. The device according to claim 9, characterized in that The scheduler obtains the scheduling parameter value of the GPU hardware cryptographic device, the scheduling parameter value of the USB hardware cryptographic device, and the scheduling parameter value of the PCIE hardware cryptographic device for: Obtaining a first initial weight, a second initial weight, and a third initial weight corresponding to the encryption task type from a configured weight table; wherein the weight table includes a correspondence between the encryption task type and the initial weight; the first initial weight represents a probability that the GPU hardware cryptographic device is assigned data to be processed, the second initial weight represents a probability that the USB hardware cryptographic device is assigned data to be processed, and the third initial weight represents a probability that the PCIE hardware cryptographic device is assigned data to be processed; Determine a first target weight based on the current performance parameters of the GPU hardware cryptographic device for the encryption task type and the first initial weight, and determine a first performance score value based on the first target weight; determine a second target weight based on the current performance parameters of the USB hardware cryptographic device for the encryption task type and the second initial weight, and determine a second performance score value based on the second target weight; determine a third target weight based on the current performance parameters of the PCIE hardware cryptographic device for the encryption task type and the third initial weight, and determine a third performance score value based on the third target weight; Determine a first hardware load score value based on the current hardware load parameters of the GPU hardware cryptographic device, determine a second hardware load score value based on the current hardware load parameters of the USB hardware cryptographic device, and determine a third hardware load score value based on the current hardware load parameters of the PCIE hardware cryptographic device; Determine an OpenSSL link load score value based on a current link load parameter of the OpenSSL link, and determine a first link load score value allocated to the GPU hardware cryptographic device, a second link load score value allocated to the USB hardware cryptographic device, and a third link load score value allocated to the PCIE hardware cryptographic device among the OpenSSL link load score values; Determining a DPDK link load score value based on a current link load parameter of the DPDK link; The scheduling parameter value of the GPU hardware cryptographic device is determined based on the first performance score value, the first hardware load score value and the first link load score value; the scheduling parameter value of the USB hardware cryptographic device is determined based on the second performance score value, the second hardware load score value and the second link load score value; the scheduling parameter value of the PCIE hardware cryptographic device is determined based on the third performance score value, the third hardware load score value, the third link load score value and the DPDK link load score value.

Citation Information

Patent Citations

  • Task processing method and device, electronic equipment and storage medium

    CN119201451A

  • Data encryption and decryption system and method

    US20230222231A1