Acceleration device and computing acceleration system
By introducing unified addressing technology and asymmetric encryption algorithm in the acceleration equipment, the problem of high data transmission delay in the prior art is solved, and a more efficient establishment of secure communication links is achieved.
Patent Information
- Application Number
- PCT/CN2024/135307
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-05
AI Technical Summary
When establishing a secure communication link, existing acceleration equipment needs to frequently transfer data, resulting in high data transmission delay.
Design an acceleration device, including a storage module, a computing high-speed link hard-core module and an acceleration computing module, through unified addressing technology, enable the host and acceleration equipment to directly access the unified address space, reduce the number of data transfers, and deploy an asymmetric encryption algorithm in the acceleration computing module for encryption calculation.
It effectively reduces the number of data transfers during the acceleration process, reduces the delay in data transmission, and improves the efficiency of establishing a secure communication link.
Smart Images

Figure CN2024135307_05062025_PF_FP_ABST
Abstract
Description
Acceleration device and computing acceleration system
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 1, 2023, with application number 202311639194.5 and application name “An Acceleration Device and Computing Acceleration System”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of accelerated computing, and in particular to an acceleration device and a computing acceleration system. Background Art
[0004] With the increasing popularity of high-speed mobile communication devices and the advancement of digitalization across industries, establishing reliable and efficient secure communication links has become a top priority. Computation based on various encryption and decryption protocols is particularly important in this process. Encryption and decryption protocols primarily utilize symmetric and asymmetric encryption technologies. Asymmetric encryption uses different keys for encryption and decryption, offering relatively high security but at the expense of high computational cost. Due to the high computational cost of asymmetric encryption and decryption algorithms, they are often used for key negotiation.
[0005] At present, the development of hardware acceleration units that establish secure communication links based on the TLS (Transport Layer Security) key negotiation protocol has become relatively mature. Generally, the entire TLS protocol and the underlying TCP (Transmission Control Protocol) / IP (Internet Protocol) protocol stack are placed in the accelerator card, and the accelerator in the accelerator card is responsible for all the work of this part of the network protocol stack. The PCIe (Peripheral Component Interconnect Express) protocol is used for communication between the accelerator and the server where the CPU (Central Processing Unit) is located. When the server sends the data needed to establish a secure communication link to the accelerator, the source data required for the calculation needs to be moved from the host-side memory to the accelerator card's memory. After the accelerator completes the calculation, the result data also needs to be moved from the accelerator card's memory to the host-side memory, which increases the delay of data transmission.
[0006] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the Invention
[0007] The purpose of this application is to provide an acceleration device and a computing acceleration system that can reduce the number of data movements during the acceleration process and reduce data transmission delays.
[0008] To solve the above technical problems, the present application provides an acceleration device, comprising:
[0009] Storage module. When the acceleration device is connected to the host, the host uniformly addresses the address space of the host memory and the address space of the storage module, so that the host directly accesses the storage module and the acceleration device directly accesses the host memory;
[0010] Computing high-speed link hard core module, used to receive and forward computing requests sent by the host;
[0011] An accelerated computing module deployed with an asymmetric encryption algorithm is used to access the data to be calculated generated by the host in the uniformly addressed address space according to the calculation request when a calculation request is received, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the uniformly addressed address space so that the host can access the result data in the uniformly addressed address space.
[0012] In some exemplary embodiments of the present application, the accelerated computing module includes a register configuration unit and a computing unit, the register configuration unit includes at least one parameter register, and the computing request includes configuration information of the at least one parameter register;
[0013] The high-speed link computing hard core module is specifically used to receive computing requests sent by the host, parse the configuration information in the computing request, and write the configuration information into the corresponding parameter register;
[0014] The computing unit is used to determine the source address and the destination address according to the configuration information in each parameter register, access the data to be calculated generated by the host in the address space after unified addressing according to the source address, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the address space after unified addressing according to the destination address, so that the host can access the result data in the address space after unified addressing.
[0015] In some exemplary embodiments of the present application, the parameter register includes a startup register and an enable register, the computing unit includes at least one computing subunit, each computing subunit has a built-in asymmetric encryption and decryption algorithm, the parameter register also includes a destination address register and a source address register corresponding to each computing subunit, and the configuration information includes startup information and enable information;
[0016] The computing high-speed link hard core module is specifically used to parse the enable information and startup information in the computing request after receiving the computing request, write the enable information into the enable register, and write the startup information into the startup register; the enable information includes the configuration value of each configuration bit in the enable register, and the startup information includes the configuration value of the startup register;
[0017] Each computing sub-unit is used to, when it detects that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the startup preset value, access the uniformly addressed address space according to the source address stored in its one-to-one corresponding source address register and obtain the data to be calculated generated by the host, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the uniformly addressed address space according to the destination address stored in its one-to-one corresponding destination address register, so that the host can access the result data in the uniformly addressed address space.
[0018] In some exemplary embodiments of the present application, each computing subunit is further configured to stop working when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the second enable preset value and the configuration value of the startup register is the startup preset value.
[0019] In some exemplary embodiments of the present application, each computing subunit is further configured to stop working when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the shutdown preset value.
[0020] In some exemplary embodiments of the present application, the accelerated computing module further includes an arbitration unit, the arbitration unit including a request and response arbitration subunit;
[0021] The computing unit is specifically configured to determine a source address and a destination address based on configuration information in each parameter register, initiate a read data request to the arbitration unit based on the source address, so as to access the data to be calculated generated by the host in the uniformly addressed address space according to the source address, perform asymmetric encryption calculation on the data to be calculated based on an asymmetric encryption and decryption algorithm to obtain result data, and then initiate a write data request to the arbitration unit based on the destination address, so as to write the result data into the uniformly addressed address space according to the destination address;
[0022] The request and response arbitration subunit is used to poll and respond to the read data request or write data request initiated by the computing unit according to the polling cycle.
[0023] In some exemplary embodiments of the present application, the arbitration unit further includes:
[0024] A management subunit, configured to generate an alarm message when an abnormality occurs in the calculation unit and the register configuration unit;
[0025] The computing high-speed link kernel module is also used to report alarm information to the host.
[0026] In some exemplary embodiments of the present application, the parameter register further includes a completion status register;
[0027] Each computing subunit is further configured to adjust the configuration value of the configuration bit corresponding to itself in the completion status register to a completion preset value after performing asymmetric encryption calculation on the computing data based on the asymmetric encryption and decryption algorithm to obtain result data;
[0028] The computing high-speed link hard core module is further used to send an interrupt prompt signal generated according to each configuration value of the completion status register to the host, so that the host can access the result data in the unified addressing address space after receiving the interrupt prompt signal.
[0029] In some exemplary embodiments of the present application, the calculation request is a calculation request generated based on various configuration values of the completion status register.
[0030] In some exemplary embodiments of the present application, the host is also used to determine the target computing subunit according to the various configuration values of the completion status register, generate a calculation request based on the target computing subunit, and generate the data to be calculated to the source address corresponding to the target computing subunit.
[0031] In some exemplary embodiments of the present application, the parameter register further includes a first cleanup register;
[0032] The high-speed link computing hard core module is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request;
[0033] The process of generating interrupt prompt signals according to the various configuration values of the completion status register includes:
[0034] When the configuration value of the first cleaning register is the configuration value that allows reporting, and the completion preset value exists in each configuration value of the completion status register, an interrupt prompt signal is generated.
[0035] In some exemplary embodiments of the present application, the first cleanup register is further configured to adjust the configuration value to a configuration value that is allowed to be reported after a preset time period after the configuration value is adjusted to a configuration value that is prohibited from being reported.
[0036] In some exemplary embodiments of the present application, the parameter register further includes a second cleanup register;
[0037] The high-speed link computing hard core module is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request;
[0038] The completion status register is further used to adjust the configuration values of each configuration bit thereof according to the configuration value of the first cleaning register.
[0039] In some exemplary embodiments of the present application, the address space after unified addressing includes multiple sub-address spaces, and the operation mode of each sub-address space is a host bias mode or a device bias mode;
[0040] The accelerated computing module is also used to generate notification information before accessing the sub-address space of the host bias mode, and access the sub-address space of the host bias mode after receiving confirmation indication information returned by the host based on the notification information.
[0041] In some exemplary embodiments of the present application, the sub-address space used to store data to be calculated is in a device bias mode, and the sub-address space used to store result data is in a host bias mode.
[0042] In some exemplary embodiments of the present application, the asymmetric encryption and decryption algorithm is an elliptic curve algorithm.
[0043] In some exemplary embodiments of the present application, the computing high-speed link hard core module includes:
[0044] Transceiver, used to receive and forward computing requests and memory access requests sent by the host;
[0045] A function control unit, configured to receive a computing request forwarded by the transceiver and configure an accelerated computing module based on the computing request;
[0046] The consistency engine unit is used to receive memory access requests forwarded by the transceiver and / or memory access requests sent by the accelerated computing module, and access the storage module based on the memory access requests.
[0047] In some exemplary embodiments of the present application, the consistency engine unit includes:
[0048] Host cache, used to cache host memory data corresponding to memory access requests;
[0049] Device cache, used to cache device memory data corresponding to memory access requests;
[0050] The request processing subunit is used to determine whether there is data corresponding to the memory access request in the host cache or device cache after receiving the memory access request sent by the sender. If so, the data is returned to the sender, which is the host or the accelerated computing module.
[0051] In some exemplary embodiments of the present application, the high-speed link computing hard core module further includes:
[0052] An interrupt interface, used to detect an interrupt control signal generated by the accelerated computing module when a processing condition is met;
[0053] The function control unit is further used to convert the interrupt control signal into an interrupt prompt signal;
[0054] The transceiver is further used to report an interrupt prompt signal to the host so that the host can access result data in the unified addressing address space.
[0055] To solve the above technical problems, the present application also provides a computing acceleration system, including a host and an acceleration device as described above.
[0056] The present application provides an acceleration device, including a storage module, an accelerated computing module, and a computing high-speed link hard-core module that supports a computing high-speed link protocol. When the acceleration device is connected to a host, the host performs unified address space addressing on the address space of the host memory and the address space of the storage module, so that the host directly accesses the storage module and the acceleration device directly accesses the host memory. The acceleration device can directly access the address space after unified addressing to obtain the data to be calculated generated by the host, perform asymmetric encryption calculation on the data to be calculated based on an asymmetric encryption and decryption algorithm to obtain the result data, and write the result data into the address space after unified addressing, so that the host can directly access and obtain the result data after the asymmetric encryption calculation, thereby reducing the number of data movements during the acceleration process and reducing the delay of data transmission. The present application also provides a computing acceleration system with the same beneficial effects as the above-mentioned acceleration device. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0058] FIG1 is a schematic structural diagram of an acceleration device provided by the present application;
[0059] FIG2 is a schematic diagram of the structure of an accelerated computing module provided by the present application;
[0060] FIG3 is a schematic diagram of the structure of a high-speed link computing hard-core module provided in this application. DETAILED DESCRIPTION
[0061] The core of this application is to provide an acceleration device and a computing acceleration system that can reduce the number of data movements during the acceleration process and reduce data transmission delays.
[0062] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0063] In the present application, each module and each unit is a device component with specific functions formed by at least one processor, circuit, component, or a combination thereof.
[0064] In the first aspect, please refer to FIG1 , which is a schematic diagram of the structure of an acceleration device provided by the present application. In FIG1 , the host side includes a consistency bridge management module, a master agent module, and a host memory. The acceleration device includes:
[0065] Storage module 1. When the acceleration device is connected to the host, the host uniformly addresses the address space of the host memory and the address space of the storage module 1, so that the host directly accesses the storage module 1 and the acceleration device directly accesses the host memory;
[0066] Computing high-speed link hard core module 2, used to receive and forward computing requests sent by the host;
[0067] An accelerated computing module 3 is deployed with an asymmetric encryption algorithm, which is used to access the data to be calculated generated by the host in the address space after unified addressing according to the computing request when a computing request is received, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the address space after unified addressing so that the host can access the result data in the address space after unified addressing.
[0068] In this embodiment, the storage module 1 can be an HDM (Host-managed Device Memory) in the acceleration device. It can be understood that when the acceleration device is connected to the host, during the initialization process on the host side, the storage module 1 of the acceleration device can be identified. At this time, after the address space of the storage module 1 of the acceleration device and the address space of the host memory are uniformly addressed based on the CXL (Compute Express Link) protocol, the host can directly access the device-side memory of the acceleration device, and the acceleration device can also directly access the host memory. For example, assuming that the address space of the host memory is ADDR1~ADDR2 (Address, address), and the address space of the storage module 1 is ADDR3~ADDR4, before the two are uniformly addressed, the address space accessible to the host CPU is ADDR1~ADDR2, and the address space accessible to the acceleration device is ADDR3~ADDR4. After the address space of the host memory and the address space of the storage module 1 are uniformly addressed, the address space accessible to the host CPU and the acceleration device is ADDR1~ADDR4. Since unified addressing is used for device memory and host memory, both the device and the host can directly access the device-side memory, making full use of device-side resources, reducing the occupation of host-side memory, and improving resource utilization.
[0069] The compute high-speed link hard core module 2 connects to the host via the PCIe bus. It receives compute requests from the host via the CXL.io protocol link and forwards them to the accelerated computing module 3. It then receives memory access requests via the CXL.cache and CXL.mem protocol links to access device memory. It is understood that memory access requests can be sent by the host, other CXL devices connected to the accelerator device, or the accelerated computing module 3 within the accelerator device. It is understood that the CXL.io protocol is essentially a modified PCIe 5.0 protocol used for initialization, linking, device discovery and enumeration, and register access. It provides a load / store interface for I / O (Input / Output) devices. The CXL.cache protocol defines the interaction between the host and accelerator devices, allowing connected accelerator devices to efficiently cache host memory with extremely low latency. The CXL.mem protocol provides load and store commands, allowing the host CPU to access the memory connected to the accelerator device. In this case, the host CPU acts as the master device and the accelerator device acts as a slave device. Both volatile and persistent memory architectures are supported.
[0070] An asymmetric encryption and decryption algorithm is deployed in the accelerated computing module 3. Specifically, the asymmetric encryption and decryption algorithm can be an elliptic curve algorithm. It can be understood that, under the premise of the same key security, the key negotiation protocol of the asymmetric encryption algorithm based on the elliptic curve algorithm requires less computation. Under the same computation, the key negotiation protocol of the asymmetric encryption algorithm based on the elliptic curve algorithm is relatively more secure. Based on this, the accelerated computing module 3 in this embodiment is deployed with an elliptic curve algorithm. In response to the computation request, the accelerated computing module 3 accesses the data to be calculated generated by the host end in the uniformly addressed address space. That is, the accelerated computing module 3 responds to the computation request by directly accessing the host memory, obtaining the data to be calculated generated by the host end, performing asymmetric encryption calculation on the data to be calculated based on the elliptic curve algorithm to obtain the result data, and writing the result data into the uniformly addressed address space.
[0071] In this embodiment, the accelerated computing module 3 only implements the elliptic curve algorithm in the TLS protocol, which has low host CPU computational efficiency. The remaining control logic and computational logic in the TLS protocol, which the CPU excels at, are still processed by the CPU. Furthermore, because the accelerated computing module 3 only accelerates the elliptic curve algorithm, the size of the data blocks it transmits is significantly reduced, fully utilizing the CXL protocol's lower latency compared to the PCIe protocol for small data blocks. Because the acceleration device in this embodiment does not deploy the full TLS protocol, the resource usage of the acceleration device is reduced.
[0072] It can be seen that in this embodiment, it includes a storage module 1, an accelerated computing module 3 and a computing high-speed link hard core module 2 that supports the computing high-speed link protocol. When the acceleration device is connected to the host, the host performs unified address space addressing on the address space of the host memory and the address space of the storage module 1, so that the host can directly access the storage module 1 and the acceleration device can directly access the host memory. The acceleration device can directly access the address space after unified addressing to obtain the data to be calculated generated by the host, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain the result data, and write it into the address space after unified addressing, so that the host can directly access and obtain the result data after the asymmetric encryption calculation, thereby reducing the number of data movements during the acceleration process and reducing the delay in data transmission.
[0073] Based on the above embodiment:
[0074] In some exemplary embodiments of the present application, as shown in FIG2 , the accelerated computing module includes a register configuration unit 31 and a computing unit 32 , the register configuration unit 31 includes at least one parameter register, and the computing request includes configuration information of the at least one parameter register;
[0075] The high-speed link computing hard core module 2 is specifically used to receive a computing request sent by the host, parse the configuration information in the computing request, and write the configuration information into the corresponding parameter register;
[0076] The computing unit 32 is used to determine the source address and the destination address according to the configuration information in each parameter register, access the data to be calculated generated by the host in the address space after unified addressing according to the source address, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the address space after unified addressing according to the destination address, so that the host can access the result data in the address space after unified addressing.
[0077] In this embodiment, the accelerated computing module 3 includes a computing unit 32 and a register configuration unit 31. The register configuration unit 31 includes at least one parameter register. After receiving a computing request through CXL.io, the computing high-speed link hard core module 2 parses the computing request, parses out the configuration information of at least one parameter register included in the computing request, and writes the configuration information into the corresponding parameter register. It can be understood that the computing unit 32 is connected to each parameter register, and the computing task that the computing unit 32 needs to complete is determined according to the configuration value of each parameter register.
[0078] In some exemplary embodiments of the present application, the parameter register includes a startup register and an enable register, the computing unit 32 includes at least one computing subunit, each computing subunit has a built-in asymmetric encryption and decryption algorithm, the parameter register also includes a destination address register and a source address register corresponding to each computing subunit, and the configuration information includes startup information and enable information;
[0079] The high-speed link computing hard core module 2 is specifically configured to, upon receiving a computing request, parse the enable information and startup information in the computing request, write the enable information into the enable register, and write the startup information into the startup register; the enable information includes the configuration value of each configuration bit in the enable register, and the startup information includes the configuration value of the startup register;
[0080] Each computing sub-unit is used to, when it detects that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the startup preset value, access the uniformly addressed address space according to the source address stored in its one-to-one corresponding source address register and obtain the data to be calculated generated by the host, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the uniformly addressed address space according to the destination address stored in its one-to-one corresponding destination address register, so that the host can access the result data in the uniformly addressed address space.
[0081] In some exemplary embodiments of the present application, each computing subunit is further configured to stop working when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the shutdown preset value.
[0082] In some exemplary embodiments of the present application, each computing subunit is further configured to stop working when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the second enable preset value and the configuration value of the startup register is the startup preset value.
[0083] In this embodiment, the register configuration unit 31 includes but is not limited to a startup register and an enable register, a source address register and a destination address register. The enable register includes multiple configuration bits, and the multiple configuration bits are connected one-to-one with multiple computing sub-units. Each computing sub-unit is connected one-to-one with a source address register and a destination address register. The bit width of each parameter register can be adjusted according to actual engineering needs. When the bit width of a parameter register is not sufficient to meet the needs of the computing sub-units in the computing unit 32, multiple registers of the same category can be set, such as setting multiple enable registers.
[0084] For example, assuming that the computing unit 32 includes 64 computing sub-units, it may include two 32-bit enable registers, namely params_enable_low and params_enable_high, where params_enable_low is used to control whether computing sub-units 0 to 31 are enabled, and params_enable_high is used to control whether computing sub-units 32 to 63 are enabled. The source address register params_csr_xx_0 and the destination address register params_csr_xx_1 are set in a one-to-one correspondence with the computing sub-units, where xx represents the number of the corresponding computing sub-unit. The start register params_ecc_start is used to control the start of the enabled computing sub-unit to perform asymmetric encryption calculations. When the host rewrites the params_csr_02_0 register to 0x60000 through the CXL.io protocol, it means that the computing subunit numbered 2 will obtain the source data required for its calculation from address 0x60000, that is, obtain the data to be calculated. When the params_csr_02_1 register is rewritten to 0x70000, it means that the result data calculated by the computing subunit will be written to the address space of 0x70000. params_enable_low, params_enable_high and params_ecc_start work together to determine which computing subunits to start for calculation. For example, assuming that all bits in the two 32-bit registers params_enable_low and params_enable_high are 1, it can be understood that bit is the configuration bit in this embodiment, 1 is the first enable preset value, and 0 is the second enable preset value. At this time, if params_ecc_start is written to 1, all 64 computing subunits will be started for calculation. If params_enable_low, params_enable_high, and params_ecc_start are written to 1, all 64 computing subunits will be started for calculation. Set le_low to 'b0000_0000_0000_0000_0000_0000_0000_1111 (i.e., enable computing subunits No. 0 to 3), set params_enable_high to 'b0000_0000_0000_0000_0000_0000_1111_0000 (i.e., enable computing subunits No. 36 to 39), and then write params_ecc_start to 1. Here, 1 can be set to the startup preset value, and 0 can be set to the shutdown preset value. Writing params_ecc_start to 1 will start computing subunits No. 0 to 3 and No. 36 to 39 for calculation.
[0085] In some exemplary embodiments of the present application, the accelerated computing module 3 further includes an arbitration unit 33 , which includes request and response arbitration subunits;
[0086] The calculation unit 32 is specifically configured to determine a source address and a destination address based on the configuration information in each parameter register, initiate a read data request to the arbitration unit 33 based on the source address, so as to access the data to be calculated generated by the host in the uniformly addressed address space according to the source address, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and then initiate a write data request to the arbitration unit 33 based on the destination address, so as to write the result data into the uniformly addressed address space according to the destination address;
[0087] The request and response arbitration subunit is used to poll and respond to the read data request or write data request initiated by the computing unit 32 according to the polling cycle.
[0088] It can be understood that the activated calculation sub-unit will obtain the source address according to the value of its respective params_csr_xx_0 source address register, and initiate a request to read data to the request and response arbitration sub-unit. After completing the calculation, it will obtain the destination address according to the value of its respective params_csr_xx_1 register, and initiate a request to write data to the request and response arbitration sub-unit, and write the calculation result to the destination address.
[0089] In some exemplary embodiments of the present application, the arbitration unit 33 further includes:
[0090] A management subunit, configured to generate an alarm message when an abnormality occurs in the calculation unit 32 and the register configuration unit 31;
[0091] The computing high-speed link kernel module is also used to report alarm information to the host.
[0092] Considering the low efficiency of arbitration polling response, a management subunit is also provided here. When an abnormality occurs in the calculation unit 32 and the register configuration unit 31, an alarm message is immediately generated and reported to the host so that the abnormality can be handled in time.
[0093] In some exemplary embodiments of the present application, the parameter register further includes a completion status register;
[0094] Each computing subunit is further configured to adjust the configuration value of the configuration bit corresponding to itself in the completion status register to a completion preset value after performing asymmetric encryption calculation on the computing data based on the asymmetric encryption and decryption algorithm to obtain result data;
[0095] The high-speed link computing hard core module 2 is further configured to send an interrupt prompt signal generated according to various configuration values of the completion status register to the host, so that the host can access the result data in the unified addressing address space after receiving the interrupt prompt signal.
[0096] In some exemplary embodiments of the present application, the calculation request is a calculation request generated based on various configuration values of the completion status register.
[0097] In some exemplary embodiments of the present application, the host is also used to determine the target computing subunit according to the various configuration values of the completion status register, generate a calculation request based on the target computing subunit, and generate the data to be calculated to the source address corresponding to the target computing subunit.
[0098] In some exemplary embodiments of the present application, the parameter register further includes a first cleanup register;
[0099] The high-speed link calculation hard core module 2 is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request;
[0100] The process of generating interrupt prompt signals according to the various configuration values of the completion status register includes:
[0101] When the configuration value of the first cleaning register is the configuration value that allows reporting, and the completion preset value exists in each configuration value of the completion status register, an interrupt prompt signal is generated.
[0102] In some exemplary embodiments of the present application, the first cleanup register is further configured to adjust the configuration value to a configuration value that is allowed to be reported after a preset time period after the configuration value is adjusted to a configuration value that is prohibited from being reported.
[0103] In some exemplary embodiments of the present application, the parameter register further includes a second cleanup register;
[0104] The high-speed link calculation hard core module 2 is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request;
[0105] The completion status register is further used to adjust the configuration value of each configuration bit thereof according to the configuration value of the first cleaning register.
[0106] It can be understood that the number of completion status registers can also be determined according to their bit width and the number of computing subunits in the computing unit 32. Assuming that the computing unit 32 includes 64 computing subunits, each with a 1-bit completion signal, these 64 1-bit completion signals will be connected to two 32-bit registers, params_done_low and params_done_high, namely the completion status registers. The former is used to store the completion signals of computing subunits No. 0 to No. 31, and the latter is used to store the completion signals of computing subunits No. 32 to No. 63. The first cleanup register params_irq_ The configuration value of clear is set by the host. The configuration value of the first clear register can be used to determine whether the host allows reporting of the interrupt prompt signal. The configuration value for allowing reporting can be set to 0, and the configuration value for prohibiting reporting can be set to 1. When the 64 bits of the params_done_low and params_done_high registers are not all 0 and params_irq_clear is not 1, the interrupt control signal IRQ is output. After the high-speed link hard core module 2 detects a valid interrupt control signal, it converts it into an interrupt prompt signal and reports it to the host. Here, the interrupt control signal IRQ can be set to be valid at a high level. The parameter register also includes a second clearing register params_done_low_clear and params_done_high_clear for clearing the completion signals stored in the corresponding bits of the two 32-bit registers params_done_low and params_done_high. For example, assuming that the current values of the two registers params_done_low and params_done_high are both 0xffffffff, that is, all 64 computing sub-units have completed the calculation, the host sets params_done_low_clear and params_done_high_clear to 0xffffffff and 0x0 respectively, which means that the completion signals of computing sub-units No. 0 to No. 63 are cleared, which actually means that the host has obtained the results of computing sub-units No. 0 to No. 32. After clearing them, computing sub-units No. 0 to No. 32 can execute new computing tasks.In addition, each time the host uses the params_done_low_clear and params_done_high_clear registers to clear some or all of the completion signals, it also needs to set params_irq_clear to 1. This register will be automatically set to 0 after being set to 1 for one clock cycle. Its function is to notify the computing device that the upper-level software has completed the processing of some or all of the result data, and can continue to trigger interrupt notifications upward according to the status of the interrupt control signal IRQ. It can be understood that the host may not process all completion signals at once, and can process part of it first, that is, partially clear the bits that are 1 in the params_done_low and params_done_high registers. Since there is still a bit that is 1, IRQ is still 1, so the host can continue to receive interrupt signals after setting params_irq_clear to 1.
[0107] In some exemplary embodiments of the present application, the address space after unified addressing includes multiple sub-address spaces, and the operation mode of each sub-address space is a host bias mode or a device bias mode;
[0108] The accelerated computing module 3 is further configured to generate notification information before accessing the sub-address space of the host bias mode, and access the sub-address space of the host bias mode after receiving confirmation indication information returned by the host based on the notification information.
[0109] In some exemplary embodiments of the present application, the sub-address space used to store data to be calculated is in a device bias mode, and the sub-address space used to store result data is in a host bias mode.
[0110] The biggest difference between the CXL protocol and the PCIe protocol is that it has a unified address space, allowing both the device and host to directly access host and device memory, and achieving cache coherence between the device and host. To minimize the impact of cache coherence on the latency of host or device access to the HDM, the HDM operation mode is divided into two types: host bias mode and device bias mode. Host bias mode optimizes the latency of host access to the HDM and is typically used when the host sends task data to the device before starting a task and when the host reads the data after the device completes its work. In this mode, if the device wants to access the HDM, it must first notify the host's coherence bridge management module and obtain permission before accessing the HDM. Device bias mode optimizes the latency of device access to the HDM and is typically used after the device starts executing an acceleration task and before the host reads the result data. In this mode, the device can directly access the HDM without notifying the host's coherence bridge management module, ensuring that the host does not have cache of the HDM memory area in device bias mode. In this mode, the host can still access the HDM, but access performance may be reduced compared to host bias mode.
[0111] The host-side master agent software module can control and manage which sub-address spaces in the HDM belong to the host bias mode and which belong to the device bias mode. In the consistency engine, the bias table tracks the bias status of each page (i.e., 4kB) in the HDM. This means that the bias mode of the memory in the HDM can be set and managed at a 4kB granularity.
[0112] Assume that the data to be calculated of all computing subunits are stored in the address space of ADDR1~ADDR2 (and the size is an integer multiple of 4kB), and the result data of all computing subunits are stored in the address space of ADDR3~ADDR4 (and the size is an integer multiple of 4kB). Because the operating mode of elliptic curve accelerated multiplication is simple enough, it simply reads data from the address space ADDR1-ADDR2 and writes it back to the address space ADDR3-ADDR4 after calculation. For address space ADDR1-ADDR2, the host, as the data writer, does not need to use the data in this address space, so there is no need to cache the data in this address space. Setting this portion of the address space to device bias mode can minimize performance loss. Because the host no longer needs to obtain this data, there is no increased latency due to the lack of cache. The device can directly access this address space without host consent and cache some data in the engine coherence unit to accelerate access. For address space ADDR3-ADDR4, the calculation subunit, as the result data writer, no longer needs to obtain the result data. By setting ADDR3-ADDR4 to host bias mode, when the device uses the result data to update this address space, it must first notify the host's coherence bridge management module, so that the host is notified of the data update in this address space and updates the host CPU cache with the latest data.
[0113] In some exemplary embodiments of the present application, as shown in FIG3 , the high-speed link computing hard core module 2 includes:
[0114] Transceiver 21, used to receive and forward computing requests and memory access requests sent by the host;
[0115] The function control unit 22 is configured to receive the computing request forwarded by the transceiver 21 and configure the accelerated computing module 3 based on the computing request;
[0116] The consistency engine unit 23 is used to receive the memory access request forwarded by the transceiver 21 and / or the memory access request sent by the accelerated computing module 3, and access the storage module 1 based on the memory access request.
[0117] In some exemplary embodiments of the present application, the consistency engine unit 23 includes:
[0118] Host cache, used to cache host memory data corresponding to memory access requests;
[0119] Device cache, used to cache device memory data corresponding to memory access requests;
[0120] The request processing subunit is used to determine whether there is data corresponding to the memory access request in the host cache or device cache after receiving the memory access request sent by the sender. If so, the data is returned to the sender, which is the host or the accelerated computing module 3.
[0121] In this embodiment, the host cache and device cache in the consistency engine unit 23 are used to store data read from the host memory and data read from the device memory, respectively, and cache the data in the next preset address segment in the memory access request in advance, so that after receiving the memory access request subsequently, it is first determined from its own cache whether it hits. If the cache hits, it is directly returned to the sending end to improve data transmission efficiency. If it does not hit, data is requested from the host memory or device memory.
[0122] In some exemplary embodiments of the present application, the high-speed link computing hard core module 2 further includes:
[0123] An interrupt interface 24, used to detect an interrupt control signal generated by the accelerated computing module 3 when a processing condition is met;
[0124] The function control unit 22 is further configured to convert the interrupt control signal into an interrupt prompt signal;
[0125] The transceiver 21 is further configured to report an interrupt prompt signal to the host so that the host can access the result data in the unified addressing address space.
[0126] In summary, this application targets the computing power bottleneck faced by the CPU in asymmetric encryption and decryption calculations in the TLS protocol. By compiling an elliptic curve multiplication unit array module and its supporting parameter registers, arbitrators and other peripheral facilities, and combining the IRQ interface, memory interface and CSR interface provided by the computing high-speed link hard core, this application has for the first time realized a heterogeneous computing system that supports direct access of devices and hosts to a unified addressable memory space and performs elliptic curve multiplication acceleration with low latency, thereby reducing the number of data movements. By rationally designing the register configuration unit 31, the hardware foundation for starting and using the results of any number of elliptic curve multiplication units is realized. By setting operation modes of different address spaces for different stages of the elliptic curve accelerator's work, the access delay is optimized while ensuring data consistency.
[0127] In a second aspect, the present application also provides a computing acceleration system, comprising a host and an acceleration device as described above.
[0128] Among them, acceleration equipment includes:
[0129] Storage module. When the acceleration device is connected to the host, the host uniformly addresses the address space of the host memory and the address space of the storage module, so that the host directly accesses the storage module and the acceleration device directly accesses the host memory;
[0130] Computing high-speed link hard core module, used to receive and forward computing requests sent by the host;
[0131] An accelerated computing module deployed with an asymmetric encryption algorithm is used to access the data to be calculated generated by the host in the uniformly addressed address space according to the calculation request when a calculation request is received, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the uniformly addressed address space so that the host can access the result data in the uniformly addressed address space.
[0132] In some exemplary embodiments of the present application, the accelerated computing module includes a register configuration unit and a computing unit, the register configuration unit includes at least one parameter register, and the computing request includes configuration information of the at least one parameter register;
[0133] The high-speed link computing hard core module is specifically used to receive computing requests sent by the host, parse the configuration information in the computing request, and write the configuration information into the corresponding parameter register;
[0134] The computing unit is used to determine the source address and the destination address according to the configuration information in each parameter register, access the data to be calculated generated by the host in the address space after unified addressing according to the source address, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the address space after unified addressing according to the destination address, so that the host can access the result data in the address space after unified addressing.
[0135] In some exemplary embodiments of the present application, the parameter register includes a startup register and an enable register, the computing unit includes at least one computing subunit, each computing subunit has a built-in asymmetric encryption and decryption algorithm, the parameter register also includes a destination address register and a source address register corresponding to each computing subunit, and the configuration information includes startup information and enable information;
[0136] The computing high-speed link hard core module is specifically used to parse the enable information and startup information in the computing request after receiving the computing request, write the enable information into the enable register, and write the startup information into the startup register; the enable information includes the configuration value of each configuration bit in the enable register, and the startup information includes the configuration value of the startup register;
[0137] Each computing sub-unit is used to, when it detects that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the startup preset value, access the uniformly addressed address space according to the source address stored in its one-to-one corresponding source address register and obtain the data to be calculated generated by the host, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the uniformly addressed address space according to the destination address stored in its one-to-one corresponding destination address register, so that the host can access the result data in the uniformly addressed address space.
[0138] In some exemplary embodiments of the present application, each computing subunit is further configured to stop working when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the second enable preset value and the configuration value of the startup register is the startup preset value.
[0139] In some exemplary embodiments of the present application, each computing subunit is further configured to stop working when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the shutdown preset value.
[0140] In some exemplary embodiments of the present application, the accelerated computing module further includes an arbitration unit, the arbitration unit including a request and response arbitration subunit;
[0141] The computing unit is specifically configured to determine a source address and a destination address based on configuration information in each parameter register, initiate a read data request to the arbitration unit based on the source address, so as to access the data to be calculated generated by the host in the uniformly addressed address space according to the source address, perform asymmetric encryption calculation on the data to be calculated based on an asymmetric encryption and decryption algorithm to obtain result data, and then initiate a write data request to the arbitration unit based on the destination address, so as to write the result data into the uniformly addressed address space according to the destination address;
[0142] The request and response arbitration subunit is used to poll and respond to the read data request or write data request initiated by the computing unit according to the polling cycle.
[0143] In some exemplary embodiments of the present application, the arbitration unit further includes:
[0144] A management subunit, configured to generate an alarm message when an abnormality occurs in the calculation unit and the register configuration unit;
[0145] The computing high-speed link kernel module is also used to report alarm information to the host.
[0146] In some exemplary embodiments of the present application, the parameter register further includes a completion status register;
[0147] Each computing subunit is further configured to adjust the configuration value of the configuration bit corresponding to itself in the completion status register to a completion preset value after performing asymmetric encryption calculation on the computing data based on the asymmetric encryption and decryption algorithm to obtain result data;
[0148] The computing high-speed link hard core module is further used to send an interrupt prompt signal generated according to each configuration value of the completion status register to the host, so that the host can access the result data in the unified addressing address space after receiving the interrupt prompt signal.
[0149] In some exemplary embodiments of the present application, the calculation request is a calculation request generated based on various configuration values of the completion status register.
[0150] In some exemplary embodiments of the present application, the host is also used to determine the target computing subunit according to the various configuration values of the completion status register, generate a calculation request based on the target computing subunit, and generate the data to be calculated to the source address corresponding to the target computing subunit.
[0151] In some exemplary embodiments of the present application, the parameter register further includes a first cleanup register;
[0152] The high-speed link computing hard core module is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request;
[0153] The process of generating interrupt prompt signals according to the various configuration values of the completion status register includes:
[0154] When the configuration value of the first cleaning register is the configuration value that allows reporting, and the completion preset value exists in each configuration value of the completion status register, an interrupt prompt signal is generated.
[0155] In some exemplary embodiments of the present application, the first cleanup register is further configured to adjust the configuration value to a configuration value that is allowed to be reported after a preset time period after the configuration value is adjusted to a configuration value that is prohibited from being reported.
[0156] In some exemplary embodiments of the present application, the parameter register further includes a second cleanup register;
[0157] The high-speed link computing hard core module is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request;
[0158] The completion status register is further used to adjust the configuration value of each configuration bit thereof according to the configuration value of the first cleaning register.
[0159] In some exemplary embodiments of the present application, the address space after unified addressing includes multiple sub-address spaces, and the operation mode of each sub-address space is a host bias mode or a device bias mode;
[0160] The accelerated computing module is also used to generate notification information before accessing the sub-address space of the host bias mode, and access the sub-address space of the host bias mode after receiving confirmation indication information returned by the host based on the notification information.
[0161] In some exemplary embodiments of the present application, the sub-address space used to store data to be calculated is in a device bias mode, and the sub-address space used to store result data is in a host bias mode.
[0162] In some exemplary embodiments of the present application, the asymmetric encryption and decryption algorithm is an elliptic curve algorithm.
[0163] In some exemplary embodiments of the present application, the computing high-speed link hard core module includes:
[0164] Transceiver, used to receive and forward computing requests and memory access requests sent by the host;
[0165] A function control unit, configured to receive a computing request forwarded by the transceiver and configure an accelerated computing module based on the computing request;
[0166] The consistency engine unit is used to receive memory access requests forwarded by the transceiver and / or memory access requests sent by the accelerated computing module, and access the storage module based on the memory access requests.
[0167] In some exemplary embodiments of the present application, the consistency engine unit includes:
[0168] Host cache, used to cache host memory data corresponding to memory access requests;
[0169] Device cache, used to cache device memory data corresponding to memory access requests;
[0170] The request processing subunit is used to determine whether there is data corresponding to the memory access request in the host cache or device cache after receiving the memory access request sent by the sender. If so, the data is returned to the sender, which is the host or the accelerated computing module.
[0171] In some exemplary embodiments of the present application, the high-speed link computing hard core module further includes:
[0172] An interrupt interface, used to detect an interrupt control signal generated by the accelerated computing module when a processing condition is met;
[0173] The function control unit is further used to convert the interrupt control signal into an interrupt prompt signal;
[0174] The transceiver is further used to report an interrupt prompt signal to the host so that the host can access result data in the unified addressing address space.
[0175] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0176] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An acceleration device, characterized in that: include: A storage module, when the acceleration device is connected to a host, the host uniformly addresses the address space of the host memory and the address space of the storage module, so that the host directly accesses the storage module and the acceleration device directly accesses the host memory; A high-speed link computing hard core module is configured to receive and forward computing requests sent by the host; An accelerated computing module deployed with an asymmetric encryption algorithm is configured to, upon receiving the computing request, access the data to be calculated generated by the host in the uniformly addressed address space according to the computing request, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the uniformly addressed address space so that the host can access the result data in the uniformly addressed address space.
2. The acceleration device according to claim 1, characterized in that: The accelerated computing module includes a register configuration unit and a computing unit, the register configuration unit includes at least one parameter register, and the computing request includes configuration information of at least one parameter register; The computing high-speed link hard core module is specifically configured to receive a computing request sent by the host, parse the configuration information in the computing request, and write the configuration information into a corresponding parameter register; The computing unit is configured to determine a source address and a destination address according to the configuration information in each of the parameter registers, access the data to be calculated generated by the host in the address space after unified addressing according to the source address, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the address space after unified addressing according to the destination address, so that the host can access the result data in the address space after unified addressing.
3. The acceleration device according to claim 2, characterized in that: The parameter register includes a startup register and an enable register, the computing unit includes at least one computing subunit, each of the computing subunits has the asymmetric encryption and decryption algorithm built in, the parameter register also includes a destination address register and a source address register corresponding to each of the computing subunits, and the configuration information includes startup information and enable information; The computing high-speed link hard core module is specifically configured to, after receiving the computing request, parse the enable information and the startup information in the computing request, write the enable information into the enable register, and write the startup information into the startup register; the enable information includes the configuration value of each configuration bit in the enable register, and the startup information includes the configuration value of the startup register; Each of the computing sub-units is configured to, when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the startup register is the startup preset value, access the address space after unified addressing according to the source address stored in the source address register corresponding to itself and obtain the data to be calculated generated by the host, perform asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, and write the result data into the address space after unified addressing according to the destination address stored in the destination address register corresponding to itself, so that the host can access the result data in the address space after unified addressing.
4. The acceleration device according to claim 3, characterized in that: Each of the computing subunits is further configured to not operate when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the second enable preset value and the configuration value of the start register is the start preset value.
5. The acceleration device according to claim 3, characterized in that: Each of the computing subunits is further configured to not work when it is detected that the configuration value of the configuration bit corresponding to itself in the enable register is the first enable preset value and the configuration value of the start register is the shutdown preset value.
6. The acceleration device according to claim 2, characterized in that: The accelerated computing module further includes an arbitration unit, wherein the arbitration unit includes request and response arbitration subunits; The calculation unit is specifically configured to determine a source address and a destination address according to the configuration information in each of the parameter registers, initiate a read data request to the arbitration unit according to the source address, so as to access the data to be calculated generated by the host in the address space after the unified addressing according to the source address, and after performing asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data, initiate a write data request to the arbitration unit based on the destination address, so as to write the result data into the address space after the unified addressing according to the destination address; The request and response arbitration subunit is configured to poll and respond to the read data request or the write data request initiated by the computing unit according to a polling cycle.
7. The acceleration device according to claim 6, characterized in that: The arbitration unit further includes: A management subunit, configured to generate alarm information when an abnormality occurs in the calculation unit and the register configuration unit; The computing high-speed link kernel module is also configured to report the alarm information to the host.
8. The acceleration device according to claim 2, characterized in that: The parameter register also includes a completion status register; Each of the computing subunits is further configured to adjust the configuration value of the configuration bit corresponding to itself in the completion status register to a completion preset value after performing asymmetric encryption calculation on the data to be calculated based on the asymmetric encryption and decryption algorithm to obtain result data; The computing high-speed link hard core module is also configured to send an interrupt prompt signal generated according to each configuration value of the completion status register to the host, so that the host can access the result data in the uniformly addressed address space after receiving the interrupt prompt signal.
9. The acceleration device according to claim 8, characterized in that: The calculation request is a calculation request generated based on each of the configuration values of the completion status register.
10. The acceleration device according to claim 9, characterized in that: The host is further configured to determine a target computing subunit according to each of the configuration values of the completion status register, generate the calculation request based on the target computing subunit, and generate the data to be calculated to a source address corresponding to the target computing subunit.
11. The acceleration device according to claim 8, characterized in that: The parameter register also includes a first cleanup register; The computing high-speed link hard core module is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request; The process of generating an interrupt prompt signal according to each configuration value of the completion status register includes: When the configuration value of the first cleaning register is a configuration value that allows reporting, and the completion preset value exists in each of the configuration values of the completion status register, an interrupt prompt signal is generated.
12. The acceleration device according to claim 11, characterized in that: The first cleaning register is further configured to adjust the configuration value to the reporting-allowed configuration value after a preset time period after the configuration value is adjusted to the reporting-prohibited configuration value.
13. The acceleration device according to claim 12, characterized in that: The parameter register also includes a second cleanup register; The computing high-speed link hard core module is further configured to receive a first cleaning request sent by the host, and configure a configuration value of the first cleaning register according to the first cleaning request; The completion status register is further configured to adjust the configuration values of each of its configuration bits according to the configuration value of the first cleaning register.
14. The acceleration device according to claim 1, characterized in that: The unified addressing address space includes a plurality of sub-address spaces, and the operation mode of each sub-address space is a host bias mode or a device bias mode; The accelerated computing module is also configured to generate notification information before accessing the sub-address space of the host bias mode, and access the sub-address space of the host bias mode after receiving confirmation indication information returned by the host based on the notification information.
15. The acceleration device according to claim 14, characterized in that: The sub-address space configured to store the data to be calculated is the device bias mode, and the sub-address space configured to store the result data is the host bias mode.
16. The acceleration device according to claim 1, characterized in that: The asymmetric encryption and decryption algorithm is an elliptic curve algorithm.
17. The acceleration device according to any one of claims 1 to 16, characterized in that: The computing high-speed link hard core module includes: A transceiver, configured to receive and forward computing requests and memory access requests sent by the host; a function control unit configured to receive a computing request forwarded by the transceiver and configure the accelerated computing module based on the computing request; The consistency engine unit is configured to receive a memory access request forwarded by the transceiver and / or a memory access request sent by the accelerated computing module, and access the storage module based on the memory access request.
18. The acceleration device according to claim 17, characterized in that: The consistency engine unit comprises: a host cache configured to cache host memory data corresponding to the memory access request; a device cache configured to cache device memory data corresponding to the memory access request; The request processing subunit is configured to, upon receiving the memory access request sent by the sending end, determine whether there is data corresponding to the memory access request in the host cache or the device cache, and if so, return the data to the sending end, which is the host or the accelerated computing module.
19. The acceleration device according to claim 17, characterized in that: The computing high-speed link hard core module also includes: an interrupt interface configured to detect an interrupt control signal generated by the accelerated computing module when a processing condition is met; The function control unit is further configured to convert the interrupt control signal into an interrupt prompt signal; The transceiver is further configured to report the interrupt prompt signal to the host so that the host can access the result data in the uniformly addressed address space.
20. A computing acceleration system, characterized in that: It comprises a host and an acceleration device as described in any one of claims 1-19.
Citation Information
Patent Citations
Memory management method and device, equipment and medium
CN114020454A
File encryption and decryption method and device, equipment and storage medium
CN116070239A
Computing system, method and device and acceleration equipment
CN116467245A
Acceleration device and calculation acceleration system
CN117608849A
Acceleration unit and related apparatus and method
US20220255721A1
Cited By
Accelerator device, computer system, and data processing method
CN120448344A
Accelerator device, computer system and data processing method
CN120448344B