Load balancing in data conversion accelerator

By performing load balancing and I/O virtualization on the host device, the command transmission and distribution of the data conversion accelerator are optimized, solving the problems of limited system throughput and increased latency in the prior art, and achieving more efficient data conversion operations.

CN122139335APending Publication Date: 2026-06-02MAXLINEAR INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MAXLINEAR INC
Filing Date
2024-09-03
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing data conversion accelerator systems, the command transmission method is suboptimal, resulting in limited system throughput and increased latency. Furthermore, the system fails to effectively implement the service categories associated with the execution of commands by the data conversion accelerator, thus limiting the system's priority operations.

Method used

By performing load balancing operations on the host device, selecting the appropriate container storage command address, and transmitting it to the data transformation accelerator, combined with IO virtualization technology, the transmission and distribution of commands are optimized, achieving more balanced resource sharing and processing.

Benefits of technology

It improved the throughput of the data conversion accelerator, reduced command execution latency, achieved a more optimized command submission process, and enhanced the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122139335A_ABST
    Figure CN122139335A_ABST
Patent Text Reader

Abstract

A method includes obtaining a plurality of command requests, each command request including a command address. The method also includes performing a load balancing operation to select a first container of a plurality of containers to store the first command address. The method also includes storing the first command address in the first data container. The method also includes transferring the first command address from the first data container to a data conversion accelerator. The method also includes obtaining converted data from the data conversion accelerator.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing related applications This U.S. patent application claims priority to U.S. Provisional Patent Application No. 63 / 580,083, filed September 1, 2023, entitled “LOAD BALANCING IN ADATA TRANSFORM ACCELERATOR,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure generally relates to data transformation accelerators, and more specifically, to load balancing within data transformation accelerators. Background Technology

[0003] Unless otherwise stated herein, the materials described herein do not constitute prior art for the claims of this application, and are not acknowledged as prior art by virtue of their inclusion in this section.

[0004] A data transformation accelerator is a coprocessor device used to accelerate data transformation operations for various applications, such as data analytics, big data, storage, encryption, and networking applications. For example, a data transformation accelerator can be configured as an accelerator in a storage accelerator, encryption accelerator, and / or network interface card (NIC).

[0005] The subject matter claimed in this disclosure is not limited to embodiments that address any drawbacks, nor is it limited to embodiments that operate only in the environment described above. Rather, this background is provided merely to illustrate an example technical field in which some of the embodiments described in this disclosure can be practiced. Summary of the Invention

[0006] In an example embodiment, the method may include acquiring a plurality of command requests, wherein each command request may include a command address. The method may further include performing a load balancing operation to select a first container among a plurality of containers to store the first command address. The method may further include storing the first command address in a first data container. The method may further include transferring the first command address from the first data container to a data transformation accelerator. The method may further include acquiring transformed data from the data transformation accelerator.

[0007] In another embodiment, the system may include a host device and at least one data conversion accelerator. The host device may include one or more processors and host software. The host software may be operated by the one or more processors and may be operated to generate a plurality of command requests, each command request including a plurality of command addresses. The host software may also be operated to perform load balancing operations to select a first container from a plurality of containers to store a first command address from the plurality of command addresses. The host software may further be operated to guide the storage of the first command address in a first data container. The host software may also be operated to transfer the first command address from the first data container.

[0008] The at least one data conversion accelerator can be operated to obtain a first command address transmitted from a first data container. The at least one data conversion accelerator can also be further operated to obtain a first command associated with multiple command requests and to obtain first input data using the first command address. The at least one data conversion accelerator can also be operated to perform a data conversion operation on the first input data using the first command to generate converted data. The at least one data conversion accelerator can also be further operated to transmit the converted data to a host device.

[0009] The objectives and advantages of the embodiments will be realized and achieved, at least by means of the elements, features and combinations particularly pointed out in the claims.

[0010] The preceding general description and the following detailed description are given as examples and are illustrative, and do not constitute a limitation on the claimed invention. Attached Figure Description

[0011] Exemplary embodiments will be described and explained in more detail with reference to the accompanying drawings, wherein: Figure 1A and Figure 1B A block diagram of an exemplary system for load balancing is shown, including a data transformation accelerator. Figure 2 A flowchart illustrating an exemplary method for load balancing in a system including a data transformation accelerator is shown; Figure 3 A flowchart illustrating an exemplary method for load balancing in a system including a data transformation accelerator is shown; Figure 4 A block diagram of an exemplary system for load balancing is shown, including a data transformation accelerator and a virtual environment. Figure 5 A flowchart illustrating an exemplary method for load balancing in a system including a data transformation accelerator is shown; Figure 6 An exemplary computing device is shown; and Figure 7 Equations are shown for use with load balancing in systems that include data conversion accelerators. Detailed Implementation

[0012] Data transformation accelerators can be used in conjunction with host devices as coprocessor devices to accelerate data transformation operations for various applications, such as data analytics, big data, storage, and / or networking applications. Data transformation operations may include, but are not limited to, compression, decompression, encryption, decryption, authentication tag generation, authentication, data deduplication, Non-Volatile Memory Express (NVMe) Protection Information (PI) generation, NVMe PI verification, and real-time verification.

[0013] The host device may be coupled to a data conversion accelerator (e.g., a system), and the host software may be operable to submit commands to the data conversion accelerator. Computational resources on the data conversion accelerator (e.g., a data conversion engine) may execute the commands and return the converted data to the host software after the data conversion operation is complete. Alternatively or additionally, the converted data may be directed to different devices, such as one or more network interface cards and / or storage arrays.

[0014] In some cases, system throughput may be limited and / or system latency may increase because commands sent from the host device to the data conversion accelerator may use the resources contained therein in a suboptimal manner. Alternatively or additionally, some existing methods fail to implement the service categories associated with the execution of commands by the data conversion accelerator, which may limit and / or reduce the priority of operations performed by the system (including the data conversion accelerator).

[0015] At least some aspects of this disclosure address these and other drawbacks of prior methods by including load balancing operations performed by a portion of the system, enabling commands to be transmitted from the host device to the data conversion accelerator for data conversion operations in a more optimized manner relative to prior methods. In some embodiments, the system described in this disclosure may experience improved throughput and / or reduced latency during data conversion operations. Furthermore, in some aspects of this disclosure, load balancing may be combined with one or more service classes for command execution by the data conversion accelerator. Thus, the command submission process of the host device (and its associated throughput) can scale with the number of central processing units (CPUs) included in the host device until a threshold bandwidth associated with the data conversion accelerator is reached. In such cases, command submission contention from the host device to the data conversion accelerator may be reduced, command execution latency may be lowered, the throughput of the data conversion accelerator executing commands may be increased, and / or resource sharing may be more evenly distributed among data conversion engines within the data conversion accelerator and / or among multiple data conversion accelerators within the system, compared to prior methods.

[0016] In some cases, the host software can be operated to perform load balancing operations to select containers in which at least commands are stored. The host software can be further operated to direct command addresses associated with commands, which may be stored in multiple containers, to a data transformation accelerator. The data transformation accelerator can be operated to retrieve one or more commands from the various containers and can be operated to process input data using the command address associated with that command. The data transformation accelerator can perform data transformation operations on the input data and can transfer the transformed data to the host device.

[0017] In some cases, containers may be grouped into one or more container sets to provide a service category (or quality of service) through the computing resources of at least one data transformation accelerator (e.g., where one container set may be allocated for each service category). A load balancing operation can be performed to select a first container belonging to a specific container set of the service category to store the first command address. The selected service category or container set may be based on service category requirements associated with the first command, and the load balancing operation can be used to select containers belonging to that container set. Alternatively or additionally, the first command address may be transferred from the first container of the container set to the data transformation accelerator based on the service category. The data transformation accelerator can be operated to retrieve a first command and first input data associated with the first command address using the first command address from the first container.

[0018] In another embodiment, the system may include a host device operable to run one or more virtual machines. Multiple containers, possibly associated with at least one data transformation accelerator, may be grouped into one or more sets, and each set may be made available to virtual machines using input / output (I / O) virtualization. Software in each virtual machine may be operable to perform a load balancing operation for selecting a first container from a set of containers allocated to that virtual machine to store a first command address. Software on the virtual machine may be operable to transfer the first command address from the first container to at least one data transformation accelerator associated with that virtual machine. The data transformation accelerator may be operable to obtain a first command and first input data using the first command address. The data transformation accelerator may also be operable to perform a data transformation operation on the first input data using the first command, thereby generating transformed data. The data transformation accelerator may be further operable to transfer the transformed data to the virtual machine that submitted the command. The host software may also be operable to perform a load balancing operation for selecting containers to store second and subsequent commands generated by the host software on each virtual machine. Alternatively or additionally, in some cases, as described herein, a system configured to run virtual machines may be further operable to provide service categories associated with commands, command addresses, and / or containers.

[0019] Figure 1A A block diagram of an example system 100a for load balancing in a system 100a including a data conversion accelerator 120, according to at least one embodiment of the present disclosure, is shown. System 100a may include a host device 110 and a data conversion accelerator 120. Host device 110 may include a host processor 112, host memory 114, and host software 116. Host memory 114 may include a first container 115a and a second container 115b (collectively referred to as container 115). Data conversion accelerator 120 may include an internal processor 122, internal memory 124, and a data conversion engine 126.

[0020] In some embodiments, host device 110 (e.g., host computer, host server, etc.) may communicate with data conversion accelerator 120 via a data communication interface (e.g., Peripheral Component Interconnect High Speed ​​(PCIe) interface, Universal Serial Bus (USB) interface, and / or other similar data communication interfaces). In some embodiments, when a user requests the conversion of source data that may be located in host memory 114, host software 116 (e.g., a software driver) on host device 110 and operated by host processor 112 may be directed to generate metadata related to the conversion of source data in host memory 114 (e.g., but not limited to: data conversion command pre-data including command description, a list of descriptors dereferencing different parts of metadata, a list of descriptors dereferencing source data and target data buffers, command pre-data including conversion algorithms and related parameters, source and operation tokens describing different parts of source data and the conversion operations to be applied to the different parts, and / or additional command metadata). In some embodiments, host software 116 may generate metadata in host memory 114 based on source data that may be obtained from one or more sources. For example, source data may be obtained from storage (e.g., a storage device) associated with host device 110, a buffer associated with host device 110, a data stream from another device, etc. In these and other embodiments, obtaining source data may include copying or moving the source data to host memory 114.

[0021] In some embodiments, host software 116 may instruct host processor 112 to generate metadata associated with source data. For example, host software 116 may generate and / or submit one or more command requests to host processor 112, which may be associated with data transformation commands and may include command addresses. In some embodiments, metadata may be stored in one or more input buffers. For example, in some cases where the metadata includes a data transformation command that may contain a list of source descriptors, a list of target descriptors, command preconditions, a list of source tokens and operation tokens, and additional command metadata, each individual component of the metadata may be stored in a separate input buffer (e.g., the data transformation command in a first input buffer, the preconditions in a second input buffer, the source tokens and operation tokens in a third input buffer, and so on).

[0022] In some embodiments, the input buffer associated with metadata may be located in host memory 114. Alternatively or additionally, the input buffer associated with metadata may be located in internal memory 124. Alternatively or additionally, the input buffer may be located in both host memory 114 and internal memory 124. For example, one or more input buffers associated with metadata may be located in host memory 114, while one or more input buffers associated with metadata may be located in internal memory 124. In these and other embodiments, host processor 112 may boot host software 116 to reserve one or more output buffers that can be used to store output from data conversion accelerator 120. In some embodiments, the output buffer may be located in host memory 114. In some embodiments, the output buffer may be located in internal memory 124 of data conversion accelerator 120.

[0023] When host processor 112 receives a command request from host software 116 (including a request to generate metadata and store the metadata in internal memory 124 (e.g., an input buffer located in internal memory 124) and / or in host memory 114), host processor 112 may transmit the command to data conversion accelerator 120 (e.g., to components of data conversion accelerator 120, such as internal processor 122 and / or data conversion engine 126) via a data communication interface. For example, host processor 112 may access and / or address internal memory 124 via a data communication interface, and if the data communication interface is PCIe, internal memory 124 may be mapped into the address space of host device 110 using base address registers associated with PCIe endpoints (e.g., data conversion accelerator 120).

[0024] In some embodiments, host software 116 may direct (e.g., via host processor 112) data conversion accelerator 120 to process data conversion commands. For example, host software 116 may generate one or more command requests, each command request including a command address, and store the command addresses in one or more containers, such as a first container 115a and / or a second container 115b as described herein. Data conversion accelerator 120 may obtain command addresses that can point to data conversion commands.

[0025] In some embodiments, the command address and / or data conversion command may reside in host memory 114, for example in a first container 115a and / or a second container 115b. Alternatively or additionally, the command address may be programmed into data conversion accelerator 120 (e.g., during initialization of data conversion accelerator 120). In such cases, data conversion accelerator 120 (e.g., internal processor 122 and / or data conversion engine 126) may obtain the command address and / or access the data conversion command in host memory 114 using a data communication interface. Alternatively or additionally, the command address and / or data conversion command may reside in one or more containers located in internal memory 124, and the command address may be obtained by internal processor 122 and / or data conversion engine 126.

[0026] In some embodiments, the data conversion accelerator 120 may use data conversion commands to convert source data based on data conversion operations contained in the data conversion commands. In some embodiments, the data conversion operations may be executed by the data conversion engine 126 in accordance with the data conversion commands. In some embodiments, the data conversion engine 126 may be arranged according to the data conversion commands and / or metadata (e.g., metadata stored in host memory 114 and / or stored in internal memory 124) such that the data conversion engine 126 forms a data conversion pipeline that can be configured to perform data conversion operations on the source data.

[0027] The data conversion accelerator 120 and / or the components contained therein (e.g., internal processor 122, internal memory 124, and / or data conversion engine 126) can be implemented using various systems and / or devices. For example, the data conversion accelerator 120 can be implemented in hardware, software, firmware, field-programmable gate arrays (FPGAs), graphics processing units (GPUs), and / or any combination of the above implementations.

[0028] Data conversion accelerator 120 can be operated to perform data conversion operations using one or more pipelines, which include the configuration of data conversion engine 126. The pipelines in data conversion accelerator 120 can be described as performing data conversion operations in at least two directions (encoding direction and / or decoding direction). Encoding-direction data conversion operations performed by a first pipeline in data conversion accelerator 120 may include one or more of the following: NVMe PI verification of input data, compression, deduplication hash generation, padding, encryption, encrypted hash generation, NVMe PI generation of encoded data, and / or real-time verification of encoded data. Decoding-direction data conversion operations performed by a second pipeline in data conversion accelerator 120 may include one or more of the following: decryption (e.g., generating verification or not generating verification for input data and / or converted data), depadding, decompression, deduplication hash generation, and / or NVMe PI verification of encoded data and / or decoded data for input data and / or converted data (e.g., obtained from input data).

[0029] In these and other embodiments, host device 110 may use a data communication interface to transmit metadata to data conversion accelerator 120, which may be stored in internal memory 124 by internal processor 122, and internal processor 122 may return the command address of the stored metadata to host processor 112. Alternatively or additionally, host device 110 may use a data communication interface to directly transmit metadata to internal memory 124 of data conversion accelerator 120.

[0030] As previously described, host software 116 may submit one or more command requests to host processor 112 and / or data conversion accelerator 120. In response to a command request, a command structure may be generated, which may reside in host memory 114, internal memory 124, and / or a combination of host memory 114 and internal memory 124. Subsequently, the command address associated with the command structure may be stored in a first container 115a, a second container 115b, and / or in one or more containers arranged in internal memory 124 (not in...). Figure 1A As shown in the figure, and will be combined Figure 1A This is discussed as a container arranged in host memory 114, but it should be understood that the container may also be arranged in internal memory 124. In these and other embodiments, the command address may be accessed by data conversion accelerator 120.

[0031] Container 115 may be initialized during the initialization of data conversion accelerator 120. Container 115 may be one or more command pointer rings and may be operable to store command addresses that can be generated and / or requested by one or more commands from host software 116. In some cases, multiple threads on the CPU of host processor 112 and / or one or more applications of host software 116 may submit command requests for storing command addresses, and container 115 may be locked to achieve mutual exclusion, which may reduce or eliminate the possibility of race conditions associated with storing command addresses in the first container 115a and / or the second container 115b.

[0032] Host device 110 can implement load balancing operations to determine the specific container for storing command addresses. For example, host processor 112 can obtain a first command request and an associated first command address, and host software 116 can determine, taking into account the load balancing method in host device 110, to store the first command address in a second container 115b. In the case of host device 110 obtaining a large number of command requests, the load balancing operation and locking of container 115b help establish parallelism during command submission to data transformation accelerator 120.

[0033] In some cases, multiple data transformation accelerators may communicate with host device 110, and load balancing operations performed by host device 110 may be operated to load balance command addresses among containers that may be associated with the multiple data transformation accelerators. For example, a first data transformation accelerator may be associated with a first container and a second container, and a second data transformation accelerator may be associated with a third container and a fourth container. In this example configuration, host device 110 may obtain a command with a command address, and host device 110 may perform load balancing operations to store the command address in one of the first container, the second container, the third container, or the fourth container, thereby enabling host device 110 to perform load balancing operations among the multiple data transformation accelerators (e.g., by performing load balancing operations).

[0034] As described herein, load balancing helps improve and / or maximize the number of input / output operations per second (IOPS) between host device 110 and data conversion accelerator 120. Alternatively or additionally, load balancing operations help scale the commands submitted to data conversion accelerator 120 based on the number of CPU cores(one or more) of host processor 112, the number of threads running on those cores, and / or the number of applications in host software 116. Scale between host device 110 and data conversion accelerator 120 can continue until a threshold bandwidth associated with data conversion accelerator 120 is met. Alternatively or additionally, load balancing can reduce or eliminate command contention submitted to data conversion accelerator 120, reduce latency of commands submitted to data conversion accelerator 120, increase throughput of commands submitted to data conversion accelerator 120, and / or distribute resources among multiple data conversion accelerators (e.g., when system 100a contains more than one data conversion accelerator 120). Load balancing can be performed using one or more load balancing methods, as further described herein.

[0035] In some embodiments, host software 116 may include one or more components that facilitate load balancing between host device 110 and data conversion accelerator 120. For example, host software 116 may include a resource management module operable to manage resources related to command submission and / or load balancing; a load balancing module operable to determine containers storing command addresses; and / or a command submission module for submitting command addresses to selected containers that can be used by data conversion accelerator 120. In some cases, containers 115 may each be individually associated with a specific data conversion accelerator. For example, a first container 115a may be associated with data conversion accelerator 120, while a second container 115b may be associated with a second data conversion accelerator 120.

[0036] As described, one or more load balancing operations may be implemented in host device 110 (e.g., performed by a load balancing module in host software 116). Load balancing methods (e.g., methods for implementing load balancing) may include round-robin methods, queue depth-based methods, CPU core ring-based methods, service class methods, and / or the use of I / O virtual environments (as described herein). Figure 4 (as described).

[0037] Load balancing using a round-robin method may include a load balancing module in host software 116 obtaining an index associated with the most recently submitted command to data transformation accelerator 120 (and / or any other data transformation accelerator that may be coupled to host device 110). This index may be associated with any concurrently running thread in host device 110. For example, in the index (e.g., the most recently submitted container), ... p In this case, the load balancing module uses Figure 7 Equation 702 selects the next command pointer ring to submit this command, where N It is the number of containers associated with the data conversion accelerator 120, and M It is the number of data conversion accelerators coupled to host device 110. xmoduloy Operation is available x Divide by y (In other words, that is) p +1) divided by ( N * M The remainder is then output. The index can be updated for the next command submitted to the data transformation accelerator 120 from any source (e.g., the current thread, different threads, different applications, etc.). p Alternatively or additionally, the load balancing operation can be operated to run across multiple instances, one of which is configured to run for each data transformation accelerator device connected to host device 110. In such cases, the next container can be selected locally from among containers 115 available for a particular data transformation accelerator (e.g., data transformation accelerator 120), rather than considering all containers connected to all data transformation accelerators connected to host device 110.

[0038] Because multiple threads in host device 110 (e.g., threads associated with individual cores in host processor 112 and / or applications in host software 116) may simultaneously submit commands and / or access indexes stored in host memory 114. p Therefore, when updating the index p Previously, resource locking might have been performed by the resource management module in host software 116. Alternatively or additionally, indexing... p It can be operated on as an atomic variable.

[0039] The selected container (e.g., first container 115a) may be locked by host software 116 (e.g., via a command submission module) using hardware locks, software locks, and / or hardware-software lock mechanisms to achieve mutual exclusion. Alternatively or additionally, container 115 may implement a lock-free mechanism, allowing access to the container while storing command addresses within it. Alternatively or additionally, container 115 may be automatically configured or reconfigured between implementing a locking mechanism and implementing a lock-free mechanism. For example, in a first instance, host software 116 may instruct container 115 to implement a locking mechanism for mutual exclusion, as described herein. In a second instance, host software 116 may instruct container 115 to be reconfigured to a lock-free mechanism.

[0040] In some cases, container 115 may be implemented in software and may not include hardware-assisted lock-free mechanisms. In such cases, container 115 may utilize synchronization mechanisms and / or software data structures that may not involve lock / unlock primitives. Container 115 may facilitate concurrent storage commands by multiple threads within container 115. For example, a read-copy-update mechanism may be used to implement this mechanism.

[0041] In another instance, container 115 may be implemented in software and may include a hardware-assisted lock-free mechanism. In such a case, container 115 may be implemented with the assistance of data conversion accelerator 120, wherein one or more of containers 115 may be locked, and the command submission module of host software 116 may write a command address to the next available location of the selected container (e.g., the selected container as described or the next container), while the command submission module of host software 116 may unlock the container. If the selected container is full, the next container (e.g., the second container 115b) may be selected. In such a lock-free mechanism, a register may be provided in data conversion accelerator 120 for each container 115, and the thread submitting a command to container 115 may write the command address to the register instead of updating the read / write pointer of container 115. Data conversion accelerator 120 may push the command address from the register and update a particular container at that address using a mechanism capable of making such updates atomically. Before writing to a register, the software must ensure that writing to the register does not cause read / write pointer conflicts when the data conversion accelerator 120 reads a value from the register and writes the command address to the specific container by updating the write pointer. The number of threads running in the specific container can be known in advance in the host software 116. Before attempting to write the command address to the register, each thread must ensure that there are at least as many available entries in the specific container as the number of threads. Each time a command address is stored in the specific container, the data conversion accelerator 120 can update the register to indicate to the host software 116 the amount of free space in the specific container, which helps reduce read / write pointer-based operation conflicts. In some cases, threads can check the register (which may be read-only) before writing, rather than calculating the amount of free space in the specific container by reading the read / write pointer.

[0042] The container locking and unlocking described herein can be applied to any load balancing method described in this disclosure.

[0043] Load balancing using a queue depth-based method may include: once the command structure is generated in host memory 114 and / or internal memory 124 (in response to a command retrieval request), the load balancing module in host software 116 selects a container with fewer pending operations than other containers. For example, if a first container 115a contains a first number of pending operations and a second container 115b contains a second number of pending operations less than the first number, host software 116 may use a queue depth-based load balancing method to select the second container 115b to store the new command address.

[0044] The amount of pending operations in a specific container might be the total amount of source data submitted from that specific container that the data transformation accelerator 120 needs to process. Based on this consideration, the container with the least amount of pending operations can be selected. (The last sentence appears to be incomplete and possibly refers to a different context.) Figure 7Equation 704 selects the container in container 115, where d pi It is a container p The length of the source data for the i-th command. n p It is a container p The total number of pending commands in the system. M It is the number of containers associated with a data transformation accelerator, and N This represents the total number of data transformation accelerators. In cases where additional data transformation accelerators exist, and each additional accelerator is associated with a different number of containers, the selection process described above can consider all rings of all data transformation accelerators. Alternatively or additionally, load balancing operations can be performed independently across different data transformation accelerators, where each instance of the load balancing operation can consider containers available for a particular data transformation accelerator (e.g., containers available for other data transformation accelerators may not be considered).

[0045] In some embodiments, N Data transformation accelerators can be identical or similar to each other. For example, N Each of the data transformation accelerators can be associated with the same number of containers. Alternatively or additionally, N Each data transformation accelerator can be different from the others. For example, N Each of the data conversion accelerators can be associated with a different number of containers. For example, the first data conversion accelerator can be a 100Gbps device and can be associated with eight containers; while the second data conversion accelerator can be a 200Gbps device and can be associated with ten containers.

[0046] The selection of containers can be based on the size associated with the commands to be load-balanced by the host software 116. Alternatively or additionally, in a queue depth-based load balancing method, the host software 116 can select containers based on command size depth and / or data size depth. Command size depth can refer to the total number of pending commands within a particular container, and / or the total number of pending bytes relative to the pending commands, while data size depth can refer to the total number of data bytes of pending commands within a particular container. Alternatively or additionally, the total number of pending operations in any given container can be defined as the total number of pending commands in the container. Figure 7 Equation 706 selects the index as m A specific container with a minimum amount of work to be processed (e.g., relative to other containers), wherein n p It is a container p The total number of uncompleted commands in the system. M This refers to the number of containers in a data transformation accelerator. NThis represents the total number of data transformation accelerators. The available containers are enumerated as 1, 2, ..., M*N. In the case of multiple data transformation accelerators, each associated with a different number of containers, the above selection process must consider all containers across all data transformation accelerators. Alternatively or additionally, load balancing operations can be performed independently across different data transformation accelerators, where each instance of a load balancing operation can consider the containers available for a specific data transformation accelerator.

[0047] Alternatively or additionally, the total pending operations estimation in the container may take into account the command type of each pending command within the container, and one or more weights may be applied to that command type when determining the total pending operations. The weights may be determined based on the workload (e.g., execution latency) of the command type presented to the data transformation accelerator 120. For example, a command executing an XP10 compression algorithm with a 64KB history size may have a higher weight than another command executing an XP10 compression algorithm with a 16KB history size. In another example, a command executing an XP10 compression algorithm with a 16KB history size may have a higher weight than a command executing a chain of operations involving AES-CBC encryption with a 192-bit key size, padding, and inserting NVMe protection information (PI) into the encrypted padded data.

[0048] In some cases, weights may be predetermined based on priorities established relative to host device 110 and / or data conversion accelerator 120. For example, a read operation may have a higher priority than a write operation because a write data conversion operation may take longer than a read data conversion operation. In this example, the latency associated with a write data conversion operation (e.g., a data conversion operation performed in the encoding direction) may be hidden because there may be an expectation of the converted data (e.g., the requesting system may not know the amount of time from the write request to the actual write data conversion operation), while the latency associated with a read data conversion operation (e.g., a data conversion operation performed in the decoding direction) may not be hidden (e.g., the amount of time from the read request to the actual read data conversion operation may be measured by the requesting system). Alternatively or additionally, weights may be based on one or more service level protocols that may guide the execution of specific tasks (e.g., commands, data conversion operations, service groups, command groups, etc.) with determined priorities. During the operation of system 100a (or system 100b), the weights (and / or the weight-based direction of the input data) may be adjusted, for example, in conjunction with changes in the input data (e.g., patterns associated with the input data and / or workload associated with the input data).

[0049] In the example, in the event of a power failure to system 100b, system 100b can operate under battery backup. In such a case, read operations that might otherwise have had high priority (e.g., decode-oriented traffic) can be updated to no priority (e.g., stop operations), while write operations (e.g., encode-oriented traffic) can be assigned all priorities to save all unwritten data before battery backup fails. In another example, the priorities of normal operation can be established in conjunction with the heavy workload of system 100b. In subsequent times, the active traffic of system 100b may decrease. In response, priority traffic could be reads, comparisons, and / or writes used for deduplication operations; or, priority traffic could be decode and / or re-encode operations with larger command sizes, which can be operated to convert hot data into cold data, thereby improving the compression ratio of system 100b and / or improving the overall effective capacity of system 100b. Hot data can include any data that may have been requested and / or otherwise accessed within a threshold time period, which can be defined by the associated storage device. Alternatively or additionally, cold data may include data that was once hot data, but a threshold amount of time may have passed and the hot data has not been requested and / or accessed, allowing the associated storage device to update the hot data as cold data. Cold data may be stored in the associated storage device in a different manner relative to hot data, for example, using a more efficient compression algorithm and / or employing a different block size relative to hot data.

[0050] Alternatively or additionally, weights may include pre-assigned weights based on commands and / or operations associated with commands. For example, because an encoding operation with RTV may take approximately 1.6 times longer than a decoding operation of similar size, an encoding operation with RTV may be weighted relative to a decoding operation by 1.6 times.

[0051] Use command type and associated weight, and use Figure 7 Equation 708, choose index as m A specific container that contains fewer pending operations compared to other containers, wherein d pi It is a container p The length of the source data for the i-th command. n p It is a ring p The total number of uncompleted commands in the system. w pi It is the weight of the i-th command based on the operations performed in the command (e.g., compression, encryption, hashing, NVMe PI verification and insertion, and / or chained operations including one or more operations). In some cases, w iA selection can be made from a set of weights, regardless of the container. In cases where multiple data transformation accelerators exist, each associated with a different number of containers, a selection process using command type and associated weights can consider all containers across all data transformation accelerators. Alternatively or additionally, load balancing operations can be performed independently across different data transformation accelerators, where each instance of a load balancing operation can consider containers available on a specific data transformation accelerator.

[0052] Alternatively or additionally, the length of the source data in the command may not be used for computation of the operation to be processed in a particular container. Therefore, the specific container to store the command address can be selected based on having the smallest command weighted sum among all considered containers. The index of the selected container. m It is possible Figure 7 Equation 710 in the equation is determined, where w pi It is the weight of the i-th command based on the operations performed in the command (such as compression, encryption, hashing, NVMe PI verification and insertion, and / or chained operations including one or more operations). n p It is a container p The total number of uncompleted commands in the system. w pi Selecting an elective reorganization without considering the container, M It is the number of containers in a data transformation accelerator, while N This represents the total number of data transformation accelerators. In the example, when the calculation contains only two weights { w compression , w others}hour, w compression It can be used for commands involving compression operations, and w others This can be used with any other command (e.g., all non-compression operation commands). In cases where multiple data transformation accelerators exist, each associated with a different number of containers, the weighted selection process described above can consider all containers across all data transformation accelerators. Alternatively or additionally, load balancing operations can be performed independently across different data transformation accelerators, where each load balancing operation instance can consider the containers available on a particular data transformation accelerator.

[0053] Load balancing using a CPU core ring-based approach (or CPU core ring method) may include multiple command submission threads associated with host device 110, which may submit command requests via a command submission module in host software 116. For example, host processor 112 may be more than one processing device (e.g., multiple CPUs) and / or host processor 112 may include multiple cores (e.g., multiple cores per CPU), and each core of host processor 112 may be multi-threaded, where each thread (e.g., a command submission thread) can be operated to submit command requests. In some cases, command submission threads may be operated to run concurrently on multiple cores. In some embodiments, each command submission thread may be bound to a specific CPU core, which may prevent migration between multiple CPU cores. In some cases, to reduce command submission latency and / or improve command execution performance, a subset of containers from all containers 115 associated with all data transformation accelerators may be mapped to each CPU core (or a finite subset of CPU cores from which command submission threads can be operated). For example, host processor 112 may have a first core, a second core, and a third core, each core being coupled to a first container associated with a first data conversion accelerator, a second container associated with a second data conversion accelerator, and a third container associated with a third data conversion accelerator. The first core may include a first thread for submitting command requests to the first container, a second thread for submitting command requests to the second container, and so on, for each of the first, second, and third cores. Alternatively or additionally, more than one container associated with each data conversion accelerator may be coupled to each CPU core. For example, the first core may include a first thread for submitting command requests to the first container in the first data conversion accelerator, and a second thread (in the first core) for submitting command requests to the second container in the first data conversion accelerator. In these and other embodiments, the allocation between cores and containers may be completed by a resource management module in host device 110 or host software 116 during the initialization of data conversion accelerator 120.

[0054] The CPU core-ring-based load balancing method can be used by the command submission module in host software 116. Once the command submission module receives a command request, the command structure can be constructed in host memory 114 and / or internal memory 124. The load balancing module in host software 116 can use operating system utilities / primitives to determine the specific CPU core of the host processor 112 that executes the thread that calls the command submission module. Once the specific CPU core is determined, one of the containers can be selected and mapped to that specific CPU core. If more than one container is mapped to a specific CPU core, another load balancing method (e.g., the round-robin load balancing method or the queue depth-based load balancing method described herein) can be used to select containers from a subset of containers coupled to a specific CPU core across multiple containers.

[0055] Load balancing using the service class approach can be used when a data transformation accelerator contains multiple sets of data transformation engines. Figure 1B An example of a system 100b that can be operated to implement service class methods is shown, which can be used with Figure 1A System 100b is similar to system 100a. System 100b is relative to... Figure 1A Some differences in system 100a may include a first data conversion engine 126a, a second data conversion engine 126b (collectively referred to as data conversion engine 126 or multiple sets of data conversion engines 126), a first direct memory access (DMA) controller 128a, a second DMA controller 128b (collectively referred to as DMA controller 128), a first service class queue 130a, a second service class queue 130b, a third service class queue 130c, and a fourth service class queue 130d (collectively referred to as service class queue 130). Alternatively or additionally, the internal memory 124 of the data conversion accelerator 120 may store containers (e.g., this is related to...) Figure 1A The containers stored in host memory 114 as described herein may differ from, or be supplemented thereof, the containers which may include a first container 125a, a second container 125b, a third container 125c, a fourth container 125d, and a fifth container 125e (collectively referred to as container 125).

[0056] In some embodiments, multiple sets of data conversion engines 126 can be operated to perform different operations. For example, a first data conversion engine 126a can perform data conversion operations in both the encoding and decoding directions, while a second data conversion engine 126b can perform data conversion operations in the decoding direction. In some embodiments, each data conversion engine 126 can interface with a DMA controller 128 individually. For example, a first DMA controller 128a can interface with a first data conversion engine 126a, and a second DMA controller 128b can interface with a second data conversion engine 126b. The DMA controllers 128 can be independent of each other and can be operated to control the traffic flowing to the data conversion engines 126 (e.g., data conversion operations submitted from the host device 110) such that traffic in the first data conversion engine 126a does not interfere with traffic in the second data conversion engine 126b.

[0057] In this configuration, the data conversion accelerator 120 may support at least two service categories (e.g., forwarding groups) using a first data conversion engine 126a and a second data conversion engine 126b. Service categories may include secured forwarding groups (e.g., the first group) and / or accelerated forwarding groups (e.g., the second group). Secure forwarding may include assigning data conversion operations in either the encoding or decoding direction to the first group. Data conversion operations assigned to the first group may be processed by the first data conversion engine 126a. Because a mixture of encoding and decoding direction conversion operations may be performed by the first group using a first DMA controller 128a (e.g., a single DMA channel), encoding commands may cause head-of-line blocking for decoding commands. To circumvent this blocking problem, latency-sensitive decoding direction conversion operations may be assigned to the second group.

[0058] The second group facilitates decoding direction reversal operations, which can be handled by the second data conversion engine 126b. As illustrated in the example, the second group may only perform decoding direction reversal operations. Therefore, encoding commands will not cause head-of-line blocking of decoding commands as might occur when using the first group.

[0059] Figure 1BThe illustrated system 100b includes two sets of data conversion engines, but more or fewer data conversion engines can be implemented in system 100b, which may support additional service categories (or forwarding groups). For example, another accelerated forwarding group can be used to send a power failure notification to system 100b, which can then temporarily switch to battery backup. In such examples, some low-priority commands (e.g., write commands) may be cached in memory so that system 100b can execute higher-priority commands (e.g., read commands, decode commands, etc.), and / or some cached low-priority write commands may be flushed to storage. Furthermore, low-priority commands may become the commands processed by system 100b if system 100b is unable to handle additional commands and / or is in a timer state (e.g., due to operation using battery backup). Alternatively or additionally, flushed low-priority write commands may be stored in the highest-priority queue used for encoding-direction operations. Furthermore, to implement Quality of Service (QoS), bandwidth can be allocated to the accelerated forwarding queue when traffic is present; and bandwidth can be reduced or limited when the accelerated forwarding queue is empty.

[0060] In some embodiments, the allocation of data conversion engines 126 in the first and second groups can be hardwired into the data conversion accelerator 120. Alternatively or additionally, the allocation of data conversion engines 126 can be configured during the initialization of the data conversion accelerator 120. For example, software in the host device 110 or firmware associated with an internal processor 122 (e.g., an embedded CPU) on the data conversion accelerator 120 can configure the allocation of data conversion engines 126. A first set of containers 125 in the internal memory 124 (e.g., first container 125a and second container 125b) can be used to submit commands to the first data conversion engine 126a, while a second set of containers 125 (e.g., third container 125c, fourth container 125d, and fifth container 125e) can be used to submit commands to the second data conversion engine 126b. As described in this disclosure, containers 125 can be disposed in the internal memory 124 (e.g., Figure 1B (as shown) and / or in host memory 114 (such as Figure 1A (As shown).

[0061] In some embodiments, the first data conversion engine 126a and / or the second data conversion engine 126b may each be associated with one or more queues that can be operated to provide service categories associated with the data conversion engine 126. As shown, the first data conversion engine 126a may be associated with a first service category queue 130a and a second service category queue 130b, and the second data conversion engine 126b may be associated with a third service category queue 130c and a fourth service category queue 130d.

[0062] Service categories can be strict priority queues, weighted round-robin (WRR) queues, and / or other service category queues. Alternatively or additionally, service category queues 130 may have different priorities based on the implemented service category, and each service category queue may be assigned a peak bandwidth limit. For example, a strict priority queue may have a higher priority than a WRR queue. In another example, a first strict priority queue may have a higher priority than a second strict priority queue, and the first and second strict priority queues may each be assigned a peak bandwidth limit. Service category queues may be configured with priorities and peak bandwidth limits for strict priority arbitration, and / or configured with weights for WRR arbitration, thereby sharing the bandwidth of the associated data transformation engine 126.

[0063] Commands from the command pointer ring can be retrieved from these queues and processed by the transformation engine using the queue's priority. Commands in container 125 can be retrieved from service category queue 130 and processed by the associated data transformation engine 126 using the priority of service category queue 130. In some cases, commands may contain command tags that can be used to identify the service category to which the command belongs. For example, DMA controller 128 can retrieve the command tag associated with a command, and DMA controller 128 can be operated to classify the command into one of the service category queues 130.

[0064] Using strict priority arbitration, commands from the highest priority service category queue can be processed first by the associated data transformation engine 126. Commands from lower priority service category queues can be processed after all commands from higher priority service category queues (e.g., at least the highest priority service category queue) have been processed, or after commands from higher priority service category queues have exhausted the peak bandwidth limit associated with that higher priority service category queue. For example, if there are waiting commands in both the first service category queue 130a and the second service category queue 130b, and the first service category queue 130a has a higher priority than the second service category queue 130b, the first service category queue 130a will be served until it is empty or the peak bandwidth limit associated with it is reached, and then the second service category queue 130b will be served.

[0065] When using WRR arbitration, the number of commands that can be processed from one of the service class queues 130 within a time interval can be proportional to the weight of that queue (e.g., the higher the weight associated with that service class queue, the more commands can be processed from that service class queue). In some embodiments, because the priority of a WRR queue may be lower than that of a strict priority queue, the arbitrator may dispatch commands from one of the WRR queues (if present) once each of the strict priority queues reaches its peak bandwidth limit. If one or more service class queues 130 have no commands to process, their bandwidth can be automatically occupied by other service class queues 130 that contain commands to be processed.

[0066] If a service category is disabled and / or not yet established (e.g., no strict priority queue, WRR queue, etc.), commands from all service category queues 130 can be dispatched among the service category queues 130 using a round-robin method. As the service category queues 130 are labeled with the appropriate category (e.g., strict priority or WRR) and the bandwidth of the strict priority queue is allocated and / or the weights of the round-robin queues are set, the service categories in the first data conversion engine 126a and / or the second data conversion engine 126b can be configured.

[0067] After service category queue 130 is established, subsets of containers 125 can be mapped to service category queue 130 respectively, where containers 125 can share time slot resources allocated to each service category queue 130. For example, the first container 125a is mapped to the first service category queue 130a, the second container 125b is mapped to the second service category queue 130b, the third container 125c is mapped to the third service category queue 130c, and the fourth container 125d and the fifth container 125e are mapped to the fourth service category queue 130d. In some cases, there may be more than one data transformation accelerator attached to host device 110, and / or each attached data transformation accelerator may contain more than one service category queue and / or more than one data transformation engine group. In each service category queue and / or each transformation engine group, multiple service categories can be defined, and containers associated with each data transformation accelerator can be mapped to different service categories.

[0068] In some embodiments, the number of containers 125 that can be associated with each service category may be determined and / or established during the initialization of the data transformation accelerator 120. During the operation of the data transformation accelerator 120, the association between containers 125 and their service categories may be static. In some cases, containers 125 may be reconfigurable, and this reconfiguration may be based on user requests or determined needs (e.g., an imbalance between the number of commands assigned to a service category and the number of containers assigned to that service category). For example, a fourth container 125d may be associated with a fourth service category queue 130d at initialization (e.g., ...). Figure 1B (as shown), and can be reconfigured to be associated with the third service category queue 130c, for example, based on the first workload associated with the third service category being greater than the second workload associated with the fourth service category.

[0069] In some embodiments, host device 110 may be operated to perform multiple load balancing operations at various stages of the load balancing operations described herein. For example, host device 110 may perform a first load balancing operation to select a second data transformation engine 126b (instead of the first data transformation engine 126a), a second load balancing operation to select a fourth service class queue 130d (instead of the third service class queue 130c), and a third load balancing operation to select a fourth container 125d (instead of the fifth container 125e). At each stage of load balancing, host device 110 may employ one or more of the load balancing methods described herein to determine where a particular command might be distributed.

[0070] As part of the initialization of the data conversion accelerator 120, software (e.g., a resource management module) in the host device 110 can assign priorities (e.g., strict priority, WRR, or a combination of strict priority and WRR) to each service category in the data conversion accelerator. When a service category is assigned strict priority, the resource management module can configure peak bandwidth for the strict priority queue. When a service category is assigned WRR, the resource management module can configure shared bandwidth weights, where the bandwidth of the WRR queue may be the remaining bandwidth after the peak bandwidth configured for the strict priority queue.

[0071] After priority and bandwidth allocation are completed, service category groups can be created, with each service category served by one queue in service category queue 130. Alternatively or additionally, container 125 can be mapped to each of service category queues 130. For example, in the sample data transformation accelerator, there are eight queues with different priorities, and each data transformation engine group in this sample data transformation accelerator has 64 containers. The first group of eight containers can be assigned to the first service category queue, the second group of eight containers can be assigned to the second service category queue, and so on. Alternatively or additionally, a load balancing method can be defined for each service category container, where different load balancing methods can be defined for different container groups mapped to different service categories.

[0072] In the case where the data conversion accelerator 120 is a storage and / or encrypted data conversion accelerator, commands that can be assigned to service categories can be grouped into one or more acceleration sessions in the software running on the host device 110. An acceleration session may include command groups, and service categories in each group of the data conversion engine 126 can be provided to these commands.

[0073] For example, a user of software on host device 110 can create three acceleration sessions in the first data conversion engine 126a: a first acceleration session using strict priority and operable to receive 50 Gbps of command execution throughput on data conversion accelerator 120; a second acceleration session and a third acceleration session, each using WRR weights set to 0.6 and 0.4 respectively. In the case where the total command execution throughput in the first data conversion engine 126a of data conversion accelerator 120 is 100 Gbps, the first acceleration session can receive a maximum of 50 Gbps, and the remaining throughput (up to 50 Gbps) can be shared between the second and third acceleration sessions in a round-robin manner, with the second acceleration session receiving approximately 60% and the third acceleration session receiving approximately 40%. When a user submits a command to the command submission module, the user can specify a specific data conversion engine group to execute the command (e.g., by specifying guaranteed forwarding or accelerated forwarding) and the service class within that specific data conversion engine group (e.g., by specifying strict priority or WRR).

[0074] Upon receiving a command, the command submission module can determine the specific service category and coordinate with the load balancing module in the host device 110 to determine the specific container in the container 125 mapped to that specific service category, to which the command can be submitted.

[0075] Modifications, additions, or omissions may be made to system 100a or system 100b without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to aid in the explanation of the concepts presented herein and is not restrictive. Furthermore, system 100a or system 100b may contain any number of other elements, or may be implemented in a system or context other than those described herein. For example, Figure 1A or Figure 1B Any component can be divided into more components or merged into fewer components.

[0076] Figure 2 A flowchart of an example method 200 for submitting commands to containers using load balancing is shown. This method may begin at block 202, where the host device (e.g., Figure 1A The host software of the host device 110 (e.g., Figure 1A An application in the host software (116) can submit a command request, and a command submission module in the host software can receive the command request.

[0077] At block 204, the command submission module can generate a command based on the command request and store the command in memory, such as host memory in the host device (e.g., ...). Figure 1A (in the host memory 114) and / or associated data conversion accelerator (e.g., Figure 1A The internal memory on the data conversion accelerator 120 (e.g., Figure 1A Internal memory 124).

[0078] At block 206, the load balancing module in the host software can determine a specific container (e.g., Figure 1A The first container 115a or the second container 115b is used to store the command address associated with the command, and the system can bootstrap to store the command address in the specific container.

[0079] At block 208, the specific container can be locked by the host software (e.g., via the command submission module) using hardware locks, software locks, and / or hardware-software lock mechanisms to achieve mutual exclusion.

[0080] At block 210, the host software's command submission module can write the command address to the next available location in that particular container. If that particular container is full, the next container can be selected.

[0081] At block 212, after the command address is written to the specific container, the host software's command submission module can unlock the specific container so that additional command addresses can be written to the specific container and / or the command address can be retrieved from it to process the command, for example, by a data conversion accelerator.

[0082] After the command address is stored in that specific container, and continuing as an example operation, the data transformation accelerator can retrieve the command address from the host device and the associated command stored in memory. The data transformation accelerator can acquire input data (which may include various metadata), configure a data transformation pipeline based on the input data, and / or perform data transformation operations on at least a portion of the input data to generate transformed data.

[0083] Subsequently, the data conversion accelerator can guide the storage of the converted data to one or more output buffers (which may be established within a command structure associated with the command). The output buffers may be located in host memory, internal memory, a combination of host memory and internal memory, and / or other remote storage devices (e.g., NVMe storage arrays or network interface cards). In some embodiments, the data conversion accelerator may be operable to consume a specific command and retrieve the associated command address from that specific container after performing a data conversion operation and generating the converted data. In such cases, a second specific container (e.g., a result container) may be associated with the specific command, and the result container may store a specific tag associated with the specific command (e.g., the result container may store tags for each completed command, which may be written by the data conversion accelerator upon completion). When the data conversion operation associated with the specific command is completed, the data conversion accelerator may notify the host device, and the host device may retrieve the specific tag to identify the specific command and / or the output buffer used to store the result of the data conversion operation associated with the specific command. For example, the tag may be stored in the result container and may contain the address of the completed command. The host software on the host device can read the output buffer associated with the command by dereferencing the address in the tag, and obtain the converted data. After consuming the command, the host software can reuse the space occupied by the command for future commands.

[0084] Alternatively, or additionally, the data conversion accelerator may provide the host device with notification (e.g., interrupts and / or flags) that a data conversion operation has been performed and / or that the converted data is available in the output buffer. The host device may receive this notification (e.g., from interrupts and / or polling flags) and may access the converted data in the output buffer. Alternatively, or additionally, the host device may transfer the converted data to an application in the host software that generates a command request associated with the converted data.

[0085] Figure 3A flowchart illustrating an example method 300 for load balancing in a system including a data conversion accelerator according to at least one embodiment of the present disclosure is shown. Method 300 may be executed by processing logic, which may include hardware (circuit, dedicated logic, etc.), software (e.g., software running on a general-purpose computer system or a dedicated machine), or a combination of both, which may be contained in any computer system or device, such as host device 110 of FIG1.

[0086] For simplicity, the methods described herein are depicted and described as a series of actions. However, actions according to this disclosure can occur in various orders and / or concurrently, accompanied by other actions not presented and described herein. Furthermore, not all actions shown are applicable to implementing the methods according to the disclosed subject matter. Moreover, those skilled in the art will understand and appreciate that the method may alternatively be represented by a state diagram or events as a series of interrelated states. Furthermore, the methods disclosed in this specification can be stored on an article of art (e.g., a non-transitory computer-readable medium) to facilitate the transfer and assignment of these methods to a computing device. As used herein, the term "article of art" is intended to encompass a computer program accessible from any computer-readable device or storage medium. Although illustrated as discrete blocks, various blocks may be divided into additional blocks, merged into fewer blocks, or eliminated, depending on the desired implementation.

[0087] At block 302, multiple command requests can be obtained. Each command request may include a command address. The multiple command requests may be generated by one or more software applications. The software application may include one or more threads, and each thread may be operated to generate the command requests among the multiple command requests.

[0088] At block 304, a load balancing operation can be performed to select a first container among multiple containers. Performing the load balancing operation can scale the number of multiple command addresses transmitted to the data transformation accelerator to the bandwidth limit associated with that data transformation accelerator. Alternatively or additionally, load balancing can be performed using one or more of the following methods: round-robin, queue depth, CPU core ring, and / or service class.

[0089] In some instances, the first container may be locked to achieve mutual exclusion before the first command address is stored in it. Alternatively or additionally, the first container may be unlocked after the first command address is stored in it.

[0090] Multiple containers can be operated on to store at least the command address, and a first container can be operated on to store a first command address. In some instances, the multiple containers can be a command pointer ring, and can be operated on to store multiple command addresses.

[0091] The first command address may point to a first command and first input data. In some instances, the data transformation accelerator may use the first command to perform data transformation operations on the first input data. The first input data may include at least source data, metadata, and additional data. The additional data may include one or more of the following: an initialization vector for encryption or decryption, a message authentication code, metadata for data compression, and / or authentication data for encryption or decryption.

[0092] At block 306, the first command address may be stored in the first data container. In some instances, a first software application may store the first command address in the first data container, and a second software application may store the second command address in a second data container.

[0093] At block 308, the first command address can be transferred from the first data container to the data conversion accelerator. In some instances, a first set of containers among multiple containers can be associated with the first data conversion accelerator, such that the first set of command addresses stored in the first container set can be transferred to the first data conversion accelerator and will not be transferred to the second data conversion accelerator.

[0094] At block 310, the converted data can be obtained from the data conversion accelerator.

[0095] Method 300 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, in some embodiments, a first command address from a first container may be used in response to obtaining the transformed data. In another example, the designation of different elements in the manner described herein is intended to aid in the explanation of the concepts presented herein and is not restrictive. Furthermore, method 300 may include any number of other elements or may be implemented in a system or context other than the system or context described herein.

[0096] Figure 4A block diagram of an example system 400 according to at least one embodiment of the present disclosure is shown. This example system 400 is used for load balancing in a system 400 including a data transformation accelerator 420 and input / output (I / O) virtualization. System 400 may include a host device 410 and a data transformation accelerator 420. Host device 410 may include a host processor 412, host memory 414, and host software 416. Host memory 414 may include a first container 415a and a second container 415b, collectively referred to as container 415. Host software 416 may include a host operating system 440, a hypervisor 442, a first virtual machine 444a, and a second virtual machine 444b, collectively referred to as virtual machine 444. First virtual machine 444a may include first virtual machine software 446a, and second virtual machine 444b may include second virtual machine software 446b, wherein first virtual machine software 446a and second virtual machine software 446b may be collectively referred to as virtual machine software 446. Data transformation accelerator 420 may include an internal processor 422, internal memory 424, and a data transformation engine 426.

[0097] System 400 and / or components of system 400 may be identical or similar to components of system 100 and / or system 100. For example, host device 410, host processor 412, host memory 414, container 415, host software 416, data conversion accelerator 420, internal processor 422, internal memory 424, and data conversion engine 426 may be identical or similar to host device 110, host processor 112, host memory 114, container 115, host software 116, data conversion accelerator 120, internal processor 122, internal memory 124, and data conversion engine 126 in FIG. 1, respectively. Alternatively or additionally, as described herein, system 400 and / or components of system 400 may be operable to perform the same or similar operations as components of system 100 and / or system 100.

[0098] In some instances, as described herein, system 400 may be identical to system 100 and may include support for I / O virtualization. In some instances, hypervisor 442 may be operable to manage virtual machine 444 on host device 410. Hypervisor 442 may be software, firmware, and / or hardware included in host device 410 for creating and / or running virtual machine 444. Virtual machine software 446 may be operable to perform operations involving host device 410 (e.g., relative to container 415) and / or involving data transformation accelerator 420. Thus, virtual machine 444 may be operable to communicate at least with host device 410. For example, data may be transferred between virtual machine 444 and host device 410, load balancing may be performed by virtual machine software 446, and / or commands and / or command addresses may be submitted from virtual machine 444 to data transformation accelerator 420 (e.g., such as via host device 410 and / or host software 416). In another example, data in a buffer (e.g., an output buffer that can be manipulated to store the output from the data conversion accelerator 420) can be accessed by the host device 410 and / or the virtual machine 444.

[0099] In some instances, data transformation accelerator 420 may be used with host device 410, which may include I / O virtualization (e.g., single root I / O virtualization). Each instance of virtual machine 444 may be operated to communicate with data transformation accelerator 420 via host device 410 (e.g., virtual machine software 446, such as communicating with host software 416), such that virtual machine 444 can utilize data transformation accelerator 420 to perform data transformation acceleration operations. In some embodiments, containers 415 may be divided into multiple groups, wherein each group of containers 415 may be associated with a group of data transformation accelerators (e.g., such as data transformation accelerator 420) and / or with a group of data transformation engines 426 operated by the data transformation accelerators. For example, a first container 415a may be associated with a first group of data transformation engines 426, and a second container 415b may be associated with a second group of data transformation engines 426, wherein both the first and second groups are associated with data transformation accelerators. In another example, a first container 415a may be associated with a first data transformation accelerator, and a second container 415b may be associated with a second data transformation accelerator.

[0100] exist Figure 4 In this context, the data transformation engine 426 is shown as a single group; however, the data transformation engine 426 can be combined with... Figure 1B The data conversion engine 426 is the same as or similar to the data conversion engine 126 in the data conversion engine 426, which may include multiple groups. Alternatively or additionally, container 415 may be arranged in host memory 414 and / or internal memory 424, and container 415 may be a single container or may represent multiple containers, such as a group of containers.

[0101] In some embodiments, dividing container 415 into multiple groups of containers may be hardwired in data conversion accelerator 420, may be configured by host software 416, and / or may be configured by firmware on an embedded CPU (e.g., internal processor 422, which may coordinate with host software 416) in data conversion accelerator 420, for example during the initialization of system 400 and / or the initialization of data conversion accelerator 420.

[0102] Alternatively or additionally, for each group of data conversion engines 426, one or more service categories may be implemented during the initialization of the data conversion accelerator 420 by firmware in the host software 416 and / or the internal processor 422. Service categories may be related to... Figure 1B The system described in system 100b has the same or similar service categories, but the system may not include I / O virtualization, as per the description. Figure 4 As stated above.

[0103] In response to the setting of service categories, as described herein, it may include establishing categories, such as strict priority, WRR, etc., with the peak bandwidth of the category (e.g., for a strict priority category) and / or the weight of sharing bandwidth among multiple categories (e.g., for a WRR category) configured by host software 416 and / or firmware on internal processor 422 guided by host software 416.

[0104] Containers 415 associated with each group of data conversion engines 426 may be divided into different groups, and each group of containers 415 may be assigned to each service category within each group of data conversion engines 426. In some instances, the partitioning of containers 415 may be hardwired on the data conversion accelerator 420. Alternatively or additionally, the partitioning of containers 415 may be configured by host software 416 and / or by firmware on the internal processor 422, wherein the partitioning may occur during the initialization of system 400 and / or the initialization of data conversion accelerator 420. Alternatively or additionally, data conversion engines 426 may not include service categories, and accordingly, containers 415 associated with data conversion engines 426 may not be partitioned into groups, as previously described.

[0105] In some embodiments, containers 415 in each service category may be subdivided into one or more subsets, where each subset may be assigned to a virtual function of virtual machine 444. A virtual function may be an operation performed and / or requested by virtual machine 444 relative to data transformation accelerator 420, and / or an operation performed by data transformation accelerator 420 (e.g., a data transformation operation). Each virtual machine 444 using one or more virtual functions (e.g., first virtual machine 444a and second virtual machine 444b) may obtain a subset of containers 415, which may be mapped to different service categories in each group of data transformation engine 426. In some instances, host device 410 may determine the total number of containers 415 available to work with virtual machine 444, and host device 410 may assign containers 415 to virtual machine 444 based on the number of virtual machines 444 communicating with host device 410.

[0106] In system 400, host software 416 may be operated to perform service class configuration and / or to allocate containers 415 between service classes and / or between virtual functions executed by virtual machine 444. Alternatively or additionally, host software 416 may boot firmware in internal processor 422 to perform service class configuration and / or allocate containers 415 between service classes. For example, in a PCIe environment, service class configuration and container 415 allocation may be performed by physical functions (PFs).

[0107] After the service category configuration and / or container 415 allocation are completed, container 415 becomes available to virtual machine software 446, where commands can be submitted from virtual machines (e.g., one of virtual machines 444) to data transformation accelerator 420. Virtual machine software 446 within virtual machine 444 can submit commands to containers 415 that can be allocated to virtual machines 444. For example, in an instance where a first container 415a is allocated to a first virtual machine 444a and a second container 415b is allocated to a second virtual machine 444b, first virtual machine software 446a can submit commands to first container 415a, and second virtual machine software 446b can submit commands to second container 415b. The load balancing method described herein can be configured by virtual machine software 446 within each virtual machine 444.

[0108] Virtual machine software 446 running on one of the virtual machines 444 can access the container 415 allocated to it. For example, host software running on a first virtual machine 444a can access the container 415 allocated to the first virtual machine 444a (e.g., first container 415a), and a second virtual machine 444b can access the container allocated to the second virtual machine 444b (e.g., second container 415b). Alternatively or additionally, the load balancing method described herein can be configured and / or operated by virtual machine software 446 running on one of the virtual machines 444. As described herein, the established load balancing method can be applied to the containers 415 allocated to the respective virtual machines.

[0109] In these and other embodiments, the virtual machine software 446 in each virtual machine 444 can be operated to perform load balancing operations using containers 415 that may be allocated to the virtual machine 444. For example, in an instance where a first container 415a represents a plurality of first containers allocated to a first virtual machine 444a and a second container 415b represents a plurality of second containers allocated to a second virtual machine 444b, the first virtual machine software 446a can perform load balancing operations involving the first container 415a, and the second virtual machine software 446b can perform load balancing operations involving the second container 415b. Alternatively or additionally, the virtual machine 444 has limited visibility to containers not allocated to itself. Referring to the foregoing examples, the first virtual machine 444a (and / or the first virtual machine software 446a) may not have visibility to the second container 415b, and the second virtual machine 444b (and / or the second virtual machine software 446b) may not have visibility to the first container 415a.

[0110] When virtual machine software 446 submits a command to data transformation accelerator 420, virtual machine software 446 can perform load balancing operations to determine the specific container 415 from which the command can be submitted. The selection of the load balancing module on the virtual machine may be limited to container 415 that may be visible to a particular virtual machine.

[0111] In the example, for the first data transformation engine group (data transformation engine 426), three service categories can be established, where the first category is a strict priority category, and the second and third categories are weighted round-robin (WRR) categories. Four containers (e.g., container 415) can be assigned to the first category, six containers can be assigned to the second category, and six containers can be assigned to the third category. Furthermore, two of the four containers assigned to the first category can be assigned to virtual functions mapped to the first virtual machine (e.g., first virtual machine 444a), while the remaining two containers assigned to the first category can be assigned to virtual functions mapped to the second virtual machine (e.g., second virtual machine 444b).

[0112] Three of the first six containers assigned to the second category may be assigned to the first virtual function mapped to the first virtual machine. The remaining three containers assigned to the second category may be assigned to the second virtual function mapped to the second virtual machine. Alternatively or additionally, three of the last six containers assigned to the third category may be assigned to the first virtual function mapped to the first virtual machine. The remaining three containers assigned to the third category may be assigned to the second virtual function mapped to the second virtual machine.

[0113] For the second data transformation engine group, three WRR categories can be created. Four containers can be assigned to the first WRR category, six containers to the second WRR category, and six containers to the third WRR category. Two of the four containers in the first WRR category can be assigned to virtual functions mapped to the first virtual machine. The remaining two containers in the first WRR category can be assigned to virtual functions mapped to the second virtual machine.

[0114] Three of the six containers in the second WRR category may be assigned to virtual functions mapped to the first virtual machine, while the remaining three containers in the second WRR category may be assigned to virtual functions mapped to the second virtual machine. Alternatively or additionally, three of the six containers in the third WRR category may be assigned to virtual functions mapped to the first virtual machine, while the remaining three containers in the third WRR category may be assigned to virtual functions mapped to the second virtual machine.

[0115] The foregoing examples are provided for illustrative purposes only. In a system implementing I / O virtualization, there may be two or more virtual machines. Alternatively or additionally, there may be more or fewer containers and / or more or fewer data transformation engine groups. Furthermore, the number of containers associated with virtual machines may be more or less than the number described. In some cases, containers may be located in the memory of host device 410 and / or in data transformation accelerator 420.

[0116] System 400 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, system 400 may include any number of other elements, or may be implemented in a system or context other than those described herein. For example, Figure 4 Any component can be divided into more components or merged into fewer components.

[0117] Figure 5A flowchart illustrating an example method 500 for load balancing in a system including a data conversion accelerator according to at least one embodiment of the present disclosure is shown. Method 500 may be executed by processing logic, which may include hardware (circuit, dedicated logic, etc.), software (e.g., software running on a general-purpose computer system or a dedicated machine), or a combination of both. This processing logic may be contained in any computer system or device, such as… Figure 4 The host device 410.

[0118] For simplicity, the methods described herein are depicted and described as a series of actions. However, actions according to this disclosure can occur in various orders and / or concurrently, accompanied by other actions not presented and described herein. Furthermore, not all actions shown are applicable to implementing the methods according to the disclosed subject matter. Moreover, those skilled in the art will understand and appreciate that the method may alternatively be represented by a state diagram or events as a series of interrelated states. Furthermore, the methods disclosed in this specification can be stored on an article of art (e.g., a non-transitory computer-readable medium) to facilitate the transfer and assignment of these methods to a computing device. As used herein, the term "article of art" is intended to encompass a computer program accessible from any computer-readable device or storage medium. Although illustrated as discrete blocks, various blocks may be divided into additional blocks, merged into fewer blocks, or eliminated, depending on the desired implementation.

[0119] At block 502, a command request can be obtained from the virtual machine. This command request may include a command address. In some cases, the command address may point to a first command and first input data. The first input data may include source data, metadata, and / or additional data. Additional data may include one or more of an initialization vector for encryption or decryption, a message authentication code, and / or authentication data for encryption or decryption.

[0120] At block 504, a load balancing operation can be performed to select a first container among multiple containers allocated to a virtual machine. Based on the load balancing operation, it can be determined which of the multiple containers (e.g., the first container) can be used to store command addresses. Load balancing can be performed using at least one of the following methods: round-robin, queue depth, CPU core ring, and / or service class.

[0121] In some cases, a first subset of multiple containers may be associated with a first service category that can be assigned to a virtual machine. Alternatively or additionally, a second subset of multiple containers may be associated with a second service category that can be assigned to a virtual machine.

[0122] At block 506, the command address may be stored in a first container. The first container may be locked to achieve mutual exclusion before the command address is stored therein. Alternatively or additionally, the first container may be unlocked after the command address is stored therein. In some cases, the first container may implement a lock-free mechanism, allowing it to be accessed while the command address is stored within it. Alternatively or additionally, the first container may be automatically configured or reconfigured between implementing a locking mechanism and implementing a lock-free mechanism.

[0123] At block 508, the first command address can be transmitted from the first container to the data transformation accelerator. In some instances, the data transformation accelerator can use the first command to perform a data transformation operation on the first input data. The load balancing described herein can be performed to scale the number of command addresses transmitted to the data transformation accelerator to the bandwidth limit associated with that data transformation accelerator.

[0124] In some instances, a first set of containers among multiple containers may be associated with a first data conversion accelerator, such that a first set of command addresses stored in the first set of containers can be transmitted to the first data conversion accelerator but not to a second data conversion accelerator.

[0125] At block 510, the converted data can be obtained from the data conversion accelerator.

[0126] At block 512, the virtual machine can facilitate access to the transformed data. For example, the transformed data can be set in one or more output buffers accessible to the virtual machine, allowing the virtual machine to retrieve the transformed data from the output buffers.

[0127] Method 500 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, in some embodiments, the command address may be removed from the first container in response to obtaining the converted data.

[0128] In another example, a second command request can be obtained from a second virtual machine. The second command request may include a second command address. A second load balancing operation can be performed to select a second container from a second plurality of containers that can be allocated to the second virtual machine. The second command address may be stored in the second container. In some cases, the plurality of containers may be associated with a first data transformation engine group in a data transformation accelerator, while the second plurality of containers may be associated with a second data transformation engine group in the data transformation accelerator.

[0129] In another example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, method 500 may include any number of other elements, or may be implemented in a system or context other than those described.

[0130] Figure 6 An example computing device 600 is illustrated, in which a set of instructions can be executed to cause a machine to perform any or more methods discussed herein. The computing device 600 may include a mobile phone, smartphone, netbook, rack server, router computer, server computer, personal computer, mainframe computer, laptop computer, tablet computer, desktop computer, or any computing device having at least one processor, etc., in which a set of instructions can be executed to cause a machine to perform any or more methods discussed herein. In alternative implementations, the machine may be connected (e.g., networked) to other machines in a local area network, intranet, extranet, or the Internet. The machine may operate as a server machine in a client-server network environment. The machine may include a personal computer (PC), set-top box (STB), server, network router, switch, or bridge, or any machine capable of executing a set of instructions (orderly or otherwise) to specify the operations to be performed by the machine. Furthermore, although only a single machine is shown, the term "machine" may also include any collection of machines that individually or collectively execute a set (or more) of instructions to perform any or more methods discussed herein.

[0131] The computing device 600 includes a processing device 602 (e.g., a processor), a main memory 604 (e.g., a read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM)), a static memory 606 (e.g., flash memory, static random access memory (SRAM)), and a data storage device 616, which communicate with each other via a bus 608.

[0132] Processing device 602 represents one or more general-purpose processing devices (e.g., microprocessors, central processing units, etc.). More specifically, processing device 602 may include a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets or combinations thereof. Processing device 602 may also include one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. Processing device 602 is configured to execute instructions 626 to perform the operations and steps discussed herein.

[0133] The computing device 600 may also include a network interface device 622 capable of communicating with the network 618. The computing device 600 may also include a display device 610 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and a signal generation device 620 (e.g., a speaker). In at least one implementation, the display device 610, the alphanumeric input device 612, and the cursor control device 614 may be combined into a single component or device (e.g., a liquid crystal touchscreen).

[0134] Data storage device 616 may include computer-readable storage medium 624 on which one or more instruction sets 626 embodying any one or more methods or functions described herein are stored. During execution of the instructions 626 by computing device 600, the instructions 626 may reside wholly or at least partially in main memory 604 and / or processing device 602, which also constitute computer-readable media. The instructions may also be transmitted or received on network 618 via network interface device 622.

[0135] Although computer-readable storage medium 624 is shown as a single medium in the example implementation, the term "computer-readable storage medium" can include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) storing one or more sets of instructions. The term "computer-readable storage medium" can also include any medium capable of storing, encoding, or carrying sets of instructions for machine execution and causing the machine to perform any one or more methods of this disclosure. Therefore, the term "computer-readable storage medium" should be understood to include, but is not limited to, solid-state memory, optical media, and magnetic media.

[0136] The terms used in this disclosure, especially in the appended claims (e.g., the text of the appended claims), are generally intended to be “open-ended terms” (e.g., the term “comprising” should be interpreted as “including but not limited to”).

[0137] Furthermore, if there is an intent to describe a specific number of claims in an introduced claim, such intent will be explicitly stated in the claim, and if no such statement is made, such intent does not exist. For example, to aid understanding, the appended claims may include the use of the introductory phrases “at least one” and “one or more” to introduce the claim reference. However, the use of such phrases should not be construed as implying that a claim reference introduced by the indefinite article “a” or “an” would limit any particular claim containing such an introduced claim reference to only one implementation of such a claim, even if the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted as meaning “at least one” or “one or more”); the same applies to the use of definite articles used to introduce claim references.

[0138] Furthermore, even when a specific number of claims are explicitly cited, those skilled in the art will recognize that such citations should be interpreted as referring to at least the number cited (e.g., a simple citation of "two citations," without further embellishment, implies at least two citations, or two or more citations). Additionally, in the use of conventions such as "at least one of A, B, and C" or "one or more of A, B, and C," such constructions are generally intended to include A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.

[0139] Furthermore, any extractive word or phrase preceding two or more alternative terms in the description, claims, or drawings should be understood to include the possibility of including one, any, or both terms. For example, the phrase "A or B" should be understood to include the possibility of including "A" or "B" or "A and B".

[0140] All examples and conditional language cited in this disclosure are intended for pedagogical purposes to aid the reader in understanding this disclosure and the concepts contributed by the inventors to advance the technology, and should be construed as not being limited to these specifically cited examples and conditions. Although implementations of this disclosure have been described in detail, various changes, substitutions, and alterations may be made without departing from the spirit and scope of this disclosure.

Claims

1. A method comprising: Obtain multiple command requests, each of which includes a command address from multiple command addresses; Perform a load balancing operation to select the first container out of multiple containers to store the first command address; Store the first command address in the first container; Transmit the first command address from the first container to the data conversion accelerator; and The converted data is obtained from the data conversion accelerator.

2. The method of claim 1, further comprising removing the first command address from the first container in response to obtaining the converted data.

3. The method according to claim 1, wherein, The plurality of command requests are generated by one or more software applications, and each software application includes one or more threads, each thread being operable to generate a command request among the plurality of command requests.

4. The method according to claim 1, wherein, The first command address points to a first command and first input data, and the data conversion accelerator uses the first command to perform a data conversion operation on the first input data.

5. The method according to claim 4, wherein, The first input data includes source data, metadata, and additional data, wherein the additional data includes one or more of the following: an initialization vector for encryption or decryption, a method authentication code, and authentication data for encryption or decryption.

6. The method according to claim 1, wherein, The plurality of containers are command pointer rings used to store the plurality of command addresses.

7. The method according to claim 1, wherein, The first software application stores the first command address in the first container, and the second software application stores the second command address in the second container.

8. The method according to claim 1, wherein, The load balancing is performed to scale the number of multiple command addresses transmitted to the data conversion accelerator to the bandwidth limit associated with the data conversion accelerator.

9. The method according to claim 1, wherein, The load balancing is performed based on the service category, and the load balancing uses at least one of the following methods: round-robin, queue depth, or CPU core ring.

10. The method according to claim 1, wherein, The first container is locked to achieve mutual exclusion before the first command address is stored in the first container, and the first container is unlocked after the first command address is stored in the first container.

11. The method according to claim 1, wherein, The first container set of the plurality of containers is associated with a first data conversion accelerator, such that a first command address set stored in the first container set is transmitted to the first data conversion accelerator, but not to the second data conversion accelerator.

12. A system comprising: Host equipment, including: One or more processors; and Host software, operated by the one or more processors, to: Generate multiple command requests, each of which includes multiple command addresses; Perform a load balancing operation to select a first container out of a plurality of containers to store the first command address out of the plurality of command addresses; The bootstrap stores the address of the first command in the first container; and Transmit the first command address from the first container; and At least one data transformation accelerator is used for: Obtain the address of the first command transmitted from the first container; and The converted data is then transmitted to the host device.

13. The system according to claim 12, wherein, In response to transmitting the converted data to the host device, the at least one data conversion accelerator also consumes a first command address from the first container.

14. The system according to claim 12, wherein, The load balancing is performed based on the service category, and the load balancing uses one of the following methods: round-robin, queue depth, or CPU core ring.

15. The system according to claim 12, wherein, The at least one data conversion accelerator is also used for: Use the first command address to obtain the first command and first input data associated with the plurality of command requests; and The first command is used to perform a data transformation operation on the first input data to generate the transformed data.

16. The system according to claim 15, wherein, The first input data includes source data, metadata, and additional data, wherein the additional data includes one or more of the following: an initialization vector for encryption or decryption, a method authentication code, and authentication data for encryption or decryption.

17. The system according to claim 12, wherein, Each of the one or more processors includes one or more cores and one or more threads in each of the one or more cores, and each of the one or more threads is operable to generate a command request in the plurality of command requests.

18. The system according to claim 17, wherein, The first thread bootloader in the one or more processors stores the first command address in a first container, and the second thread bootloader in the one or more processors stores the second command address in a second container.

19. The system according to claim 12, wherein, The first container is locked to achieve mutual exclusion before the first command address is stored in the first container, and the first container is unlocked after the first command address is stored in the first container.

20. The system according to claim 12, wherein, Transmitting the converted data to the host device includes outputting the converted data to one or more output buffers accessible by the host device.