Computing device and data processing method
By introducing auxiliary processors and special application processors into computer systems, offloading part of data processing work and freeing up the burden from general-purpose processors, solving the problem of inefficient data processing in traditional computer architectures and achieving more efficient data processing.
Patent Information
- Application Number
- CN202110219103.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-26
- Filing Date
- 2021-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-05-23
AI Technical Summary
In traditional computer architectures, the central processor needs to spend a lot of time on data transfer and processing, resulting in inefficient data processing and an innovative method is needed to accelerate data processing.
The auxiliary processor processes control flows, and the special application processor processes data flows, and the special application processor processes data flows, enabling hardware acceleration.
This method can perform data processing without the intervention of a general-purpose processor, reduce the load of the general-purpose processor, save power consumption, reduce delay, and improve data processing efficiency.
Smart Images

Figure CN112882984B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing, and more particularly to a data processing method and a related computing device for offloading at least a portion of data processing work from at least one general purpose processor through at least one auxiliary processor and at least one special application processor. Background Art
[0002] According to the traditional computer architecture, the storage device can transmit and receive data with the central processor through the bus. For example, a solid-state drive (SSD) can be connected to a PCIe (Peripheral Component Interconnect Express) bus or a SATA (Serial Advanced Technology Attachment) bus. In this way, the central processor of the host can write data to the solid-state drive of the host through the PCIe bus / SATA bus, and the solid-state drive of the host can also transmit the stored data to the central processor of the host through the PCIe bus / SATA bus. In addition, with the development of network technology, the storage device can also be set remotely and connected to the host through the network. In this way, the central processor of the host can write data to the remote storage device through the network, and the remote storage device can also transmit the stored data to the central processor of the host through the network.
[0003] Regardless of whether the storage device is installed on the host side or is set up in a remote storage device, the application program executed on the central processing unit needs to read data from the storage device for processing based on the traditional computer architecture. Since data movement through the central processing unit consumes a lot of time, in order to speed up the efficiency of data processing, an innovative data processing method and related computing device are urgently needed. Summary of the invention
[0004] Therefore, one of the objectives of the present invention is to provide a data processing method and a related computing device for offloading at least a portion of data processing work from at least one general-purpose processor through at least one auxiliary processor and at least one special application processor.
[0005] In one embodiment of the present invention, a computing device is disclosed. The computing device includes at least one general-purpose processor, at least one auxiliary processor, and at least one special application processor. The at least one general-purpose processor is used to execute an application, wherein at least a portion of data processing of a data processing task is offloaded from the at least one general-purpose processor. The at least one auxiliary processor is used to process a control flow of the data processing without the intervention of the application executed on the at least one general-purpose processor. The at least one special application processor is used to process a data flow of the data processing without the intervention of the application executed on the at least one general-purpose processor.
[0006] In another embodiment of the present invention, a data processing method is disclosed, which includes: executing an application program through at least one general-purpose processor, wherein at least a portion of data processing of a data processing task is offloaded from the at least one general-purpose processor; and processing a control flow of the data processing through at least one auxiliary processor and processing a data flow of the data processing through at least one special application processor without intervention of the application program executed on the at least one general-purpose processor.
[0007] The computing device of the present invention may be equipped with a network subsystem to connect to the network and may perform relevant data processing for object storage, and therefore has extremely high scalability. In one application, the computing device of the present invention may be compatible with existing object storage services (such as Amazon S3 or other cloud storage services), and thus may perform relevant data processing for data acquisition on the object storage device to which the computing device is connected based on object storage instructions (such as Amazon S3 Select) from the network. In another application, the computing device of the present invention may receive NVMe / TCP instructions from the network, and perform relevant data processing operations on the storage device to which the computing device is connected based on the NVMe / TCP instructions. If the storage device to which the computing device is connected is part of a distributed storage system (such as part of a key-value database), the NVMe / TCP instructions received from the network may include key-value instructions, and the computing device of the present invention may perform relevant key-value database data processing operations on the storage device based on the key-value instructions. In addition, data can be processed through the hardware acceleration circuit during the movement process, and there is no need for the general-purpose processor executing the application to intervene in the data movement and the communication between the software and the hardware. Therefore, internal network operations and / or storage internal operations can be realized, thereby saving power consumption, reducing delays and reducing the load of the general-purpose processor; furthermore, the computing device of the present invention can be implemented using a multi-processing system chip. For example, the multi-processing system chip can include a field programmable logic gate array and a general-purpose processor core using an ARM architecture, so it has high design flexibility. Designers can design the application / program code to be executed by the general-purpose processor core and the hardware data processing acceleration function to be possessed by the field programmable logic gate array according to their needs. For example, the computing device of the present invention can be applied to a data center, and can be customized to support various data types and unit formats and obtain optimal performance. Since a single multi-processing system chip can replace a high-end server, a data center using the computing device of the present invention can have a lower construction cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A schematic diagram of a computer system using an accelerator card.
[0009] Figure 2 is a schematic diagram of a computing device according to an embodiment of the present invention.
[0010] Figure 3 A schematic diagram of the functional correspondence between the computing device of the present invention and a computer system using an accelerator card.
[0011] Figure 4 FIG. 1 is a schematic diagram of a computing device that uses virtual storage memory technology to perform object data processing according to an embodiment of the present invention.
[0012] Figure 5 FIG. 4 is a schematic diagram of a programmable circuit according to an embodiment of the present invention.
[0013] Figure 6 FIG. 4 is a schematic diagram of a programmable circuit according to another embodiment of the present invention. DETAILED DESCRIPTION
[0014] Figure 1Schematic diagram of a computer system using an accelerator card. The computer system 100 includes a computer host 102, a storage device 104, and an accelerator card 105. The computer host 102 includes a central processing unit 106, a system memory 108, a network interface 110, an input / output interface 112, and other components (not shown). The storage device 104 can be connected to the computer host 102 through the network interface 110 or the input / output interface 112. The accelerator card 105 can be connected to the computer host 102 through the input / output interface 112. For example, the network interface 110 can provide wired network access or wireless network access, and the input / output interface 112 can be a PCIe interface. In this example, the accelerator card 105 is a PCIe adapter card that can be installed in the slot of the input / output interface 112. For example, the accelerator card 105 can be an adapter card based on a field programmable gate array (FPGA), which can be used for data processing acceleration and other applications. Compared with the calculation by the central processor 106, the accelerator card 105 has the advantages of larger data throughput, lower processing delay and lower power consumption. When the computer system 100 is running, the central processor 106 will first move the storage data to be processed in the storage device 104 to the system memory 108, and then move the storage data to be processed in the system memory 108 to the memory 114 in the accelerator card 105. Then, the accelerator card 105 will read the data in the memory 114 to perform calculations and return the calculation results to the central processor 106. In addition, the central processor 106 may need to first convert the format of the storage data to be processed in the storage device 104, and then move the data that meets the data format required by the accelerator card 105 to the memory 114 on the accelerator card 105. Although the accelerator card 105 can provide a larger data throughput, lower processing latency and lower power consumption, the accelerator card 105 is connected to the input / output interface 112 (e.g., PCIe interface) of the computer host 102. Therefore, the central processing unit 106 still needs to intervene in the data movement and / or data format conversion between the storage device 104 and the accelerator card 105. For example, according to the traditional computer architecture, the central processing unit 106 needs to process multiple layers of the input / output stack, and based on the von Neumann architecture, the central processing unit 106 needs to perform frequent load / store operations. In this way, even if the computer system 100 is additionally installed with the accelerator card 105, the overall data processing performance of the computer system 100 will not be significantly improved due to the influence of these factors.Furthermore, the accelerator card 105 is directly connected to the input / output interface 112 (eg, PCIe interface) of the computer host 102 , so the accelerator card 105 itself lacks scalability.
[0015] In order to improve the above-mentioned shortcomings of the computer system 100 using the acceleration card 105, the present invention proposes a new hardware acceleration architecture. Figure 2 FIG. 2 is a schematic diagram of a computing device according to an embodiment of the present invention. The computing device 200 includes at least one general purpose processor 202, at least one coprocessor 204, and at least one application specific processor 206. For the sake of simplicity, Figure 2Only one general purpose processor 202, one auxiliary processor 204, and one special application processor 206 are shown. However, in actual application, the number of general purpose processors 202, auxiliary processors 204, and special application processors 206 can be determined according to the requirements, and the present invention is not limited thereto. In the present embodiment, the general purpose processor 202, auxiliary processor 204, and special application processor 206 are all disposed in the same chip 201. For example, the chip 201 can be a multiprocessor system on a chip (MPSoC), but the present invention is not limited thereto. In addition, the general processor 202 includes at least one general processor core (not shown), the auxiliary processor 204 includes at least one general processor core 212 and a portion of the programmable circuit 208, and the special application processor 206 includes another portion of the programmable circuit 208. For example, the general processor core can adopt the x86 architecture or the ARM (Advanced RISC Machine) architecture, and the programmable circuit 208 can be a field programmable gate array. In this embodiment, the general processor 202 (consisting only of the general processor core) and the auxiliary processor 204 (consisting of the general processor core and the field programmable gate array) are heterogeneous processors. The general processor 202 is responsible for executing the application program APP, while the processing and data movement of the input / output stack are completely handed over to the auxiliary processor 204. In addition, since the special application processor 206 is implemented using a field programmable gate array, compared to data processing through the general-purpose processor 202, the special application processor 206 can provide advantages such as higher data throughput, lower processing latency and lower power consumption. For example, the special application processor 206 can escape the von Neumann architecture and thus can use other architectures such as pipelines and data flows to provide data parallel processing capabilities, thereby having better data processing performance.
[0016] In the application of using a multi-processing system chip to implement the computing device 200, the processor core of the general processor 202 can be an application processor unit (APU) implemented by ARM Cotex-A53, the general processor core 212 can be a real-time processor unit (RPU) implemented by ARM Cotex-R5, and the programmable circuit 208 is a field programmable gate array. Figure 3Schematic diagram of the functional correspondence between the computing device of the present invention and the computer system using an accelerator card. The computer system 300 can be a traditional server, which can execute the application 302, the operating system kernel 304, the file system 306 and the driver 308 through the central processing unit (not shown). In addition, the computer system 300 is provided with a network interface 310 to provide network access function, and can use the storage device 312 to provide large-capacity data access, and can provide data processing acceleration function through the accelerator card 314 (such as a PCIe adapter card). The computing device 200 can be implemented using a multi-processing system chip 320, based on Figure 2 As shown in the structure, the multi-processing system chip 320 may include an application processor unit 324, a real-time processor unit 326 and a field programmable logic gate array 328, wherein the network subsystem 330, the storage subsystem 332 and the acceleration circuit 334 are all implemented by the field programmable logic gate array 328. Figure 3 As shown, the function of the application processor unit 324 corresponds to the application 302, the function of the network subsystem 330, the real-time processor unit 326, the storage subsystem 332 and the storage device 322 corresponds to the network interface 310, the operating system kernel 304, the file system 306, the driver 308 and the storage device 312, and the function of the acceleration circuit 334 corresponds to the acceleration card 314. Different from the computer system 300, the multi-processing system chip 320 can unload at least part of the data processing work from the application processor unit 324 through the real-time processor unit 326 and the field programmable logic gate array 328 (especially the acceleration circuit 334), so as to avoid the application processor unit 324 spending a lot of time processing data movement. The details of the computing circuit 200 of the present invention are further described as follows.
[0017] Please refer to Figure 2The general purpose processor 202 is used to execute the application program APP, wherein at least a portion (i.e., part or all) of the data processing of a data processing task is offloaded from the general purpose processor 202. In other words, the general purpose processor 202 does not need to intervene in at least a portion of the data processing of the data processing task. In this way, for the data processing flow, the general purpose processor 202 does not need to perform any processing on multiple layers of the input / output stack. The auxiliary processor 204 is used to process a control flow of the data processing without the intervention of the application program APP executed on the general purpose processor 202. In addition, the special application processor 206 is used to process a data flow of the data processing without the intervention of the application program APP executed on the general purpose processor 202. In this embodiment, the application program APP executed on the general purpose processor 202 can offload at least a portion of the data processing of the data processing task to the auxiliary processor 204 and the special application processor 206 by calling the application programming interface (API) function API_F.
[0018] The general purpose processor core 212 of the auxiliary processor 204 can load and execute program code SW to execute and control the processing of multiple layers of the input / output stack. In particular, the general purpose processor core 212 communicates with the programmable circuit 208 to allow the entire data flow to be processed smoothly without the intervention of the general purpose processor 202. In addition, the auxiliary processor 204 also includes a network subsystem 214, a storage subsystem 216, and a plurality of data converter circuits 234, 236 implemented by the programmable circuit 208. The network subsystem 214 includes a transmission control protocol / internet protocol (TCP / IP) offload engine 222 and a network handler 224. The TCP / IP offload engine 222 is used to process the TCP / IP stack between the network processing circuit 224 and the network-attached device 10. For example, the network-attached device 10 can be a client or an object storage device in a distributed object storage system and is connected to the computing device 200 via the network 30. Therefore, the instructions or data of the distributed object storage system can be transmitted to the computing device 200 via the network 30. Since the TCP / IP offload engine 222 is responsible for the processing of the network layer, the general processor core 212 does not need to intervene in the processing of the TCP / IP stack. The network processing circuit 224 is used to communicate with the general processor core 212 and control the network flow.
[0019] In this embodiment, the special application processor 206 is implemented by a programmable circuit 208 and includes at least one accelerator circuit 232. For the sake of simplicity, Figure 2 Only one acceleration circuit 232 is shown. However, in actual application, the number of acceleration circuits 232 may be determined according to demand. For example, each acceleration circuit 232 is designed to execute a kernel function. Therefore, the special application processor 206 may be provided with a plurality of acceleration circuits 232 to respectively execute different kernel functions. Please note that the kernel function itself does not perform any processing on the multiple layers of the input / output stack.
[0020] The acceleration circuit 232 is designed to provide a hardware data processing acceleration function, and can receive a data input from the network processing circuit 224, and process the data stream of the data processing of at least a part of the data processing work according to the data input. If the data format of the payload data obtained from the network stream is different from the predetermined data format required by the acceleration circuit 232, the data conversion circuit 234 will be used to process the data conversion between the network processing circuit 224 and the acceleration circuit 232. For example, the payload data obtained and output by the network processing circuit 224 from the network stream will include a complete data, and the core function to be executed by the acceleration circuit 232 only needs to process a specific field in the complete data, so the data conversion circuit 234 will obtain the specific field from the complete data and transmit it to the acceleration circuit 232. In addition, if the network mounting device 10 is part of a distributed object storage system and is connected to the computing device 200 through the network 30, the network processing circuit 224 can be used to control the network stream between the acceleration circuit 232 and the network mounting device 10.
[0021] The storage subsystem 216 includes a storage handler 226 and a storage controller 228. The storage handler 226 is used to communicate with the general processor core 212 and control data access to the storage device 20. For example, the storage handler 226 can perform information transmission, synchronization processing and data flow control in response to the application programming interface function related to data access. The storage controller 228 is used to perform actual data storage on the storage device 20. For example, the storage device 20 can be a solid state drive connected to the computing device 200 through the input / output interface 40 (such as a PCIe interface or a SATA interface), and the storage controller 228 will output a write instruction, a write address and write data to the storage device 20 to write data, and will output a read instruction and a read address to the storage device 20 to read data.
[0022] The acceleration circuit 232 is designed to provide a hardware data processing acceleration function. The acceleration circuit 232 can receive a data input from the storage processing circuit 226, and process the data stream of the data processing of at least a part of the data processing work according to the data input. If the data format of the data obtained from the storage processing circuit 226 is different from the predetermined data format required by the acceleration circuit 232, the data conversion circuit 236 will be used to process the data conversion between the storage processing circuit 226 and the acceleration circuit 232. For example, the data obtained and output by the storage processing circuit 226 will include a complete data, and the core function to be executed by the acceleration circuit 232 only needs to process a specific field in the complete data, so the data conversion circuit 236 will obtain the specific field from the complete data and transmit it to the acceleration circuit 232.
[0023] For the sake of brevity, Figure 2 In the figure, only one data processing circuit 234 is shown between the acceleration circuit 232 and the network processing circuit 224, and only one data processing circuit 236 is shown between the acceleration circuit 232 and the data processing circuit 226. However, in actual application, the number of data processing circuits 234 and 236 may be determined according to requirements. For example, in another embodiment, the special application processor 206 may include a plurality of acceleration circuits 232, which respectively execute different core functions. Since different core functions may have different data format requirements, a plurality of data processing circuits 234 may be set between the special application processor 206 and the network processing circuit 224 to perform different data format conversions, and a plurality of data processing circuits 236 may be set between the special application processor 206 and the data processing circuit 226 to perform different data format conversions.
[0024] As described above, the general processor 202 can offload at least a portion of the data processing work to the auxiliary processor 204 and the special application processor 206, wherein the auxiliary processor 204 is responsible for the control flow of the data processing (at least including the processing of multiple layers of the input / output stack), and the special application processor 206 is responsible for the data flow of the data processing. In this embodiment, the computing device 200 also includes a control channel 218, which is coupled between the pins of the special application processor 206 (especially the acceleration circuit 232) and the pins of the auxiliary processor 204 (especially the general processor core 212). The control channel 218 can be used to transmit control information between the special application processor 206 (especially the acceleration circuit 232) and the auxiliary processor 204 (especially the general processor core 212).
[0025] In one application, the acceleration circuit 232 may receive a data input from the network processing circuit 224, and transmit a data output of the acceleration circuit 232 through the network processing circuit 224, that is, the data from the network mounted device 10 is processed by the acceleration circuit 232 and then written back to the network mounted device 10. Since the data is processed in the path between the network mounted device 10 and the acceleration circuit 232 without passing through the general processor 202, in-network computing can be achieved. In another application, the acceleration circuit 232 may receive a data input from the network processing circuit 224, and transmit a data output of the acceleration circuit 232 through the storage processing circuit 226. That is, the data from the network mounted device 10 is processed by the acceleration circuit 232 and then written to the storage device 20. Since the data is processed in the path between the network mounted device 10, the acceleration circuit 232 and the storage device 20 without passing through the general processor 202, in-network computing can be achieved. In another application, the acceleration circuit 232 may receive a data input from the storage processing circuit 226 and transmit a data output of the acceleration circuit 232 through the network processing circuit 224, that is, the data from the storage device 20 is processed by the acceleration circuit 232 and then written to the network mounting device 10. Since the data is processed in the path of the storage device 20, the acceleration circuit 232 and the network mounting device 10 without passing through the general processor 202, in-storage computation can be achieved. In another application, the acceleration circuit 232 may receive a data input from the storage processing circuit 226 and transmit a data output of the acceleration circuit 232 through the storage processing circuit 226. That is, the data from the storage device 20 is processed by the acceleration circuit 232 and then written back to the storage device 20. Since the data is processed in the path of the storage device 20 and the acceleration circuit 232 without passing through the general processor 202, in-storage computation can be achieved.
[0026] Unlike file storage, object storage is a non-hierarchical data storage method that does not use a directory tree. Instead, discrete data units (objects) exist at the same level in the storage area, and each object has a unique identifier for applications to access the object. Object storage is widely used in cloud storage, and the computing device 200 disclosed in the present invention can also be used for data processing in object storage devices.
[0027] In an object storage application, the application APP executed on the general processor 202 can offload at least part of the data processing work to the auxiliary processor 204 and the special application processor 206 by calling the application programming interface function API_F. For example, the special application processor 206 is designed to process a kernel function with a kernel identifier. The data processing is used to process an object located in an object storage device (object storage) and having an object identifier (object identifier), and the parameters of the application programming interface function API_F may include the kernel identifier and the object identifier, wherein the object storage device may be a storage device 20 or a network mounted device 10.For example, the storage device 20 is a solid state drive and is connected to the computing device 200 via a PCIe interface. Therefore, the computing device 200 and the storage device 20 can be considered as a computational storage device (CSD) as a whole. In addition, the computational storage device can be used as a part of a distributed object storage system. Therefore, the storage device 20 can be used to store a plurality of objects, and each object has its own object identifier. The application APP executed on the general processor 202 can call the application programming interface function API_F to offload the operation of object data processing to the auxiliary processor 204 and the special application processor 206. For example, the application programming interface function API_F can include csd_sts csd_put(object_id, object_data, buf_len), csd_sts csd_put_acc(object_id, object_data, acc_id, buf_len), csd_sts csd_get(object_id, object_data, buf_len) and csd_sts csd_get_acc(object_id,object_data,acc_id,buf_len), wherein csd_sts csd_put(object_id,object_data,buf_len) is used to write the object data object_data having the object identifier object_id to the storage device 20, csd_sts csd_put_acc(object_id,object_data,acc_id,buf_len) is used to use the acceleration circuit 232 having the core identifier acc_id to process the object data object_data having the object identifier object_id and write the corresponding operation result to the storage device 20, csd_sts csd_get(object_id,object_data,buf_len) is used to read the object data object_data having the object identifier object_id from the storage device 20, and csd_sts csd_get_acc(object_id, object_data, acc_id, buf_len) is used to send the object data object_data with the object identifier object_id read from the storage device 20 to the acceleration circuit 232 with the core identifier acc_id for processing, and transmit the corresponding operation result.
[0028] For example, the operation of csd_sts csd_put(object_id, object_data, buf_len) can be simply represented by the following pseudo code:
[0029] struct nvme_cmd io;
[0030] io.opcode = nvme_sdcs;
[0031] io.object_id = object_id;
[0032] io.object_data=&object_data;
[0033] io.xfterlen = buf_len;
[0034] return ioctl(fd,NVME_IOCTL_SUBMIT_IO,&io)
[0035] In addition, the operation of csd_sts csd_get_acc(object_id, object_data, acc_id, buf_len) can be simply represented by the following virtual program code:
[0036] struct nvme_cmd io;
[0037] io.opcode = nvme_sdcs;
[0038] io.object_id = object_id;
[0039] io.object_data=&object_data;
[0040] io.acc_id = acc_id;
[0041] io.xfterlen = buf_len;
[0042] return ioctl(fd,NVME_IOCTL_SUBMIT_IO,&io)
[0043] Please note that the above virtual program codes are only used as examples for illustration, and the present invention is not limited thereto. In addition, the application programming interface function API_F actually used by the computing device 200 may also be determined according to actual design requirements.
[0044] In another object storage application, the network mount device 10 may be a client in a distributed object storage system and connected to the computing device 200 via the network 30. In addition, the storage device 20 may be a part of the distributed storage system (e.g., a part of a key-value store). The acceleration circuit 232 is designed to execute a core function with a core identifier, and an object with an object identifier is stored in the storage device 20. The network mount device 10 can send a specific application programming interface function through the network 30, whose parameters include the core identifier and the object identifier. Therefore, the network subsystem 214 in the computing device 200 will receive the specific application programming interface function (whose parameters include the core identifier and the object identifier) from the network 30, and then the general processor core 212 will obtain the core identifier and the object identifier from the network subsystem 214, and trigger the core function with the core identifier (that is, the acceleration circuit 232) to process the object located in an object storage device (that is, the storage device 20) and having the object identifier, wherein the acceleration circuit 232 in the special application processor 206 processes the object without the intervention of the application APP executed on the general processor 202.
[0045] As described above, the special application processor 206 is implemented by using a field programmable logic gate array. Since the internal memory capacity of the field programmable logic gate array is very small, the memory capacity that can be used by the special application processor 206 (especially the acceleration circuit 232) is limited. However, if the computing device 200 is used for data processing of an object storage device, the computing device 200 can also use the virtual storage memory technology of the present invention, so that the on-chip memory / embedded memory (such as Block RAM (BRAM) or UltraRAM (URAM)) can be equivalently regarded as having a large capacity like a storage device. Further, the general processor core 212 triggers a core function (i.e., the acceleration circuit 232) with the core identifier to process an object located in an object storage device (i.e., the storage device 20) and having the object identifier according to the core identifier and the object identifier. Based on the characteristics of object storage, the continuous data of the object with the object identifier can be continuously read to the on-chip memory / embedded memory used by the special application processor 206 (especially the acceleration circuit 232) through a data stream according to the capacity of the on-chip memory / embedded memory used by the special application processor 206 (especially the acceleration circuit 232). memory for the special application processor 206 (especially the acceleration circuit 232) to process until the data processing of the object with the object identifier is completed. In addition, during the object data processing, the data movement between the storage device 20 and the embedded memory used by the special application processor 206 (especially the acceleration circuit 232) and the synchronization between the core function and the application APP will be the responsibility of the general processor core 212 of the auxiliary processor 204. Therefore, the application APP executed by the general processor 202 does not need to intervene in the data movement between the storage device 20 and the on-chip memory / embedded memory used by the special application processor 206 (especially the acceleration circuit 232) and the synchronization between the core function and the application APP.
[0046] Figure 4 FIG. 2 is a schematic diagram of a computing device that uses virtual storage memory technology to process object data according to an embodiment of the present invention. Assume that the computing device 200 is implemented by a multi-processing system chip, and the storage device connected to the computing device 200 is an object storage device 412. The multi-processing system chip can be divided into a processing system (PS) 400 and a programmable logic module (PL) 401, wherein the processing system 400 includes an application processor unit 402 (for implementing Figure 2The general purpose processor 202 shown in the figure) and the real-time processor unit 404 (for implementing Figure 2 The general purpose processor core 212 shown in FIG. 4 is a general purpose processor core 212 shown in FIG. 4 , and the programmable logic module 401 includes an acceleration circuit 406 (for implementing Figure 2 The acceleration circuit 232 shown in FIG. 1 ), an on-chip memory 408 (such as BRAM or URAM), and a storage controller 410 (for implementing Figure 2 228). Please note that for the sake of brevity, Figure 4 Only a portion of the components are shown. In practice, the multi-processing system chip may include other components of the computing device 200 .
[0047] At the beginning, the application processor unit 402 sends an instruction (e.g., an application programming interface function) to the real-time processor unit 404, wherein the instruction (e.g., an application programming interface function) may include a core identifier and an object identifier. In addition, the instruction (e.g., an application programming interface function) may also include some parameters of the programmable logic module 401. Then, the real-time processor unit 404 determines the storage location of the object 414 having the object identifier, and sends an instruction to the storage controller 410 according to the capacity size (buffer size) of the on-chip memory 408. Therefore, the storage controller 410 reads a piece of data having the capacity size of the on-chip memory 408 in the object 414 from the object storage device 412 and writes it to the on-chip memory 408. Then, the real-time processor unit 404 sends an instruction to the acceleration circuit 406 to trigger the core function having the core identifier. Therefore, the acceleration circuit 406 reads the object data from the on-chip memory 408 and executes the core function having the core identifier to process the object data. After completing the processing of the object data stored in the on-chip memory 408, the acceleration circuit 406 transmits information to inform the real-time processor unit 404. Then, the real-time processor unit 404 determines whether the data processing for the object 414 is completely completed. If the data processing for the object 414 is not completely completed, the real-time processor unit 404 will send instructions to the storage controller 410 according to the capacity of the on-chip memory 408, so that the storage controller 410 will read the next data of the object 414 with the capacity of the on-chip memory 408 from the object storage device 412 and write it to the on-chip memory 408. Then, the real-time processor unit 404 sends instructions to the acceleration circuit 406 to trigger the core function with the core identifier, so that the acceleration circuit 406 reads the object data from the on-chip memory 408 and executes the core function with the core identifier to process the object data. The above steps are repeatedly executed until the real-time processor unit 404 determines that the data processing for the object 414 has been completed. In addition, when the data processing for the object 414 has been completed, the real-time processor unit 404 transmits information to inform the application processor unit 402.
[0048] exist Figure 2 In the illustrated embodiment, the programmable circuit 208 also includes a network subsystem 214 and a storage subsystem 216. However, this is merely an example and the present invention is not limited thereto. Any circuit architecture that utilizes auxiliary processors and special application processors to offload data processing from a general-purpose processor falls within the scope of the present invention.
[0049] Figure 5 FIG. 4 is a schematic diagram of a programmable circuit according to an embodiment of the present invention. Figure 2 The computing device 200 shown may be modified to employ Figure 5 The programmable circuit 500 shown in FIG. 5 may replace Figure 2 Compared with the programmable circuit 208, the programmable circuit 500 does not include the network subsystem 214 and the data conversion circuit 234. Therefore, for applications that do not need to be connected to the network mounting device 10 through the network 30, the programmable circuit in the computing device of the present invention can be used. Figure 5 Design shown.
[0050] Figure 6 FIG. 4 is a schematic diagram of a programmable circuit according to another embodiment of the present invention. Figure 2 The computing device 200 shown may be modified to adopt Figure 6 The programmable circuit 600 shown in FIG. 600 may replace Figure 2 Compared with the programmable circuit 208, the programmable circuit 600 does not include the storage subsystem 216 and the data conversion circuit 236. Therefore, for applications that do not require connection to the storage device 20 via the input / output interface 40, the programmable circuit in the computing device of the present invention can be used. Figure 6 Design shown.
[0051] Please note, Figure 2 , Figure 5 and Figure 6 The data conversion circuits 234 and 236 in the embodiment may be optional components, that is, whether the programmable circuits 208, 500, and 600 need the data conversion circuits 234 and 236 may be determined according to different design requirements. For example, if the predetermined data format required by the acceleration circuit 232 can cover the data format of the payload data obtained from the network flow, the data conversion circuit 234 may be omitted. Similarly, if the predetermined data format required by the acceleration circuit 232 can cover the data format of the data obtained from the storage processing circuit 226, the data conversion circuit 236 may be omitted. All these design changes are within the scope of the present invention.
[0052] In summary, the computing device of the present invention may be equipped with a network subsystem to connect to the network and may perform relevant data processing for object storage, and therefore has extremely high scalability. In one application, the computing device of the present invention may be compatible with existing object storage services (such as Amazon S3 or other cloud storage services), and thus may perform relevant data processing for data acquisition on the object storage device to which the computing device is connected based on object storage instructions (such as Amazon S3 Select) from the network. In another application, the computing device of the present invention may receive NVMe / TCP instructions from the network, and perform relevant data processing operations on the storage device to which the computing device is connected based on the NVMe / TCP instructions. If the storage device to which the computing device is connected is part of a distributed storage system (such as part of a key-value database), the NVMe / TCP instructions received from the network may include key-value instructions, and the computing device of the present invention may perform relevant key-value database data processing operations on the storage device based on the key-value instructions. In addition, data can be processed by the hardware acceleration circuit during the movement process, and there is no need for the general-purpose processor executing the application to intervene in the data movement and the communication between the software and the hardware. Therefore, internal network operations and / or storage internal operations can be realized, thereby saving power consumption, reducing delays and reducing the load of the general-purpose processor; furthermore, the computing device of the present invention can be implemented using a multi-processing system chip. For example, the multi-processing system chip can include a field programmable logic gate array and a general-purpose processor core using an ARM architecture, so it has high design flexibility. Designers can design the application / program code to be executed by the general-purpose processor core and the hardware data processing acceleration function to be possessed by the field programmable logic gate array according to their needs. For example, the computing device of the present invention can be applied to a data center, and can be customized to support various data types and unit formats and obtain optimal performance. Since a single multi-processing system chip can replace a high-end server, a data center using the computing device of the present invention can have a lower construction cost.
[0053] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the claims of the present invention should fall within the scope of the present invention.
[0054] [Explanation of symbols]
[0055] 10: Network mount device
[0056] 20,104,312,322: Storage device
[0057] 30: Network
[0058] 40,112: Input / output interface
[0059] 100,300:Computer Systems
[0060] 102: Computer host
[0061] 105,314: Accelerator card
[0062] 106:CPU
[0063] 108: System memory
[0064] 110,310: Network interface
[0065] 114: Memory
[0066] 200: Computing device
[0067] 201: Chip
[0068] 202: General Processor
[0069] 204: Auxiliary processor
[0070] 206: Special Application Processor
[0071] 208,500,600: Programmable circuit
[0072] 212: General purpose processor core
[0073] 214,330: Network subsystem
[0074] 216,332: Storage subsystem
[0075] 222:TCP / IP Offload Engine
[0076] 224: Network processing circuit
[0077] 226: Storage processing circuit
[0078] 228,410: Storage Controller
[0079] 232,334,406: Acceleration circuit
[0080] 234,236: Data conversion circuit
[0081] 302: Application
[0082] 304: Operating system core
[0083] 306: File System
[0084] 308: Driver
[0085] 320: Multi-processing system chip
[0086] 324,402: Application Processor Unit
[0087] 326,404: Real-time processor unit
[0088] 328: Field Programmable Gate Array
[0089] 400: Processing system
[0090] 401: Programmable Logic Module
[0091] 408: On-chip memory
[0092] 412: Object Storage Device
[0093] 414: Object
[0094] APP: Application
[0095] API_F: Application Programming Interface Function
[0096] SW: Program code
Claims
1. A computing device, comprising: at least one general purpose processor for executing an application program, wherein at least a portion of data processing of a data processing task is offloaded from the at least one general purpose processor; at least one auxiliary processor for handling a control flow of the data processing without intervention of the application program executed on the at least one general purpose processor; and at least one application specific processor for processing a data stream of the data processing without intervention of the application program executed on the at least one general purpose processor, The application executed on the at least one general purpose processor offloads the data processing by calling an application programming interface function, The at least one special application processor is used to process a core function with a core identifier, the data processor is used to process an object located in an object storage device and having an object identifier, and the parameters of the application programming interface function include the core identifier and the object identifier.
2. The computing device of claim 1, wherein the control flow executed on the at least one auxiliary processor comprises a plurality of layers of an input / output stack. 3 . The computing device as claimed in claim 1 , wherein the at least one general purpose processor and the at least one auxiliary processor are heterogeneous processors. 4 . The computing device as claimed in claim 1 , wherein the at least one special application processor is a programmable circuit. 5 . The computing device as claimed in claim 4 , wherein the programmable circuit is a field programmable gate array. 6 . The computing device as claimed in claim 5 , wherein the at least one general purpose processor, the at least one auxiliary processor and the at least one special application processor are all integrated into a same chip.
7. The computing device of claim 1, wherein the at least one auxiliary processor comprises: A programmable circuit comprising: a network subsystem for receiving the core identifier and an object identifier from a network; and At least one general-purpose processor core is used to obtain the core identifier and the object identifier from the programmable circuit, and trigger the core function with the core identifier to process an object located in an object storage device and having the object identifier, wherein when the at least one special application processor processes the object, intervention of the application executed on the at least one general-purpose processor is not required.
8. The computing device of claim 1 , wherein the at least one auxiliary processor comprises at least one general purpose processor core, and the computing device further comprises: A control channel is coupled between the pin of the at least one special application processor and the pin of the at least one general processor core, wherein the control channel is used to transmit control information between the at least one special application processor and the at least one general processor core.
9. The computing device of claim 1, wherein the at least one auxiliary processor comprises: At least one general-purpose processor core; and A programmable circuit includes a network subsystem, wherein the network subsystem includes: a network processing circuit for communicating with the at least one general purpose processor core and controlling a network flow; and The at least one special application processor comprises: At least one acceleration circuit is used to receive a data input from the network processing circuit and process the data stream of the data processing according to the data input. 10 . The computing device of claim 9 , wherein the at least one acceleration circuit is further configured to transmit a data output of the at least one acceleration circuit through the network processing circuit.
11. The computing device of claim 9, wherein the network subsystem further comprises: A TCP / IP offload engine is used to process a TCP / IP stack between the network processing circuit and a network mounting device.
12. The computing device of claim 9, wherein the programmable circuit further comprises: At least one data conversion circuit is used to process data conversion between the network processing circuit and the at least one acceleration circuit, wherein the data format of the payload data obtained from the network flow is different from the predetermined data format required by the at least one acceleration circuit.
13. The computing device of claim 9, wherein the network processing circuit is used to control the network flow between the at least one acceleration circuit and a portion of a distributed object storage system.
14. The computing device of claim 9, wherein the programmable circuit further comprises: A storage subsystem, comprising: a storage processing circuit for communicating with the at least one general purpose processor core and controlling data access to a storage device; The at least one acceleration circuit is further used to transmit a data output of the at least one acceleration circuit through the storage processing circuit.
15. The computing device of claim 1, wherein the at least one auxiliary processor comprises: At least one general-purpose processor core; and A programmable circuit includes a storage subsystem, wherein the storage subsystem includes: a storage processing circuit for communicating with the at least one general purpose processor core and controlling data access to a storage device; and The at least one special application processor comprises: At least one acceleration circuit is used to receive a data input from the storage processing circuit and process the data stream of the data processing according to the data input. 16 . The computing device of claim 15 , wherein the at least one acceleration circuit is further configured to transmit a data output of the at least one acceleration circuit through the storage processing circuit.
17. The computing device of claim 15, wherein the storage subsystem further comprises: A storage controller is used to perform actual data access to the storage device.
18. The computing device of claim 15, wherein the programmable circuit further comprises: At least one data conversion circuit is used to process data conversion between the storage processing circuit and the at least one acceleration circuit, wherein the data format of the data obtained from the storage processing circuit is different from the predetermined data format required by the at least one acceleration circuit.
19. The computing device of claim 15, wherein the programmable circuit further comprises: A network subsystem, including: a network processing circuit for communicating with the at least one general purpose processor core and controlling a network flow; The at least one acceleration circuit is further used to transmit a data output of the at least one acceleration circuit through the network processing circuit.
20. A data processing method, comprising: executing an application program by at least one general purpose processor, wherein at least a portion of data processing of a data processing task is offloaded from the at least one general purpose processor; and processing a control flow of the data processing by at least one auxiliary processor and processing a data flow of the data processing by at least one special application processor without intervention of the application program executed on the at least one general-purpose processor, The application executed on the at least one general purpose processor calls an application programming interface function to offload the data processing, A core function with a core identifier is processed by the at least one special application processor, an object located in an object storage device and having an object identifier is processed by the data processor, and the parameters of the application programming interface function include the core identifier and the object identifier.
Citation Information
Patent Citations
Resource Efficient Acceleration of Datastream Analytics Processing Using an Analytics Accelerator
US20180052708A1
System, Apparatus And Method For Multi-Kernel Performance Monitoring In A Field Programmable Gate Array
US20180267878A1
Technologies for hybrid field-programmable gate array-application-specific integrated circuit code acceleration
WO2018176238A1