Data processing method and device, equipment and equipment resource pool
By using a separate architecture and FPGA in the data processing equipment to complete the processing tasks independently, the problems of high cost and low efficiency in the prior art are solved, and lower costs and higher efficiency are achieved.
Patent Information
- Application Number
- CN202410458074.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-04-16
- Publication Date
- 2025-06-27
AI Technical Summary
The existing data processing methods have high server costs and low efficiency due to the high data transmission delay between the central processor and heterogeneous operators.
The data processing device adopts a separate architecture uses FPGA to independently complete multiple processing tasks that need to be executed in the specified order, and directly obtains the pending data in the FPGA's storage unit to avoid data transmission delay.
It reduces data processing costs, improves data processing efficiency, simplifies scheduling logic, and enhances the scalability and parallel processing capabilities of the device.
Smart Images

Figure CN120216112A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, device, and device resource pool. Background Art
[0002] Currently, data processing requests are usually executed by a server. When the server processes multiple processing tasks related to a data processing request, it is usually completed by a central processing unit (CPU) and heterogeneous computing power in cooperation. For example, the central processing unit executes task scheduling and processing tasks with less time consumption, and the heterogeneous computing power executes processing tasks with more time consumption.
[0003] However, in this data processing method, due to the high cost of the server, the data processing cost is also high. In addition, it increases the data transmission delay between the central processing unit and the heterogeneous operators, resulting in low data processing efficiency. Summary of the Invention
[0004] This application provides a data processing method, apparatus, device, and device resource pool, which can not only reduce the data processing cost, but also improve the data processing efficiency.
[0005] To achieve the above object, this application adopts the following technical solutions:
[0006] In a first aspect, a data processing method is provided, which is applied to a data processing device. The field programmable gate array (FPGA) of the data processing device includes a scheduling unit, a processing unit, and a storage unit. The method includes: The scheduling unit determines M processing tasks and the to-be-processed data of the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the to-be-processed data is stored in the storage unit; the scheduling unit sends M request commands to the processing unit; wherein, each request command is used to request to execute each processing task according to the to-be-processed data of each processing task; the processing unit executes the M processing tasks in the target order based on the received M request commands, and obtains the output data of the M processing tasks; wherein, the output data of the Nth processing task is the to-be-processed data of the (N + 1)th processing task in the target order; N is a positive integer less than M; the processing unit stores the output data of the M processing tasks in the storage unit.
[0007] In this solution, a data processing device provides data processing services. The data processing device adopts a split architecture. Since the data processing device does not need to deploy hardware resources based on the X86 server architecture in related technologies, the hardware resources used by the data processing device are reduced. In this way, not only the data processing cost is reduced, but also the utilization rate of the hardware resources of the data processing device is improved. In addition, the FPGA independently completes M processing tasks that need to be executed in a specified order, and the output data of each processing task is stored in the storage unit of the FPGA. In this way, when the processing unit executes each processing task, it can directly obtain the data to be processed from the storage unit in hardware, thereby avoiding the end-to-end data transmission delay during the execution of the M processing tasks, and further improving the data processing efficiency.
[0008] In addition, since the data processing device adopts a numerically controlled separation architecture, that is, when executing M processing tasks, the task scheduling is performed by the scheduling unit of the FPGA. For example, a request command is sent to the processing unit to instruct the processing unit to execute the processing task on the data to be processed, and the processing unit of the FPGA executes the processing task on the data to be processed, thereby realizing the decoupling / separation of the control flow (such as task scheduling) and the data flow (such as the data to be processed generated by the processing task). In this way, it helps to simplify the scheduling logic of the scheduling unit and improve the scalability of the data processing device. For example, when adding an executable processing task to the processing unit, there is no need to change the scheduling logic (i.e., software code) of the scheduling unit. In addition, due to the adoption of the numerically controlled separation architecture, the data processing device can also execute task scheduling and processing tasks in parallel, thereby improving the parallel processing ability of the data processing device, and further helping to improve the data processing efficiency.
[0009] In a possible implementation, the data to be processed is image data; the processing unit includes a decoding operator unit, and the decoding operator unit is used to execute an image decoding task. The decoding operator unit includes multiple decoding modules, and the decoding operator unit decodes 8Q pixels of the image data in parallel through the multiple decoding modules, where K is a positive integer greater than 1. In this way, it helps to improve the decoding efficiency.
[0010] In another possible implementation, the processing unit includes an encoding operator unit, and the encoding operator unit is used to execute an image encoding task. The encoding operator unit includes multiple encoding modules, and the encoding operator unit encodes 8Q pixels of the image data in parallel through the multiple encoding modules. In this way, it helps to improve the encoding efficiency.
[0011] In another possible implementation, the data processing device further includes a processor, which is configured to convert the received data processing request into a task request recognizable by the FPGA; the scheduling unit determines M processing tasks and the data to be processed for the M processing tasks, including: the processor determines a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks; the processor sends the target task request to the scheduling unit; the target task request is used to indicate the original data and the target operation; the scheduling unit determines the first-sorted processing task among the M processing tasks and the data to be processed for the first-sorted processing task based on the received target task request; the data to be processed for the first-sorted processing task is the original data.
[0012] In this implementation, by using a separate processor to receive the data processing request sent by the electronic device, the software development cost of the FPGA can be reduced. Among them, the processor can be a DPU / CPU.
[0013] In another possible implementation, the processing unit includes multiple operator units; the FPGA stores an operator form, which is used to indicate the processing tasks performed by each operator unit among the multiple operator units; the scheduling unit sends M request commands to the processing unit, including: the scheduling unit determines a target operator unit for the target processing task among the M processing tasks from the multiple operator units indicated by the operator form, and the operator form is used to indicate that the target operator unit is used to perform the target processing task; the scheduling unit sends a target request command to the target operator unit, and the target request command is used to request to perform the target processing task on the target data to be processed.
[0014] In this implementation, the operator unit for each processing task is determined according to the content indicated by the operator form. In this way, when a new operator unit is added to the processing unit, the scheduling logic of the scheduling unit does not need to be changed, which helps to reduce the development cost.
[0015] In another possible implementation, the operator form is further used to indicate the working state of each operator unit, and the working state includes an idle state or a busy state; the operator form indicates that the working state of the target operator unit is the idle state. In this way, the operator unit in the idle state can be allocated for the processing task, thereby improving the processing efficiency.
[0016] In another possible implementation, after the scheduling unit determines a target operator unit for the target processing task among the M processing tasks from the multiple operator units indicated by the operator form, the method further includes: the scheduling unit updates the working state of the target operator unit indicated by the operator form to the busy state. In this way, it helps to ensure the accuracy of the working state indicated by the operator form.
[0017] In another possible implementation, when the target operator unit has completed the target processing task, the scheduling unit updates the working state of the target operator unit indicated in the operator form to the idle state. This helps to ensure the accuracy of the working state indicated in the operator form.
[0018] In another possible implementation, the method further includes: when the target operator unit executes the target processing task based on the target request command, the target operator unit returns target response information to the scheduling unit, and the target response information is used to indicate the storage address of the output data of the target processing task. In this way, the scheduling unit can determine the storage address of the data to be processed for the next processing task according to the response information, which helps to ensure the accuracy of the data to be processed for the execution of the request command.
[0019] In another possible implementation, the target response information is further used to indicate the execution result of the target processing task, and the execution result includes execution success or execution failure. This helps the scheduling unit to understand the execution situation of the processing task.
[0020] In another possible implementation, the FPGA stores a task form, and the task form is used to indicate the task states of M processing tasks, and the task states include the executed state or the unexecuted state; the method further includes: the scheduling unit updates the task form based on the received target response information; the updated task form is used to indicate that the target processing task is in the executed state. In this way, the scheduling unit can end the task scheduling after sending the request command and no longer needs to continuously monitor the execution situation of the processing task, which can reduce the scheduling time of each processing task during the task scheduling process.
[0021] In another possible implementation, the scheduling unit includes multiple sub-scheduling units, where each sub-scheduling unit is used to send a request command to the processing unit. In this way, the scheduling unit can execute the task scheduling of multiple data processing requests in parallel, which helps to improve the data processing ability of the data processing device
[0022] In another possible implementation, the method further includes: when the processing unit has completed the M processing tasks, the scheduling unit outputs target data to the processor, and the target data is the output data of the processing task ranked Mth in the target order.
[0023] In a second aspect, a data processing device is provided. The device includes: functional units for performing the functions of any of the methods provided in the first aspect, and the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the data processing device is used for a data processing device, and the hardware processor of the data processing device includes a processing unit and a storage unit; the data processing device may include a scheduling module and a processing module; the scheduling module is configured to determine M processing tasks and the data to be processed for the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; the scheduling module is further configured to send M request commands to the processing unit, where each request command is used to request to execute each processing task according to the data to be processed for each processing task; the processing module is configured to execute the M processing tasks in the target order based on the received M request commands to obtain the output data of the M processing tasks; where the output data of the Nth processing task is the data to be processed for the (N + 1)th processing task in the target order; N is a positive integer less than M; the processing module is further configured to store the output data of the M processing tasks in the storage unit.
[0024] In a possible implementation, the processing module includes multiple operators; the FPGA stores an operator form, and the operator form is used to indicate the processing tasks performed by each operator among the multiple operators; the scheduling module is specifically configured to: determine a target operator for a target processing task among the M processing tasks from the multiple operators indicated by the operator form, and the operator form is used to indicate that the target operator is used to perform the target processing task; send a target request command to the target operator, and the target request command is used to request to perform the target processing task on the target data to be processed.
[0025] In another possible implementation, the operator form is further used to indicate the working state of each operator unit, and the working state includes an idle state or a busy state; the operator form indicates that the working state of the target operator unit is the idle state.
[0026] In another possible implementation, after the scheduling module determines a target operator unit for a target processing task among the M processing tasks from the multiple operator units indicated by the operator form, the scheduling module is further configured to: update the working state of the target operator unit indicated by the operator form to the busy state.
[0027] In another possible implementation, when the target operator finishes executing the target processing task, the scheduling module is further configured to: update the working state of the target operator indicated by the operator form to the idle state.
[0028] In another possible implementation, when the target operator has completed the target processing task based on the target request command, the target operator is further configured to return target response information to the scheduling module, and the target response information is used to indicate the storage address of the output data of the target processing task.
[0029] In another possible implementation, the target response information is further used to indicate the execution result of the target processing task, and the execution result includes successful execution or failed execution.
[0030] In another possible implementation, the FPGA stores a task form, and the task form is used to indicate the task status of M processing tasks, where the task status includes an executed status or an unexecuted status; the scheduling module is further configured to: update the task form based on the received target response information; the updated task form is used to indicate that the target processing task is in an executed status.
[0031] In another possible implementation, the scheduling module may include multiple sub-scheduling modules, where each sub-scheduling module is used to send a request command to the processing unit.
[0032] In another possible implementation, the data processing device further includes a software interface module; the software interface module is configured to: determine a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks; the software interface module is further configured to: send the target task request to the scheduling module; the target task request is used to indicate the original data and the target operation; specifically, the scheduling module is configured to: determine the first-sorted processing task among the M processing tasks and the data to be processed of the first-sorted processing task based on the received target task request; the data to be processed of the first-sorted processing task is the original data.
[0033] In another possible implementation, when the processing unit has completed the M processing tasks, the scheduling module is further configured to: output target data to the processor, and the target data is the output data of the Mth-sorted processing task in the target order.
[0034] In another possible implementation, the data to be processed is image data. The processing module includes a decoding operator, and the decoding operator is used to perform an image decoding task. The decoding operator includes multiple decoding modules, and the multiple decoding modules are used to decode 8Q pixels of the image data in parallel.
[0035] In another possible implementation, the processing module includes an encoding operator, and the encoding operator is used to perform an image encoding task. The encoding operator includes multiple encoding modules, and the multiple encoding modules are used to encode 8Q pixels of the image data in parallel.
[0036] In a third aspect, a data processing device is provided, including: a scheduling unit, a processing unit, and a storage unit; the scheduling unit is configured to determine M processing tasks and the data to be processed for the M processing tasks; the execution order of the M processing tasks is the target order, where M is a positive integer greater than 1; the data to be processed is stored in the storage unit; the scheduling unit is further configured to send M request commands to the processing unit; wherein each request command is used to request to execute each processing task according to the data to be processed for each processing task; the processing unit is configured to execute the M processing tasks in the target order based on the received M request commands, to obtain the output data of the M processing tasks; wherein the output data of the Nth processing task is the data to be processed for the (N + 1)th processing task in the target order; N is a positive integer less than M; the processing unit is further configured to store the output data of the M processing tasks in the storage unit.
[0037] It should be noted that in the third aspect, the scheduling unit and the processing unit can also be used to execute any possible implementation manner provided in the first aspect above.
[0038] In a fourth aspect, a resource pool of data processing devices is provided, including: at least one data processing device; the at least one data processing device is configured to execute the steps of any one of the methods provided in the first aspect.
[0039] In a fifth aspect, a computer program product is provided, the computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a data processing device, the steps of any one of the methods provided in the first aspect are implemented.
[0040] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a data processing device, the steps of any one of the methods provided in the first aspect are implemented.
[0041] Wherein, the technical effects brought by any implementation manner in the second aspect to the sixth aspect can refer to the technical effects brought by different implementation manners in the first aspect, and will not be elaborated here. Description of the Drawings
[0042] Figure 1 A schematic diagram of a related art provided by an embodiment of the present application;
[0043] Figure 2 A schematic diagram of another related art provided by an embodiment of the present application;
[0044] Figure 3 A schematic diagram of a system architecture provided by an embodiment of the present application;
[0045] Figure 4 A schematic diagram of another system architecture provided by an embodiment of the present application;
[0046] Figure 5 Schematic diagram of another system architecture provided by an embodiment of the present application;
[0047] Figure 6 Schematic diagram of a data device provided by an embodiment of the present application;
[0048] Figure 7 Schematic diagram of an encoding operator unit provided by an embodiment of the present application;
[0049] Figure 8 Schematic diagram of a decoding operator unit provided by an embodiment of the present application;
[0050] Figure 9 Flowchart of a data processing method provided by an embodiment of the present application;
[0051] Figure 10 Schematic diagram of a data processing device provided by an embodiment of the present application. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application.
[0053] Among them, in the description of the present application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B may be singular or plural.
[0054] Moreover, in the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.
[0055] In addition, for the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0056] The following provides an exemplary introduction to the relevant terms involved in the embodiments of the present application.
[0057] Split architecture: It refers to an architecture that breaks up various hardware resources in the server architecture and uses some of the hardware resources in the server architecture. Among them, some of the hardware resources in the split architecture can communicate and interact through the network.
[0058] Operator: It refers to a module on the hardware accelerator used to execute a specified operation.
[0059] In the embodiments of the present application, an operator can be a software module on an FPGA used to execute processing tasks, such as: a module for executing decoding tasks, a module for executing encoding tasks, a module for executing image magnification tasks, etc.
[0060] Zero-layer logic architecture: It is used to define the relationships and interaction methods among various modules in the system during the software development process. The zero-layer logic architecture divides the system into multiple layers according to functions, and each layer is responsible for different functions and responsibilities.
[0061] Application Programming Interface (API): It is a pre-defined function used to provide the application and developers with the ability to access a set of routines based on a certain software or hardware, without the need to access the source code.
[0062] Direct Memory Access (DMA): It is a mechanism that allows direct data transfer between a hardware device and a memory device, that is, when the hardware device and the memory device perform data transfer, they do not need to rely on the CPU.
[0063] The following provides an exemplary introduction to the application scenarios of the embodiments of the present application.
[0064] The data processing method provided by the embodiments of the present application is applicable to scenarios such as processing image data and audio data. Hereinafter, taking image data as an example, an exemplary description of the data processing method provided by the embodiments of the present application will be given.
[0065] It should be noted that the embodiments of the present application do not limit the application scenarios of the data processing method, and the above is only an exemplary description.
[0066] In the related art, image data processing services are usually deployed on a server cluster, and the server cluster provides image processing services in the form of microservices. As Figure 1 shown, after the user stores the image data on the cloud storage device, the user can send an image processing request to the server cluster through an electronic device. After receiving the image processing request, the server cluster can specify Server 1 in the server cluster to execute the image processing request. Server 1 obtains the image data from the cloud storage device and returns the processed image to the user. Among them, an image processing request may include multiple image processing tasks. For example: a processing request for sharpening a JPEG image may include decoding the JPEG image to obtain a bitmap image, sharpening the bitmap image, and encoding the sharpened bitmap image to obtain a JPEG image, etc.
[0067] Server 1 deploys hardware resources based on the X86 server architecture. Server 1 includes a central processing unit, a network interface card (NIC), a solid state disk / drive (SSD), a dynamic random access memory (DRAM), etc. When Server 1 executes the image processing request, it is specifically executed by the CPU inside Server 1. Due to the limited processing power of the CPU, the image data processing ability provided by Server 1 is poor.
[0068] In order to improve the image processing ability of Server 1, as Figure 2As shown in the figure, a graphics processing unit (GPU) is configured for the related technology server 1, so that the CPU and the GPU (heterogeneous computing power) cooperate to complete the image processing request. Among them, the CPU is used to execute task scheduling and simple image processing tasks (such as: sharpening tasks), and the GPU is used to execute time-consuming image processing tasks (such as: image encoding tasks, image decoding tasks, etc.) to improve data processing capabilities. Specifically, after the CPU obtains the image data based on the image processing request, it sends the image data to the GPU and instructs the GPU to execute the image decoding task. The GPU sends the bitmap image obtained by executing the image decoding task to the CPU. After the CPU performs the sharpening task on the bitmap image, it sends the sharpened bitmap image to the GPU and instructs the GPU to execute the image encoding task. The GPU returns the JFEG image obtained by executing the image encoding task to the CPU, and the CPU returns it to the user.
[0069] However, since the server architecture also needs to include hardware resources such as a network interface card (NIC) and a solid state disk (SSD), therefore, the cooperation between the CPU and the GPU inside the server to complete the image processing request not only cannot make full use of the hardware resources on the server, resulting in low hardware resource utilization and high image processing costs of the server, but also increases the end-to-end data transmission delay, that is, the delay of data transmission between the CPU and the GPU, thus reducing the image processing efficiency.
[0070] In view of this, an embodiment of the present application provides a data processing method, which is applied to a data processing device. The FPGA of the data processing device includes a scheduling unit, a processing unit, and a storage unit. When the data processing device processes multiple processing tasks of a data processing request, the scheduling unit executes task scheduling. For example, the scheduling unit sends multiple request commands to the processing unit, where each of the multiple request commands is used to request the execution of a processing task. The processing unit is used to execute multiple processing tasks. For example, it executes multiple processing tasks based on the multiple request commands and stores the output data of each processing task in the storage unit.
[0071] Since the embodiments of the present application are executed by a data processing device, and the data processing device adopts a split architecture. Since the data processing device does not need to deploy hardware resources based on the X86 server architecture in the related art, the hardware resources used by the data processing device are reduced. In this way, not only the data processing cost is reduced, but also the utilization rate of the hardware resources of the data processing device is improved. In addition, by using the FPGA to independently complete multiple processing tasks, and the output data of each processing task is stored in the storage unit of the FPGA. In this way, when the processing unit executes each processing task, it can directly obtain the data to be processed from the storage unit, thus avoiding the end-to-end data transmission delay during the execution of multiple processing tasks, that is, the data transmission delay between the CPU and the GPU in the related art, and further improving the data processing efficiency.
[0072] In addition, since the data processing device adopts a numerical control separation architecture, that is, when executing multiple processing tasks, the task scheduling is performed by the scheduling unit of the FPGA. For example, a request command is sent to the processing unit to instruct the processing unit to execute the processing task on the data to be processed, and the processing unit of the FPGA executes the processing task on the data to be processed, thereby realizing the decoupling / separation of the control flow (such as task scheduling) and the data flow (such as the data to be processed generated by the processing task). In this way, it helps to simplify the scheduling logic of the scheduling unit and improve the scalability of the data processing device. For example, when adding an executable processing task to the processing unit, there is no need to change the scheduling logic (i.e., software code) of the scheduling unit. In addition, due to the adoption of the numerical control separation architecture, the data processing device can also execute task scheduling and processing tasks in parallel, thereby improving the parallel processing ability of the data processing device and further helping to improve the data processing efficiency.
[0073] Hereinafter, an exemplary introduction to the system architecture provided by the embodiments of the present application will be given.
[0074] The embodiments of the present application provide a data processing device, which can be used to execute the data processing method provided by the embodiments of the present application. Among them, the data processing device may include an FPGA, and the FPGA is used to execute task scheduling and multiple processing tasks. The FPGA may include a scheduling unit, a processing unit, and a storage unit. The scheduling unit may be used to execute the task scheduling of multiple processing tasks, the processing unit may be used to execute multiple processing tasks, and the storage unit may be used to store the output data of each processing task in the multiple processing tasks.
[0075] In one implementation, the FPGA may include a receiving unit. The receiving unit may be used to receive a data processing request sent by an electronic device, convert the data processing request into a task request recognizable by the scheduling unit, and then send the task request to the scheduling unit. In this way, it helps to reduce the hardware cost of the data processing device.
[0076] In another implementation, the data processing device may further include a processor. The processor may be configured to receive a data processing request sent by an electronic device, convert the data processing request into a task request recognizable by a scheduling unit, and then send the task request to the scheduling unit. This helps reduce the complexity of the software logic of the FPGA, thereby helping to reduce the software development cost of the FPGA.
[0077] Exemplarily, the processor may be a data processing unit (DPU), a CPU, etc.
[0078] It should be noted that the embodiments of the present application do not limit the type of the processor, and the above is only an exemplary illustration.
[0079] In the embodiments of the present application, since the data processing device adopts a split architecture, that is, a non-X86 server architecture, the data processing device may also be referred to as a split device, which will not be elaborated hereinafter.
[0080] The embodiments of the present application further provide a data processing device resource pool, which includes a plurality of data processing devices, and different data processing devices among the plurality of data processing devices can communicate with each other.
[0081] Among them, the data processing device resource pool can be used to execute the data processing method provided by the embodiments of the present application. This helps improve the data processing ability and thus improve the data processing efficiency.
[0082] Hereinafter, taking the data processing device including an FPGA and a processor as an example, the embodiments of the present application will be exemplarily described.
[0083] Figure 3 It is a schematic diagram of a system architecture provided by the embodiments of the present application.
[0084] As Figure 3 shown, the system architecture may include an electronic device and a data processing device. The data processing device may include an FPGA and a processor (such as a DPU). Among them, the electronic device can communicate with the data processing device.
[0085] In the embodiments of the present application, the electronic device sends a data processing request and original data to the data processing device. The data processing request may include multiple processing tasks. After receiving the data processing request, the data processing device may execute multiple processing tasks on the original data based on the data processing method provided by the embodiments of the present application.
[0086] It should be noted that the embodiments of the present application do not limit the manner in which the electronic device sends a data processing request to the data processing device. For example, an application program can be installed on the electronic device, and the electronic device can send a data processing request to the data processing device through the application program.
[0087] In the embodiments of the present application, the FPGA may include a scheduling unit, a processing unit, and a storage unit, and the scheduling unit, the processing unit, and the storage unit can communicate with each other.
[0088] During the process of the data processing device executing the data processing request, the scheduling unit can be used to schedule multiple processing tasks, that is, determine the operator units for executing each processing task, and send a request command to the operator units to request the operator units to execute the processing tasks. The processing unit can execute multiple processing tasks through multiple operator units, and the storage unit can store the output data of executing each processing task.
[0089] Optionally, the scheduling unit may include a task parsing unit, a task distribution unit, at least one sub-scheduling unit, at least one task recycling unit, a task output unit, etc. The task parsing unit can communicate with the task distribution unit and the task output unit, and the task distribution unit can communicate with each sub-scheduling unit in at least one sub-scheduling unit, each task recycling unit in at least one task recycling unit, and the task output unit.
[0090] In the embodiments of the present application, the task parsing unit can be used to parse the task request to determine the original data and multiple processing tasks, etc. The task distribution unit can be used to determine the sub-scheduling unit for executing task scheduling, the task recycling unit for updating the task form, etc. The sub-scheduling unit can be used to determine the operator unit for executing the processing task, send a request command, update the operator form, etc. The task recycling unit can be used to update the task form, update the operator form to release the operator unit, etc. The task output unit is used to output the target data, etc.
[0091] Optionally, the FPGA may further include a memory management unit. The memory management unit can communicate with the task output unit, the task parsing unit, etc. The memory management unit is used to manage the storage space of the storage unit. For example, allocate storage space for the task parsing task, etc.
[0092] Optionally, the scheduling unit may further include a DMA arbitration unit. The DMA arbitration unit is connected to the task parsing unit, the task output unit, each sub-scheduling unit, each task recycling unit, etc. The DMA arbitration unit can be used to arbitrate the unit using the DMA controller.
[0093] Optionally, the processing unit includes multiple operator units, and each operator unit in the multiple operator units can be used to execute the processing task.
[0094] Exemplarily, multiple operator units may include an Encoder operator unit, a Decoder operator unit, an Image Cropping operator unit, an Image Resizing operator unit, an Image Watermarking operator unit, an Image Sharpening operator unit, etc.
[0095] The Decoder operator unit is used to perform image decoding tasks, such as: decompressing an image in a specified format into a bitmap format image. The Encoder operator unit is used to perform image encoding tasks, such as: compressing a bitmap format image into a specified format image. The Image Resizing operator unit is used to resize an image. The Image Watermarking operator unit is used to process the watermark of an image. The Image Cropping operator unit is used to crop an image. The Image Sharpening operator unit is used to sharpen an image.
[0096] Exemplarily, the Encoder operator unit may include a Joint Photographic Experts Group (JPEG) Encoder operator unit, a Web Picture (WEBP) Encoder operator unit, a Portable Network Graphics (PNG) Encoder operator unit, a Graphics Interchange Format (GIF) Encoder operator unit, etc. The Decoder operator unit may include a JPEG Decoder operator unit, a WEBP Decoder operator unit, a PNG Decoder operator unit, a GIF Decoder operator unit.
[0097] Optionally, the processing unit may include a DMA arbitration unit. The DMA arbitration unit may communicate with each operator unit. The DMA arbitration unit may be used to arbitrate the operator units using the DMA controller.
[0098] Optionally, the processing unit may include a Control Flow Interconnection and Arbitration unit. The Control Flow Interconnection and Arbitration unit may communicate with each operator unit. The Control Flow Interconnection and Arbitration unit is used to implement the communication between the scheduling unit and the processing unit, such as: the sub-scheduling unit may communicate with the operator unit through the Control Flow Interconnection and Arbitration unit, and the sub-operator unit may communicate with the task distribution unit through the Control Flow Interconnection and Arbitration unit, etc.
[0099] Optionally, the processing unit may include a memory operation arbitration unit. The memory operation arbitration unit may communicate with each operator unit. The memory arbitration unit may be used to allocate storage space for the operator unit, release the storage space not used by the operator unit, etc.
[0100] Optionally, the storage unit may be a volatile storage medium, such as dynamic random access memory (DRAM), etc.
[0101] In the embodiments of the present application, the FPGA may include multiple configurable logic blocks (CLBs). Among them, different units of the FPGA are composed of different CLBs. For example, the scheduling unit, the processing unit, etc. are composed of different CLBs.
[0102] Exemplarily, the multiple configurable logic blocks may include a first configurable logic block and a second configurable logic block, and the first configurable logic block and the second configurable logic block are different configurable logic blocks.
[0103] It should be noted that the first configurable logic block / second configurable logic block refers to a type of configurable logic block among the multiple configurable logic blocks, and does not specifically refer to a certain configurable logic block. In addition, the embodiments of the present application do not limit the number of configurable logic blocks included in the first configurable logic block / second configurable logic block.
[0104] Exemplarily, the scheduling unit includes the first configurable logic block, and the processing unit includes the second configurable logic block. Among them, each operator unit among the multiple operator units is composed of a different second configurable logic block.
[0105] Optionally, the electronic device may be a terminal device or a network device.
[0106] The terminal device may include a mobile phone, a tablet computer, a handheld computer, a PC, a cellular phone, a personal digital assistant (PDA), a wearable device (such as a smart watch, a smart bracelet, etc.), a smart home device (such as a television, etc.), a car computer (such as an in-vehicle computer, etc.), a smart screen, a game console, an earphone, an AI speaker, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a notebook computer, a netbook, a desktop computer or an all-in-one computer, etc.
[0107] The network device may include a server, etc. Among them, the server may be a physical server, or may be two or more physical servers sharing different responsibilities and cooperating with each other to implement the various functions of the server. Exemplarily, the server may be a blade server, a high-density server, a rack server or a tower server, etc.
[0108] It should be noted that the device type of the electronic device in the embodiments of the present application is not limited, and the above is only an exemplary description.
[0109] Figure 4 It is a schematic diagram of another system architecture provided by the embodiments of the present application.
[0110] As Figure 4 shown, the system architecture may include an electronic device, a data processing device, and a cloud storage device. The electronic device can communicate with the data processing device and the cloud storage device, and the data processing device can communicate with the cloud storage device.
[0111] Exemplarily, the cloud storage device stores a plurality of data, such as: data 10,......, data P0, where P is a positive integer greater than 1, and the plurality of data can be image data, audio data, etc.
[0112] In the embodiments of the present application, the user can store a plurality of data in the cloud storage device through the electronic device. After that, the data processing device can execute the data processing method provided by the embodiments of the present application on the data stored in the cloud storage device.
[0113] It should be noted that the electronic device that stores a plurality of data in the cloud storage device and the electronic device that requests the data processing device to execute the data processing method provided by the present application may be the same, or they may also be different. The present application does not limit this.
[0114] Exemplarily, the electronic device sends a data processing request to the data processing device. The data processing request is used to indicate that data 10 in the cloud storage device is the original data, and the data processing request includes a plurality of processing tasks. The data processing device obtains data 10 from the cloud storage device based on the data processing request and executes a plurality of processing tasks on data 10.
[0115] It should be noted that Figure 4 For other related descriptions of the system architecture shown, reference can be made to Figure 3 the description of the system architecture shown, which will not be elaborated here.
[0116] Figure 5 It is a schematic diagram of yet another system architecture provided by the embodiments of the present application.
[0117] As Figure 5 shown, the system architecture may include an electronic device, a management device, a data processing device resource pool, and a cloud storage device. The electronic device can communicate with the management device and the cloud storage device, and the management device can communicate with each data processing device in the data processing device resource pool.
[0118] In an embodiment of the present application, an electronic device may send a data processing request to a management device, and the management device may determine a data processing device for executing the data processing request from a data processing device resource pool according to a pre-determined balancing policy.
[0119] Exemplarily, the management device may include an Elastic Load Balancer (ELB), and the management device may determine a data processing device for executing the data processing request from the data processing device resource pool through the ELB.
[0120] It should be noted that Figure 5 For other relevant descriptions of the system architecture shown, reference may be made to Figure 3 、 Figure 4 the description of the system architecture shown, which will not be elaborated here.
[0121] Figure 6 It is a schematic diagram of a data processing device provided by an embodiment of the present application.
[0122] Among them, Figure 6 the data processing device shown may be Figures 3 to 5 the data processing device in the system architecture shown. Exemplarily, the data processing device may adopt a zero-layer logic architecture.
[0123] Such as Figure 6 shown, in terms of hardware, the data processing device may include an FPGA and a DPU. Exemplarily, the FPGA and the DPU may communicate based on the PCIE protocol.
[0124] Among them, the FPGA may include a Dynamic Random Access Memory (DRAM), and the dynamic random access memory may also be referred to as memory. It should be noted that the dynamic random access memory of the FPGA may also be referred to as On-Board DRAM. The DRAM of the FPGA may be used to store raw data, intermediate output data, target data, etc.
[0125] It should be noted that the intermediate output data and the target data will be described in subsequent embodiments and will not be elaborated here.
[0126] Among them, the DPU may include a Memory, and the memory may be used to store raw data, target data, etc.
[0127] Exemplarily, the operating system of the DPU may use the Linux operating system, and the Linux operating system running on the DPU may include a Linux Userspace and a Linux Kernel.
[0128] It should be noted that the embodiments of this application do not limit the operating system used by the DPU. The above is only an exemplary illustration.
[0129] In terms of software, the data processing device may include a software interface module, an FPGA infrastructure, a scheduling module, an operator pool, etc. The operator pool may include multiple operators. Among them, the software interface module can communicate with the FPGA infrastructure, the FPGA infrastructure can communicate with the scheduling module, and the scheduling module can communicate with each operator in the operator pool.
[0130] Exemplarily, a task request (Job Request), a task response (Job Response), data (Data), etc. can be transmitted between the software interface module and the FPGA infrastructure.
[0131] It should be noted that when a software module performs an operation (such as communication), it can be considered that the hardware performs an operation during the process of running the software module, or the hardware performs an operation through the software module. This will not be elaborated further hereinafter.
[0132] It can be understood that in the case where the data processing device adopts a zero-layer logic architecture, the software interface module, the FPGA infrastructure, the scheduling module, and the operator pool can be regarded as zero-layer logic units.
[0133] It should be noted that the embodiments of this application do not limit the names of the software modules included in the data device. The above is only an exemplary illustration. For example, the scheduling module can also be called a central controller.
[0134] Hereinafter, an exemplary introduction to the software interface module will be given.
[0135] The software interface module can be deployed on the DPU. The DPU converts the received data processing request into a task request recognizable by the FPGA by running the software interface module. Exemplarily, the software interface module can run in Linux Userspace.
[0136] Optionally, the software interface module may include a call surface and a management and control surface.
[0137] The call surface is used to provide a software call interface, convert a data processing request into a task request, send the task request to the FPGA, etc. The management and control surface is used to monitor the working state of the software interface module, output an alarm message and generate a system log when the working state is abnormal, etc.
[0138] In the embodiments of the present application, the invocation plane may include a Software APIs module and a Job Manager. Among them, the Software APIs module is used to provide a unified task invocation interface externally and convert data processing requests into task requests, such as integrating multiple consecutive operations involved in a data processing request into a single task request. The Job Manager is used to provide task management capabilities, such as sending the task requests generated by the Software APIs module to the FPGA, etc.
[0139] In the embodiments of the present application, the management and control plane may include a System Logger module and an Abnormal Alarm module. Among them, the System Logger module can be used to provide log management capabilities. The Abnormal Alarm module can be used to provide abnormal alarm capabilities.
[0140] Hereinafter, an exemplary introduction to the FPGA infrastructure will be given.
[0141] The FPGA infrastructure can be deployed on the DPU and the FPGA. That is to say, the FPGA infrastructure can span across the DPU and the FPGA. The DPU and the FPGA can achieve data transmission and interaction by running the FPGA infrastructure. That is to say, the FPGA infrastructure provides an interface for the DPU and the FPGA to perform data transmission and exchange.
[0142] Exemplarily, the FPGA infrastructure can be used to write data to the memory of the DPU, read data from the memory, etc.
[0143] In the embodiments of the present application, the FPGA infrastructure may include an API module, a Drive, and a Shell module. The API module and the Drive are deployed on the DPU, and the Shell module is deployed on the FPGA. The API module is used to communicate with the software interface module, and the Shell module is used to communicate with the central controller. The API module and the Shell module communicate through the Drive. Among them, the API module can be a module implemented in the C language.
[0144] Exemplarily, a Job Request, a Job Response, Data, etc. can be transmitted between the API module and the Drive. A Job Request, a Job Response, Data, etc. can be transmitted between the Drive and the Shell module. Data can be transmitted between the Drive and the storage of the DPU. Data can be transmitted between the Shell module and the storage unit of the FPGA.
[0145] Exemplarily, the API module can run in the Linux Userspace. The driver can run in the Linux Kernel.
[0146] Exemplarily, the API module can include a Read / Write module and a User Interrupt / Poll module. Among them, the Read / Write module is used to communicate with the PCIE driver to transfer task requests, task responses, raw data, target data, etc.
[0147] Exemplarily, the Shell module can include a DRAM Controller, a DMA Controller, an AXI Data Mover, etc.
[0148] Exemplarily, the driver of the FPGA infrastructure can be used to write data to the memory of the DPU, read data from the memory, etc.
[0149] Exemplarily, the DPU and the FPGA can communicate based on the PCIE protocol. Based on this, the Drive can be a PCIE Drive, and the DMA Controller can be a PCIE DMA Controller.
[0150] The following is an exemplary introduction to the central controller.
[0151] The central controller can be deployed on the scheduling unit of the FPGA. The scheduling unit of the FPGA can achieve task scheduling by running the central controller, such as: parsing, scheduling, and managing task flows, managing the operator pool, and controlling data flows, etc.
[0152] In the embodiments of this application, the central controller can include a Job Parser, a Job Output, a Job Descriptor Dispatcher, at least one Scheduler (Stateless Scheduler), and at least one Recycler.
[0153] Exemplarily, the Job Parser, the Job Output, the Scheduler, and the Recycler are Stateless modules.
[0154] Exemplarily, a task request (Job Request) and the like can be transmitted between the task parser and the Shell module. A job descriptor (Job Descriptor) can be transmitted between the task parser, the task dispatcher, and the task output module. A job response (Job Response) can be transmitted between the task output module and the Shell module. A job descriptor (Job Descriptor) can be transmitted between the task dispatcher, the task output module, and the scheduler. A job descriptor (Job Descriptor), a command response (Cmd Response), and the like can be transmitted between the task dispatcher and the recycler. Other control flows (Other Control Flow) can be transmitted between the scheduler and the task form and the operator form. Other control flows (Other Control Flow) can be transmitted between the recycler and the task form and the operator form.
[0155] Among them, the job descriptor can also be referred to as job description information, and the command response can also be referred to as response information, which will not be elaborated hereinafter.
[0156] The task parser can be used to parse a task request to determine the original data and multiple processing tasks, etc. The task parser is deployed on the task parsing unit, and the task parsing unit realizes parsing the task request to determine the original data and multiple processing tasks, etc. by running the task parser. The task output module is used to output target data, etc. The task output module is deployed on the task output unit, and the task output unit realizes outputting the target data, etc. by running the task output module.
[0157] The task dispatcher can be used to determine the sub-dispatching unit for executing task scheduling, determine the task recycling unit for updating the task form / operator form, etc. The task dispatcher is deployed on the task dispatching unit, and the task dispatching unit realizes determining the sub-dispatching unit for executing task scheduling, determining the task recycling unit for updating the task form / operator form, etc. by running the task dispatcher.
[0158] The scheduler can be used to execute task scheduling, such as: determining the processing tasks to be scheduled, determining the operator unit for the processing tasks, scheduling the operator unit to execute the processing tasks, etc. Among them, the scheduler corresponds one-to-one with the sub-dispatching unit. One scheduler is deployed on the sub-dispatching unit, and one sub-dispatching unit realizes executing task scheduling by running one scheduler.
[0159] The recycler can be used to release the operator unit, check the execution result of the processing task, synchronize the execution result of the processing task, etc. Among them, the recycler corresponds one-to-one with the task recycling unit. One recycler is deployed on one task recycling unit, and one task recycling unit realizes releasing the operator unit, checking the execution result of the processing task, synchronizing the execution result of the processing task, etc. by running one recycler.
[0160] Optionally, the central controller may further include a Memory Manager. The Memory Manager is used to manage the storage space of the storage unit. The Memory Manager may be deployed in the memory management unit, and the memory management unit manages the storage space of the storage unit by running the Memory Manager. Exemplarily, the Memory Manager is a Stateful module.
[0161] Exemplarily, other control flows may be transmitted between the Memory Manager, the task parser, and the task output module.
[0162] Optionally, the central controller may further include a DMA Arbiter Stateless. The DMA Arbiter may be used to arbitrate the units using the DMA controller. The DMA Arbiter is deployed in the DMA arbitration unit, and the DMA arbitration unit arbitrates the units using the DMA controller by running the DMA Arbiter.
[0163] Exemplarily, data may be transmitted between the DMA controller, the task parser, and the task output module. Other data flows may be transmitted between the DMA controller, the scheduler, and the recycler.
[0164] Hereinafter, an exemplary introduction to the operator pool is given.
[0165] The operator pool is deployed on the processing unit of the FPGA. The operator pool may include multiple operators. The processing unit of the FPGA may execute multiple processing tasks by running multiple operators. Among them, one operator may be deployed on one operator unit, and one operator unit may implement one processing task by running one operator.
[0166] Optionally, the multiple operators may include an encoding operator, a decoding operator, an image scaling operator, an image watermarking operator, an image cropping operator, etc.
[0167] The decoding operator is used to decompress an image in a specified format into a bitmap format image. The decoding operator is deployed on the decoding operator unit, and the decoding operator unit decompresses an image in a specified format into a bitmap format image by running the decoding operator. The encoding operator is used to compress a bitmap format image into a specified format image. The encoding operator is deployed on the encoding operator unit, and the encoding operator unit compresses a bitmap format image into a specified format image by running the encoding operator.
[0168] The image scaling operator is used to scale an image. The image scaling operator is deployed on an image scaling operator unit, and the image scaling operator unit scales the image by running the image scaling operator. The image watermarking operator is used to process the watermark of an image. The image watermarking operator is deployed on an image watermarking operator unit, and the image watermarking operator unit processes the watermark of the image by running the image watermarking operator. The image cropping operator is used to crop an image. The image cropping operator is deployed on an image cropping operator unit, and the image cropping operator unit crops the image by running the image cropping operator.
[0169] Exemplarily, the encoding operators may include a JPEG Encoder operator, a WEBP Encoder operator, a PNG Encoder operator, a GIF Encoder operator, etc. The decoding operators may include a JPEG Decoder operator, a WEBP Decoder operator, a PNG Decoder operator, a GIF Decoder operator, etc.
[0170] Optionally, the decoding operator may include multiple run-length decoding modules, and the decoding operator unit decodes 8Q pixels of the image data in parallel through the multiple run-length decoding modules; where Q is a positive integer greater than 1. In this way, it helps to improve the decoding efficiency.
[0171] Optionally, the encoding operator unit may include multiple run-length encoding modules, and the encoding operator unit encodes 8Q pixels of the image data in parallel through the multiple run-length encoding modules. In this way, it helps to improve the encoding efficiency.
[0172] Optionally, the operator pool may further include a control flow arbitration and interconnection module. The control flow arbitration and interconnection module is deployed on a control flow arbitration and interconnection unit, and the control flow arbitration and interconnection unit realizes the specified function by running the control flow arbitration and interconnection module.
[0173] Exemplarily, a response command (CmdResponse) etc. may be transmitted between the control flow arbitration and interconnection module and the task dispatcher. A request command (Cmd Request) may be transmitted between the control flow arbitration and interconnection module and the scheduler. A request command (Cmd Request), a response command (CmdResponse), etc. may be transmitted between the control flow arbitration and interconnection module and the operator.
[0174] Optionally, the operator pool may further include a DMA arbiter. The DMA arbiter is deployed on a DMA arbitration unit, and the DMA arbitration unit realizes the specified function by running the DMA arbiter.
[0175] Exemplarily, data can be transmitted between the DMA arbiter of the operator pool and the DMA arbiter of the central controller. Data can be transmitted between the DMA arbiter of the operator pool and the operator.
[0176] Optionally, the operator pool may further include a memory operation arbitration module. The memory operation arbitration module is deployed on the memory operation arbitration unit, and the memory operation arbitration unit implements specified functions by running the memory operation arbitration module.
[0177] Exemplarily, other control flows can be transmitted between the memory operation arbitration module and the operator.
[0178] Figure 7 Schematic diagram of a JPEG encoding operator provided by an embodiment of this application.
[0179] The JPEG encoding operator can compress a bitmap image (RGB pixels) into an encoded bitstream according to the JPEG baseline standard.
[0180] As Figure 7 shown, the JPEG encoding operator may include a controller (JPEG controller), an encoding chain, and a JPEG header generator (JFIF Generator). Among them, the controller is used to receive a request command (commandRequest), parse the request command to obtain a control signal, obtain uncompressed bitmap pixels from the bitmap stream (bitmapStream) according to the control signal, output a response command (commandResponse), output a working state (such as: busy signal), etc. The encoding chain is used to perform streaming encoding on the uncompressed bitmap pixels to obtain compressed data, synthesize the compressed data and JPEG header data into a compressed data stream (JPEGStream), and output the compressed data stream, etc.
[0181] Exemplarily, the request command can be input 512 bits (bit) at a time. The response command can be output 512 bit at a time.
[0182] Optionally, the encoding chain includes a pre-encoder, an entropy encoder (Entropy Code), an encoding output generator (Output Generator), etc.
[0183] In the embodiment of this application, the pre-encoder may include: a pixel conversion module (RGB TO YCbCr), a segmentation module (BlockSplit), a discrete cosine transform module (DCT 2D), a zigzag conversion module (ZIG ZAG), a color component (AddColor Component), a quantization module (Quantize), etc.
[0184] Exemplarily, RGB TO YCbCr is used to convert RGB pixels into YUV pixels, and BlockSplit is used to split YUV pixels into Y pixels, Cb pixels, and Cr pixels and output them in sequence. DCT 2D is used to perform parallel discrete cosine transform calculations. ZIGZAG is used for zigzag conversion. Quantize is used to perform parallel quantization calculations. Add Color Component is used to identify the type of pixel (e.g., Y, Cb, or Cr).
[0185] In an embodiment of the present application, the entropy encoder may include a color component (Add Color Component), a differential code module (DiffCode), a data scheduling module (Data Dispatch), multiple run-length encoding modules (RunLengthEncode1 - RunLength Encode8), a data integration module (Data Integrate), a Huffman encoding module (Huffman Encode), an alignment module (AlignCode), a byte stuffing module (ByteStuff), etc.
[0186] Exemplarily, DiffCode is used to perform parallel differential encoding. Data Dispatch is used to input the received pixels of the same type into the same RunLength Encode. RunLength Encode is used to filter out the high-frequency information of the image data to reduce the data volume of the image data. The Data Integrate module is used to sort the pixels output by RunLength Encode according to the type and output them in sequence. Huffman Encode is used to perform Huffman encoding according to the Huffman encoding table (Huffman EncodeTable). AlignCode is used to byte-align the compressed image data. ByteStuff is used to perform byte stuffing on the compressed image data.
[0187] In an embodiment of the present application, the encoding output generator may include an addition module (EOI Adder). EOI Adder is used to synthesize the compressed data output by the entropy encoder and the JPEG header data into a compressed data stream and output the compressed data stream.
[0188] Figure 8 This is a schematic diagram of a JPEG decoding operator unit provided in an embodiment of the present application.
[0189] The JPEG decoding operator can decode the compressed bitstream into uncompressed pixels (RGB) according to the JPEG baseline standard.
[0190] Such as Figure 8As shown, the JPEG decoding operator may include a controller, a decoding chain, and a JPEG header parser (JFIFParser). Among them, the controller is used to receive a request command (commandRequest), parse the request command to obtain basic parameters, output a response command (commandResponse), output a working state (such as: busy signal), etc. The decoding chain is used to perform streaming decoding on the compressed data in the compressed data stream to obtain uncompressed bitmap data and output the uncompressed bitmap data through the bitmap stream, etc. The JPEG header parser is used to parse the header information in the compressed data stream (JPEGStream) to obtain decoding information, such as: image size, Huffman coding table, quantization table, and other information required during the decoding process.
[0191] Optionally, the decoding chain may include a pre-decoder, an entropy decoder (Entropy Decode), a data processing module, and a decoding output generator (Output Generator).
[0192] In an embodiment of the present application, the pre-decoder may include a width converter (Width Converter). Exemplarily, the Width Converter is used to split the compressed data.
[0193] In an embodiment of the present application, the entropy encoder may include: a Huffman decoding module (Huffman Decode), a data scheduling module (Data Dispatch), multiple run-length decoding modules (RunLength Decode1 - RunLength Decode8), a data integration module (Data Integrate), and a differential decoding module (DiffDecode).
[0194] Exemplarily, Huffman Decode is used to perform Huffman decoding according to the Huffman decoding table (Huffman Decode Table) to restore the run-length encoded compressed information. The run-length decoding module is used to restore the compressed information to pixels in the frequency domain. DiffDecode is used to perform parallel differential decoding.
[0195] In an embodiment of the present application, the data processing module may include: an inverse quantization module (DeQuantize), an inverse zigzag module (DeZigzag), a two-dimensional inverse discrete cosine transform module (IDCT 2D), YUV upsampling (YUVUpSampling), a block merging module (BlockMerge), and a pixel conversion module (YCbCr TO RGB).
[0196] Exemplarily, DeQuantize is used to perform parallel dequantization calculations. YUVUpSampling is used to perform upsampling to obtain YUV pixels.
[0197] It should be noted that for other descriptions of the JPEG decoding operator, reference can be made to the description of the JPEG encoding operator, which will not be elaborated here. For example, the descriptions of request commands and response commands, etc.
[0198] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0199] For ease of understanding, the data processing method provided by the embodiments of the present application will be introduced below in conjunction with the above system architecture and the accompanying drawings.
[0200] Figure 9 It is a flowchart of a data processing method provided by the embodiments of the present application. Exemplarily, the data processing method may include the following steps 901 to step 904.
[0201] It should be noted that the "step" in the embodiments of the present application can be abbreviated as "S", which will not be elaborated further.
[0202] In the embodiments of the present application, Figure 9 the data processing method shown can be executed by Figures 3 to 6 the data processing device shown.
[0203] The data processing device may include an FPGA, and the FPGA may include a scheduling unit, a processing unit, and a storage unit. Among them, the scheduling unit is used to execute task scheduling, such as: steps 901 - step 902, the processing unit is used to execute processing tasks, such as: steps 903 - step 904, and the storage unit is used to store the data to be processed for each processing task, the output data of each processing task, etc.
[0204] Step 901: The scheduling unit of the FPGA determines M processing tasks and the data to be processed for the M processing tasks.
[0205] Among them, the execution order of the M processing tasks is the target order, M is a positive integer greater than 1, and the data to be processed for the M processing tasks is stored in the storage unit.
[0206] In the embodiments of the present application, the scheduling unit can implement the determination of M processing tasks and the data to be processed for the M processing tasks by running a central controller (such as Figure 6 the central controller shown).
[0207] It should be noted that the processing task ranked first in the target order is called the first processing task / the first processing task,......, and the processing task ranked M can be called the Mth processing task / the Mth processing task. In addition, the data to be processed of the first processing task can be called the first data to be processed,......, and the data to be processed of the Mth processing task can be called the Mth data to be processed, which will not be elaborated hereinafter.
[0208] Exemplarily, the M processing tasks include an image decoding task, an image magnification task, and an image encoding task. Among them, the target order is the image decoding task, the image magnification task, and the image encoding task. That is to say, the image decoding task is executed first, then the image magnification task is executed, and finally the image encoding task is executed.
[0209] Hereinafter, through Method 1 to Method 2, the implementation manners of determining the first data to be processed will be exemplarily described.
[0210] Optionally, Method 1 may include the following S1 - S3.
[0211] In the embodiments of the present application, the data processing device may include a DPU, and the DPU may be used to receive a data processing request sent by an electronic device and convert the data processing request into a task request recognizable by an FPGA.
[0212] S1: The DPU determines a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks.
[0213] In the embodiments of the present application, the DPU may receive the data processing request and determine the target task request by running a software interface module. Exemplarily, when the electronic device needs to perform a target operation on the original image (i.e., the original data), such as: when the user instructs the electronic device to perform a target operation on the original image or when the electronic device needs to perform a target operation on the original image when performing a certain task, the electronic device may send a data processing request to the data processing device, and the data processing request is used to request to perform a target operation on the original image. After receiving the data processing request, the DPU of the data processing device converts the data processing request into a task request recognizable by the FPGA, such as: the target task request.
[0214] Example 1, the data processing request may include the original data. After receiving the original data, the DPU stores the original data in a local storage medium (such as Figure 6in the memory of the DPU shown). Example 2, the data processing request is used to indicate the identification, storage address, etc. of the original data. After receiving the data processing request, the DPU can obtain the original data based on the identification, storage address, etc. of the original data. For example, obtain the original data from a cloud storage device and store the obtained original data in a local storage medium. Exemplarily, the target operation may include sharpening, enlarging, color correction, shrinking, watermarking, matting, compositing, light and shade modification, modification of chroma and hue, adding special effects, editing, restoration, etc.
[0215] It should be noted that the embodiments of the present application do not limit the type of the target operation, and the above is only an exemplary illustration. Hereinafter, taking the target operation of enlarging as an example, the embodiments of the present application will be exemplarily described.
[0216] S2: The DPU sends a target task request to the scheduling unit; the target task request is used to indicate the original data and the target operation.
[0217] In the embodiments of the present application, the DPU can send a target task request to the scheduling unit of the FPGA by running the FPGA infrastructure, so as to request the FPGA to perform the target operation on the original data.
[0218] Exemplarily, the target task request can indicate the original data by indicating the storage address of the original data in the memory of the DPU.
[0219] S3: The scheduling unit determines M processing tasks and the first data to be processed based on the received target task request. Wherein, the first data to be processed is the original data, and the first data to be processed is the data to be processed of the processing task ranked first among the M processing tasks.
[0220] In the embodiments of the present application, the scheduling unit can receive the target task request sent by the DPU by running the FPGA infrastructure. On this basis, the scheduling unit can determine M processing tasks and the first data to be processed by running the central controller.
[0221] Optionally, the scheduling unit can obtain the original data from the local storage medium of the DPU based on the target task request.
[0222] Exemplarily, the target task request indicates the storage address of the original data. The scheduling unit obtains the storage address of the original data by parsing the task request. After that, the scheduling unit obtains the original data from the DPU through the storage address of the original data.
[0223] Optionally, the scheduling unit can obtain multiple processing tasks based on the target task request.
[0224] Exemplarily, the target task request is used to request to perform a target operation on the original image. The target operation is a zoom operation, and the target image is a JPEG image. Based on this, the scheduling unit determines, by parsing the task request, that the multiple processing tasks include a JPEG image decoding task, an image zooming task, and a JPEG image encoding task.
[0225] In this implementation manner, the processor receives a data processing request sent by the electronic device and converts the data processing request into a task request recognizable by the FPGA. In this way, it helps to reduce the software complexity of the FPGA, and thus helps to reduce the software development cost of the FPGA.
[0226] Optionally, the FPGA stores a task form, which can be used to indicate the execution order of multiple processing tasks, the task status of each processing task among the multiple processing tasks, the data to be processed of each processing task, etc. The task status includes an executed status or an unexecuted status.
[0227] In the embodiment of the present application, the scheduling unit can generate task description information based on the target task request. The task description information can be used to indicate the execution order of M processing tasks, the task status of each processing task among the M processing tasks, the storage address of the original data (i.e., the data to be processed of the processing task ranked first), etc. On this basis, the scheduling unit can store the task description information in the task form, so as to indicate the relevant information of the M processing tasks through the task identifier, such as: execution order, task status, storage address of the data to be processed, etc.
[0228] It should be noted that generating task description information based on the target task request can be referred to as initial task description information. The task status of each processing task indicated by the initial task description information is the initial task status, and the initial task status is the unexecuted status.
[0229] Exemplarily, the task form can be stored in the display look-up-table (LUT), DRMA, etc. of the FPGA.
[0230] In this embodiment, by configuring the task form to indicate the task status of each processing task, in this way, the scheduling unit does not need to perceive the processing task itself, such as: does not need to know what specific processing task is executed, the task status of the processing task, the life cycle of the processing task, the execution status of the processing task, etc., which helps to improve the applicable range of the data processing method, such as: can be used to process image data, audio data, etc., and further helps to improve the service migration ability of the data processing device. In addition, since the scheduling unit does not need to perceive the processing task itself, it helps to shorten the time occupied by the processing task on the scheduling unit, and thus helps to improve the task scheduling efficiency.
[0231] Optionally, Mode 2 may include S4 - S6.
[0232] In the embodiments of the present application, the FPGA may include a receiving unit, and the receiving unit may be configured to receive a data processing request sent by an electronic device and convert the data processing request into a task request recognizable by the FPGA.
[0233] S4: The receiving unit determines a target task request based on the received data processing request.
[0234] S5: The receiving unit sends the target task request to the scheduling unit.
[0235] S6: The receiving unit determines M processing tasks and first data to be processed based on the received target task request.
[0236] It should be noted that for the relevant descriptions of S4 - S6, reference can be made to the descriptions of S1 - S3, which will not be elaborated here.
[0237] In this implementation manner, the receiving unit of the FPGA receives the data processing request sent by the electronic device and converts the data processing request into a task request recognizable by the processing unit, which helps to reduce the hardware cost of the data processing device.
[0238] Optionally, the data processing method may further include: the scheduling unit stores the original data in the storage unit of the FPGA.
[0239] In the embodiments of the present application, the scheduling unit may apply for a storage space from the memory management unit of the FPGA. For example, the memory management unit allocates a first storage space to the scheduling unit. Based on this, the scheduling unit stores the original data in the first storage space.
[0240] Hereinafter, taking the second data to be processed as an example, through S7 - S8, an exemplary description of the implementation manner for determining the second data to be processed to the Mth data to be processed will be given.
[0241] S7: The scheduling unit receives the first response information returned by the processing unit. The first response information is used to indicate a first storage address, and the output data of the first processing task is stored at the first storage address. The output data of the first processing task is the data to be processed for the second processing task.
[0242] In the embodiments of the present application, after the processing unit finishes executing the first processing task, it returns the first response information to the scheduling unit. The first response information is used to indicate the output data of the first processing task. For example, it is used to indicate the storage address of the output data of the first processing task in the storage unit of the FPGA.
[0243] Exemplarily, the first processing task is a JPEG image decoding task. After the decoding operator unit completes the JPEG image decoding task by running the decoding operator, it returns the first response information to the scheduling unit. The scheduling unit receives the first response information by running the central controller.
[0244] S8: The scheduling unit determines the data stored at the first storage address as the second data to be processed.
[0245] In the embodiment of the present application, the storage address indicated by the first response information stores the output data of the first processing task. Based on this, the scheduling unit uses the output data of the first processing task as the second data to be processed.
[0246] Exemplarily, the scheduling unit determines the data stored at the first storage address as the second data to be processed by running the central controller.
[0247] In the above embodiment, the data to be processed for the next processing task is determined through the response information returned by the processing unit. In this way, it helps to ensure the accuracy of the determined data to be processed.
[0248] In the embodiment of the present application, the scheduling unit can update the task description information based on the received first response information. The updated task description information can be used to indicate that the task status of the first processing task is the executed state, the storage address of the output data of the first processing task, etc. In this way, it helps to improve the accuracy of the task description information, and thus realizes recording the latest status of different processing tasks and the data to be processed of different processing tasks through the task description information.
[0249] Step 902: The scheduling unit of the FPGA sends M request commands to the processing unit of the FPGA.
[0250] Wherein, each of the M request commands is used to request to execute each processing task according to the data to be processed of each processing task.
[0251] That is to say, one request command is used to request to execute one processing task according to the data to be processed of one processing task. Wherein, one processing task can be any one of the M processing tasks.
[0252] In the embodiment of the present application, the scheduling unit can send M request commands to the processing unit by running the central controller.
[0253] Optionally, the scheduling unit includes multiple sub-scheduling units. Wherein, each of the multiple sub-scheduling units can be used to send a request command to the processing unit. Exemplarily, step 902 can be that one scheduling unit sends one request command to the processing unit.
[0254] In this embodiment, by setting the scheduling unit to include multiple sub-scheduling units, it helps to schedule the processing tasks of different data processing requests simultaneously, thereby helping to improve the scheduling efficiency of multiple data processing requests. In addition, by combining the task form to maintain the task status, metadata, etc. of the processing tasks, each sub-scheduling unit no longer needs to maintain the life cycle, task status, etc. of the processing tasks. It only needs to allocate operator units for the processing tasks and send request commands, without waiting for the operator units to complete the processing tasks. In this way, it helps to achieve elastic scheduling of processing tasks, avoid processing tasks from occupying sub-scheduling units for a long time, and further helps to improve the scheduling efficiency.
[0255] Hereinafter, through Method A to Method B, an exemplary introduction to the implementation process of sending M request commands will be given.
[0256] Optionally, the processing unit includes multiple operator units. Among them, one operator unit can be used to execute one processing task.
[0257] It should be noted that multiple in the embodiments of the present application means two or more, which will not be elaborated hereinafter.
[0258] In one example, different operator units among multiple operator units are used to execute different processing tasks. In this way, it helps to increase the types of processing tasks that the processing unit can execute.
[0259] In another example, some of the multiple operator units are used to execute the same processing task. In this way, the processing unit can process multiple identical processing tasks simultaneously, such as: multiple identical processing tasks of multiple different data processing requests, thereby achieving parallel processing of multiple different data processing requests, and further being able to improve the data processing efficiency and data processing capacity.
[0260] In yet another example, some of the multiple operator units can be used to execute the same processing task, and another part of the operator units are used to execute different processing tasks. In this way, not only can the types of processing tasks executed by the processing unit be increased, but also parallel processing of different data processing requests can be achieved, which further helps to improve the data processing efficiency and data processing capacity of the data processing device.
[0261] Exemplarily, the M processing tasks include a target processing task, and the target processing task can be any one of the M processing tasks. Among the multiple operator units, there may be a target operator unit, and the target operator unit is the operator unit used to execute the target processing task.
[0262] Hereinafter, taking the target processing task and the target operator unit as an example, an exemplary description of step 902 will be given.
[0263] Optionally, Method A may include the following S9 - S10.
[0264] In the embodiments of the present application, the FPGA stores an operator form, and the operator form is used to indicate the processing tasks executed by each operator unit among multiple operator units. Among them, the operator form is used to indicate that the target operator unit is used to execute the target processing task.
[0265] Exemplarily, the operator form can be stored in the display look-up-table (LUT), DRMA, etc. of the FPGA.
[0266] Among them, the operator form and the task label form can be stored in different LUTs. In this way, it helps to improve the accuracy of obtaining the operator form / task form and the accuracy of updating the operator form / task form.
[0267] S9: The scheduling unit determines a target operator unit for the target processing task from among the multiple operator units indicated by the operator form.
[0268] In the embodiments of the present application, the operator form is used to indicate the processing tasks that each operator unit can execute. Based on this, when the scheduling unit needs to determine the operator unit that executes the target processing task, it can determine the operator unit that can execute the target processing task according to the content indicated by the operator form, that is, the target operator unit. For example, when the target processing task is an image decoding task, the scheduling unit can determine the decoding operator unit as the target operator unit.
[0269] Exemplarily, the target sub-scheduling unit among multiple sub-scheduling units determines a target operator to execute the target processing task from among the multiple operators indicated by the operator form by running a scheduler, where the target operator is run by the target operator unit, thereby realizing determining that the target operator unit is used to execute the target processing task.
[0270] S10: The scheduling unit sends a target request command to the target operator unit, and the target request command is used to request to execute the target processing task on the target data to be processed.
[0271] Exemplarily, the operator form includes the identifier of each operator unit, and different operator units have different identifiers. For example, the identifier of the target operator unit is a. Based on this, the scheduling unit can send a target request command to the operator unit with the identifier a among the multiple operator units (i.e., the target operator unit) to request the target operator unit to execute the target processing task. For example, it can be that the target sub-scheduling unit realizes sending a target request command to the target operator unit running the target operator by running a scheduler.
[0272] In the embodiments of the present application, the target request command is further used to indicate the target data to be processed, and the target data to be processed is the data to be processed for the target processing task. For example, the target request command may include the storage address of the target data to be processed, so as to realize indicating the target data to be processed, where the storage address is the storage address of the target data to be processed in the storage unit.
[0273] In the embodiments of the present application, when the target processing task is the processing task ranked first in the target order, for example: the target processing task is an image decoding task, and the target data to be processed is the original data. In addition, the output data of the Nth processing task is the data to be processed for the (N + 1)th processing task in the target order; N is a positive integer less than M. That is to say, when the target processing task is a processing task that is not ranked first in the target order, the target data to be processed is the output data of the processing task ranked one position before the target processing task in the target order. For example: when the target processing task is an image magnification task, the target data to be processed is the output data of the image decoding task; when the target processing task is an image encoding task, the target data to be processed is the output data of the image magnification task.
[0274] In this implementation manner, by configuring an operator form for the FPGA to indicate the processing tasks that each operator unit can execute, the scheduling logic of the scheduling unit is to determine the operator unit used to execute each processing task according to the operator form. In this way, when expanding the operator unit, it is only necessary to add the newly expanded operator unit to the operator form to realize scheduling the newly added operator unit to implement the newly added processing task, which helps to simplify the difficulty of expanding the operator unit, and further helps to improve the scalability of the FPGA.
[0275] Optionally, the operator form is further used to indicate the working state of each operator unit, and the working state includes an idle state or a busy state. The operator form indicates that the working state of the target operator unit is the idle state.
[0276] Among them, the idle state means that the operator unit has not been assigned a processing task, that is, the operator unit is not currently executing a processing task, and the busy state means that the operator unit has been assigned a processing task, that is, the operator unit is currently executing a processing task.
[0277] In this embodiment, by setting the operator form, the working state of each operator unit can be indicated. In this way, when determining the operator unit for executing each processing task, the operator units in the idle state can be screened out to execute the processing task, which helps to improve the execution efficiency of each processing task, and further helps to improve the working efficiency of the data processing device.
[0278] Optionally, after the scheduling unit determines a target operator unit for a target processing task among multiple operator units indicated by an operator form, the data processing method may include: the scheduling unit updates the working state of the target operator unit indicated by the operator form to a busy state.
[0279] In an embodiment of the present application, the target sub-scheduling unit updates the working state of the target operator unit to a busy state by running a scheduler.
[0280] In this embodiment, after allocating a processing task to the target operator unit, the working state of the target operator unit is updated to a busy state. In this way, it not only helps to ensure the accuracy of the working state indicated by the operator form, but also helps to avoid the target operator unit being assigned to other processing tasks during the execution of the target processing task, which affects the execution efficiency of other processing tasks.
[0281] Optionally, Mode B may include the following S11 - S13.
[0282] S11: The scheduling unit sends a query command to each operator unit among the multiple operator units, and the query command is used to query the processing tasks executed by each operator unit.
[0283] Exemplarily, the scheduling unit sends a query command to each operator unit to request querying the processing tasks executed by each operator unit, so as to determine the operator unit executing each processing task. For example, the target sub-scheduling unit updates the working state of the target operator unit to a busy state by running a scheduler.
[0284] S12: The scheduling unit determines a target operator unit for the target processing task based on the received multiple operator information.
[0285] Wherein, one piece of operator information is used to indicate the processing task executed by one operator unit.
[0286] In an embodiment of the present application, each operator unit among the multiple operator units responds to the query command by running an operator and returns its own operator information to the scheduling unit. Exemplarily, the target operator unit returns target operator information to the scheduling unit by running a target operator, and the target operator information is used to indicate that the target operator unit is used to execute the target processing unit.
[0287] Exemplarily, the target operator unit is used to execute an image decoding task. The target operator unit responds to the query command and returns target operator information to the scheduling unit, and the target operator information indicates that the target operator unit executes the image decoding task.
[0288] Exemplarily, the target processing task is an image decoding task. Based on this, the scheduling unit determines that the target operator unit executes the target processing task based on the target operator information among the received multiple operator information.
[0289] S13: The scheduling unit sends a target request command to the target operator unit, which is used to request the target operator unit to execute the target processing task on the target data to be processed.
[0290] It should be noted that for the relevant description of S13, reference can be made to the relevant description of S10, which will not be elaborated here.
[0291] In this implementation, by sending query commands to multiple operator units, the processing tasks executed by each operator unit are determined. In this way, when it is necessary to expand the operator units, only the newly expanded operator units need to be used as the objects for the scheduling unit to send query commands, which helps to simplify the difficulty of expanding the operator units and thus helps to improve the reliability of the FPGA.
[0292] In the embodiments of the present application, the query command can also be used to query the working status of each operator unit. An operator information can also be used to indicate the working status of an operator unit. For example, the target operator information can be used to execute the working status of the target operator unit. Based on this, when the target operator information indicates that the target operator unit is in an idle state, the scheduling unit determines that the target operator unit executes the target processing task.
[0293] It should be noted that for other relevant descriptions of Method B, reference can be made to the description of Method A above, which will not be elaborated here.
[0294] It should be noted that the embodiments of the present application do not limit how to send the request command. The above is only an exemplary description. Hereinafter, taking the above Method A as an example, the embodiments of the present application will be introduced exemplarily.
[0295] Step 903: The processing unit of the FPGA executes M processing tasks in the target order based on the received M request commands, and obtains the output data of the M processing tasks.
[0296] In the embodiments of the present application, the processing unit realizes receiving the request command by running operators. For example, the target operator unit of the processing unit realizes receiving the target request command by running the target operator. Exemplarily, after the target operator unit receives the target request command, based on the storage address of the target data to be processed indicated by the target request command, it obtains the target data to be processed from the storage unit of the FPGA. Then, the target operator unit executes the target processing task on the target data to be processed and obtains the output data of the target processing task.
[0297] Exemplarily, the target processing task is an image decoding task. After the target operator unit obtains the target data to be processed (i.e., the original data), it performs the target decoding task on the original data to obtain the output data of the image decoding task. Among them, the output data of the image decoding task is the data to be processed for the image magnification task.
[0298] Next, in combination with Figure 8 , an exemplary introduction to the process of performing the image decoding task will be given.
[0299] Exemplarily, the data to be processed for the image decoding task is the original data. After the decoding operator unit receives the image decoding request (i.e., commandRequest), it obtains the original data from the storage unit of the FPGA. For example, 512 bits of compressed data are obtained per clock cycle. Next, taking the processing process of 512 bits of compressed data as an example, an exemplary description of the decoding process will be given.
[0300] The Width Converter of the decoding operator unit splits the 512-bit compressed data and transmits the obtained Data 1 to the JFIF Parser. Among them, Data 1 includes 64 pixels, and each pixel is 8 bits. The JFIF Parser parses Data 1 to obtain the header information of the original data, the decoding mode (decodeMode), the quantization tables (Quantization Tables), the Huffman coding table, etc., and transmits Data 1 to the entropy decoder.
[0301] The Huffman Decode of the entropy decoder decodes Data 1 according to the Huffman decoding table to restore the compressed information of the run-length encoding, and transmits the obtained Data 2 to the Data Dispatch. Among them, each pixel of Data 2 is 19 bits. The Data Dispatch inputs the 64 pixels of Data 2 into multiple RunLengthDecodes for parallel decoding to improve the decoding efficiency. For example, input RunLengthDecode1 - RunLengthDecode8, and multiple RunLengthDecodes parallelly decode Data 2 to restore the compressed information to the pixels in the frequency domain, and transmit the obtained Data 3 to the Data Integrate. Among them, each pixel of Data 3 is 14 bits. The Data Integrate sorts the received Data 3 and transmits the obtained Data 4 to the DiffDecode.
[0302] DiffDecode performs parallel inverse quantization on the received data 4 according to the quantization table, and transmits the obtained data 5 to DeZigzag. DeZigzag processes the received data 5 and transmits the obtained data 6 to IDCT 2D. IDCT 2D processes the received data 6 and transmits the obtained data 7 to YUVUpSampling. YUVUpSampling performs parallel upsampling on the received data 7 and transmits the obtained data 8 to YCbCr TO RGB. YCbCr TO RGB performs parallel conversion on the received data 8 and transmits the RGB pixel data 9 of the obtained data to Output Generator, where the data 9 is 512 - bit data. Output Generator outputs the received data 9 through bitMapStream.
[0303] Next, in conjunction with Figure 7 , an exemplary introduction to the process of performing an image encoding task will be given.
[0304] Exemplarily, the data to be processed in the image encoding task is the output data of an image magnification task, where the data to be processed in the image encoding task is bitmap pixel data. The data to be processed includes multiple minimum coded units (MCUs). Each MCU includes 8 rows of images, each row of images includes 8 pixels, and each pixel is 24 bits. Next, taking the processing process of a single MCU as an example, an exemplary description of the execution process of the image encoding task will be given.
[0305] After the encoding operator unit receives an image encoding request, it obtains 1 row of images from the storage unit every clock cycle. RGB TO YCbCr can convert the image of 1 row of bitmap pixels into an image of YCbCr pixels. Among them, 1 bitmap pixel can be converted into 1 Y pixel, 1 Cb pixel, and 1 Cr pixel. Based on this, 1 row of bitmap pixels can be converted into 1 row of Y pixels, 1 row of Cb pixels, and 1 row of Cr pixels. 1 row of Y pixels / Cb pixels / Cr pixels includes 8 Y pixels / Cb pixels / Cr pixels. After 8 clock cycles, RGB TO YCbCr can output data 10 to BlockSplit. The data 10 includes 8 rows of YCbCr pixels, and each row of YCbCr pixels includes 1 row of Y pixels, 1 row of Cb pixels, and 1 row of Cr pixels. Among them, each Y pixel / Cb pixel / Cr pixel of the data 10 is 8 bits.
[0306] BlockSplit splits data 10 to obtain data 11. Data 11 includes 8 rows of Y pixels, 8 rows of Cb pixels, and 8 rows of Cr pixels. Each Y pixel / Cb pixel / Cr pixel in data 11 is 8 bits. BlockSplit sequentially outputs the 8 rows of Y pixels, 8 rows of Cb pixels, and 8 rows of Cr pixels in data 11 to DCT 2D. DCT 2D processes the received data 11 in parallel. For example, it receives 8Q pixels per clock cycle, converts the 8Q pixels from the spatial domain to the frequency domain, and transmits the obtained data 12 to ZIG ZAG. Here, each Y pixel / Cb pixel / Cr pixel in data 12 is 12 bits.
[0307] ZIG ZAG processes the received data 12. For example, it converts the two-dimensional image matrix into a one-dimensional vector in Zig-Zag order and transmits the obtained data 13 to Add Color Component. Here, each Y pixel / Cb pixel / Cr pixel in data 13 is 12 bits. Add Color Component adds type information to each row of pixels in data 13 to obtain data 14. The type information is used to indicate the pixel type, and the pixel types include Y pixels, Cb pixels, or Cr pixels. Each row of Y pixel / Cb pixel / Cr pixel in data 14 is 8 * 12 + 2 bits, where 8 represents 8 pixels, 12 represents 12 bits per pixel, and 2 bits are used to record the type information. Add Color Component transmits data 14 to Quantize. Quantize performs parallel quantization calculations on the received data 14. For example, it receives 8Q pixels per clock cycle, compresses the high-frequency components of the 8Q pixels, and transmits the obtained data 15 to the entropy encoder. Here, each row of Y pixel / Cb pixel / Cr pixel in data 15 is 8 * 12 bits.
[0308] The Add Color Component of the entropy encoder adds type information to each row of pixels in the received data 15 and transmits the obtained data 16 to DiffCode. Here, each row of Y pixel / Cb pixel / Cr pixel in data 16 is 8 * 12 + 2 bits. DiffCode performs parallel differential encoding on the received data 16. For example, it performs differential encoding on 8Q pixels per clock cycle and transmits the obtained data 17 to Data Dispatch. Data 17 includes 8 rows of Y pixels (i.e., 64 Y pixels), 8 rows of Cb pixels (i.e., 64 Cb pixels), and 8 rows of Cr pixels (i.e., 64 Cr pixels). Here, each row of Y pixel / Cb pixel / Cr pixel includes 8 Y pixels / Cb pixels / Cr pixels, and each Y pixel / Cb pixel / Cr pixel is 14 bits.
[0309] Data Dispatch transfers data 17 to multiple RunLength Encodes for parallel compression to improve compression efficiency, and transfers the resulting data 18 to Data Integrate. Among them, Data Dispatch can transfer the same type of pixels to the same RunLength Encode. For example, it transfers 8 rows of Y pixels to RunLength Encode1, 8 rows of Cb pixels to RunLength Encode2, and 8 rows of Cr to RunLength Encode3. Among them, data 18 includes A Y pixels output by RunLength Encode1, B Cb pixels output by RunLength Encode2, and C Cr pixels output by RunLength Encode3. A / B / C is less than 64, and each Y pixel / Cb pixel / Cr pixel is 19 bits.
[0310] Data Integrate sorts the received data 18 to obtain data 19, and sequentially outputs A rows of Y pixels, B rows of Cb pixels, and C rows of Cr pixels in data 19 to Huffman Encode in order. Each Y pixel / Cb pixel / Cr pixel is 19 bits. Huffman Encode encodes the received data 19 according to the Huffman coding table, and transfers the resulting data 20 to AlignCode. Each Y pixel / Cb pixel / Cr pixel in data 20 is 33 bits. AlignCode processes the received data 20 and transfers the resulting data 21 to ByteStuff. ByteStuff processes the received data 21 and transfers the resulting data 22 to Output Generator. Output Generator synthesizes the received data 22 and JPEG header data, and outputs the resulting data 23 through JPEGStream.
[0311] Optionally, the data processing method may further include: after the target operator unit finishes executing the target processing task, it returns target response information to the scheduling unit. Among them, the target response information is used to indicate the storage address of the output data of the target processing task, the execution result of the target processing task, etc. The execution result may include execution success or failure.
[0312] In this embodiment, after the operator unit finishes executing the processing task, it returns response information to the scheduling unit. In this way, the scheduling unit can not only determine the data to be processed for the next processing task according to the response information, but also determine the final data obtained by executing M processing tasks.
[0313] In the embodiment of the present application, after the target operator unit finishes executing the target processing task, it returns target response information to the scheduling unit, and the scheduling unit can receive the target response information returned by the target operator unit. Exemplarily, the target operator unit returns the target response information to the scheduling unit by running the target operator, and the scheduling unit receives the target response information by running the central controller.
[0314] In the embodiment of the present application, based on the received target response information, the scheduling unit can update the task description information, and the updated task description information can be used to indicate that the task status of the target processing task is the executed status, the storage address of the output data of the target processing task, etc. On this basis, the scheduling unit can write the updated task description information into the task form, so as to update the task form, and the updated task form can be used to indicate the content indicated by the updated task description information.
[0315] Exemplarily, when the target processing task is an image decoding task, the updated task description information can be used to indicate that the task status of the image decoding task is the executed status and the storage address of the output data of the image decoding task (i.e., the data to be processed by the image magnification task).
[0316] Optionally, the data processing method may include: in response to the target response information, the scheduling unit updates the working status of the target operator unit indicated in the operator form to the idle status. Exemplarily, the scheduling unit updates the working status of the target operator unit to the idle status by running the central controller. In this way, it not only helps to ensure the accuracy of the working status indicated in the operator form, but also enables the scheduling unit to allocate the operator unit for other processing tasks, thus helping to improve the execution efficiency of other processing tasks.
[0317] Step 904: The processing unit stores the output data of the M processing tasks in the storage unit.
[0318] Exemplarily, the target operator unit can apply for storage space from the storage unit, such as: allocating the third storage space to the target operator unit. Based on this, the target operator unit stores the output data of the target processing task in the third storage space.
[0319] Optionally, the data processing method may further include: when the processing unit finishes executing the M processing tasks, the scheduling unit outputs target data to the DPU, and the target data is the output data of the processing task ranked Mth in the target order.
[0320] Exemplarily, the scheduling unit may determine the storage address of the target data according to the response information returned by the processing unit and obtain the target data. After that, the scheduling unit may output the target data to the DPU so that the DPU returns the target data to the electronic device, thereby completing the content requested by the data processing request sent by the electronic device.
[0321] In the embodiments of the present application, the execution order of steps 901 to 904 is not limited. Hereinafter, taking the target operation as an image magnification operation as an example, the execution process of steps 901-904 will be introduced exemplarily.
[0322] After receiving the image magnification operation request sent by the user through the electronic device, the DPU converts the image magnification operation request into an image magnification task request recognizable by the FPGA and sends the image magnification task request to the scheduling unit of the FPGA. In response to receiving the image magnification task request, the scheduling unit stores the original data obtained from the DPU in the first storage space of the storage unit of the FPGA and generates the first task description information of the image magnification task request. The first task description information is used to indicate a plurality of processing tasks and the storage address of the original data on the storage unit, wherein the execution order of the plurality of processing tasks is an image decoding task, an image magnification task, and an image encoding task.
[0323] On this basis, the scheduling unit may store the first task description information in the task form and determine the first decoding operator unit from the plurality of operator units indicated by the operator form to execute the image decoding task. After that, the scheduling unit sends a decoding request to the first decoding operator unit. The decoding request indicates the storage address of the original data and is used to request the first decoding operator unit to execute the decoding task on the original data.
[0324] Exemplarily, the task parsing unit of the scheduling unit, in response to receiving the image magnification task request, applies to the memory management unit for a storage space to store the original data. The memory management unit allocates the first storage space of the storage unit to the task parsing unit. The task parsing unit stores the original data obtained from the DPU in the first storage space of the storage unit and generates the first task description information. After that, the task parsing unit sends the first task description information to the task distribution unit. In response to receiving the task description information, the task distribution unit determines the first sub-scheduling unit from at least one sub-scheduling unit for task scheduling of the image decoding task. After receiving the first task description information forwarded by the task distribution unit, the first sub-scheduling unit stores the first task description information in the task form and determines the first decoding operator unit to execute the image decoding task. After that, the first sub-scheduling unit sends a decoding request to the first decoding operator unit through the control flow arbitration and the interconnection unit.
[0325] Exemplarily, the first sub-scheduling unit is any one of the sub-scheduling units in the idle state. The operator form can also be used to indicate that the working state of the first decoding operator unit is the idle state. After the first sub-scheduling unit determines that the first decoding operator unit is used to execute the image decoding task, it can update the working state of the first decoding operator unit indicated by the operator form to the busy state.
[0326] In response to the received decoding request, the first decoding operator unit obtains the data to be processed (i.e., the original data) from the first storage space of the storage unit, and executes the decoding task on the data to be processed. After the first decoding operator unit finishes executing the decoding task on the original data, it obtains the first output data. The first decoding operator unit applies to the memory operation arbitration unit for a storage space to store the first output data. The memory operation arbitration unit allocates the second storage space of the storage unit to the first decoding operator unit. The first decoding operator unit stores the first output data in the second storage space of the storage unit, and returns the first response information to the scheduling unit through the control flow arbitration and the interconnection unit. The first response information is used to indicate that the image decoding task has been completed, the storage address of the first output data of the image decoding task, etc.
[0327] Based on the received first response information, the scheduling unit updates the first task description information to obtain the second task description information. The second task description information is used to indicate that the task state of the image decoding task is the executed state, the storage address of the first output data of the image decoding task, etc.
[0328] Exemplarily, in response to the received first response information, the task distribution unit of the scheduling unit determines the first task recovery unit from at least one task recovery unit to perform the task recovery work of the image decoding task, and forwards the first response information to the first task recovery unit. In response to the received first response information, the first task recovery unit obtains the first task description information from the task form and updates the first task description information, thereby obtaining the second task description information.
[0329] Exemplarily, the first task recovery unit is any one of the task recovery units in the idle state. The first task recovery unit can also, in response to the first response information, update the working state of the first decoding operator unit indicated by the operator form to the idle state to release the first decoding operator unit.
[0330] After the scheduling unit obtains the second task description information, it stores the second task description information in the task form, and determines that the first image magnification operator unit executes the image magnification task. Then, the scheduling unit sends an image magnification request to the first image magnification operator unit to request the first image magnification operator unit to perform the image magnification task on the first output data.
[0331] It should be noted that for the determination of the first image magnification operator unit by the scheduling unit and other relevant descriptions of sending the image magnification request, reference can be made to the description of the scheduling unit determining the first decoding operator unit and sending the image decoding request, which will not be elaborated here.
[0332] Exemplarily, the first task recycling unit may send the obtained second task description information to the task distribution unit, and the task distribution unit determines a sub-scheduling unit for the next processing task (i.e., the image magnification task). After receiving the second task description information, the task distribution unit determines a second sub-scheduling unit from at least the first sub-scheduling unit for task scheduling of the image magnification task. The second sub-scheduling unit updates the task form based on the second task description information so that the updated task form can execute the content indicated by the second task description information, determines the first image magnification operator unit for executing the image magnification task, and sends an image magnification request to the first image magnification operator unit.
[0333] It should be noted that for other relevant descriptions of the second sub-scheduling unit, the first image magnification operator unit, and the image magnification request, reference can be made to the description of the first sub-scheduling unit, the first decoding operator unit, and the image decoding request, which will not be elaborated here.
[0334] After the first image magnification operator unit finishes executing the image magnification task, it returns second response information to the scheduling unit and stores the second output data of the image magnification task in the third storage space of the storage unit. In response to the received second response information, the scheduling unit updates the second task description information to obtain third task description information, which is used to indicate that the task status of the image magnification task is the executed state, the storage address of the second output data of the image magnification task, etc. Then, the scheduling unit determines the first encoding operator unit for executing the image encoding task and sends an encoding request to the first encoding operator unit.
[0335] It should be noted that for other relevant descriptions of the scheduling unit determining the first decoding operator unit and sending the encoding request, reference can be made to the description of the scheduling unit determining the first decoding operator unit and sending the decoding request, which will not be elaborated here.
[0336] Exemplarily, the task distribution unit determines the second sub-scheduling unit for task scheduling of the image magnification task. The task distribution unit determines that the second task recycling unit performs the task recycling work of the image magnification task.
[0337] It should be noted that for other relevant descriptions of the second sub-scheduling unit and the second task recycling unit, reference can be made to the description of the first sub-scheduling unit and the first task recycling unit, which will not be elaborated here.
[0338] After the first encoding operator unit finishes the image encoding task, it returns the third response information to the scheduling unit and stores the third output data of the image encoding task in the fourth storage space of the storage unit. In response to the received third response information, the scheduling unit returns a task response information to the DPU, and the task response information is used to indicate the storage address of the third output data, etc. In response to the received task response information, the DPU obtains the third output data in the storage unit of the FPGA and sends the third output data to the electronic device to return the target data (i.e., the third output data) obtained by executing the image magnification operation request to the user.
[0339] Exemplarily, in response to the received third response information, the task distribution unit determines that the third task recovery unit is used to perform the task recovery work of the image encoding task and forwards the third response information to the third task recovery unit. Based on the received third response information, the third task recovery unit updates the third task description information to obtain the fourth task description information, and the fourth task description information is used to indicate that the task status of the image encoding task is the executed state, the storage address of the third output data of the image encoding task, etc. Then, the third task recovery unit sends the fourth task description information to the task distribution unit, and the task distribution unit determines that the image magnification task request has been completed based on the received fourth task description information. Then, the task distribution unit sends the fourth description information to the task output unit, and the task output unit sends the received task response information to the DPU and sends the third output data stored in the fourth storage space of the storage unit to the DPU.
[0340] So far, the data processing device has completed the image magnification operation request for the original data.
[0341] Exemplarily, combined with Figure 6 , the operations performed by the task parsing unit are implemented by running the task parser. The operations performed by the task output module are implemented by running the task output module. The operations performed by the task distribution unit are implemented by running the task dispatcher. The operations performed by the sub-scheduling units (such as: the first sub-scheduling unit, the second sub-scheduling unit, the third sub-scheduling unit, etc.) are implemented by running the scheduler. The operations performed by the task recovery units (such as: the first task recovery unit, the second task recovery unit, the third task recovery unit, etc.) are implemented by running the recovery agent. The operations performed by the memory management unit are implemented by running the memory manager. The operations performed by the control flow arbitration and interconnection unit are implemented by running the control flow arbitration and interconnection module. The operations performed by the memory operation arbitration unit are implemented by the memory operation arbitration. The operations performed by the operator units (such as: the decoding operator unit, the image magnification operator unit, the encoding operator unit, etc.) are implemented by running the operators (such as: the decoding operator, the image magnification operator, the encoding operator, etc.).
[0342] In the above embodiments, data processing services are provided by a data processing device. Since the data processing device does not need to deploy hardware resources based on the X86 server architecture in related technologies, the hardware resources used by the data processing device are reduced. In this way, not only the data processing cost is reduced, but also the utilization rate of the hardware resources of the data processing device is improved. In addition, the FPGA independently completes M processing tasks that need to be executed in a specified order, and the output data of each processing task is stored in the storage unit of the FPGA. In this way, when the processing unit executes each processing task, it can directly obtain the data to be processed from the storage unit of the accelerator in hardware, thereby avoiding the end-to-end data transmission delay during the execution of the M processing tasks, and further improving the data processing efficiency.
[0343] In addition, since the data processing device adopts a numerically controlled separation architecture, that is, when executing M processing tasks, task scheduling is performed by the scheduling unit of the FPGA. For example, a request command is sent to the processing unit to instruct the processing unit to execute a processing task on the data to be processed, and the processing unit of the FPGA executes the processing task on the data to be processed, thereby realizing the decoupling / separation of the control flow (such as task scheduling) and the data flow (such as the data to be processed generated by the processing task). In this way, it helps to simplify the scheduling logic of the scheduling unit and improve the scalability of the data processing device. For example, when adding an executable processing task to the processing unit, there is no need to change the scheduling logic (i.e., software code) of the scheduling unit. In addition, due to the adoption of the numerically controlled separation architecture, the data processing device can also perform task scheduling and processing tasks in parallel, thereby improving the parallel processing ability of the data processing device, and further helping to improve the data processing efficiency.
[0344] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, the data processing device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0345] Embodiments of the present application can, according to the above method, exemplarily divide a data processing device into functional modules. For example, the data processing device may include respective functional modules corresponding to each functional division, or two or more functions may be integrated into one processing module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, merely a logical functional division, and there may be other division methods in actual implementation.
[0346] Exemplarily, Figure 10 FIG. shows a possible structural schematic diagram of the data processing device (denoted as data processing device 1000) involved in the above embodiments. The actions performed by the data processing device are implemented by a data processing device or by the data processing device executing corresponding software. The data processing device may include an FPGA, and the FPGA may include a scheduling unit, a processing unit, and a storage unit. The data processing device 1000 may include a scheduling module 1001 and a processing module 1002. The scheduling module 1001 is configured to determine M processing tasks and the data to be processed for the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit. For example, as Figure 9 shown in S901. The scheduling module 1001 is further configured to send M request commands to the processing unit; wherein each request command is used to request to execute each processing task according to the data to be processed for each processing task. For example, as Figure 9 shown in S902. The processing module 1002 is configured to, based on the received M request commands, execute the M processing tasks in the target order to obtain the output data of the M processing tasks; wherein the output data of the Nth processing task is the data to be processed for the (N + 1)th processing task in the target order; N is a positive integer less than M. For example, as Figure 9 shown in S903. The processing module 1002 is further configured to store the output data of the M processing tasks in the storage unit. For example, as Figure 9 shown in S904.
[0347] Optionally, the data processing device 1000 further includes a software interface module 1003; the software interface module 1003 is configured to: determine a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks; the software interface module 1003 is further configured to: send the target task request to the scheduling module 1001; the target task request is used to indicate the original data and the target operation; the scheduling module 1001 is specifically configured to: based on the received target task request, determine the first-ranked processing task among the M processing tasks and the data to be processed of the first-ranked processing task; the data to be processed of the first-ranked processing task is the original data.
[0348] Optionally, the processing module includes a plurality of operators; the FPGA stores an operator form, and the operator form is used to indicate the processing tasks performed by each operator among the plurality of operators; the scheduling module 1001 is specifically configured to: determine a target operator for the target processing task among the M processing tasks from the plurality of operators indicated by the operator form, and the operator form is used to indicate that the target operator is used to perform the target processing task; send a target request command to the target operator, and the target request command is used to request to perform the target processing task on the target data to be processed.
[0349] Optionally, the operator form is further used to indicate the working state of each operator unit, and the working state includes an idle state or a busy state; the operator form indicates that the working state of the target operator unit is the idle state.
[0350] Optionally, after the scheduling module 1001 determines a target operator unit for the target processing task among the M processing tasks from the plurality of operator units indicated by the operator form, the scheduling module 1001 is further configured to: update the working state of the target operator unit indicated by the operator form to the busy state.
[0351] Optionally, when the target operator finishes executing the target processing task, the scheduling module 1001 is further configured to: update the working state of the target operator indicated by the operator form to the idle state.
[0352] Optionally, when the target operator is used to perform the target processing task based on the target request command, the target operator is configured to return target response information to the scheduling module, and the target response information is used to indicate the storage address of the output data of the target processing task.
[0353] Optionally, the target response information is used to indicate the execution result of the target processing task, and the execution result includes execution success or execution failure.
[0354] Optionally, the FPGA stores a task form, which is used to indicate the task status of M processing tasks. The task status includes an executed status or an unexecuted status. The scheduling module 1001 is further configured to: update the task form based on the received target response information; the updated task form is used to indicate that the target processing task is in an executed status.
[0355] Optionally, the scheduling module 1001 may include multiple sub-scheduling modules, where each sub-scheduling module is configured to send a request command to the processing unit.
[0356] Optionally, when the processing unit has completed the M processing tasks, the scheduling module 1001 is further configured to: output target data to the processor, where the target data is the output data of the processing task ranked Mth in the target order.
[0357] Optionally, the data to be processed is image data. The processing module 1002 includes a decoding operator, which is configured to perform an image decoding task. The decoding operator includes multiple decoding modules, and the multiple decoding modules are configured to decode 8Q pixels of the image data in parallel.
[0358] Optionally, the processing module 1002 includes an encoding operator, which is configured to perform an image encoding task. The encoding operator includes multiple encoding modules, and the multiple encoding modules are configured to encode 8Q pixels of the image data in parallel.
[0359] For the specific descriptions of the above optional manners, reference may be made to the foregoing method embodiments, which will not be elaborated herein. In addition, the explanations and descriptions of the beneficial effects of any of the data processing devices 1000 provided above may refer to the corresponding method embodiments above, which will not be elaborated.
[0360] An embodiment of the present application further provides a computer program product, which includes computer programs / instructions. When the computer programs / instructions are executed by a data processing device, the data processing device executes the steps of the above data processing method.
[0361] An embodiment of the present application further provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a data processing device, the data processing device executes the steps of the above data processing method.
[0362] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method, characterized in that: The method is applied to a data processing device, wherein the data processing device includes a field programmable gate array (FPGA), and the FPGA includes a scheduling unit, a processing unit, and a storage unit; the method includes: The scheduling unit determines M processing tasks and data to be processed of the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; The scheduling unit sends M request commands to the processing unit, wherein each request command is used to request execution of each processing task according to the to-be-processed data of each processing task; The processing unit executes the M processing tasks according to the target sequence based on the received M request commands to obtain output data of the M processing tasks; wherein the output data of the Nth processing task is the to-be-processed data of the N+1th processing task in the target sequence; and N is a positive integer less than M; The processing unit stores the output data of the M processing tasks in the storage unit.
2. The method according to claim 1, characterized in that The processing unit includes a plurality of operator units; the FPGA stores an operator table, and the operator table is used to indicate the processing task performed by each of the plurality of operator units; The scheduling unit sends M request commands to the processing unit, including: The scheduling unit determines a target operator unit for a target processing task among the M processing tasks from a plurality of operator units indicated by the operator table, wherein the operator table is used to indicate that the target operator unit is used to execute the target processing task; The scheduling unit sends a target request command to the target operator unit, where the target request command is used to request execution of the target processing task on the target data to be processed.
3. The method according to claim 2, characterized in that The operator table is further used to indicate the working state of each operator unit, where the working state includes an idle state or a busy state; the operator table is used to indicate that the working state of the target operator unit is an idle state.
4. The method according to claim 2 or 3, characterized in that: The method further comprises: When the target operator unit completes executing the target processing task based on the target request command, the target operator unit returns target response information to the scheduling unit, where the target response information is used to indicate a storage address of output data of the target processing task.
5. The method according to claim 4, characterized in that The target response information is also used to indicate the execution result of the target processing task, and the execution result includes execution success or execution failure.
6. The method according to claim 4 or 5, characterized in that: The FPGA stores a task form, and the task form is used to indicate the task status of the M processing tasks, and the task status includes an executed state or an unexecuted state; the method further includes: The scheduling unit updates the task list based on the received target response information; the updated task list is used to indicate that the target processing task is in an executed state.
7. The method according to any one of claims 1 to 6, characterized in that The data processing device further includes a processor, the processor being configured to convert a received data processing request into a task request recognizable by the FPGA; The scheduling unit determines M processing tasks and the to-be-processed data of the M processing tasks, including: The processor determines a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes the M processing tasks; The processor sends the target task request to the scheduling unit; the target task request is used to indicate the original data and the target operation; The scheduling unit determines the M processing tasks and the to-be-processed data of the processing task ranked first among the M processing tasks based on the received target task request; the to-be-processed data of the processing task ranked first among the M processing tasks is the original data.
8. The method according to claim 7, characterized in that The method further comprises: When the processing unit completes executing the M processing tasks, the scheduling unit outputs target data to the processor, where the target data is output data of the processing task ranked Mth in the target sequence.
9. The method according to any one of claims 1 to 8, characterized in that The scheduling unit includes a plurality of sub-scheduling units, wherein each sub-scheduling unit is used to send a request command to the processing unit.
10. The method according to any one of claims 1 to 9, characterized in that The data to be processed is image data; the processing unit includes a decoding operator unit, the decoding operator unit is used to perform an image decoding task, the decoding operator unit includes a plurality of run-length decoding modules, and the decoding operator unit decodes 8Q pixels of the image data in parallel through the plurality of run-length decoding modules; Q is a positive integer greater than 1; and / or The processing unit includes a coding operator unit, which is used to perform image coding tasks. The coding operator unit includes multiple run-length coding modules. The coding operator unit encodes 8Q pixels of the image data in parallel through the multiple run-length coding modules.
11. A data processing device, characterized in that: Used in a data processing device, wherein the hardware processor of the data processing device includes a processing unit and a storage unit; the data processing device includes: A scheduling module, used for determining M processing tasks and data to be processed of the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; The scheduling module is further used to send M request commands to the processing unit, wherein each request command is used to request execution of each processing task according to the to-be-processed data of each processing task; A processing module, configured to execute the M processing tasks in the target order based on the received M request commands, and obtain output data of the M processing tasks; wherein the output data of the Nth processing task is the to-be-processed data of the N+1th processing task in the target order; and N is a positive integer less than M; The processing module is further used to store the output data of the M processing tasks in the storage unit.
12. A data processing device, characterized in that: comprising an FPGA, the FPGA comprising: A scheduling unit is used to determine M processing tasks and data to be processed of the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; The scheduling unit is further used to send M request commands to the processing unit; wherein each request command is used to request execution of each processing task according to the to-be-processed data of each processing task; A processing unit, configured to execute the M processing tasks in the target order based on the received M request commands, and obtain output data of the M processing tasks; wherein the output data of the Nth processing task is the to-be-processed data of the N+1th processing task in the target order; and N is a positive integer less than M; The processing unit is further used to store the output data of the M processing tasks in the storage unit.
13. A data processing equipment resource pool, characterized in that: The data processing device resource pool includes at least one data processing device, and the at least one data processing device is used to execute the steps of the method according to any one of claims 1-10.
14. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a data processing device, implements the steps of the method according to any one of claims 1 to 10.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a data processing device, the steps of the method according to any one of claims 1 to 10 are implemented.