Data processing method, apparatus and device, and device resource pool

Through the separate architecture of FPGA, data processing tasks are independently completed, which solves the problems of high cost and low efficiency in the existing technology, and realizes the efficient utilization of hardware resources and the improvement of data processing efficiency.

WO2025139007A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/116583
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2024-09-03
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing data processing methods rely on high-cost server architecture, resulting in low hardware resource utilization and high data transmission delay, thereby reducing data processing efficiency.

Method used

The separate architecture of FPGA is adopted, including scheduling units, processing units and storage units. Multiple processing tasks are completed independently through FPGA, and the control flow and data flow are decoupled to achieve the separation of task scheduling and data storage, reducing hardware resource requirements and improving parallel processing capabilities.

Benefits of technology

It reduces data processing costs, improves hardware resource utilization and data processing efficiency, simplifies scheduling logic, and enhances the scalability and parallel processing capabilities of data processing equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116583_03072025_PF_FP_ABST
    Figure CN2024116583_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a data processing method, apparatus and device, and a device resource pool, which relate to the technical field of data processing. The method is applied to a data processing device, wherein an FPGA of the data processing device comprises a scheduling unit, a processing unit and a storage unit. During data processing, the scheduling unit executes task scheduling, and the processing unit executes a plurality of processing tasks and stores output data of each processing task in the storage unit to serve as data to be processed of the next processing task. In the method, an FPGA independently completes a plurality of processing tasks, and output data of each processing task is stored in a storage unit. In this way, when a processing unit executes each processing task, data to be processed can be directly acquired from the storage unit, so as to avoid an end-to-end data transmission delay during the execution of M processing tasks, thereby improving the data processing efficiency. In addition, the FPGA uses a data-control separation architecture, so as to realize the decoupling of a control flow and a data flow, thereby facilitating an improvement to the scalability of data processing devices.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device, equipment and equipment resource pool

[0001] This application claims priority to the Chinese patent applications filed with the State Intellectual Property Office on December 26, 2023, with application number 202311811962.0 and application name “A method, device and other equipment for image processing”, and filed with the State Intellectual Property Office on April 16, 2024, with application number 202410458074.3 and application name “Data processing method, device, equipment and equipment resource pool”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of data processing technology, and in particular to a data processing method, apparatus, equipment, and equipment resource pool. Background Art

[0003] Currently, data processing requests are usually executed by servers. When servers handle multiple processing tasks related to data processing requests, they are usually completed by the central processing unit (CPU) and heterogeneous computing power in collaboration. For example, the CPU performs task scheduling and processing tasks that take less time, while heterogeneous computing power performs processing tasks that take more time.

[0004] However, this data processing approach results in high data processing costs due to the high cost of servers. Furthermore, it increases the data transmission latency between the central processing unit and heterogeneous operators, resulting in lower data processing efficiency.

[0005] Summary of the Invention

[0006] The present application provides a data processing method, apparatus, equipment, and equipment resource pool, which can not only reduce data processing costs but also improve data processing efficiency.

[0007] In order to achieve the above objectives, this application adopts the following technical solutions:

[0008] In a first aspect, a data processing method is provided, which is applied to a data processing device, wherein a field programmable gate array (FPGA) of the data processing device includes a scheduling unit, a processing unit, and a storage unit; the method includes: the scheduling unit determines M processing tasks and the to-be-processed data of the M processing tasks; the execution order of the M processing tasks is a target order, M is a positive integer greater than 1; the to-be-processed data is stored in the storage unit; the scheduling unit sends M request commands to the processing unit; wherein each request command is used to request execution of each processing task according to the to-be-processed data of each processing task; based on the received M request commands, the processing unit executes the M processing tasks in the target order to obtain output data of the M processing tasks; wherein the output data of the Nth processing task is the to-be-processed data of the N+1th processing task in the target order; N is a positive integer less than M; and the processing unit stores the output data of the M processing tasks in the storage unit.

[0009] In this solution, data processing services are provided by a data processing device that utilizes a decoupled architecture. Because the data processing device does not require hardware resources deployed based on the X86 server architecture used in related technologies, the hardware resources used by the data processing device are reduced. This not only reduces data processing costs but also improves the hardware resource utilization of the data processing device. Furthermore, the FPGA independently completes M processing tasks that must be executed in a specified order, and the output data of each processing task is stored in the FPGA's memory unit. As the processing unit executes each processing task, it can directly retrieve the data to be processed from the memory unit through hardware, thus avoiding end-to-end data transmission delays during the execution of the M processing tasks and thereby improving data processing efficiency.

[0010] In addition, since the data processing device adopts a numerical control separation architecture, that is, when executing M processing tasks, task scheduling is performed through the FPGA's scheduling unit, such as sending a request command to the processing unit to instruct the processing unit to perform a processing task on the processing data, and the processing unit of the FPGA performs the processing task on the processing data, thereby achieving the decoupling / separation of the control flow (such as task scheduling) and the data flow (such as the processing data generated by the processing task). This helps to simplify the scheduling logic of the scheduling unit and improve the scalability of the data processing device. For example, when adding an executable processing task to the processing unit, there is no need to change the scheduling logic (i.e., software code) of the scheduling unit. In addition, due to the adoption of the numerical control separation architecture, the data processing device can also perform task scheduling and processing tasks in parallel, thereby improving the parallel processing capability of the data processing device, and thus helping to improve data processing efficiency.

[0011] In one possible implementation, the data to be processed is image data; the processing unit includes a decoding operator unit, which is configured to perform image decoding tasks. The decoding operator unit includes multiple decoding modules, and the decoding operator unit decodes 8Q pixels of the image data in parallel using the multiple decoding modules, where K is a positive integer greater than 1. This helps improve decoding efficiency.

[0012] In another possible implementation, the processing unit includes an encoding operator unit, which is configured to perform image encoding tasks. The encoding operator unit includes multiple encoding modules, and the encoding operator unit encodes 8Q pixels of the image data in parallel using the multiple encoding modules. This helps improve encoding efficiency.

[0013] In another possible implementation, the data processing device also includes a processor, which is used to convert a received data processing request into a task request recognizable by the FPGA; the scheduling unit determines M processing tasks and the data to be processed of the M processing tasks, including: the processor determines a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks; the processor sends a target task request to the scheduling unit; the target task request is used to indicate the original data and the target operation; the scheduling unit determines the processing task ranked first among the M processing tasks and the data to be processed of the processing task ranked first based on the received target task request; the data to be processed of the processing task ranked first is the original data.

[0014] In this implementation, a separate processor receives data processing requests sent by an electronic device, which can reduce the software development cost of the FPGA. The processor can be a DPU / CPU.

[0015] In another possible implementation, the processing unit includes multiple operator units; the FPGA stores an operator table, which is used to indicate the processing tasks performed by each operator unit in the multiple operator units; the scheduling unit sends M request commands to the processing unit, including: the scheduling unit determines a target operator unit for a target processing task in the M processing tasks from the multiple operator units indicated by the operator table, and the operator table is used to indicate that the target operator unit is used to perform the target processing task; the scheduling unit sends a target request command to the target operator unit, and the target request command is used to request to perform the target processing task on the target data to be processed.

[0016] In this implementation, the operator unit of each processing task is determined by the content indicated by the operator form. In this way, when a new operator unit is added to the processing unit, the scheduling logic of the scheduling unit does not need to be changed, which helps to reduce development costs.

[0017] In another possible implementation, the operator table also indicates the operating status of each operator unit, including idle or busy. The operator table indicates that the target operator unit's operating status is idle. This allows idle operators to be assigned to processing tasks, thereby improving processing efficiency.

[0018] In another possible implementation, after the scheduling unit determines a target operator unit for a target processing task among the M processing tasks from the multiple operator units indicated in the operator table, the method further includes: the scheduling unit updating the working state of the target operator unit indicated in the operator table to a busy state. This helps ensure the accuracy of the working state indicated by the operator table.

[0019] In another possible implementation, when the target operator unit completes the target processing task, the scheduling unit updates the working state of the target operator unit indicated by the operator table to an idle state. This helps ensure the accuracy of the working state indicated by the operator table.

[0020] In another possible implementation, the method further includes: when the target operator unit executes the target processing task based on the target request command, returning target response information to the scheduling unit, where the target response information indicates the storage address of the output data of the target processing task. In this way, the scheduling unit can determine the storage address of the to-be-processed data of the next processing task based on the response information, thereby helping to ensure the accuracy of the to-be-processed data executed by the request command.

[0021] In another possible implementation, the target response information is also used to indicate the execution result of the target processing task, where the execution result includes execution success or execution failure. This helps the scheduling unit understand the execution status of the processing task.

[0022] In another possible implementation, the FPGA stores a task table indicating the status of M processing tasks, including whether the task status is executed or not. The method further includes: the scheduling unit updating the task table based on the received target response information; the updated task table indicates that the target processing task is in the executed state. In this way, the scheduling unit can terminate task scheduling after sending the request command, eliminating the need to constantly monitor the execution status of the processing tasks, thereby reducing the scheduling time for each processing task during the task scheduling process.

[0023] In another possible implementation, the scheduling unit includes multiple sub-scheduling units, wherein each sub-scheduling unit is used to send a request command to the processing unit. In this way, the scheduling unit can execute task scheduling of multiple data processing requests in parallel, thereby helping to improve the data processing capability of the data processing device.

[0024] In another possible implementation, the method further includes: when the processing unit completes executing M processing tasks, the scheduling unit outputs target data to the processor, where the target data is output data of the processing task ranked Mth in the target sequence.

[0025] In a second aspect, a data processing device is provided, comprising: a functional unit for executing any one of the methods provided in the first aspect, wherein the actions executed by each functional unit are implemented by hardware or by hardware executing corresponding software implementations. For example, the data processing device is used in a data processing device, wherein the hardware processor of the data processing device includes a processing unit and a storage unit; the data processing device may include a scheduling module and a processing module; the scheduling module is configured to determine M processing tasks and data to be processed for the M processing tasks; the execution order of the M processing tasks is a target order, where M is a positive integer greater than 1; the data to be processed is stored in the storage unit; the scheduling module is further configured to send M request commands to the processing unit, wherein each request command is configured to request execution of each processing task based on the data to be processed of each processing task; the processing module is configured to execute the M processing tasks in the target order based on the received M request commands, and obtain output data of the M processing tasks; wherein the output data of the Nth processing task is the data to be processed of the N+1th processing task in the target order, where N is a positive integer less than M; and the processing module is further configured to store the output data of the M processing tasks in the storage unit.

[0026] In one possible implementation, the processing module includes multiple operators; the FPGA stores an operator list, which is used to indicate the processing tasks performed by each of the multiple operators; the scheduling module is specifically used to: determine a target operator for a target processing task among the M processing tasks from the multiple operators indicated by the operator list, the operator list is used to indicate the target operator for performing the target processing task; and send a target request command to the target operator, the target request command being used to request execution of the target processing task on the target data to be processed.

[0027] In another possible implementation, the operator table is further used to indicate the working state of each operator unit, where the working state includes an idle state or a busy state; the operator table indicates that the working state of the target operator unit is an idle state.

[0028] In another possible implementation, after the scheduling module determines the target operator unit for the target processing task in the M processing tasks from the multiple operator units indicated by the operator table, the scheduling module is also used to: update the working status of the target operator unit indicated by the operator table to a busy state.

[0029] In another possible implementation, when the target operator completes executing the target processing task, the scheduling module is further configured to update the working state of the target operator indicated by the operator table to an idle state.

[0030] In another possible implementation, when the target operator completes the target processing task based on the target request command, the target operator is also used to return target response information to the scheduling module, and the target response information is used to indicate the storage address of the output data of the target processing task.

[0031] In another possible implementation, the target response information is further used to indicate the execution result of the target processing task, where the execution result includes execution success or execution failure.

[0032] In another possible implementation, the FPGA stores a task form, which is used to indicate the task status of M processing tasks, including the executed status or the unexecuted status. The scheduling module is also used to: update the task form based on the received target response information; the updated task form is used to indicate that the target processing task is in the executed state.

[0033] In another possible implementation, the scheduling module may include multiple sub-scheduling modules, wherein each sub-scheduling module is used to send a request command to the processing unit.

[0034] In another possible implementation, the data processing device also includes a software interface module; the software interface module is used to: determine a target task request based on a received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks; the software interface module is also used to: send a target task request to the scheduling module; the target task request is used to indicate the original data and the target operation; the scheduling module is specifically used to: determine the first-ranked processing task among the M processing tasks and the data to be processed of the first-ranked processing task based on the received target task request; the data to be processed of the first-ranked processing task is the original data.

[0035] In another possible implementation, when the processing unit completes executing M processing tasks, the scheduling module is further configured to: output target data to the processor, where the target data is output data of the processing task ranked Mth in the target sequence.

[0036] In another possible implementation, the data to be processed is image data. The processing module includes a decoding operator, which is used to perform image decoding tasks. The decoding operator includes multiple decoding modules, which are used to decode 8Q pixels of the image data in parallel.

[0037] In another possible implementation, the processing module includes a coding operator, which is used to perform an image coding task. The coding operator includes multiple coding modules, which are used to encode 8Q pixels of the image data in parallel.

[0038] According to a third aspect, a data processing device is provided, including: a scheduling unit, a processing unit and a storage unit; the scheduling unit is used to determine M processing tasks and the data to be processed of the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; the scheduling unit is also used to send M request commands to the processing unit; wherein each request command is used to request the execution of each processing task according to the data to be processed of each processing task; the processing unit is used to execute the M processing tasks in the target order based on the received M request commands to obtain the output data of the M processing tasks; wherein the output data of the Nth processing task is the data to be processed of the N+1th processing task in the target order; N is a positive integer less than M; the processing unit is also used to store the output data of the M processing tasks in the storage unit.

[0039] It should be noted that in the third aspect, the scheduling unit and the processing unit can also be used to execute any possible implementation method provided in the first aspect above.

[0040] In a fourth aspect, a data processing device resource pool is provided, comprising: at least one data processing device; and at least one data processing device configured to execute the steps of any one of the methods provided in the first aspect.

[0041] In a fifth aspect, a computer program product is provided, which includes a computer program / instructions, and when the computer program / instructions are executed by a data processing device, the steps of any one of the methods provided in the first aspect are implemented.

[0042] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a data processing device, the steps of any one of the methods provided in the first aspect are implemented.

[0043] Among them, the technical effects brought about by any implementation method in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] FIG1 is a schematic diagram of a related technology provided by an embodiment of the present application;

[0045] FIG2 is a schematic diagram of another related technology provided by an embodiment of the present application;

[0046] FIG3 is a schematic diagram of a system architecture provided in an embodiment of the present application;

[0047] FIG4 is a schematic diagram of another system architecture provided in an embodiment of the present application;

[0048] FIG5 is a schematic diagram of another system architecture provided in an embodiment of the present application;

[0049] FIG6 is a schematic diagram of a data device provided in an embodiment of the present application;

[0050] FIG7 is a schematic diagram of an encoding operator unit provided in an embodiment of the present application;

[0051] FIG8 is a schematic diagram of a decoding operator unit provided in an embodiment of the present application;

[0052] FIG9 is a flow chart of a data processing method provided in an embodiment of the present application;

[0053] FIG10 is a schematic diagram of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0055] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, among which A and B can be singular or plural.

[0056] Furthermore, in the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0057] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0058] The following is an illustrative introduction to the relevant terms involved in the embodiments of the present application.

[0059] Decoupled architecture: This refers to an architecture that breaks down the various hardware resources in a server architecture and uses a subset of them. Some of these hardware resources can communicate and interact over the network.

[0060] Operator: refers to the module on the hardware accelerator used to perform specified operations.

[0061] In an embodiment of the present application, the operator can be a software module on the FPGA for performing processing tasks, such as: a module for performing a decoding task, a module for performing an encoding task, a module for performing an image magnification task, etc.

[0062] Zero-layer logical architecture: This is used during software development to define the relationships and interactions between modules within a system. It divides the system into multiple layers based on functionality, with each layer responsible for different functions and responsibilities.

[0063] Application programming interface (API): A predefined function that provides applications and developers with access to a set of routines based on a piece of software or hardware without having access to the source code.

[0064] Direct Memory Access (DMA): is a mechanism that allows data to be transferred directly between hardware devices and memory devices. In other words, data can be transferred between hardware devices and memory devices without relying on the CPU.

[0065] The following is an illustrative introduction to the application scenarios of the embodiments of the present application.

[0066] The data processing method provided in the embodiment of the present application is applicable to scenarios such as processing image data, audio data, etc. Below, the data processing method provided in the embodiment of the present application is exemplarily described using image data as an example.

[0067] It should be noted that the embodiments of the present application do not limit the application scenarios of the data processing method, and the above is only an exemplary description.

[0068] In related technologies, image data processing services are usually deployed on server clusters, which provide image processing services in the form of microservices. As shown in Figure 1, after a user stores image data on a cloud storage device, he or she can send an image processing request to the server cluster through an electronic device. After the server cluster receives the image processing request, it can specify server 1 in the server cluster to execute the image processing request. Server 1 obtains the image data from the cloud storage device and returns the processed image to the user. Among them, an image processing request can include multiple image processing tasks. For example, a processing request to sharpen a JPEG image can include decoding the JPEG image to obtain a bitmap image, sharpening the bitmap image, encoding the sharpened bitmap image to obtain a JPEG image, etc.

[0069] Server 1 deploys hardware resources based on an X86 server architecture, including a central processing unit (CPU), a network interface card (NIC), a solid-state disk / drive (SSD), and dynamic random access memory (DRAM). When server 1 executes an image processing request, it is specifically executed by the CPU within server 1. Due to the limited processing power of the CPU, the image data processing capabilities provided by server 1 are relatively poor.

[0070] In order to improve the image processing capability of server 1, as shown in FIG2 , the related art configures a graphics processing unit (GPU) for server 1 so that the CPU and GPU (heterogeneous computing power) can collaboratively complete image processing requests. Among them, the CPU is used to perform task scheduling and simple image processing tasks (such as sharpening tasks), and the GPU is used to perform time-consuming image processing tasks (such as image encoding tasks, image decoding tasks, etc.) to improve data processing capabilities. Specifically, after the CPU obtains image data based on the image processing request, it sends the image data to the GPU and instructs the GPU to perform the image decoding task. The GPU sends the bitmap image obtained by performing the image decoding task to the CPU. After the CPU performs the sharpening task on the bitmap image, it sends the sharpened bitmap image to the GPU and instructs the GPU to perform the image encoding task. The GPU returns the JFEG image obtained by performing the image encoding task to the CPU, which is then returned to the user by the CPU.

[0071] However, since the server architecture also requires hardware resources such as network interface cards (NICs) and solid-state disks (SSDs), having the CPU and GPU inside the server collaborate to complete image processing requests not only fails to fully utilize the server's hardware resources, resulting in low server hardware resource utilization and high image processing costs, but also increases end-to-end data transmission latency, that is, the latency in transmitting data between the CPU and GPU, thereby reducing image processing efficiency.

[0072] In view of this, an embodiment of the present application provides a data processing method, which is applied to a data processing device, wherein the FPGA of the data processing device includes a scheduling unit, a processing unit, and a storage unit. When the data processing device processes multiple processing tasks requested by data processing, the scheduling unit performs task scheduling, such as sending multiple request commands to the processing unit, wherein each of the multiple request commands is used to request the execution of a processing task, and the processing unit is used to execute the multiple processing tasks, such as executing the multiple processing tasks based on the multiple request commands, and storing output data of each processing task in the storage unit.

[0073] Since the data processing method is executed by a data processing device in the embodiment of the present application, the data processing device adopts a split architecture. Since the data processing device does not need to deploy hardware resources based on the X86 server architecture in the related art, the hardware resources used by the data processing device are reduced. This not only reduces the data processing cost but also improves the hardware resource utilization of the data processing device. In addition, multiple processing tasks are independently completed by the FPGA, and the output data of each processing task is stored in the storage unit of the FPGA. In this way, when the processing unit executes each processing task, it can directly obtain the data to be processed from the storage unit, thereby avoiding the end-to-end data transmission delay during the execution of multiple processing tasks, that is, the data transmission delay between the CPU and GPU in the related art, thereby improving data processing efficiency.

[0074] In addition, since the data processing device adopts a numerical control separation architecture, that is, when executing multiple processing tasks, task scheduling is performed through the FPGA's scheduling unit, such as sending a request command to the processing unit to instruct the processing unit to perform a processing task on the processing data, and the processing unit of the FPGA performs the processing task on the processing data, thereby achieving the decoupling / separation of the control flow (such as task scheduling) and the data flow (such as the processing data generated by the processing task). This helps to simplify the scheduling logic of the scheduling unit and improve the scalability of the data processing device. For example, when adding an executable processing task to the processing unit, there is no need to change the scheduling logic (i.e., software code) of the scheduling unit. In addition, due to the adoption of the numerical control separation architecture, the data processing device can also perform task scheduling and processing tasks in parallel, thereby improving the parallel processing capability of the data processing device, and thus helping to improve data processing efficiency.

[0075] The following is an exemplary introduction to the system architecture provided in the embodiments of the present application.

[0076] An embodiment of the present application provides a data processing device that can be used to perform the data processing method provided in the embodiment of the present application. The data processing device may include an FPGA, which is used to perform task scheduling and multiple processing tasks. The FPGA may include a scheduling unit, a processing unit, and a storage unit. The scheduling unit may be used to perform task scheduling of multiple processing tasks, the processing unit may be used to perform multiple processing tasks, and the storage unit may be used to store output data of each of the multiple processing tasks.

[0077] In one implementation, the FPGA may include a receiving unit. The receiving unit may be configured to receive data processing requests sent by an electronic device, convert the data processing requests into task requests recognizable by a scheduling unit, and then send the task requests to the scheduling unit. This helps reduce the hardware cost of the data processing device.

[0078] In another implementation, the data processing device may further include a processor. The processor may be configured to receive data processing requests from the electronic device, convert the data processing requests into task requests recognizable by the scheduling unit, and then transmit the task requests to the scheduling unit. This helps reduce the complexity of the FPGA's software logic, thereby reducing the cost of FPGA software development.

[0079] Exemplarily, the processor may be a data processing unit (DPU), a CPU, or the like.

[0080] It should be noted that the embodiments of the present application do not limit the type of processor, and the above is only an exemplary description.

[0081] In the embodiment of the present application, since the data processing device adopts a separate architecture, that is, a non-X86 server architecture, the data processing device can also be called a separate device, which will not be described in detail later.

[0082] An embodiment of the present application further provides a data processing device resource pool, which includes multiple data processing devices, and different data processing devices in the multiple data processing devices can communicate with each other.

[0083] Among them, the data processing device resource pool can be used to execute the data processing method provided in the embodiment of the present application, which helps to improve data processing capabilities and thus improve data processing efficiency.

[0084] Hereinafter, the embodiments of the present application are exemplarily described by taking a data processing device including an FPGA and a processor as an example.

[0085] FIG3 is a schematic diagram of a system architecture provided in an embodiment of the present application.

[0086] As shown in FIG3 , the system architecture may include an electronic device and a data processing device, wherein the data processing device may include an FPGA and a processor (eg, a DPU). The electronic device may communicate with the data processing device.

[0087] In an embodiment of the present application, an electronic device sends a data processing request and original data to a data processing device. The data processing request may include multiple processing tasks. After receiving the data processing request, the data processing device may perform multiple processing tasks on the original data based on the data processing method provided in the embodiment of the present application.

[0088] It should be noted that the embodiments of the present application do not limit the manner in which the electronic device sends a data processing request to the data processing device. For example, an application may be installed on the electronic device, and the electronic device may send a data processing request to the data processing device through the application.

[0089] In an embodiment of the present application, the FPGA may include a scheduling unit, a processing unit, and a storage unit, and the scheduling unit, the processing unit, and the storage unit may communicate with each other.

[0090] During the execution of a data processing request by a data processing device, the scheduling unit can be used to schedule multiple processing tasks. Specifically, the scheduling unit determines the operator unit that will execute each processing task and sends a request command to the operator unit to request the operator unit to execute the processing task. The processing unit can execute multiple processing tasks using multiple operator units, and the storage unit can store output data from executing each processing task.

[0091] Optionally, the scheduling unit may include a task parsing unit, a task dispatching unit, at least one sub-scheduling unit, at least one task recycling unit, a task output unit, etc. The task parsing unit may communicate with the task dispatching unit and the task output unit, and the task dispatching unit may communicate with each sub-scheduling unit in the at least one sub-scheduling unit, each task recycling unit in the at least one task recycling unit, and the task output unit.

[0092] In embodiments of the present application, the task parsing unit can be used to parse task requests to determine raw data and multiple processing tasks. The task dispatching unit can be used to determine a sub-scheduling unit to perform task scheduling, a task recycling unit to update a task table, and the like. The sub-scheduling unit can be used to determine an operator unit to perform processing tasks, send request commands, update an operator table, and the like. The task recycling unit can be used to update the task table and update the operator table to release an operator unit. The task output unit is used to output target data, and the like.

[0093] Optionally, the FPGA may also include a memory management unit. The memory management unit can communicate with the task output unit, the task parsing unit, and the like. The memory management unit is used to manage the storage space of the storage unit, such as allocating storage space for task parsing tasks.

[0094] Optionally, the scheduling unit may further include a DMA arbitration unit. The DMA arbitration unit is connected to the task parsing unit, the task output unit, each sub-scheduling unit, each task recycling unit, etc. The DMA arbitration unit may be used to arbitrate units using the DMA controller.

[0095] Optionally, the processing unit includes multiple operator units, and each of the multiple operator units can be used to perform processing tasks.

[0096] Exemplarily, the multiple operator units may include an encoder operator unit, a decoder operator unit, an image cropping operator unit, an image resizing operator unit, an image watermarking operator unit, an image sharpening operator unit, and the like.

[0097] The decoding operator unit is used to perform image decoding tasks, such as decompressing an image in a specified format into a bitmap format. The encoding operator unit is used to perform image encoding tasks, such as compressing a bitmap format into an image in a specified format. The image scaling operator unit is used to scale images. The image watermarking operator unit is used to apply watermarks to images. The image cropping operator unit is used to crop images. The image sharpening operator unit is used to sharpen images.

[0098] Exemplarily, the encoding operator unit may include a Joint Photographic Experts Group (JPEG) encoding operator unit, a Web Picture (WEBP) encoding operator unit, a Portable Network Graphics (PNG) encoding operator unit, a Graphics Interchange Format (GIF) encoding operator unit, etc. The decoding operator unit may include a JPEG decoding operator unit, a WEBP decoding operator unit, a PNG decoding operator unit, and a GIF decoding operator unit.

[0099] Optionally, the processing unit may include a DMA arbitration unit. The DMA arbitration unit may communicate with each operator unit and may be used to arbitrate between operator units using the DMA controller.

[0100] Optionally, the processing unit may include a control flow arbitration and interconnection unit. The control flow arbitration and interconnection unit can communicate with each operator unit. The control flow arbitration and interconnection unit is used to implement communication between the scheduling unit and the processing unit. For example, the sub-scheduling unit can communicate with the operator unit through the control flow arbitration and interconnection unit, and the sub-operator unit can communicate with the task dispatch unit through the control flow arbitration and interconnection unit.

[0101] Optionally, the processing unit may include a memory operation arbitration unit. The memory operation arbitration unit may communicate with each operator unit. The memory arbitration unit may be used to allocate storage space to the operator unit, release storage space not used by the operator unit, etc.

[0102] Optionally, the storage unit may be a volatile storage medium, such as a dynamic random access memory (DRAM).

[0103] In the embodiment of the present application, the FPGA may include multiple configurable logic blocks (CLBs), wherein different units of the FPGA are composed of different CLBs, such as a scheduling unit and a processing unit, which are composed of different CLBs.

[0104] Exemplarily, the plurality of configurable logic blocks may include a first configurable logic block and a second configurable logic block, and the first configurable logic block and the second configurable logic block are different configurable logic blocks.

[0105] It should be noted that the first configurable logic block / second configurable logic block refers to a type of configurable logic block among multiple configurable logic blocks, and does not refer to a specific configurable logic block. In addition, the embodiments of the present application do not limit the number of configurable logic blocks included in the first configurable logic block / second configurable logic block.

[0106] Exemplarily, the scheduling unit includes a first configurable logic block, and the processing unit includes a second configurable logic block, wherein each of the plurality of operator units is composed of a different second configurable logic block.

[0107] Optionally, the electronic device may be a terminal device or a network device.

[0108] Terminal devices may include mobile phones, tablet computers, handheld computers, PCs, cellular phones, personal digital assistants (PDAs), wearable devices (such as smart watches and smart bracelets), smart home devices (such as televisions), car computers (such as on-board computers), smart screens, game consoles, headphones, AI speakers, augmented reality (AR) / virtual reality (VR) devices, ultra-mobile personal computers (UMPCs), laptops, netbooks, desktop computers or all-in-one computers, etc.

[0109] Network devices may include servers, etc. A server may be a single physical server, or two or more physical servers sharing different responsibilities and working together to implement various server functions. For example, the server may be a blade server, a high-density server, a rack server, or a tower server.

[0110] It should be noted that the embodiments of the present application do not limit the device type of the electronic device, and the above description is only an example.

[0111] FIG4 is a schematic diagram of another system architecture provided in an embodiment of the present application.

[0112] As shown in Figure 4, the system architecture may include an electronic device, a data processing device, and a cloud storage device. The electronic device may communicate with the data processing device and the cloud storage device, and the data processing device may communicate with the cloud storage device.

[0113] Exemplarily, the cloud storage device stores a plurality of data, such as data 10, ..., data P0, where P is a positive integer greater than 1, and the plurality of data may be image data, audio data, and the like.

[0114] In an embodiment of the present application, a user can store multiple data on a cloud storage device through an electronic device. Afterwards, a data processing device can execute the data processing method provided in an embodiment of the present application on the data stored in the cloud storage device.

[0115] It should be noted that the electronic device that stores multiple data on the cloud storage device may be the same as or different from the electronic device that requests the data processing device to execute the data processing method provided in this application, and this application does not impose any restrictions on this.

[0116] Exemplarily, an electronic device sends a data processing request to a data processing device. The data processing request indicates that data 10 in a cloud storage device is original data, and the data processing request includes multiple processing tasks. Based on the data processing request, the data processing device obtains data 10 from the cloud storage device and performs the multiple processing tasks on the data 10.

[0117] It should be noted that for other relevant descriptions of the system architecture shown in FIG4 , reference can be made to the description of the system architecture shown in FIG3 , which will not be repeated here.

[0118] FIG5 is a schematic diagram of another system architecture provided in an embodiment of the present application.

[0119] As shown in Figure 5, the system architecture may include electronic devices, management devices, a data processing device resource pool, and cloud storage devices. The electronic devices may communicate with the management device and the cloud storage device, and the management device may communicate with each data processing device in the data processing device resource pool.

[0120] In an embodiment of the present application, the electronic device may send a data processing request to the management device, and the management device may determine a data processing device to execute the data processing request from a data processing device resource pool according to a predetermined balancing strategy.

[0121] Exemplarily, the management device may include an elastic load balance (ELB), and the management device may determine the data processing device that executes the data processing request from a data processing device resource pool through the ELB.

[0122] It should be noted that for other relevant descriptions of the system architecture shown in FIG5 , reference can be made to the descriptions of the system architecture shown in FIG3 and FIG4 , which will not be repeated here.

[0123] FIG6 is a schematic diagram of a data processing device provided in an embodiment of the present application.

[0124] The data processing device shown in Figure 6 may be the data processing device in the system architecture shown in Figures 3 to 5. Exemplarily, the data processing device may adopt a zero-layer logical architecture.

[0125] As shown in Figure 6, in terms of hardware, the data processing device may include an FPGA and a DPU. Exemplarily, the FPGA and the DPU may communicate based on the PCIE protocol.

[0126] FPGAs can include dynamic random access memory (DRAM), also known as internal memory. It should be noted that FPGA DRAM is also known as on-board DRAM. FPGA DRAM can be used to store raw data, intermediate output data, target data, and more.

[0127] It should be noted that the intermediate output data and target data will be described in subsequent embodiments and will not be described in detail here.

[0128] The DPU may include a memory, which may be used to store original data, target data, and the like.

[0129] Exemplarily, the operating system of the DPU may use a Linux operating system, and the Linux operating system running on the DPU may include a Linux user space (Userspace) and a Linux kernel (Kernel).

[0130] It should be noted that the embodiments of the present application do not limit the operating system used by the DPU, and the above is only an exemplary description.

[0131] In terms of software, a data processing device may include a software interface module, FPGA infrastructure, a scheduling module, and an operator pool. The operator pool may include multiple operators. The software interface module can communicate with the FPGA infrastructure, which in turn can communicate with the scheduling module, which in turn can communicate with each operator in the operator pool.

[0132] For example, job requests (Job Request), job responses (Job Response), data (Data), etc. can be transmitted between the software interface module and the FPGA infrastructure.

[0133] It should be noted that when a software module performs an operation (such as communication), it can be considered that the hardware performs an operation during the process of running the software module, or the hardware performs an operation through the software module, which will not be further described later.

[0134] It can be understood that when the data processing device adopts a zero-layer logic architecture, the software interface module, FPGA infrastructure, scheduling module and operator pool can be considered as zero-layer logic units.

[0135] It should be noted that the embodiments of the present application do not limit the names of the software modules included in the data device. The above is only an exemplary description. For example, the scheduling module can also be called a central controller.

[0136] The following is an exemplary introduction to the software interface module.

[0137] The software interface module can be deployed on the DPU. The DPU converts received data processing requests into task requests recognizable by the FPGA by running the software interface module. For example, the software interface module can run in Linux Userspace.

[0138] Optionally, the software interface module may include a calling surface and a management and control surface.

[0139] The call plane is used to provide a software call interface, convert data processing requests into task requests, and send task requests to the FPGA. The control plane is used to monitor the working status of the software interface module, output alarm information when the working status is abnormal, and generate system logs.

[0140] In embodiments of the present application, the call interface may include a call interface (Software APIs) module and a job manager (Job Manager). The call interface module is used to provide a unified task call interface and convert data processing requests into task requests, such as consolidating multiple consecutive operations involved in a data processing request into a single task request. The task manager is used to provide task management capabilities, such as sending task requests generated by the call interface module to the FPGA.

[0141] In an embodiment of the present application, the control plane may include a system logger module and an abnormal alarm module. The system logger module may be used to provide log management capabilities, and the abnormal alarm module may be used to provide abnormal alarm capabilities.

[0142] The following is an example introduction to the FPGA infrastructure.

[0143] FPGA infrastructure can be deployed on both the DPU and the FPGA. In other words, the FPGA infrastructure spans the DPU and FPGA. The DPU and FPGA can communicate and exchange data by running the FPGA infrastructure. In other words, the FPGA infrastructure provides an interface for data transmission and exchange between the DPU and FPGA.

[0144] For example, the FPGA infrastructure can be used to write data to the DPU's memory, read data from the memory, etc.

[0145] In an embodiment of the present application, the FPGA infrastructure may include an API module, a driver (Drive), and a shell (Shell) module. The API module and driver are deployed on the DPU, while the Shell module is deployed on the FPGA. The API module is used to communicate with the software interface module, and the Shell module is used to communicate with the central controller. The API module and the Shell module communicate through the driver. The API module can be a module implemented in C language.

[0146] For example, job requests, job responses, and data can be transmitted between the API module and the driver. Job requests, job responses, and data can be transmitted between the driver and the Shell module. Data can be transmitted between the driver and the DPU's storage. Data can be transmitted between the Shell module and the FPGA's storage units.

[0147] For example, the API module can run in Linux Userspace, and the driver can run in Linux Kernel.

[0148] Exemplarily, the API module may include a Read / Write module and a User Interrupt / Poll module. The Read / Write module is used to communicate with the PCIE driver to transmit task requests, task responses, original data, target data, etc.

[0149] Exemplarily, the Shell module may include a memory controller (DRAM Controller), a direct memory access controller (DMA Controller), an advanced extensible interface data move engine (AXI Data Mover), and the like.

[0150] For example, the driver of the FPGA infrastructure can be used to write data to the memory of the DPU, read data from the memory, etc.

[0151] For example, the DPU and FPGA can communicate based on the PCIE protocol. Based on this, the drive can be a PCIE drive and the DMA controller can be a PCIE DMA controller.

[0152] The following is an exemplary introduction to the central controller.

[0153] The central controller can be deployed on the FPGA's scheduling unit. The FPGA's scheduling unit can implement task scheduling by running the central controller, such as parsing, scheduling, and managing task flows, managing operator pools, and controlling data flows.

[0154] In an embodiment of the present application, the central controller may include a task parser (Job Parser), a task output module (Job Output), a task dispatcher (Job Descriptor Dispatcher), at least one scheduler (Scheduler Stateless Scheduler), and at least one recycler (Recycler).

[0155] Exemplarily, the task parser, task output module, scheduler, and recycler are stateless modules.

[0156] For example, a task request (Job Request) and the like can be transmitted between the task parser and the Shell module. A task descriptor (Job Descriptor) can be transmitted between the task parser and the task dispatcher and the task output module. A task response (Job Response) can be transmitted between the task output module and the Shell module. A task descriptor (Job Descriptor) can be transmitted between the task dispatcher and the task output module and the scheduler. A task descriptor (Job Descriptor), a response command (Cmd Response), and the like can be transmitted between the task dispatcher and the recycler. Other control flows (Other Control Flow) can be transmitted between the scheduler and the task form and the operator form. Other control flows (Other Control Flow) can be transmitted between the recycler and the task form and the operator form.

[0157] Among them, the task descriptor can also be called task description information, and the response command can also be called response information, which will not be repeated later.

[0158] The task resolver can be used to parse task requests to determine the raw data and multiple processing tasks. The task resolver is deployed on the task resolution unit, which runs the task resolver to parse task requests to determine the raw data and multiple processing tasks. The task output module is used to output target data. The task output module is deployed on the task output unit, which runs the task output module to output the target data.

[0159] The task dispatcher can be used to determine the sub-scheduling unit that executes task scheduling and the task recovery unit that updates the task form / operator form. The task dispatcher is deployed on the task dispatch unit, which runs the task dispatcher to determine the sub-scheduling unit that executes task scheduling and the task recovery unit that updates the task form / operator form.

[0160] The scheduler can be used to perform task scheduling, such as determining the processing tasks to be scheduled, determining the operator units for the processing tasks, and scheduling the operator units to execute the processing tasks. There is a one-to-one correspondence between schedulers and sub-scheduling units. A scheduler is deployed on a sub-scheduling unit, and a sub-scheduling unit executes task scheduling by running a scheduler.

[0161] The collector can be used to release operator units, check the execution results of processing tasks, synchronize the execution results of processing tasks, etc. There is a one-to-one correspondence between collectors and task recycling units. One collector is deployed on one task recycling unit. A task recycling unit releases operator units, checks the execution results of processing tasks, synchronizes the execution results of processing tasks, etc. by running one collector.

[0162] Optionally, the central controller may further include a memory manager. The memory manager is used to manage the storage space of the storage unit. The memory manager may be deployed in the memory management unit, which manages the storage space of the storage unit by running the memory manager. Exemplarily, the memory manager is a stateful module.

[0163] Exemplarily, other control flows (Other Control Flow) can be transmitted between the memory manager and the task parser and task output module.

[0164] Optionally, the central controller may further include a DMA arbiter (DMA Arbiter Stateless). The DMA arbiter may be used to arbitrate between units using the DMA controller. The DMA arbiter is deployed in the DMA arbitration unit, which arbitrates between units using the DMA controller by running the DMA arbiter.

[0165] For example, data can be transmitted between the DMA controller and the task parser and task output module, and other data flows can be transmitted between the DMA controller and the scheduler and recycler.

[0166] The following is an exemplary introduction to the operator pool.

[0167] An operator pool is deployed on the FPGA's processing unit and can include multiple operators. An FPGA's processing unit can execute multiple processing tasks by running multiple operators. Each operator can be deployed on an operator unit, and each operator unit can execute one processing task by running one operator.

[0168] Optionally, the multiple operators may include an encoding operator, a decoding operator, an image scaling operator, an image watermarking operator, an image cropping operator, and the like.

[0169] The decoding operator is used to decompress an image in a specified format into a bitmap format. The decoding operator is deployed on the decoding operator unit, which runs the decoding operator to decompress the image in the specified format into a bitmap format. The encoding operator is used to compress a bitmap format into an image in a specified format. The encoding operator is deployed on the encoding operator unit, which runs the encoding operator to compress the bitmap format into an image in the specified format.

[0170] The image scaling operator is used to scale images. It is deployed on the image scaling operator unit, which scales images by running the image scaling operator. The image watermark operator is used to process image watermarks. It is deployed on the image watermark operator unit, which processes image watermarks by running the image watermark operator. The image cropping operator is used to crop images. It is deployed on the image cropping operator unit, which crops images by running the image cropping operator.

[0171] For example, encoding operators may include a JPEG Encoder operator, a WEBP Encoder operator, a PNG Encoder operator, a GIF Encoder operator, etc. Decoding operators may include a JPEG Decoder operator, a WEBP Decoder operator, a PNG Decoder operator, a GIF Encoder operator, etc.

[0172] Optionally, the decoding operator may include multiple run-length decoding modules, and the decoding operator unit decodes 8Q pixels of the image data in parallel through the multiple run-length decoding modules, where Q is a positive integer greater than 1. This helps improve decoding efficiency.

[0173] Optionally, the encoding operator unit may include multiple run-length encoding modules, and the encoding operator unit encodes 8Q pixels of the image data in parallel through the multiple run-length encoding modules, which helps to improve encoding efficiency.

[0174] Optionally, the operator pool may further include a control flow arbitration and interconnection module. The control flow arbitration and interconnection module is deployed on the control flow arbitration and interconnection unit, and the control flow arbitration and interconnection unit implements designated functions by running the control flow arbitration and interconnection module.

[0175] For example, the control flow arbitration and interconnection module can transmit response commands (Cmd Response) and the task dispatcher, request commands (Cmd Request) and the scheduler, and request commands (Cmd Request) and response commands (Cmd Response) and the operators.

[0176] Optionally, the operator pool may further include a DMA arbiter. The DMA arbiter is deployed on a DMA arbitration unit, and the DMA arbitration unit implements a specified function by running the DMA arbiter.

[0177] For example, data can be transmitted between the DMA arbiter of the operator pool and the DMA arbiter of the central controller. Data can also be transmitted between the DMA arbiter of the operator pool and the operators.

[0178] Optionally, the operator pool may further include a memory operation arbitration module. The memory operation arbitration module is deployed on the memory operation arbitration unit, and the memory operation arbitration unit implements a specified function by running the memory operation arbitration module.

[0179] Exemplarily, other control flows (Other Control Flow) can be transmitted between the memory operation arbitration module and the operator.

[0180] FIG7 is a schematic diagram of a JPEG encoding operator provided in an embodiment of the present application.

[0181] The JPEG encoding operator can compress a bitmap image (RGB pixels) into a coded bitstream according to the JPEG baseline standard.

[0182] As shown in Figure 7, the JPEG encoding operator can include a controller (JPEG controller), an encoding chain, and a JPEG header generator (JFIF Generator). The controller is used to receive a request command (commandRequest), parse the request command to obtain a control signal, obtain uncompressed bitmap pixels from the bitmap stream (bitmapStream) based on the control signal, output a response command (commandResponse), and output the working status (such as a busy signal). The encoding chain is used to stream-encode the uncompressed bitmap pixels to obtain compressed data, combine the compressed data and JPEG header data into a compressed data stream (JPEGStream), and output the compressed data stream.

[0183] For example, a request command may input 512 bits at a time, and a response command may output 512 bits at a time.

[0184] Optionally, the coding chain includes a precoder, an entropy coder (Entropy Code), a coding output generator (Output Generator), etc.

[0185] In an embodiment of the present application, the precoder may include: a pixel conversion module (RGB TO YCbCr), a segmentation module (BlockSplit), a discrete cosine transform module (DCT 2D), a zigzag transform module (ZIG ZAG), a color component (Add Color Component), a quantization module (Quantize), etc.

[0186] For example, RGB TO YCbCr is used to convert RGB pixels into YUV pixels, BlockSplit is used to split YUV pixels into Y pixels, Cb pixels, and Cr pixels and output them in sequence. DCT 2D is used to parallelize discrete sine and cosine transforms. ZIG ZAG is used for zigzag transforms. Quantize is used to perform parallel quantization calculations. Add Color Component is used to identify the pixel type (e.g., Y, Cb, or Cr).

[0187] In an embodiment of the present application, the entropy encoder may include a color component (Add Color Component), a difference code module (DiffCode), a data dispatch module (Data Dispatch), multiple run-length encoding modules (RunLength Encode1-RunLength Encode8), a data integration module (Data Integrate), a Huffman encoding module (Huffman Encode), an alignment module (AlignCode), a byte stuffing module (ByteStuff), etc.

[0188] Exemplarily, DiffCode is used to perform parallel differential encoding. Data Dispatch is used to input received pixels of the same type into the same RunLength Encode. RunLength Encode is used to filter high-frequency information of image data to reduce the amount of image data. The Data Integrate module is used to sort the pixels output by RunLength Encode by type and output them in sequence. Huffman Encode is used to perform Huffman encoding according to the Huffman Encode Table. AlignCode is used to align the bytes of compressed image data. ByteStuff is used to perform byte padding on the compressed image data.

[0189] In the embodiment of the present application, the encoding output generator may include an addition module (EOI Adder). The EOI Adder is used to combine the compressed data output by the entropy encoder and the JPEG header data into a compressed data stream, and output the compressed data stream.

[0190] FIG8 is a schematic diagram of a JPEG decoding operator unit provided in an embodiment of the present application.

[0191] The JPEG decoding operator can decode the compressed bitstream into uncompressed pixels (RGB) according to the JPEG baseline standard.

[0192] As shown in Figure 8, the JPEG decoding operator may include a controller, a decoding chain, and a JPEG header parser (JFIF Parser). The controller is used to receive a request command (commandRequest), parse the request command to obtain basic parameters, output a response command (commandResponse), and output the working status (e.g., a busy signal). The decoding chain is used to stream-decode the compressed data in the compressed data stream to obtain uncompressed bitmap data and output the uncompressed bitmap data via the bitmap stream. The JPEG header parser is used to parse the header information in the compressed data stream (JPEGStream) to obtain decoding information, such as image size, Huffman coding table, quantization table, and other information required for the decoding process.

[0193] Optionally, the decoding chain may include a pre-decoder, an entropy decoder (Entropy Decode), a data processing module, and a decoding output generator (Output Generator).

[0194] In the embodiment of the present application, the pre-decoder may include a width converter. Exemplarily, the width converter is used to split the compressed data.

[0195] In an embodiment of the present application, the entropy encoder may include: a Huffman decoding module (Huffman Decode), a data scheduling module (Data Dispatch), multiple run-length decoding modules (RunLength Decode1-RunLength Decode8), a data integration module (Data Integrate), and a differential decoding module (DiffDecode).

[0196] For example, Huffman Decode is used to perform Huffman decoding according to a Huffman Decode Table to restore the run-length coded compressed information. The run-length decoding module is used to restore the compressed information to pixels in the frequency domain. DiffDecode is used to perform parallel difference decoding.

[0197] In an embodiment of the present application, the data processing module may include: an inverse quantization module (DeQuantize), an inverse zigzag module (DeZigzag), an inverse discrete cosine transform module (IDCT 2D), a YUV upsampling (YUVUpSampling), a block merging module (BlockMerge), and a pixel conversion module (YCbCr TO RGB).

[0198] For example, DeQuantize is used to perform parallel dequantization calculations, and YUVUpSampling is used to perform upsampling to obtain YUV pixels.

[0199] For other descriptions of the JPEG decoding operator, please refer to the description of the JPEG encoding operator, which will not be repeated here, such as the description of the request command and response command.

[0200] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0201] For ease of understanding, the data processing method provided in the embodiment of the present application is introduced below in combination with the above system architecture and accompanying drawings.

[0202] FIG9 is a flow chart of a data processing method provided in an embodiment of the present application. Exemplarily, the data processing method may include the following steps 901 to 904.

[0203] It should be noted that the "step" in the embodiment of the present application can be abbreviated as "S" and will not be repeated later.

[0204] In the embodiment of the present application, the data processing method shown in FIG. 9 can be executed by the data processing device shown in FIG. 3 to FIG. 6 .

[0205] The data processing device may include an FPGA, which may include a scheduling unit, a processing unit, and a storage unit. The scheduling unit is used to perform task scheduling, such as steps 901 and 902; the processing unit is used to perform processing tasks, such as steps 903 and 904; and the storage unit is used to store data to be processed for each processing task, output data of each processing task, and the like.

[0206] Step 901: The scheduling unit of the FPGA determines M processing tasks and data to be processed of the M processing tasks.

[0207] The execution order of the M processing tasks is the target order, M is a positive integer greater than 1, and the to-be-processed data of the M processing tasks are stored in the storage unit.

[0208] In the embodiment of the present application, the scheduling unit can determine M processing tasks and the data to be processed of the M processing tasks by running a central controller (the central controller as shown in FIG6 ).

[0209] It should be noted that the processing task ranked first in the target order is referred to as the first processing task, ..., and the processing task ranked Mth can be referred to as the Mth processing task. Furthermore, the to-be-processed data of the first processing task can be referred to as the first to-be-processed data, ..., and the to-be-processed data of the Mth processing task can be referred to as the Mth to-be-processed data. This will not be further elaborated.

[0210] Exemplarily, the M processing tasks include an image decoding task, an image enlargement task, and an image encoding task, wherein the target order is image decoding task, image enlargement task, and image encoding task, that is, the image decoding task is executed first, followed by the image enlargement task, and finally the image encoding task.

[0211] Hereinafter, through Method 1 and Method 2, an implementation method of determining the first data to be processed is exemplarily described.

[0212] Optionally, mode 1 may include the following S1-S3.

[0213] In an embodiment of the present application, the data processing device may include a DPU, which may be used to receive a data processing request sent by an electronic device and convert the data processing request into a task request that can be recognized by the FPGA.

[0214] S1: The DPU determines a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks.

[0215] In an embodiment of the present application, the DPU can receive a data processing request and determine a target task request based on the data processing request by running a software interface module. For example, when an electronic device needs to perform a target operation on an original image (i.e., original data), such as when a user instructs the electronic device to perform a target operation on an original image or when the electronic device needs to perform a target operation on an original image when performing a certain task, the electronic device can send a data processing request to a data processing device, and the data processing request is used to request that the target operation be performed on the original image. After receiving the data processing request, the DPU of the data processing device converts the data processing request into a task request that can be recognized by the FPGA, such as a target task request.

[0216] In Example 1, a data processing request may include raw data. After receiving the raw data, the DPU stores the raw data in a local storage medium (in the memory of the DPU as shown in FIG6 ). In Example 2, a data processing request is used to indicate the identifier, storage address, etc. of the raw data. After receiving the data processing request, the DPU may obtain the raw data based on the identifier, storage address, etc. of the raw data, such as obtaining the raw data from a cloud storage device and storing the obtained raw data in a local storage medium. Exemplary target operations may include sharpening, amplification, color adjustment, reduction, watermarking, cutout, synthesis, light and dark modification, color and chroma modification, adding special effects, editing, repairing, etc.

[0217] It should be noted that the embodiment of the present application does not limit the type of target operation, and the above is only an exemplary description. Below, the embodiment of the present application is exemplified by taking the target operation as amplification as an example.

[0218] S2: The DPU sends a target task request to the scheduling unit; the target task request is used to indicate the original data and the target operation.

[0219] In an embodiment of the present application, the DPU can send a target task request to the scheduling unit of the FPGA by running the FPGA infrastructure to request the FPGA to perform a target operation on the original data.

[0220] Exemplarily, the target task request may indicate the original data by indicating the storage address of the original data in the memory of the DPU.

[0221] S3: The scheduling unit determines M processing tasks and first data to be processed based on the received target task request, wherein the first data to be processed is original data, that is, the data to be processed of the processing task ranked first among the M processing tasks.

[0222] In the embodiment of the present application, the scheduling unit can receive the target task request sent by the DPU by running the FPGA infrastructure. On this basis, the scheduling unit can determine the M processing tasks and the first data to be processed by running the central controller.

[0223] Optionally, the scheduling unit may obtain original data from a local storage medium of the DPU based on the target task request.

[0224] Exemplarily, the target task request indicates the storage address of the original data, and the scheduling unit obtains the storage address of the original data by parsing the task request. Thereafter, the scheduling unit obtains the original data from the DPU through the storage address of the original data.

[0225] Optionally, the scheduling unit may obtain multiple processing tasks based on the target task request.

[0226] For example, the target task request is used to request to perform a target operation on the original image, the target operation is an enlargement operation, and the target image is a JPEG image. Based on this, the scheduling unit parses the task request and determines that the multiple processing tasks include a JPEG image decoding task, an image enlargement task, and a JPEG image encoding task.

[0227] In this implementation, a processor receives a data processing request sent by an electronic device and converts the data processing request into a task request recognizable by the FPGA. This helps reduce the software complexity of the FPGA and thus helps reduce the software development cost of the FPGA.

[0228] Optionally, the FPGA stores a task form, which can be used to indicate the execution order of multiple processing tasks, the task status of each processing task in the multiple processing tasks, the data to be processed of each processing task, etc. The task status includes an executed state or an unexecuted state.

[0229] In an embodiment of the present application, the scheduling unit may generate task description information based on the target task request. The task description information may be used to indicate the execution order of the M processing tasks, the task status of each of the M processing tasks, the storage address of the original data (i.e., the data to be processed of the first-ranked processing task), etc. On this basis, the scheduling unit may store the task description information in a task form, thereby indicating relevant information of the M processing tasks, such as the execution order, task status, storage address of the data to be processed, etc., through the task identifier.

[0230] It should be noted that the task description information generated based on the target task request can be called initial task description information. The task status of each processing task indicated by the initial task description information is the initial task status, and the initial task status is the unexecuted status.

[0231] For example, the task list may be stored in a display look-up table (LUT), DRMA, etc. of the FPGA.

[0232] In this embodiment, by configuring a task form to indicate the task status of each processing task, the scheduling unit does not need to be aware of the processing task itself, such as the specific processing task being executed, the task status of the processing task, the lifecycle of the processing task, the execution status of the processing task, etc. This helps to expand the applicability of the data processing method, such as enabling it to process image data, audio data, etc., thereby helping to improve the business migration capabilities of the data processing equipment. Furthermore, since the scheduling unit does not need to be aware of the processing task itself, it helps to shorten the time that the processing task occupies the scheduling unit, thereby helping to improve task scheduling efficiency.

[0233] Optionally, mode 2 may include S4-S6.

[0234] In an embodiment of the present application, the FPGA may include a receiving unit, which may be used to receive a data processing request sent by an electronic device and convert the data processing request into a task request that can be recognized by the FPGA.

[0235] S4: The receiving unit determines a target task request based on the received data processing request.

[0236] S5: The receiving unit sends a target task request to the scheduling unit.

[0237] S6: The receiving unit determines M processing tasks and first data to be processed based on the received target task request.

[0238] It should be noted that for the relevant instructions of S4-S6, please refer to the instructions of S1-S3, which will not be repeated here.

[0239] In this implementation, a data processing request sent by an electronic device is received by a receiving unit of the FPGA, and the data processing request is converted into a task request that can be recognized by the processing unit, which helps to reduce the hardware cost of the data processing device.

[0240] Optionally, the data processing method may further include: the scheduling unit storing the original data in a storage unit of the FPGA.

[0241] In the embodiment of the present application, the scheduling unit may request storage space from the memory management unit of the FPGA, such as the memory management unit assigning the first storage space to the scheduling unit. Based on this, the scheduling unit stores the original data in the first storage space.

[0242] In the following, taking the second data to be processed as an example, an implementation method for determining the second data to be processed to the Mth data to be processed is exemplarily described through S7-S8.

[0243] S7: The scheduling unit receives the first response information returned by the processing unit. The first response information is used to indicate a first storage address, and the first storage address stores the output data of the first processing task. The output data of the first processing task is the data to be processed by the second processing task.

[0244] In an embodiment of the present application, after the processing unit executes the first processing task, it returns a first response message to the scheduling unit. The first response message is used to indicate the output data of the first processing task, such as: used to indicate the storage address of the output data of the first processing task in the storage unit of the FPGA.

[0245] Exemplarily, the first processing task is a JPEG image decoding task. After the decoding operator completes the JPEG image decoding task by running the decoding operator, the decoding operator returns a first response message to the scheduling unit. The scheduling unit receives the first response message by running the central controller.

[0246] S8: The scheduling unit determines the data stored in the first storage address as the second data to be processed.

[0247] In the embodiment of the present application, the storage address indicated by the first response information stores the output data of the first processing task. Based on this, the scheduling unit uses the output data of the first processing task as the second data to be processed.

[0248] Exemplarily, the scheduling unit determines the data stored at the first storage address as the second data to be processed by running the central controller.

[0249] In the above embodiment, the data to be processed of the next processing task is determined by the response information returned by the processing unit, which helps to ensure the accuracy of the determined data to be processed.

[0250] In the embodiment of the present application, the scheduling unit can update the task description information based on receiving the first response information. The updated task description information can be used to indicate that the task status of the first processing task is an executed state, the storage address of the output data of the first processing task, etc. This helps to improve the accuracy of the task description information, thereby enabling the task description information to record the latest status of different processing tasks and the pending data of different processing tasks.

[0251] Step 902: The scheduling unit of the FPGA sends M request commands to the processing unit of the FPGA.

[0252] Each of the M request commands is used to request execution of each processing task according to the to-be-processed data of each processing task.

[0253] That is, a request command is used to request execution of a processing task based on the to-be-processed data of a processing task, wherein a processing task can be any one of the M processing tasks.

[0254] In the embodiment of the present application, the scheduling unit can send M request commands to the processing unit by running the central controller.

[0255] Optionally, the scheduling unit includes multiple sub-scheduling units, wherein each of the multiple sub-scheduling units can be used to send a request command to the processing unit. Exemplarily, step 902 can be that a scheduling unit sends a request command to the processing unit.

[0256] In this embodiment, by configuring a scheduling unit to include multiple sub-scheduling units, it is possible to simultaneously schedule processing tasks for different data processing requests, thereby improving the scheduling efficiency of multiple data processing requests. Furthermore, by integrating a task form to maintain the task status, metadata, and other aspects of the processing task, each sub-scheduling unit no longer needs to maintain the lifecycle and task status of the processing task. Instead, it only needs to assign an operator unit to the processing task and send a request command, without having to wait for the operator unit to complete the processing task. This facilitates flexible scheduling of processing tasks, prevents processing tasks from occupying sub-scheduling units for extended periods, and further improves scheduling efficiency.

[0257] The following is an illustrative introduction to the implementation process of sending M request commands through method A to method B.

[0258] Optionally, the processing unit includes multiple operator units, wherein one operator unit can be used to execute one processing task.

[0259] It should be noted that the term "multiple" in the embodiments of the present application refers to two or more, which will not be further elaborated later.

[0260] In one example, different operator units among the multiple operator units are used to perform different processing tasks, which helps to increase the types of processing tasks that can be performed by the processing units.

[0261] In another example, some of the multiple operator units are used to perform the same processing task. In this way, the processing unit can simultaneously process multiple identical processing tasks, such as multiple identical processing tasks for multiple different data processing requests, thereby achieving parallel processing of multiple different data processing requests and improving data processing efficiency and data processing capabilities.

[0262] In another example, some of the multiple operator units can be used to perform the same processing task, while other operator units can be used to perform different processing tasks. This not only increases the types of processing tasks performed by the processing units, but also enables parallel processing of different data processing requests, thereby helping to improve the data processing efficiency and data processing capabilities of the data processing device.

[0263] Exemplarily, the M processing tasks include a target processing task, which may be any one of the M processing tasks. The multiple operator units may include a target operator unit, which is an operator unit for executing the target processing task.

[0264] In the following, step 902 is exemplarily described by taking the target processing task and the target operator unit as an example.

[0265] Optionally, method A may include the following S9-S10.

[0266] In the embodiment of the present application, the FPGA stores an operator table, which is used to indicate the processing tasks performed by each operator unit in the plurality of operator units, wherein the operator table is used to indicate that a target operator unit is used to perform a target processing task.

[0267] For example, the operator table may be stored in a display look-up table (LUT), DRMA, etc. of an FPGA.

[0268] The operator form and task form can be stored in different LUTs, which helps improve the accuracy of obtaining the operator form / task form and updating the operator form / task form.

[0269] S9: The scheduling unit determines a target operator unit for the target processing task from the multiple operator units indicated in the operator table.

[0270] In the embodiment of the present application, the operator table is used to indicate the processing tasks that each operator unit can perform. Based on this, when the scheduling unit needs to determine the operator unit that performs the target processing task, it can determine the operator unit that can perform the target processing task based on the content indicated in the operator table, that is, the target operator unit. For example, when the target processing task is an image decoding task, the scheduling unit can determine the decoding operator unit as the target operator unit.

[0271] Exemplarily, the target sub-scheduling unit among the multiple sub-scheduling units determines the target operator to execute the target processing task from the multiple operators indicated in the operator list by running the scheduler, wherein the target operator is run by the target operator unit, thereby determining that the target operator unit is used to execute the target processing task.

[0272] S10: The scheduling unit sends a target request command to the target operator unit, where the target request command is used to request execution of a target processing task on the target data to be processed.

[0273] Exemplarily, the operator table includes the identifier of each operator unit, and different operator units have different identifiers. For example, the identifier of the target operator unit is a. Based on this, the scheduling unit can send a target request command to the operator unit identified as a (i.e., the target operator unit) among the multiple operator units to request the target operator unit to perform the target processing task. For example, the target sub-scheduling unit can send the target request command to the target operator unit running the target operator by running the scheduler.

[0274] In an embodiment of the present application, the target request command is further used to indicate target data to be processed, where the target data to be processed is the data to be processed for the target processing task. For example, the target request command may include the storage address of the target data to be processed, thereby indicating the target data to be processed, where the storage address is the storage address of the target data to be processed in the storage unit.

[0275] In an embodiment of the present application, when the target processing task is the processing task ranked first in the target sequence, such as the image decoding task, the target data to be processed is the original data. In addition, the output data of the Nth processing task is the data to be processed of the N+1th processing task in the target sequence; N is a positive integer less than M. In other words, when the target processing task is a processing task that is not ranked first in the target sequence, the target data to be processed is the output data of the processing task ranked before the target processing task in the target sequence. For example, when the target processing task is an image magnification task, the target data to be processed is the output data of the image decoding task; when the target processing task is an image encoding task, the target data to be processed is the output data of the image magnification task.

[0276] In this implementation, an operator form is configured for the FPGA to indicate the processing tasks that each operator unit can perform, so that the scheduling logic of the scheduling unit determines the operator unit used to perform each processing task based on the operator form. In this way, when expanding the operator unit, it is only necessary to add the newly expanded operator unit to the operator form to schedule the newly added operator unit to perform the newly added processing task, which helps to simplify the difficulty of expanding the operator unit and further helps to improve the scalability of the FPGA.

[0277] Optionally, the operator table is further used to indicate the working state of each operator unit, where the working state includes an idle state or a busy state. The operator table indicates that the working state of the target operator unit is an idle state.

[0278] The idle state means that the operator unit has not been assigned a processing task, that is, the operator unit is not currently executing a processing task, and the busy state means that the operator unit has been assigned a processing task, that is, the operator unit is currently executing a processing task.

[0279] In this embodiment, the working status of each operator unit can be indicated by setting an operator form. In this way, when determining the operator unit to execute each processing task, the operator units in the idle state can be screened for executing the processing task, which helps to improve the execution efficiency of each processing task, and further helps to improve the working efficiency of the data processing equipment.

[0280] Optionally, after the scheduling unit determines the target operator unit for the target processing task in the multiple processing tasks from the multiple operator units indicated by the operator table, the data processing method may include: the scheduling unit updates the working status of the target operator unit indicated by the operator table to a busy state.

[0281] In an embodiment of the present application, the target sub-scheduling unit updates the working state of the target operator unit to a busy state by running the scheduler.

[0282] In this embodiment, after the processing task is assigned to the target operator unit, the working status of the target operator unit is updated to a busy status. This not only helps to ensure the accuracy of the working status indicated by the operator form, but also helps to avoid the target operator unit being assigned to other processing tasks during the execution of the target processing task, thereby affecting the execution efficiency of other processing tasks.

[0283] Optionally, method B may include the following S11-S13.

[0284] S11: The scheduling unit sends a query command to each of the multiple operator units, where the query command is used to query the processing tasks performed by each operator unit.

[0285] Exemplarily, the scheduling unit sends a query command to each operator unit to request to query the processing tasks executed by each operator unit, thereby determining the operator unit executing each processing task. For example, the target sub-scheduling unit sends the query command to each operator unit running each operator by running the scheduler.

[0286] S12: The scheduling unit determines a target operator unit for the target processing task based on the received information of the multiple operators.

[0287] Among them, an operator information is used to indicate a processing task performed by an operator unit.

[0288] In the embodiment of the present application, each of the multiple operator units responds to the query command by running the operator and returns its own operator information to the scheduling unit. For example, the target operator unit returns the target operator information to the scheduling unit by running the target operator. The target operator information is used to instruct the target operator unit to execute the target processing unit.

[0289] Exemplarily, the target operator unit is used to perform an image decoding task. In response to a query command, the target operator unit returns target operator information to the scheduling unit, where the target operator information instructs the target operator unit to perform the image decoding task.

[0290] Exemplarily, the target processing task is an image decoding task. Based on this, the scheduling unit determines that the target operator unit performs the target processing task based on the target operator information among the received multiple operator information.

[0291] S13: The scheduling unit sends a target request command to the target operator unit, requesting to execute a target processing task on the target data to be processed.

[0292] It should be noted that for the relevant instructions of S13, please refer to the relevant instructions of S10, which will not be repeated here.

[0293] In this implementation, query commands are sent to multiple operator units to determine the processing tasks performed by each operator unit. In this way, when it is necessary to expand the operator unit, it is only necessary to send the query command to the newly expanded operator unit as the object of the scheduling unit. This helps to simplify the difficulty of expanding the operator unit and thus helps to improve the reliability of the FPGA.

[0294] In embodiments of the present application, a query command can also be used to query the operating status of each operator unit, and an operator information can also be used to indicate the operating status of an operator unit. For example, target operator information can be used to determine the operating status of a target operator unit. Based on this, if the target operator information indicates that the target operator unit is idle, the scheduling unit determines that the target operator unit will execute the target processing task.

[0295] It should be noted that for other relevant instructions of method B, please refer to the instructions of method A above, which will not be repeated here.

[0296] It should be noted that the embodiment of the present application does not limit how to send the request command, and the above is only an exemplary description. Below, the embodiment of the present application is exemplarily introduced by taking the above method A as an example.

[0297] Step 903: The processing unit of the FPGA executes the M processing tasks in a target order based on the received M request commands, and obtains output data of the M processing tasks.

[0298] In an embodiment of the present application, a processing unit receives a request command by running an operator. For example, a target operator unit of a processing unit receives a target request command by running a target operator. Exemplarily, after receiving the target request command, the target operator unit retrieves the target data to be processed from a storage unit of an FPGA based on the storage address of the target data to be processed indicated by the target request command. The target operator unit then executes a target processing task on the target data to be processed, obtaining output data of the target processing task.

[0299] For example, the target processing task is an image decoding task. After the target operator unit obtains the target data to be processed (i.e., the original data), it performs the target decoding task on the original data to obtain the output data of the image decoding task. The output data of the image decoding task is the data to be processed for the image magnification task.

[0300] The following is an exemplary introduction to the process of performing the image decoding task with reference to FIG8 .

[0301] For example, the data to be processed in an image decoding task is raw data. Upon receiving an image decoding request (i.e., commandRequest), the decoding operator unit retrieves the raw data from the FPGA's storage unit, for example, 512 bits of compressed data per clock cycle. The following describes the decoding process using 512 bits of compressed data as an example.

[0302] The Width Converter in the decoding operator unit segments the 512-bit compressed data and transmits the resulting data 1 to the JFIF Parser. Data 1 consists of 64 pixels, each containing 8 bits. The JFIF Parser parses Data 1, obtaining the original data header information, decoding mode, quantization tables, Huffman coding tables, and other information. It then transmits Data 1 to the entropy decoder.

[0303] The Huffman Decoder in the entropy decoder decodes Data 1 according to the Huffman decoding table to restore the run-length encoded compressed information and transmits the resulting Data 2 to Data Dispatch, where each pixel of Data 2 is 19 bits. Data Dispatch inputs the 64 pixels of Data 2 into multiple RunLengthDecodes for parallel decoding to improve decoding efficiency. For example, multiple RunLengthDecodes decode Data 2 in parallel, restoring the compressed information to pixels in the frequency domain. The resulting Data 3 is transmitted to Data Integrate, where each pixel of Data 3 is 14 bits. Data Integrate sorts the received Data 3 and transmits the resulting Data 4 to DiffDecode.

[0304] DiffDecode performs parallel dequantization on the received data 4 according to the quantization table and transmits the resulting data 5 to DeZigzag. DeZigzag processes the received data 5 and transmits the resulting data 6 to IDCT 2D. IDCT 2D processes the received data 6 and transmits the resulting data 7 to YUVUpSampling. YUVUpSampling performs parallel upsampling on the received data 7 and transmits the resulting data 8 to YCbCr TO RGB. YCbCr TO RGB performs parallel conversion on the received data 8 and transmits the resulting RGB pixel data 9 to the Output Generator, where data 9 is 512 bits. The Output Generator outputs the received data 9 through bitMapStream.

[0305] The following is an exemplary introduction to the process of performing the image coding task with reference to FIG7 .

[0306] Exemplarily, the data to be processed for the image encoding task is the output data of the image magnification task, wherein the data to be processed for the image encoding task is bitmap pixel data. The data to be processed includes multiple minimum coded units (MCUs). Each MCU includes 8 rows of images, each row of images includes 8 pixels, and each pixel is 24 bits. The following describes the execution process of the image encoding task using the processing of a single MCU as an example.

[0307] After receiving the image encoding request, the encoding operator unit retrieves one row of image from the storage unit during each clock cycle. RGB TO YCbCr can convert one row of bitmap pixels into one row of YCbCr pixels, where one bitmap pixel can be converted into one Y pixel, one Cb pixel, and one Cr pixel. Based on this, one row of bitmap pixels can be converted into one row of Y pixels, one row of Cb pixels, and one row of Cr pixels. One row of Y pixels, Cb pixels, and Cr pixels includes eight Y pixels, Cb pixels, and Cr pixels. After eight clock cycles, RGB TO YCbCr can output data 10 to BlockSplit. Data 10 includes eight rows of YCbCr pixels, each row of YCbCr pixels includes one row of Y pixels, one row of Cb pixels, and one row of Cr pixels. Each Y pixel, Cb pixel, and Cr pixel in data 10 is 8 bits.

[0308] BlockSplit splits data 10 into data 11, which consists of 8 rows of Y pixels, 8 rows of Cb pixels, and 8 rows of Cr pixels. Each Y pixel, Cb pixel, and Cr pixel in data 11 is 8 bits. BlockSplit sequentially outputs the 8 rows of Y pixels, 8 rows of Cb pixels, and 8 rows of Cr pixels in data 11 to DCT 2D. DCT 2D processes data 11 in parallel, for example, receiving 8Q pixels per clock cycle, converting these 8Q pixels from the spatial domain to the frequency domain, and transmitting the resulting data 12 to ZIG ZAG. Each Y pixel, Cb pixel, and Cr pixel in data 12 is 12 bits.

[0309] ZIG ZAG processes the received data 12, such as converting the two-dimensional image matrix into a one-dimensional vector using the Zig-Zag sequence, and transmits the resulting data 13 to the Add Color Component, where each Y pixel, Cb pixel, or Cr pixel in data 13 is 12 bits. The Add Color Component adds type information to each row of pixels in data 13 to generate data 14. The type information indicates the pixel type, which can be Y, Cb, or Cr. Each row of Y pixels, Cb pixels, or Cr pixels in data 14 is 8*12+2 bits, where 8 represents 8 pixels, 12 represents 12 bits per pixel, and 2 bits are used to record the type information. The Add Color Component transmits data 14 to Quantize, which performs parallel quantization calculations on the received data 14, such as receiving 8Q pixels per clock cycle, compressing the high-frequency components of the 8Q pixels, and transmitting the resulting data 15 to the entropy encoder, where each row of Y pixels, Cb pixels, or Cr pixels in data 15 is 8*12 bits.

[0310] The Add Color Component of the entropy encoder adds type information to each row of pixels in the received data 15 and transmits the resulting data 16 to the DiffCode, where each row of Y pixels / Cb pixels / Cr pixels in the data 16 is 8*12+2 bits. The DiffCode performs parallel differential encoding on the received data 16, such as differential encoding 8Q pixels per clock cycle, and transmits the resulting data 17 to the Data Dispatch, where the data 17 includes 8 rows of Y pixels (i.e., 64 Y pixels), 8 rows of Cb pixels (i.e., 64 Cb pixels), and 8 rows of Cr pixels (i.e., 64 Cr pixels). Each row of Y pixels / Cb pixels / Cr pixels includes 8 Y pixels / Cb pixels / Cr pixels, and each Y pixel / Cb pixel / Cr pixel is 14 bits.

[0311] Data Dispatch transmits data 17 to multiple RunLength Encoders for parallel compression to improve compression efficiency, and transmits the resulting data 18 to Data Integrate. Data Dispatch can transmit pixels of the same type to the same RunLength Encoder, for example, transmitting 8 rows of Y pixels to RunLength Encode1, 8 rows of Cb pixels to RunLength Encode2, and 8 rows of Cr pixels to RunLength Encode3. Data 18 includes A Y pixels output by RunLength Encode1, B Cb pixels output by RunLength Encode2, and C Cr pixels output by RunLength Encode3. The ratio A / B / C is less than 64, and each Y pixel, Cb pixel, and Cr pixel is 19 bits.

[0312] Data Integrate sorts the received data 18 to obtain data 19, and outputs the Y pixels in row A, the Cb pixels in row B, and the Cr pixels in row C of data 19 to the Huffman Encoder in sequence, with each Y pixel, Cb pixel, and Cr pixel being 19 bits. Huffman Encode encodes the received data 19 according to the Huffman coding table and transmits the resulting data 20 to AlignCode, with each Y pixel, Cb pixel, and Cr pixel in data 20 being 33 bits. AlignCode processes the received data 20 and transmits the resulting data 21 to ByteStuff. ByteStuff processes the received data 21 and transmits the resulting data 22 to the Output Generator. The Output Generator combines the received data 22 with the JPEG header data and outputs the resulting data 23 through the JPEGStream.

[0313] Optionally, the data processing method may further include: after the target operator unit completes the target processing task, returning target response information to the scheduling unit. The target response information is used to indicate the storage address of the output data of the target processing task, the execution result of the target processing task, etc. The execution result may include execution success or execution failure.

[0314] In this embodiment, after the operator unit completes the processing task, it returns response information to the scheduling unit. In this way, the scheduling unit can not only determine the data to be processed for the next processing task based on the response information, but also determine the final data obtained by executing M processing tasks.

[0315] In an embodiment of the present application, after the target operator unit completes the target processing task, it returns target response information to the scheduling unit, and the scheduling unit can receive the target response information returned by the target operator unit. For example, the target operator unit returns the target response information to the scheduling unit by running the target operator, and the scheduling unit receives the target response information by running the central controller.

[0316] In an embodiment of the present application, the scheduling unit can update the task description information based on the received target response information. The updated task description information can be used to indicate that the task status of the target processing task is an executed state, the storage address of the output data of the target processing task, etc. On this basis, the scheduling unit can write the updated task description information into the task form, thereby updating the task form. The updated task form can be used to indicate the content indicated by the updated task description information.

[0317] Exemplarily, when the target processing task is an image decoding task, the updated task description information can be used to indicate that the task status of the image decoding task is the executed state and the storage address of the output data of the image decoding task (i.e., the data to be processed of the image magnification task).

[0318] Optionally, the data processing method may include: the scheduling unit updating the operating state of the target operator unit indicated by the operator table to an idle state in response to the target response information. Exemplarily, the scheduling unit updates the operating state of the target operator unit to an idle state by running the central controller. This not only helps ensure the accuracy of the operating state indicated by the operator table, but also ensures that the scheduling unit can allocate operator units to other processing tasks, thereby helping to improve the execution efficiency of other processing tasks.

[0319] Step 904: The processing unit stores the output data of executing the M processing tasks in the storage unit.

[0320] For example, the target operator unit may request storage space from the storage unit, such as allocating the third storage space to the target operator unit. Based on this, the target operator unit stores the output data of the target processing task in the third storage space.

[0321] Optionally, the data processing method may further include: when the processing unit completes executing M processing tasks, the scheduling unit outputs target data to the DPU, where the target data is output data of the processing task ranked Mth in the target sequence.

[0322] For example, the scheduling unit may determine the storage address of the target data based on the response information returned by the processing unit and obtain the target data. The scheduling unit may then output the target data to the DPU, which then returns the target data to the electronic device, thereby completing the content requested by the data processing request sent by the electronic device.

[0323] In the embodiment of the present application, there is no restriction on the execution order of step 901 to step 904. In the following, the execution process of step 901 to step 904 is exemplarily introduced by taking the target operation as an image zoom operation as an example.

[0324] After receiving an image magnification operation request from a user via an electronic device, the DPU converts the image magnification operation request into an image magnification task request recognizable by the FPGA and sends the image magnification task request to the FPGA's scheduling unit. In response to receiving the image magnification task request, the scheduling unit stores the raw data obtained from the DPU in a first storage space of the FPGA's storage unit and generates first task description information for the image magnification task request. The first task description information is used to indicate multiple processing tasks and the storage addresses of the raw data in the storage unit. The multiple processing tasks are executed in the following order: image decoding task, image magnification task, and image encoding task.

[0325] On this basis, the scheduling unit can store the first task description information in the task form, and determine the first decoding operator unit from the multiple operator units indicated by the operator form to perform the image decoding task. After that, the scheduling unit sends a decoding request to the first decoding operator unit. The decoding request indicates the storage address of the original data, which is used to request the first decoding operator unit to perform the decoding task on the original data.

[0326] Exemplarily, in response to the received image magnification task request, the task parsing unit of the scheduling unit applies to the memory management unit for storage space for storing the original data. The memory management unit allocates the first storage space of the storage unit to the task parsing unit. The task parsing unit stores the original data obtained from the DPU in the first storage space of the storage unit and generates first task description information. Afterwards, the task parsing unit sends the first task description information to the task distribution unit. In response to the received task description information, the task distribution unit determines the task scheduling of the first sub-scheduling unit from at least one sub-scheduling unit for performing the image decoding task. After receiving the first task description information forwarded by the task distribution unit, the first sub-scheduling unit stores the first task description information in the task form and determines the first decoding operator unit for performing the image decoding task. Afterwards, the first sub-scheduling unit sends a decoding request to the first decoding operator unit through the control flow arbitration and interconnection unit.

[0327] Exemplarily, the first sub-scheduling unit is any sub-scheduling unit in an idle state. The operator table may also be used to indicate that the working state of the first decoding operator unit is an idle state. After the first sub-scheduling unit determines that the first decoding operator unit is used to perform the image decoding task, it may update the working state of the first decoding operator unit indicated in the operator table to a busy state.

[0328] In response to the received decoding request, the first decoding operator unit obtains the data to be processed (i.e., the original data) from the first storage space of the storage unit and performs a decoding task on the data to be processed. After the first decoding operator unit completes the decoding task on the original data, it obtains the first output data. The first decoding operator unit applies to the memory operation arbitration unit for storage space for storing the first output data. The memory operation arbitration unit allocates the second storage space of the storage unit to the first decoding operator unit. The first decoding operator unit stores the first output data in the second storage space of the storage unit and returns a first response message to the scheduling unit through the control flow arbitration and interconnection unit. The first response message is used to indicate the completion of the image decoding task, the storage address of the first output data of the image decoding task, etc.

[0329] The scheduling unit updates the first task description information based on the received first response information to obtain second task description information, wherein the second task description information is used to indicate that the task status of the image decoding task is the executed state, the storage address of the first output data of the image decoding task, etc.

[0330] Exemplarily, in response to the received first response information, the task dispatching unit of the scheduling unit determines, from the at least one task recycling unit, a first task recycling unit for performing task recycling work for the image decoding task, and forwards the first response information to the first task recycling unit. In response to the received first response information, the first task recycling unit obtains the first task description information from the task form and updates the first task description information, thereby obtaining the second task description information.

[0331] Exemplarily, the first task recycling unit is any task recycling unit in an idle state. The first task recycling unit may also update the working state of the first decoding operator unit indicated by the operator table to an idle state in response to the first response information to release the first decoding operator unit.

[0332] After obtaining the second task description information, the scheduling unit stores the second task description information in a task form and determines that the first image enlargement operator unit performs the image enlargement task. Thereafter, the scheduling unit sends an image enlargement request to the first image enlargement operator unit to request the first image enlargement operator unit to perform the image enlargement task on the first output data.

[0333] It should be noted that for other related instructions on the scheduling unit determining the first image enlargement operator unit and sending the image enlargement request, please refer to the instructions on the scheduling unit determining the first decoding operator unit and sending the image decoding request, which will not be repeated here.

[0334] Exemplarily, the first task recycling unit may transmit the obtained second task description information to the task dispatching unit, which then determines a sub-scheduling unit for the next processing task (i.e., the image magnification task). Upon receiving the second task description information, the task dispatching unit determines a second sub-scheduling unit from at least the first sub-scheduling unit to perform task scheduling for the image magnification task. The second sub-scheduling unit updates the task form based on the second task description information so that the updated task form can execute the content indicated by the second task description information, determines the first image magnification operator unit to execute the image magnification task, and sends an image magnification request to the first image magnification operator unit.

[0335] It should be noted that for other related descriptions of the second sub-scheduling unit, the first image enlargement operator unit, and the image enlargement request, reference can be made to the descriptions of the first sub-scheduling unit, the first decoding operator unit, and the image decoding request, which will not be repeated here.

[0336] After completing the image enlargement task, the first image enlargement operator unit returns a second response message to the scheduling unit and stores the second output data of the image enlargement task in the third storage space of the storage unit. In response to the received second response message, the scheduling unit updates the second task description information to obtain third task description information. The third task description information indicates that the task status of the image enlargement task is "executed" and the storage address of the second output data of the image enlargement task. The scheduling unit then determines that the first encoding operator unit is used to perform the image encoding task and sends an encoding request to the first encoding operator unit.

[0337] It should be noted that for other relevant instructions on the scheduling unit determining the first decoding operator unit and sending the encoding request, please refer to the instructions on the scheduling unit determining the first decoding operator unit and sending the decoding request, which will not be repeated here.

[0338] Exemplarily, the task distribution unit determines that the second sub-scheduling unit is used to perform task scheduling of the image magnification task. The task distribution unit determines that the second task recycling unit is used to perform task recycling of the image magnification task.

[0339] It should be noted that for other relevant descriptions of the second sub-scheduling unit and the second task recovery unit, reference can be made to the descriptions of the first sub-scheduling unit and the first task recovery unit, and they will not be repeated here.

[0340] After completing the image encoding task, the first encoding operator unit returns a third response message to the scheduling unit and stores the third output data of the image encoding task in the fourth storage space of the storage unit. In response to receiving the third response message, the scheduling unit returns task response information to the DPU, where the task response information indicates, among other things, the storage address of the third output data. In response to receiving the task response information, the DPU retrieves the third output data from the storage unit of the FPGA and transmits the third output data to the electronic device, thereby returning the target data (i.e., the third output data) obtained from executing the image magnification operation request to the user.

[0341] Exemplarily, in response to the received third response information, the task distribution unit determines that the third task recovery unit is used to perform the task recovery work of the image coding task, and forwards the third response information to the third task recovery unit. The third task recovery unit updates the third task description information based on the received third response information to obtain the fourth task description information, and the fourth task description information is used to indicate that the task status of the image coding task is the executed state, the storage address of the third output data of the image coding task, etc. Afterwards, the third task recovery unit sends the fourth task description information to the task distribution unit, and the task distribution unit determines that the image magnification task request is completed based on the received fourth task description information. Afterwards, the task distribution unit sends the fourth description information to the task output unit, and the task output unit sends the received task response information to the DPU, and sends the third output data stored in the fourth storage space of the storage unit to the DPU.

[0342] At this point, the data processing device completes the image magnification operation request on the original data.

[0343] For example, in conjunction with Figure 6, the operations performed by the task parsing unit are implemented by running the task parser. The operations performed by the task output module are implemented by running the task output module. The operations performed by the task dispatching unit are implemented by running the task dispatcher. The operations performed by the sub-scheduling units (such as the first sub-scheduling unit, the second sub-scheduling unit, the third sub-scheduling unit, etc.) are implemented by running the scheduler. The operations performed by the task recycling units (such as the first task recycling unit, the second task recycling unit, the third task recycling unit, etc.) are implemented by running the recycler. The operations performed by the memory management unit are implemented by running the memory manager. The operations performed by the control flow arbitration and interconnection unit are implemented by running the control flow arbitration and interconnection module. The operations performed by the memory operation arbitration unit are implemented by memory operation arbitration. The operations performed by the operator units (such as the decoding operator unit, the image magnification operator unit, the encoding operator unit, etc.) are implemented by running the operators (such as the decoding operator, the image magnification operator, the encoding operator, etc.).

[0344] In the above embodiment, data processing services are provided by a data processing device. Since the data processing device does not require hardware resources deployed based on the X86 server architecture used in related technologies, the hardware resources used by the data processing device are reduced. This not only reduces data processing costs but also improves the hardware resource utilization of the data processing device. Furthermore, the FPGA independently completes M processing tasks that need to be executed in a specified order, and the output data of each processing task is stored in the FPGA's memory unit. As the processing unit executes each processing task, it can directly retrieve the data to be processed from the accelerator's memory unit through hardware, thereby avoiding end-to-end data transmission delays during the execution of the M processing tasks and thereby improving data processing efficiency.

[0345] In addition, since the data processing device adopts a numerical control separation architecture, that is, when executing M processing tasks, task scheduling is performed through the FPGA's scheduling unit, such as sending a request command to the processing unit to instruct the processing unit to perform a processing task on the processing data, and the processing unit of the FPGA performs the processing task on the processing data, thereby achieving the decoupling / separation of the control flow (such as task scheduling) and the data flow (such as the processing data generated by the processing task). This helps to simplify the scheduling logic of the scheduling unit and improve the scalability of the data processing device. For example, when adding an executable processing task to the processing unit, there is no need to change the scheduling logic (i.e., software code) of the scheduling unit. In addition, due to the adoption of the numerical control separation architecture, the data processing device can also perform task scheduling and processing tasks in parallel, thereby improving the parallel processing capability of the data processing device, and thus helping to improve data processing efficiency.

[0346] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, the data processing device includes a hardware structure and / or software module corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0347] In the embodiment of the present application, the data processing device can be divided into functional modules according to the above method. For example, the data processing device can include functional modules corresponding to the functional divisions, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0348] For example, FIG10 shows a possible structural diagram of the data processing device involved in the above embodiment (denoted as data processing device 1000). The actions performed by the data processing device are implemented by a data processing device or by the data processing device executing corresponding software. The data processing device may include an FPGA, and the FPGA may include a scheduling unit, a processing unit, and a storage unit. The data processing device 1000 may include a scheduling module 1001 and a processing module 1002. The scheduling module 1001 is used to determine M processing tasks and the data to be processed of the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit. For example, as shown in S901 in FIG9. The scheduling module 1001 is also used to send M request commands to the processing unit; wherein each request command is used to request the execution of each processing task based on the data to be processed of each processing task. For example, as shown in S902 in FIG9. Processing module 1002 is configured to execute M processing tasks in a target order based on the received M request commands, obtaining output data of the M processing tasks; wherein the output data of the Nth processing task is the to-be-processed data of the N+1th processing task in the target order; N is a positive integer less than M. For example, see S903 in FIG9 . Processing module 1002 is further configured to store the output data of the M processing tasks in a storage unit. For example, see S904 in FIG9 .

[0349] Optionally, the data processing device 1000 also includes a software interface module 1003; the software interface module 1003 is used to: determine a target task request based on a received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes M processing tasks; the software interface module 1003 is also used to: send a target task request to the scheduling module 1001; the target task request is used to indicate the original data and the target operation; the scheduling module 1001 is specifically used to: determine the first-ranked processing task among the M processing tasks and the data to be processed of the first-ranked processing task based on the received target task request; the data to be processed of the first-ranked processing task is the original data.

[0350] Optionally, the processing module includes multiple operators; the FPGA stores an operator form, which is used to indicate the processing tasks performed by each operator in the multiple operators; the scheduling module 1001 is specifically used to: determine a target operator for a target processing task in M ​​processing tasks from the multiple operators indicated by the operator form, the operator form is used to indicate that the target operator is used to perform the target processing task; and send a target request command to the target operator, the target request command is used to request execution of the target processing task on the target data to be processed.

[0351] Optionally, the operator table is further used to indicate the working state of each operator unit, where the working state includes an idle state or a busy state; the operator table indicates that the working state of the target operator unit is an idle state.

[0352] Optionally, after the scheduling module 1001 determines the target operator unit for the target processing task in the M processing tasks from the multiple operator units indicated by the operator table, the scheduling module 1001 is further used to: update the working state of the target operator unit indicated by the operator table to a busy state.

[0353] Optionally, when the target operator completes executing the target processing task, the scheduling module 1001 is further configured to update the working state of the target operator indicated by the operator table to an idle state.

[0354] Optionally, the target operator is used to: when executing the target processing task based on the target request command, return target response information to the scheduling module, where the target response information is used to indicate the storage address of the output data of the target processing task.

[0355] Optionally, the target response information is used to indicate the execution result of the target processing task, where the execution result includes execution success or execution failure.

[0356] Optionally, the FPGA stores a task form, which is used to indicate the task status of M processing tasks, including the executed status or the unexecuted status; the scheduling module 1001 is also used to: update the task form based on the received target response information; the updated task form is used to indicate that the target processing task is in the executed state.

[0357] Optionally, the scheduling module 1001 may include multiple sub-scheduling modules, wherein each sub-scheduling module is used to send a request command to the processing unit.

[0358] Optionally, when the processing unit completes executing M processing tasks, the scheduling module 1001 is further configured to: output target data to the processor, where the target data is output data of the processing task ranked Mth in the target sequence.

[0359] Optionally, the data to be processed is image data. The processing module 1002 includes a decoding operator, which is used to perform image decoding tasks. The decoding operator includes multiple decoding modules, which are used to decode 8Q pixels of the image data in parallel.

[0360] Optionally, the processing module 1002 includes a coding operator, which is used to perform an image coding task. The coding operator includes multiple coding modules, and the multiple coding modules are used to encode 8Q pixels of the image data in parallel.

[0361] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation of any of the above data processing devices 1000 and the description of the beneficial effects can refer to the above corresponding method embodiments, which will not be repeated here.

[0362] An embodiment of the present application further provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a data processing device, the data processing device performs the steps of the above-mentioned data processing method.

[0363] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a data processing device, the data processing device performs the steps of the above-mentioned data processing method.

[0364] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data processing method, characterized in that, Applied to a data processing device, the data processing device includes a field programmable gate array (FPGA), and the FPGA includes a scheduling unit, a processing unit, and a storage unit; the method includes: The scheduling unit determines M processing tasks and the data to be processed for the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; The scheduling unit sends M request commands to the processing unit, where each request command is used to request to execute each processing task according to the data to be processed for each processing task; The processing unit executes the M processing tasks in the target order based on the received M request commands to obtain the output data of the M processing tasks; among them, the output data of the Nth processing task is the data to be processed for the (N + 1)th processing task in the target order; N is a positive integer less than M; The processing unit stores the output data of the M processing tasks in the storage unit.

2. The method according to claim 1, wherein The processing unit includes a plurality of operator units; the FPGA stores an operator form, and the operator form is used to indicate the processing tasks executed by each operator unit among the plurality of operator units; The scheduling unit sending M request commands to the processing unit includes: The scheduling unit determines a target operator unit for a target processing task among the plurality of operator units indicated by the operator form, and the operator form is used to indicate that the target operator unit is used to execute the target processing task; The scheduling unit sends a target request command to the target operator unit, and the target request command is used to request to execute the target processing task on the target data to be processed.

3. The method according to claim 2, wherein The operator form is further used to indicate the working state of each operator unit, and the working state includes an idle state or a busy state; the operator form is used to indicate that the working state of the target operator unit is the idle state.

4. The method according to claim 2 or 3, characterized in that, The method further includes: When the target operator unit finishes executing the target processing task based on the target request command, the target operator unit returns target response information to the scheduling unit, and the target response information is used to indicate the storage address of the output data of the target processing task.

5. The method according to claim 4, wherein The target response information is further used to indicate the execution result of the target processing task, and the execution result includes execution success or execution failure.

6. The method according to claim 4 or 5, characterized in that, The FPGA stores a task form, and the task form is used to indicate the task status of the M processing tasks, and the task status includes an executed state or an unexecuted state; the method further includes: The scheduling unit updates the task form based on the received target response information; the updated task form is used to indicate that the target processing task is in the executed state.

7. The method according to any one of claims 1-6, characterized in that, The data processing device further includes a processor, and the processor is used to convert the received data processing request into a task request recognizable by the FPGA; The scheduling unit determines M processing tasks and the data to be processed for the M processing tasks, including: The processor determines a target task request based on the received data processing request; the data processing request is used to request to perform a target operation on the original data, and the target operation includes the M processing tasks; The processor sends the target task request to the scheduling unit; the target task request is used to indicate the original data and the target operation; The scheduling unit determines the M processing tasks and the data to be processed for the first-sorted processing task among the M processing tasks based on the received target task request; the data to be processed for the first-sorted processing task is the original data.

8. The method according to claim 7, characterized in that, The method further includes: When the processing unit finishes executing the M processing tasks, the scheduling unit outputs target data to the processor, and the target data is the output data of the Mth-sorted processing task in the target order.

9. The method according to any one of claims 1-8, wherein The scheduling unit includes a plurality of sub-scheduling units, wherein each sub-scheduling unit is used to send a request command to the processing unit.

10. The method according to any one of claims 1-9, wherein The data to be processed is image data; the processing unit includes a decoding operator unit, and the decoding operator unit is used to perform an image decoding task. The decoding operator unit includes a plurality of run-length decoding modules, and the decoding operator unit decodes 8Q pixels of the image data in parallel through the plurality of run-length decoding modules; Q is a positive integer greater than 1; and / or The processing unit includes an encoding operator unit, and the encoding operator unit is used to perform an image encoding task. The encoding operator unit includes A plurality of run-length encoding modules, and the encoding operator unit encodes 8Q pixels of the image data in parallel through the plurality of run-length encoding modules.

11. A data processing device, characterized in that, For a data processing device, the hardware processor of the data processing device includes a processing unit and a storage unit; the data processing device includes: A scheduling module, configured to determine M processing tasks and the data to be processed for the M processing tasks; the execution order of the M processing tasks is the target order, and M is a positive integer greater than 1; the data to be processed is stored in the storage unit; The scheduling module is further configured to send M request commands to the processing unit, wherein each request command is used to request to execute each processing task according to the data to be processed for each processing task; A processing module, configured to execute the M processing tasks in the target order based on the received M request commands to obtain the output data of the M processing tasks; wherein the output data of the Nth processing task is the data to be processed for the (N + 1)th processing task in the target order; N is a positive integer less than M; The processing module is further configured to store the output data of the M processing tasks in the storage unit.

12. A data processing device, characterized in that, Including an FPGA, the FPGA includes: A scheduling unit, configured to determine M processing tasks and the to-be-processed data of the M processing tasks; the execution order of the M processing tasks is the target order, where M is a positive integer greater than 1; the to-be-processed data is stored in the storage unit; The scheduling unit is further configured to send M request commands to the processing unit; wherein, each request command is used to request to execute each processing task according to the to-be-processed data of each processing task; A processing unit, configured to execute the M processing tasks in the target order based on the received M request commands, to obtain the output data of the M processing tasks; wherein, the output data of the Nth processing task is the to-be-processed data of the (N + 1)th processing task in the target order; N is a positive integer less than M; The processing unit is further configured to store the output data of the M processing tasks in the storage unit.

13. A data processing device resource pool, characterized in that The data processing device resource pool includes at least one data processing device, and the at least one data processing device is configured to execute the steps of the method according to any one of claims 1-10.

14. A computer program product, characterized in that, Including computer programs / instructions, which, when executed by a data processing device, implement the steps of the method according to any one of claims 1-10.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer programs / instructions, which, when executed by a data processing device, implement the steps of the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Heterogeneous cluster-orientated task processing method and device

    CN107678752A

  • Data processing method, system and terminal equipment

    CN110399221A

  • Data processing method and device, electronic equipment and medium

    CN112035258A

  • Task execution method and storage device

    CN113821311A