Computing system, video encoding method and data processing module

By adding a hardware processing unit to the data processing module, the video encoding function is offloaded to the data processing module, and the problem of video encoding consumes CPU resources in the prior art is solved, and a more efficient video encoding process is realized.

CN116347094BActive Publication Date: 2025-05-06ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310315385.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-05-06
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

The prior art consumes more CPU resources and memory resources in the video encoding process, which can easily lead to a bottleneck in the CPU performance of computing devices.

Method used

By adding a hardware processing unit to the data processing module, including a DMA engine and a hardware encoder, the video encoding function is offloaded from the host's CPU to the data processing module, and the DMA engine is used to access the host memory for video encoding.

Benefits of technology

It reduces the computing pressure of the host CPU, reduces the probability of the host CPU reaching a performance bottleneck, and improves the video encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116347094B_ABST
    Figure CN116347094B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a computing system, a video encoding method and a data processing module. In the embodiment of the present application, the video encoding function is unloaded from the CPU of the host to the data processing module, which reduces the computing pressure of the CPU of the host and reduces the probability of the CPU of the host reaching a performance bottleneck. On the other hand, the hardware processing unit in the data processing module uses the DMA engine to access the host memory, which can increase the speed at which the hardware processing unit obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a computing system, a video encoding method, and a data processing module. Background Art

[0002] Video encoding refers to converting the original video format file into another video format file through compression technology. In this way, redundant video information can be reduced and the video transmission bandwidth can be reduced. The video encoding algorithm is complex. If the central processing unit (CPU) of the computing device is directly used for software encoding, it will consume more CPU resources and memory resources, which is easy to cause the CPU performance bottleneck of the computing device. Summary of the invention

[0003] Multiple aspects of the present application provide a computing system, a data processing method, and a data processing module for offloading video encoding from a host processor to reduce the probability of a CPU of a computing device reaching a performance bottleneck.

[0004] In a first aspect, an embodiment of the present application provides a computing system, comprising: a host and a data processing module; the host and the data processing module are in communication connection; the data processing module comprises: a hardware processing unit and a central processing unit CPU; the hardware processing unit comprises: a direct memory access DMA engine and a control plane transmission channel; the hardware processing unit is additionally provided with a hardware encoder;

[0005] The host is used to send a video encoding request to the hardware processing unit;

[0006] The hardware processing unit is used to transparently transmit the video encoding request to the CPU through the control plane transmission channel;

[0007] The CPU runs the firmware of the hardware encoder and drives the hardware encoder to operate based on the video encoding request;

[0008] The hardware encoder is used to obtain the video to be encoded corresponding to the video encoding request from the memory of the host using the DMA engine under the drive of the CPU; and encode the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

[0009] In a second aspect, an embodiment of the present application further provides a computing system, comprising: a host and a data processing module; the host and the data processing module are communicatively connected; the data processing module comprises: a hardware processing unit and a central processing unit CPU; the hardware processing unit comprises: a direct memory access DMA engine and a control plane transmission channel;

[0010] The host is used to send a video encoding request to the hardware processing unit;

[0011] The hardware processing unit is used to transparently transmit the video encoding request to the CPU through the control plane transmission channel;

[0012] The CPU, based on the video encoding request, controls the hardware processing unit to use the DMA engine to obtain the video to be encoded from the memory of the host, and uses the DMA engine to store the video to be encoded in the memory of the CPU; reads the video to be encoded from the memory of the CPU, and encodes the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

[0013] In a third aspect, an embodiment of the present application further provides a video encoding method, which is applicable to a data processing module, wherein the data processing module is connected to a host; the data processing module comprises: a hardware processing unit and a CPU; the hardware processing unit comprises: a direct memory access DMA engine and a control plane transmission channel; the hardware processing unit is additionally provided with a hardware encoder;

[0014] The method comprises:

[0015] Obtaining a video encoding request sent by the host;

[0016] Transmitting the video encoding request to the CPU through the control plane transmission channel;

[0017] The CPU runs the firmware of the hardware encoder and drives the hardware encoder to operate based on the video encoding request;

[0018] The hardware encoder, driven by the CPU, uses the DMA engine to obtain the video to be encoded corresponding to the video encoding request from the memory of the host; and encodes the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

[0019] In a fourth aspect, the embodiment of the present application also provides a data processing module, the data processing module is connected to a host; the data processing module includes: a hardware processing unit and a CPU; the hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; the method includes:

[0020] Obtaining a video encoding request sent by the host;

[0021] Transmitting the video encoding request to the CPU through the control plane transmission channel;

[0022] The CPU runs the firmware of the hardware encoder and drives the hardware encoder to operate based on the video encoding request;

[0023] The hardware encoder, under the drive of the CPU, uses the DMA engine to obtain the video to be encoded corresponding to the video encoding request from the memory of the host; and uses the DMA engine to store the video to be encoded in the memory of the CPU;

[0024] The CPU encodes the video to be encoded in the memory to obtain video encoding data corresponding to the video to be encoded.

[0025] In a fifth aspect, an embodiment of the present application further provides a data processing module, comprising: a hardware processing unit and a CPU; the hardware processing unit and the CPU are in communication connection; the hardware processing unit comprises: a direct memory access DMA engine and a control plane transmission channel; the hardware processing unit is additionally provided with a hardware encoder; when the data processing module is in communication connection with the host, the CPU is used to execute the steps in the method executed by the CPU in the video encoding method provided in the third aspect above;

[0026] The hardware encoder is used to execute the steps in the hardware encoder execution method in the video encoding method provided in the third aspect above.

[0027] In a sixth aspect, an embodiment of the present application further provides a data processing module, comprising: a hardware processing unit and a CPU; the hardware processing unit and the CPU are in communication connection; the hardware processing unit comprises: a direct memory access DMA engine and a control plane transmission channel; when the data processing module is in communication connection with the host, the CPU is used to execute the steps in the method executed by the CPU in the video encoding method provided in the fourth aspect above;

[0028] The hardware processing unit is used to execute the steps in the hardware processing unit execution method in the video encoding method provided in the fourth aspect above.

[0029] In a seventh aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, causes the one or more processors to execute the steps in the above-mentioned video encoding method.

[0030] In the embodiment of the present application, the video encoding function is unloaded from the host CPU to the data processing module, which reduces the computing pressure of the host CPU and reduces the probability of the host CPU reaching a performance bottleneck. On the other hand, the hardware processing unit in the data processing module uses the DMA engine to access the host memory, which can increase the speed at which the hardware processing unit obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0032] Figure 1-Figure 3 A schematic diagram of the structure of a computing system provided in an embodiment of the present application;

[0033] Figure 4 A schematic diagram of a flow chart of a computing system according to an embodiment of the present application using a pipelined task execution method to perform video encoding;

[0034] Figure 5 A schematic diagram of an interactive process for performing video encoding on a computing system provided in an embodiment of the present application;

[0035] Figure 6 A schematic diagram of the structure of another computing system provided in an embodiment of the present application;

[0036] Figure 7a and Figure 7b A schematic diagram of a video encoding method according to an embodiment of the present invention;

[0037] Figure 8a and Figure 8b A schematic diagram of the structure of a data processing module provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0039] In some embodiments of the present application, the video encoding function is unloaded from the host CPU to the data processing module, which reduces the computing pressure of the host CPU and reduces the probability of the host CPU reaching a performance bottleneck. On the other hand, the hardware processing unit in the data processing module uses the DMA engine to access the host memory, which can increase the speed at which the hardware processing unit obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency.

[0040] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0041] It should be noted that the same reference numerals denote the same objects in the following drawings and embodiments, and therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in the subsequent drawings and embodiments.

[0042] Figure 1-Figure 3 A schematic diagram of the structure of a computing system provided in an embodiment of the present application. Figure 1-Figure 3 The computing system includes: a host 10 and a data processing module 20.

[0043] In this embodiment, the host 10 refers to any computer device with computing, storage and communication functions. For example, the host 10 can be a server, a computer, a mobile phone, etc. In this embodiment, the host can include: a general processing unit, etc. In this embodiment, the number of general processing units is not limited. The general processing unit can be at least one, that is, one or more; each general processing unit can be a single-core processing unit or a multi-core processing unit.

[0044] In this embodiment, the general processing unit is generally a processing chip arranged on the mainboard of the host 10, such as the central processing unit (CPU) 101 of the host, etc., and single-machine expansion cannot be achieved. The general processing unit can be any processing device with processing and computing capabilities. The general processing unit can be a serial processing unit or a parallel processing unit. For example, the general processing unit can be a general processor, such as a CPU. A parallel processing unit refers to a processing device that can perform parallel computing and processing. For example, the parallel processing unit can be a graphics processing unit (GPU) or a field programmable gate array (FPGA). Optionally, the memory of the general processing unit is greater than the memory of the parallel processing unit. Figure 1 The general processing unit is only illustrated as a CPU as an example, but it does not constitute a limitation.

[0045] The data processing module 20 is any programmable hardware device or component with computing and communication functions. The data processing module 20 may include a hardware processing unit 201 and a CPU 202. The CPU 202 may be an independent processor or an on-chip processor in a system on chip (SoC).

[0046] The hardware processing unit 201 may be a hardware processor built by an electronic device, or a hardware processor that uses a hardware description language (HDL) for data processing. The hardware description language may be a very-high-speed integrated circuit hardware description language (VHDL), Verilog HDL, System Verilog or System C, etc. The hardware processing unit 201 may be an FPGA, a programmable array logic device (PAL), a general array logic device (GAL), a complex programmable logic device (CPLD), etc. Alternatively, the hardware processing unit 201 may also be an application-specific integrated circuit (ASIC), etc.

[0047] In the present embodiment, the data processing module 20 is connected to the host 10 by communication, and more specifically, the data processing module 20 is connected to the CPU 101 of the host 10 by communication. Specifically, the data processing module 20 and the host 10 can be connected to each other by communication through a bus interface. Among them, the bus interface can be a serial bus interface, such as a Peripheral Component Interconnect Express (PCIe) bus interface, a Peripheral Component Interconnect (PCI) bus interface, an Ultra Path Interconnect (UPI) bus interface, a Universal Serial Bus (USB) serial interface, an RS485 interface or an RS232 interface, etc. Preferably, the bus interface is a PCIe interface, which can improve the data transmission rate between the data processing module 20 and the host 10.

[0048] The bus interface of the host 10 can be expanded according to the specifications of the host 10. Generally, the host 10 has multiple communication interfaces. In the embodiments of the present application, "multiple" means more than one, that is, two or more. When the data processing module 20 is connected to the host 10 through the bus interface, there can be multiple data processing modules 20 to achieve the expansion of the data processing module 20.

[0049] In some embodiments, the data processing module 20 and the host 10 may be arranged in different physical machines, and the host 10 and the data processing module 20 may be connected via network communication. For example, the host 10 and the data processing module may be arranged in different cloud servers, connected via network communication; and so on. Figure 1-Figure 3 The figure only illustrates that the host 10 and the data processing module 20 are arranged on the same physical machine, but this does not constitute a limitation.

[0050] In this embodiment, in order to reduce the pressure on the CPU of the virtual machine in the host 10 and reduce the probability of the CPU reaching a performance bottleneck, the virtual machine in the host 10 can map the data processing module 20 to a virtual device of the virtual machine through virtualization technology; and through the mapped virtual device, the data processing function is offloaded to the data processing module 20.

[0051] In the embodiment of the present application, the specific implementation form of the virtual machine mapping the data processing module 20 to the virtual device of the host 10 through the virtualization technology is not limited. In some embodiments, the virtual machine can map the data processing module 20 to the virtual device of the virtual machine through the semi-virtualization technology. Among them, the semi-virtualization technology is compared with full virtualization. In a fully virtualized solution, the client virtual machine (Virtual Machine, VM) needs to use the underlying host resources, and a virtual machine monitor (Virtual Machine Monitor, VMM) is required to intercept all request instructions and then simulate the behavior of these instructions, which is bound to bring a lot of performance overhead. Semi-virtualization completes the partially virtualized instructions through hardware through the assistance of the underlying hardware, and the VMM is only responsible for completing the virtualization of some instructions. To do this, the client VM needs to cooperate, the client VM completes the front-end drivers of different devices, and the VMM cooperates with the client VM to complete the corresponding back-end drivers, so that the two can achieve an efficient virtualization process through a certain interaction mechanism.

[0052] In some embodiments, the paravirtualization technology may be an input / output virtualization (VirtIO) technology. Among them, VirtIO is an I / O paravirtualization solution, which is a lubricant for communication between the client (Guest) and the host (Host). It provides a set of general frameworks, as well as standard interfaces or protocols to complete the interaction process between the two, which greatly solves the adaptation problem between various drivers and different virtualization solutions. VirtIO technology provides a communication framework and programming interface between upper-layer applications and various VMM virtualization devices (such as kernel-based virtual machines (KVM), Xen, VMware, etc.). Generally speaking, VirtIO can be divided into four layers, including various driver modules in the front-end client, processing program modules on the back-end VMM, and the middle VirtIO layer and VirtIO ring layer (i.e., VirtIO Ring) layer for front-end and back-end communication. Among them, the VirtIO layer implements a virtual queue interface, which can be regarded as a bridge for front-end and back-end communication, and the VirtIO Ring layer is the specific implementation of the bridge. It implements two ring buffers, which are used to store information executed by the front-end driver and the back-end processing program respectively.

[0053] Accordingly, if Figure 2 As shown, the data processing module 20 can support single root I / O virtualization (SR-IOV). SR-IOV can virtualize a data processing module into multiple lightweight virtual data processing modules, which are allocated to multiple virtual machines deployed by the host. The SR-IOV protocol introduces the concept of two types of functions: physical function (PF) and virtual function (VF). PF is used to support PCI functions of SR-IOV functions, as defined in the SR-IOV specification. PF contains an SR-IOV function configuration structure for managing SR-IOV functions. PF is a full-featured PCIe function that can be discovered, managed and processed like a PCIe device. PF has fully configured resources and can be used to configure or control PCIe devices. VF is a function associated with PF, a lightweight PCIe function that can share one or more physical resources with a physical function and other VFs associated with the same physical function. VF is only allowed to have configuration resources for its own behavior.

[0054] Based on the principle of the above-mentioned semi-virtualization technology, the virtual machine can virtualize the data processing module 20 into multiple virtual video devices through the SR-IOV technology, and distribute them to multiple virtual machines. Specifically, the virtual machine can initialize the data processing module 20, and in the process of initializing the data processing module 20, configure the identity of the data processing module 20 to virtualize the data processing module 20 as a virtual device of the virtual machine, that is, a virtual video module. Among them, the identity of the data processing module 20 refers to the identifier of the virtual device on the bus that uniquely identifies a host. For the embodiment in which the data processing module 20 is connected to the bus of the host 10 through the PCIe bus interface and communicates with the host 10, the identity of the data processing module 20 can be represented by the bus number (Bus Number), the device number (Device Number) and the function number (Function Number) (abbreviated as BDF).

[0055] BDF is a unique identifier for each function in a PCIe device. Each PCIe device can have only one function or multiple functions, up to a maximum of 8 functions. No matter how many functions a PCIe device has, each of its functions has a unique and independent configuration space. The host 10 can obtain some information about the PCIe device through this space, and can also configure the PCIe device through this space. This space is called the PCIe configuration space. The PCIe configuration software (such as the PCIe root complex, etc.) can identify the topological logic of the PCIe bus system, as well as each bus, each PCIe device and each function of the PCIe device, that is, it can identify BDF.

[0056] In an embodiment of the present application, in order to map the data processing module 20 to a virtual video (Video) device of the virtual machine, the function number in the identity identifier of the virtual machine configurable data processing module 20 includes the identifier of the video device. For the host 10 to identify the PCIe device, it is also necessary to perform device enumeration on the PCIe device. The PCIe architecture generally includes a root component (Root Complex), a switch (Switch) and various PCIe devices. Among them, the PCIe device is the terminal device (Endpoint) of the PCIe architecture. For so many devices, after the CPU of the virtual machine is started, PCIe device enumeration is required to identify these devices. The root complex usually uses a specific algorithm, such as a depth-first algorithm, to access every possible branch path until it can no longer be accessed in depth, and each PCIe device is only accessed once. This process is called PCIe device enumeration.

[0057] In this embodiment, for the host 10, during the device enumeration process, the identity of the data processing module 20 can be obtained from the register of the virtual device to determine that the function of the virtual device is a video device. Further, the host 10 can be equipped with a driver corresponding to the function of the virtual device. In this way, for the host 10, the data processing module 20 is a video device of the virtual machine, that is, the data processing module 20 is virtualized as a virtual video device of the virtual machine. Among them, for the VirtIO technology, the virtual device driver can be a VirtIO video (VirtIO-video) driver. Accordingly, the communication protocol between the hardware processing unit and the host can be an input / output semi-virtualized video (VirtIO-video) protocol.

[0058] In the embodiment of the present application, in order to realize the offloading of the video encoding function, a hardware encoder 201a can be added to the hardware processing unit 201 to realize the hardware encoding function of the video. The hardware encoder 201a can be a video encoder constructed by electronic devices, or a video encoder encoded by a hardware description language.

[0059] In this embodiment, the hardware processing unit 201 in the data processing module 20 includes: a direct memory access (DMA) engine 201b and a control plane transmission channel 201c. The control plane transmission channel is a channel for transmitting control plane information (such as control plane messages or data packets, etc.). The control plane transmission channel 201c is used to transparently transmit data between the host 10 and the CPU 202. The DMA engine 201b can access the host memory using the DMA method, and can also access the memory of the CPU 202 using the DMA method.

[0060] Since the data processing module 20 has a DMA engine and a control plane transmission channel, there is no need to develop a corresponding DMA engine and a control plane transmission channel for the hardware encoder 201a added to the data processing module 20. The data transmission and the DMA engine can reuse the existing paths in the data processing module 20, thereby reducing the development cost of video encoding.

[0061] In this embodiment, the data processing module 20 can be implemented as any form of offload card. In some embodiments, the data processing module 20 can be implemented as a data processing unit (DPU). Among them, the DPU is a dedicated electronic circuit with hardware acceleration function, which is used for data-centric computing. Data is transmitted and transmitted in the form of multiplexed packets of messages. The DPU usually includes: a CPU, a network card, and a DMA engine. This makes the DPU have the versatility and programmability of the CPU, and can efficiently operate network data packets, storage requests, or analysis requests. Among them, the network card refers to computer hardware with network communication functions.

[0062] The virtualization of the encoder mainly involves three parts: device enumeration, VirtIO-video message transmission, and direct memory access (DMA) of data. Since the DPU at least offloads the network and / or storage functions and supports the virtualization of the corresponding devices, the hardware encoder 201a can directly reuse these frameworks. For example, for SR-IOV device enumeration, the DPU has implemented the simulation of the device space through software and hardware, and the hardware encoder can be virtualized as a VirtIO-video device based on a small amount of expansion. The transmission of VirtIO-video messages and the DMA engine can reuse the existing paths in the DPU, and can even directly reuse the virtual queues of the network core and / or storage.

[0063] like Figure 2 As shown, the CPU 202 may also integrate the firmware corresponding to the hardware encoder 201a. In this embodiment, the firmware corresponding to the hardware encoder refers to the driver of the hardware encoder. The CPU 202 may drive the hardware programmer 201a through the firmware to implement operations such as video encoding.

[0064] Based on the hardware processing unit 201 with video encoding function, the host 10 can offload the video encoding function to the data processing module 20. In this embodiment, in order to improve the versatility and programmable flexibility of the CPU, the control plane program can be deployed on the CPU 202 to facilitate the update of the control plane program. The CPU 202 may also have its own CPU memory, that is, Figure 3 The CPU 202 is connected to the memory 203 via a communication interface. The memory is generally a dynamic random access memory (DRAM). Accordingly, the communication interface can be implemented as a communication interface supported by DRAM, referred to as a DRAM interface. The memory 203 can be used to store the control plane program of the CPU 202.

[0065] The hardware processing unit 201 is used for video encoding, and the CPU 202 is used to control the hardware processing unit 201 to achieve separation of the data plane and the control plane. In the embodiment of the present application, the control plane program includes but is not limited to: message parsing and programs (such as firmware) for controlling the hardware processing unit 201.

[0066] In this embodiment, when the host 10 has a video encoding requirement, the host 10 can load the video to be encoded into the host memory 102. Figure 1-Figure 3 , sends a video encoding request to the data processing module 20. The video encoding request can request the data processing module 20 to allocate encoder resources and perform encoding processing on the corresponding video.

[0067] The hardware processing unit 201 in the data processing module 20 receives the video encoding request, and transparently transmits the video encoding request to the CPU 202 through the control plane transmission channel 201c. For the VirtIO technology, the video encoding request complies with the VirtIO-video protocol.

[0068] Accordingly, the CPU 202 may run the firmware of the hardware encoder 201a and drive the hardware encoder 201a to operate (corresponding to Figure 1 "Controlling Data Reading and Video Encoding" in .

[0069] The hardware encoder 201a can obtain the video to be encoded from the host memory 102 using the DMA engine 201b under the drive of the CPU 202. Further, the hardware encoder 201a can encode the video to be encoded to obtain video encoding data corresponding to the video to be encoded, such as video code stream data. The video encoding data can be an elementary stream (ES). The elementary stream is a data stream directly output from the hardware encoder 201a, that is, an encoded video data stream.

[0070] In this embodiment, the video encoding function is unloaded from the host CPU to the data processing module, which reduces the computing pressure of the host CPU and reduces the probability of the host CPU reaching a performance bottleneck. On the other hand, the hardware encoder in the hardware processing unit uses the DMA engine to access the host memory, which can increase the speed at which the hardware encoder obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency. Moreover, unloading the video encoding to the hardware encoder for processing can achieve hardware acceleration of video encoding, which helps to improve the efficiency of video encoding.

[0071] In addition, due to the flexibility of CPU software update, in this embodiment, the control plane program is executed by the CPU in the data processing module, which facilitates the update of the control plane program.

[0072] For an embodiment in which a data processing module is pre-integrated with a DMA engine and a control plane transmission channel, a hardware encoder added to the hardware processing unit can directly reuse the DMA engine and the control plane transmission channel of the data processing module, thereby reducing the development cost of the hardware encoder.

[0073] In some embodiments, in combination Figure 2 and Figure 3 , the hardware processing unit 201 may also include: a memory controller 201d and a memory 201e. The memory controller 201d is mainly used to access the memory 201e. Among them, the memory 201e can be a memory integrated in the hardware processing unit 201 itself, or it can be an external memory of the encoding hardware unit 201. Accordingly, the memory 201e can be connected to the hardware processing unit 201 through a communication interface. The memory 201e can be a DRAM. Accordingly, the communication interface can be implemented as a DRAM interface. In some embodiments, the memory 201e can be implemented as a double data rate synchronous dynamic random access memory (Double Data Rate Synchronous Dynamic Random Access Memory, DDR SDRAM). Accordingly, the memory controller 201d can be a DDR memory controller.

[0074] In this embodiment, the memory controller 201d can write the video to be encoded read by the DMA engine 201b into the memory 201e, and the hardware encoder 201a can read the video to be encoded from the memory 201e and encode the video to be encoded.

[0075] In the embodiment of the present application, the specific algorithm used by the hardware encoder 201a to encode the video to be encoded is not limited. In some embodiments, the video encoding algorithm used by the hardware encoder 201a may be: H.264 algorithm, H.265 algorithm, H.266 algorithm, technology, i.e., motion still image (or frame by frame) compression algorithm (Motion Joint Photographic Experts Group, MJPEG) or compression algorithm based on discrete cosine transform (Discrete Cosine Transform, DCT), etc., but is not limited thereto.

[0076] In some embodiments of the present application, Figure 2 and Figure 3 , the host 10 is deployed with a virtual machine 30. The number of virtual machines 30 may be 1 or more. More means 2 or more. The resources of the virtual machine 30 may include: the CPU and memory of the virtual machine, etc. The CPU and memory of the virtual machine are connected via a communication interface. The memory may also be DRAM. Accordingly, the communication interface may be implemented as a DRAM interface. Figure 3 The communication interface is only illustrated as a DRAM interface as an example, but it does not constitute a limitation.

[0077] The CPU of the virtual machine may be referred to as a virtual CPU (Virtual CPU, vCPU). The virtual machine 30 may send a video encoding request to the data processing module 20 according to the video encoding requirement. Specifically, the virtual machine 30 deploys an application (Application, APP) and the like. The virtual machine 30 may send a video encoding request to the data processing module 20 according to the video encoding requirement of the application.

[0078] The user of the virtual machine 30 can purchase the corresponding encoding capacity according to actual needs. The scheduling node corresponding to the host 10 can allocate encoding capacity to the virtual machine 30 according to the encoding capacity requirements of the user of the virtual machine 30. Among them, the encoding capacity is used to reflect the encoding capability of the resources allocated to the virtual machine, and the encoding capacity limits the amount of video data that the hardware encoder 201a can encode. Generally, the encoding capacity is equal to the product of the resolution and the frame rate. The frame rate refers to the number of frames processed per second (Frame PerSecond, FPS). The encoding capacity available to the virtual machine 30 shall not exceed the encoding capacity allocated to the virtual machine.

[0079] Based on this, when the CPU 202 drives the hardware encoder 201a to operate based on the video encoding request, it can determine whether to accept the video encoding request according to the encoding capacity corresponding to the virtual machine (ie, the encoding capacity allocated to the virtual machine) and the video encoding request.

[0080] Specifically, the CPU 202 may parse the video encoding request (corresponding to Figure 2 ). Optionally, the CPU 202 may obtain the encoding capacity requirement information of the virtual machine 30 based on the video encoding request. Specifically, the CPU 202 may obtain the resolution and frame rate of the video frame to be encoded from the video encoding request; and calculate the encoding capacity requirement information of the virtual machine according to the resolution and frame rate of the video frame to be encoded. For example, the resolution of the video frame of the encoded video and the product of the frame rate may be calculated to obtain the encoding capacity requirement information of the virtual machine.

[0081] In actual applications, a single virtual machine may send a single video encoding request, or may send multiple video encoding requests concurrently. Multiple refers to 2 or more. For an embodiment in which a single virtual machine sends multiple video encoding requests concurrently, the CPU 202 may obtain encoding capacity requirement information corresponding to each video encoding request based on the multiple video encoding requests; and calculate the encoding capacity requirement information of the virtual machine according to the encoding capacity requirement information corresponding to each of the multiple video encoding requests. For example, the sum of the encoding capacity requirement information corresponding to each of the multiple video encoding requests may be calculated to obtain the encoding capacity requirement information of the virtual machine.

[0082] Further, the CPU 202 may determine whether the remaining encoding capacity of the virtual machine 30 meets the requirements of the encoding capacity requirement information according to the encoding capacity corresponding to the virtual machine (i.e., the encoding capacity allocated to the virtual machine) and the encoding capacity currently consumed by the virtual machine 30. If the remaining encoding capacity of the virtual machine 30 is greater than or equal to the encoding capacity requirement information, it is determined that the remaining encoding capacity of the virtual machine 30 meets the requirements of the encoding capacity requirement information. If the remaining encoding capacity of the virtual machine 30 is less than the encoding capacity requirement information, it is determined that the remaining encoding capacity of the virtual machine 30 does not meet the requirements of the encoding capacity requirement information.

[0083] Further, if the remaining encoding capacity of the virtual machine 30 meets the requirements of the encoding capacity requirement information, it is determined to accept the video encoding request. Further, the CPU 202 may return a request acceptance message to the virtual machine 30. The hardware processing unit 201 may transparently transmit the request acceptance message to the virtual machine 30 through the control plane transmission channel 201c. Of course, if the remaining encoding capacity of the virtual machine 30 does not meet the requirements of the encoding capacity requirement information, a request rejection message may be returned to the virtual machine 30. The hardware processing unit 201 may transparently transmit the request rejection message to the virtual machine 30 through the control plane transmission channel 201c.

[0084] For the embodiment of returning the request acceptance message, the virtual machine 30 may determine the video frame to be encoded from the video to be encoded in response to the request acceptance message; and provide the metadata of the video frame to be encoded to the CPU 202. Specifically, the virtual machine 30 may provide the metadata of the video frame to be encoded to the data processing module 20. The hardware processing unit 201 in the data processing module 20 transparently transmits the metadata of the video frame to be encoded to the CPU 202 through the control plane transmission channel 201c. The metadata of the video frame to be encoded is information describing the attributes of the video frame, and may include: the storage location of the video frame, the identification and resolution of the video frame, etc. In the VirtIO technology, the metadata of the video frame to be encoded can be encapsulated in a data packet that complies with the VirtIO-video protocol, and transparently transmitted to the CPU 202 in the form of a data packet.

[0085] The CPU 202 may obtain metadata of the video frame to be encoded. Specifically, the CPU 202 may parse a data packet containing metadata of the video frame to be encoded to obtain the metadata of the video frame to be encoded carried by the data packet.

[0086] Furthermore, the CPU 202 can drive the hardware encoder 201a to operate based on the storage location information in the metadata of the video frame to be encoded. Accordingly, the hardware encoder 201a, driven by the CPU 202, uses the DMA engine 201b to obtain the video frame to be encoded from the memory 102 of the virtual machine, and encodes the video frame to be encoded.

[0087] In order to prevent the virtual machine from using the encoding resources excessively, the CPU 202 may monitor the actual usage of the encoding capacity by the virtual machine 30 in real time during the encoding of the video to be encoded, when determining to accept the video encoding request of the virtual machine. If it is detected that the actual usage is greater than the encoding capacity allocated to the virtual machine 30, the frame rate of the video encoding of the hardware processing unit 201 may be reduced so that the actual usage of the encoding capacity by the virtual machine 30 does not exceed the encoding capacity allocated to the virtual machine.

[0088] In some embodiments, the frame rate of the video encoding of the hardware encoder 201a can be reduced so that the actual usage of the encoding capacity by the virtual machine 30 is equal to the encoding capacity allocated to the virtual machine. Alternatively, the frame rate of the video encoding of the hardware encoder 201a can be reduced so that the actual usage of the encoding capacity by the virtual machine 30 is less than the encoding capacity allocated to the virtual machine, etc.

[0089] For the above-mentioned embodiment in which a single virtual machine concurrently sends multiple video encoding requests, when the CPU 202 determines to accept the multiple video encoding requests concurrently sent by the single virtual machine, the CPU 202 may drive the hardware encoder 201a to execute the encoding tasks corresponding to the multiple video encoding requests in parallel in a time-division multiplexing manner (corresponding to Figure 2 Accordingly, the hardware encoder 201a, driven by the CPU 202, executes encoding tasks corresponding to multiple video encoding requests in parallel in a time-division multiplexing manner.

[0090] Specifically, the CPU 202 may allocate 1 / N processing time to each video encoding request based on a fair scheduling strategy. Wherein, N is the number of multiple video encoding requests, N≥2, and is an integer, and controls the hardware encoder 201a to perform the encoding task of the video encoding request corresponding to the current 1 / N processing time. For example, the CPU 202 may receive metadata of the video frame to be encoded of the video encoding request corresponding to the current 1 / N processing time sent by the virtual machine 30. Further, the CPU 202 may drive the hardware encoder 201a to act based on the metadata of the video frame to be encoded. Accordingly, the hardware encoder 201a uses the DMA engine 201b to obtain the metadata of the video frame to be encoded of the video encoding request corresponding to the current 1 / N processing time from the memory of the virtual machine 30, and encodes the video frame to be encoded read by the DMA engine 201b until the 1 / N processing time is reached.

[0091] For an embodiment in which a plurality of virtual machines 30 are deployed on the host 10 , there may be at least two virtual machines among the plurality of virtual machines 30 , which may concurrently send video encoding requests to the data processing module 20 .

[0092] Although the hardware encoder 201a can be virtualized into multiple virtual video devices and allocated to multiple virtual machines 30 through the above-mentioned paravirtualization technology, the physical hardware encoder 201a is limited. Therefore, it is necessary to perform resource scheduling for video encoding requests sent concurrently by at least two virtual machines.

[0093] For the embodiment in which at least two virtual machines concurrently send video encoding requests, the CPU 202 may also determine whether to accept the video encoding request of each virtual machine according to the encoding capacity corresponding to each virtual machine and the video encoding request sent by the virtual machine. For any virtual machine A among the at least two virtual machines that concurrently send video encoding requests, it may be determined whether to accept the video encoding request of virtual machine A according to the encoding capacity of virtual machine A and the video encoding request sent by virtual machine A. For the specific implementation of determining whether to accept the video encoding request of virtual machine A, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0094] In the case of determining to accept the video encoding requests sent concurrently by the at least two virtual machines, the CPU 202 may schedule the encoding tasks of the at least two virtual machines based on the set encoding scheduling strategy; and control the hardware processing unit 201 to execute the encoding task of the currently scheduled virtual machine (corresponding to Figure 2 in “Scheduling”).

[0095] The set encoding scheduling strategy refers to a scheduling strategy for concurrent video encoding requests, and the strategy can determine which virtual machine is selected from the concurrent video encoding requests to process the video encoding request.

[0096] In the embodiment of the present application, the specific implementation form of the encoding scheduling strategy is not limited. In some embodiments, the users corresponding to the virtual machines have different priorities, and the encoding scheduling strategy can be implemented as a priority scheduling strategy. Accordingly, the CPU 202 can control the hardware encoder 201a to preferentially execute the encoding task of the virtual machine with the highest user priority according to the user priorities corresponding to at least two virtual machines providing concurrent video encoding requests.

[0097] In other embodiments, the encoding scheduling strategy may be implemented as a time-division multiplexing scheduling strategy. Accordingly, the CPU 202 may schedule the encoding tasks of at least two virtual machines providing concurrent video encoding requests based on the time-division multiplexing scheduling strategy, and drive the hardware encoder 201a to adopt a time-division multiplexing method to execute the encoding tasks of at least two virtual machines providing concurrent video encoding requests. Accordingly, the hardware encoder 201a, driven by the CPU 202, adopts a time-division multiplexing method to execute the encoding tasks of at least two virtual machines providing concurrent video encoding requests, thereby realizing time-division multiplexing of the hardware encoder between the encoding tasks of different virtual machines, thereby realizing the parallel execution of the encoding tasks of multiple virtual machines.

[0098] Optionally, the time-sharing multiplexing scheduling strategy may be a fair scheduling strategy, that is, users of virtual machines have the same priority. Accordingly, the CPU 202 may schedule the encoding tasks of at least two virtual machines providing concurrent video encoding requests based on the fair scheduling strategy, thereby implementing time-sharing multiplexing of the hardware processing unit between the encoding tasks of different virtual machines, thereby implementing the parallel execution of the encoding tasks of multiple virtual machines, which can be called temporal parallelism.

[0099] Specifically, the CPU 202 may allocate 1 / M of the processing time for the above-mentioned two virtual machines, and drive the hardware encoder 201a to execute the encoding task of the target virtual machine corresponding to the current 1 / M processing time. Wherein, M is the number of at least two virtual machines providing concurrent video encoding requests. M≥2 and is an integer. Accordingly, the hardware encoder 201a, driven by the CPU 202, executes the encoding task of the target virtual machine corresponding to the current 1 / M processing time until the 1 / M processing time is reached.

[0100] Specifically, the CPU 202 may receive metadata of the video frame to be encoded sent by the target virtual machine, and drive the hardware encoder 201a to operate based on the storage location information in the metadata of the video frame to be encoded. Under the drive of the CPU 202, the hardware encoder 201a uses the DMA engine 201c to read the video frame to be encoded from the memory of the target virtual machine based on the storage location information of the video frame to be encoded, and performs video encoding on the video frame to be encoded to obtain video encoding data of the video frame to be encoded.

[0101] In some embodiments, the fair scheduling strategy may be a Completely Fair Scheduler (CFS) strategy. The CFS strategy is a scheduling algorithm based on the idea of ​​weighted fair queuing. The CFS strategy introduces the concept of weight, using weight to represent the user priority of the virtual machine, and the video encoding task of each virtual machine is allocated processing time according to the proportion of the weight. Among them, the higher the priority of the user, the greater the weight of its virtual machine. For example: virtual machines A and B. The weight of virtual machine A is 1024, and the weight of virtual machine B is 2048. The proportion of processing time of the encoding task obtained by virtual machine A is 1024 / (1024+2048)=33.3%. The proportion of processing time of the encoding task obtained by virtual machine B is 2048 / (1024+2048)=66.7%.

[0102] Accordingly, the CPU 202 may determine the weights of the at least two virtual machines according to the user priorities of the at least two virtual machines providing concurrent video encoding requests; and determine the processing time of the encoding tasks of the at least two virtual machines according to the weights of the at least two virtual machines; and according to the processing time of the encoding tasks of the at least two virtual machines, drive the hardware encoder 201a to execute the encoding task of the target virtual machine corresponding to the processing time at the corresponding processing time. Accordingly, the hardware encoder 201a, driven by the CPU 202, executes the encoding task of the target virtual machine corresponding to the processing time at the corresponding processing time, realizing time-division multiplexing of the hardware encoder between the encoding tasks of different virtual machines, thereby realizing the parallel execution of the encoding tasks of multiple virtual machines.

[0103] The scheduling method of the encoding task of the virtual machine shown in the above embodiment is only an example and does not constitute a limitation.

[0104] In addition to achieving temporal parallelism, the embodiment of the present application can also achieve spatial parallel processing of video encoding. Specifically, the hardware encoder 201a can be configured as a multi-level encoding unit according to the video encoding process of the video encoding algorithm. Multi-level refers to 2 levels or more. Figure 4 In the figure, only the coding unit of level R is used as an example for illustration. R is the number of levels of the coding unit, R≥2, and is an integer.

[0105] The multi-level encoding units can be connected according to the logical relationship between the video encoding processes of the video encoding algorithm. The user can write the hardware description language corresponding to the multi-level encoding unit through the development system corresponding to the hardware processing unit 201, and burn the video encoding primitive into the hardware processing unit 201 to obtain the hardware encoder 201a. Alternatively, the multi-level encoding unit can also be a hardware module built according to the electronic device corresponding to the video encoding process of the video encoding algorithm.

[0106] Among them, different video coding algorithms have different video coding processes, and accordingly the number and function of multi-level coding units are also different. For example, for the H.264 algorithm, the video coding process may include five processes: inter-frame and intra-frame prediction (Estimation), transform and inverse transform, quantization and inverse quantization, loop filtering (LoopFilter) and entropy coding (Entropy Coding). Accordingly, the multi-level coding unit may include: inter-frame and intra-frame prediction unit, transform and inverse transform unit, quantization and inverse quantization unit, loop filtering unit and entropy coding unit.

[0107] Based on the above multi-level encoding units, the CPU 202 can drive the multi-level encoding units to execute the encoding tasks of different video frames in the video to be encoded in a pipelined task execution manner, so as to encode different video frames. Accordingly, under the drive of the CPU 202, the multi-level encoding units execute the encoding tasks of different video frames in the video to be encoded in a pipelined task execution manner, thereby achieving parallelism of the encoding tasks in space, which helps to improve the throughput rate of the hardware processing unit 201. That is, by controlling the encoding tasks of the multi-level encoding units to be carried out simultaneously and separately processing the encoding tasks of different video frames, the time when the hardware encoder 201a is in an idle waiting state can be reduced, and the video encoding efficiency can be improved.

[0108] Specifically, as Figure 4 shown, the video encoding process of each video frame can be divided into multiple sub-processes, and the encoding tasks of each sub-process are the same as the functions of the encoding units corresponding to the processes. Accordingly, for any two adjacent encoding units A and B in the multi-level encoding units, where encoding unit B is the next process of the encoding process of encoding unit A, after the encoding task corresponding to the Kth video frame executed by encoding unit A is completed, the processed data corresponding to the Kth video frame can be transmitted to encoding unit B. For encoding unit A, it can then execute the encoding task corresponding to the (K + 1)th video frame without waiting for other encoding units to complete the encoding task of the Kth video frame, and can start the encoding task of the next frame, which can reduce the idle waiting time of each level of encoding units in the multi-level encoding units and help improve the video encoding efficiency. Wherein, K is a positive integer, and 1 ≤ K < Q. Q is the total number of video frames in the video to be encoded.

[0109] For example, for the H.264 algorithm, after the inter-frame and intra-frame prediction units complete the inter-frame and intra-frame prediction of the Kth video frame, the inter-frame and intra-frame prediction results of the Kth video frame can be transmitted to the transform and inverse transform units; then, the inter-frame and intra-frame prediction units can then perform the inter-frame and intra-frame prediction of the (K + 1)th video frame without waiting for other encoding units to complete the encoding of the Kth video frame, and so on, to achieve pipelined video encoding, and realize parallel processing of multiple video frames on the multi-level encoding units, which helps to improve the video encoding efficiency.

[0110] Based on the above pipelined task execution manner, for the case where the number of video frames in the video to be encoded that have not been encoded is greater than or equal to the number of levels of the encoding units, R video frames can be encoded in parallel in the hardware encoder 201a, which helps to improve the encoding efficiency. Figure 4 Only the first to Rth video frames in the video to be encoded that are currently being encoded are taken as examples for illustration, but it does not constitute a limitation.

[0111] In the process of the above-mentioned multi-level encoding unit executing the encoding tasks of different video frames in the video to be encoded in a pipelined task execution manner, the CPU 202 can drive the hardware encoder 201a to use the DMA engine to read the (K+1)th video frame from the memory of the virtual machine, and drive the first-level encoding unit in the hardware encoder 201a to perform the encoding task of the (K+1)th video frame after the first-level encoding unit in the multi-level encoding unit completes the encoding task of the Kth video frame.

[0112] Combination Figure 4 and Figure 5 After each video frame is encoded, the CPU 202 may also drive the hardware encoder 201a to use the DMA engine to transmit the video encoding data (such as Figure 5 The video encoding data of the video frames 1-Q in the video encoding process is written into the memory of the virtual machine 30 to realize DMA access of the memory, which can improve the data transmission efficiency. For the above-mentioned embodiment of video encoding in the pipeline task execution mode, the CPU 202 drives the hardware encoder 201a to alternately read the video frames and send the video encoding data.

[0113] In some embodiments, the host 10 may also offload the network to the data processing module 20. The data processing module 20 may include: a network card (not shown in the drawings). Accordingly, after each video frame is encoded, the video encoding data corresponding to the video frame may be sent to the destination end of the video request through the network card, etc.

[0114] like Figure 4 As shown, after the encoding of each video frame is completed, the CPU 202 may also send an encoding success message to the virtual machine 30. The message may include metadata of the successfully encoded video frame, such as identification information and numbering information of the video frame. The hardware processing unit 201 may transparently transmit the encoding success message to the virtual machine 30 through the control plane transmission channel 201c. The virtual machine 30 may determine which video frame in the video to be encoded is successfully encoded based on the programming success message. Accordingly, the virtual machine 30 may determine whether the encoding of the video to be encoded is completed based on the metadata of the video frame in the encoding success message; if the virtual machine 30 receives encoding success messages corresponding to all frames in the video to be encoded, it is determined that the encoding of the video to be encoded is completed.

[0115] Further, if Figure 5As shown, the virtual machine 30 may send an encoder destruction request to the data processing module 20 after the programming of the video to be encoded is completed. The hardware processing unit 201 may transparently transmit the encoder destruction request to the CPU 202 via the control plane transmission channel 201c. The CPU 202 may recycle the encoder resources corresponding to the virtual machine in response to the encoder destruction request. After reclaiming the encoder resources corresponding to the virtual machine, the CPU 202 may return an encoder destruction success message to the virtual machine 30, and the hardware processing unit 201 may transparently transmit the encoder destruction success message to the virtual machine 30 via the control plane transmission channel 201c, so that the virtual machine 30 can be informed of the destruction of its encoder resources. About Figure 5 For the specific implementation of the VM 30 sending the video encoding request and the CPU 202 determining whether to accept or reject the video encoding request, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0116] In the embodiments of the present application, the specific implementation form of the data processing module 20 is not limited. In some embodiments, the data processing module 20 can be implemented as a DPU, a hardware encoder can be added to the DPU, and the firmware of the hardware encoder can be integrated into the CPU of the DPU to achieve multiplexing of the CPU of the DPU without the need for the hardware encoder to deploy the CPU separately. Since the SR-IOV software and hardware framework, DMA framework, and hot upgrade and hot migration framework of the existing DPU are mature technologies, the hardware encoder is integrated into the DPU, and the SR-IOV software and hardware framework, DMA framework, and hot upgrade and hot migration framework of the DPU can be reused to achieve the integration of the hardware encoder and the DPU, and to achieve resource reuse of the DPU.

[0117] Since DPU generally has communication components such as network cards, in some cloud service scenarios, such as cloud desktop, cloud application, cloud game or live broadcast, the encoded data of the image to be transmitted can be sent to the client through the DPU's network card, thereby realizing the offloading of the network protocol on the host side.

[0118] Since the DPU has realized the uninstallation of the control client, the hardware encoder can be controlled directly in the DPU, and the encoder firmware can be hot-upgraded and other operation and maintenance operations can be performed through the DPU control client. At the same time, the hardware encoder can also reuse the hot migration channel of the DPU's network devices and / or storage devices to copy the context of the virtual encoder to the destination end, and realize the hot migration of the virtual encoder.

[0119] The above embodiment only takes the example of unloading video encoding to the hardware encoder to realize hardware acceleration to exemplify the video encoding process. Of course, in some embodiments, the video encoding process of the host can also be unloaded to the CPU 202 of the data processing module 20 to realize the software unloading of video encoding. The software unloading process of video encoding is exemplified below.

[0120] Figure 6 A schematic diagram of the structure of a computing system for offloading video encoding software provided in an embodiment of the present application. Figure 6 As shown, the computing system includes: a host 10 and a data processing module 20. For the description of the implementation form and connection mode of the host 10 and the data processing module 20, please refer to the relevant content of the above embodiment, which will not be repeated here. Figure 6 As shown, the data processing module 20 may include: a hardware processing unit 201 and a CPU 202. For the description of the implementation forms of the hardware processing unit 201 and the CPU 202, reference may be made to the relevant contents of the above embodiments, which will not be repeated here.

[0121] In this embodiment, in order to reduce the pressure on the CPU of the host 10 and reduce the probability of the CPU reaching a performance bottleneck, the host 10 can map the data processing module 20 to a virtual device of the host 10 through virtualization technology; and offload the data processing function to the data processing module 20 through the mapped virtual device. For the specific implementation of virtualizing the data processing module 20, please refer to the above Figure 1-Figure 3 The relevant contents of the embodiment will not be repeated here.

[0122] In the embodiment of the present application, in order to realize the offloading of the video encoding function, a software module, plug-in or container with the video encoding function may be configured in the CPU 202 .

[0123] In this embodiment, when the host 10 has a video encoding requirement, the host 10 can load the video to be encoded into the host memory 102. Figure 6 As shown, the host 10 sends a video encoding request to the data processing module 20. The hardware processing unit 201 in the data processing module 20 receives the video encoding request and transparently transmits the video encoding request to the CPU 202 through the control plane transmission channel 201c. For VirtIO technology, the video encoding request follows the VirtIO-video protocol.

[0124] Accordingly, the CPU 202 runs the firmware of the hardware processing unit 201, and drives the hardware processing unit 201 to act based on the video encoding request. Accordingly, the hardware processing unit 201, driven by the CPU 202, uses the DMA engine 201b to obtain the video to be encoded from the memory 102 of the host, and uses the DMA engine 201b to store the video to be encoded in the memory 203 of the CPU 202. Further, the CPU 202 can encode the video to be encoded in the memory 203 to obtain video encoding data corresponding to the video to be encoded, such as video code stream data.

[0125] In the embodiment of the present application, the specific algorithm of the CPU 202 for encoding the video to be encoded can be found in the relevant content of the above embodiment, which will not be repeated here.

[0126] In this embodiment, the video encoding function is unloaded from the CPU of the host to the data processing module, which reduces the computing pressure of the host CPU and reduces the probability of the CPU of the host reaching a performance bottleneck. On the other hand, the hardware processing unit uses DMA to access the host memory, which can increase the speed at which the hardware processing unit obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency. In addition, due to the flexibility of CPU software updates, the CPU in the data processing module performs the video encoding operation in this embodiment, which facilitates the update of the program corresponding to the video encoding operation.

[0127] In some embodiments of the present application, the host 10 is deployed with a virtual machine 30. The number of virtual machines 30 may be 1 or more. More than 2 means 2 or more. The virtual machine 30 may send a video encoding request to the data processing module 20 according to the video encoding requirement. The hardware processing unit 201 may transparently transmit the video encoding request to the CPU 202.

[0128] When the CPU 202 controls the hardware processing unit 201 to obtain the video to be encoded from the memory of the host 10 in a DMA manner based on the video encoding request, it can determine whether to accept the video encoding request according to the encoding capacity corresponding to the virtual machine (i.e., the encoding capacity allocated to the virtual machine) and the video encoding request. For the specific implementation of determining whether to accept the video encoding request, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0129] If it is determined to accept the video encoding request, the CPU 202 may return a request acceptance message to the virtual machine 30. The hardware processing unit 201 may transparently transmit the request acceptance message to the virtual machine 30 through the control plane transmission channel 201c. Of course, if the remaining encoding capacity of the virtual machine 30 does not meet the requirements of the encoding capacity demand information, a request rejection message may be returned to the virtual machine 30. The hardware processing unit 201 may transparently transmit the request rejection message to the virtual machine 30 through the control plane transmission channel 201c.

[0130] For the embodiment of returning the request acceptance message, the virtual machine 30 may determine the video frame to be encoded from the video to be encoded in response to the request acceptance message, and provide the metadata of the video frame to be encoded to the CPU 202. The metadata of the video frame to be encoded is information describing the attributes of the video frame, and may include: the storage location of the video frame, the identification and resolution of the video frame, etc.

[0131] The CPU 202 may drive the hardware processing unit 201 to act based on the storage location information in the metadata of the video frame to be encoded. Under the drive of the CPU 202, the hardware processing unit 201 uses the DMA engine 201b to obtain the video frame to be encoded from the memory 102 of the virtual machine in a DMA manner; and uses the DMA engine 201b to write the video frame to be encoded into the memory 203 corresponding to the CPU 202 in a DMA manner. The CPU 202 may encode the video frame to be encoded in the memory 203. The CPU 202 may start at least one process or at least one thread to encode the video frame to be encoded, etc.

[0132] For the above-mentioned embodiment in which a single virtual machine concurrently sends multiple video encoding requests, when the CPU 202 determines to accept multiple video encoding requests concurrently sent by a single virtual machine, the CPU 202 can start multiple processes or multiple threads to execute the encoding tasks corresponding to the multiple video encoding requests in parallel in a time-sharing multiplexing manner.

[0133] In other embodiments, the encoding scheduling strategy may be implemented as a time-division multiplexing scheduling strategy. Accordingly, the CPU 202 may schedule the encoding tasks of at least two virtual machines providing concurrent video encoding requests based on the time-division multiplexing scheduling strategy, thereby implementing time-division multiplexing between encoding tasks of different virtual machines, thereby implementing parallel execution of encoding tasks of multiple virtual machines.

[0134] Specifically, the CPU 202 may receive metadata of the video frame to be encoded sent by the target virtual machine, and based on the storage location information in the metadata of the video frame to be encoded, drive the hardware processing unit 201. Under the drive of the CPU 202, the hardware processing unit 201 may use the DMA engine 201b to read the video frame to be encoded from the memory of the target virtual machine based on the storage location information of the video frame to be encoded, and write the video frame to be encoded into the memory 203 of the CPU 202. The CPU 202 may perform video encoding on the video frame to be encoded in the memory 203 to obtain video encoding data of the video frame to be encoded.

[0135] In addition to achieving temporal parallelism, the embodiment of the present application can also achieve spatial parallel processing of video encoding. Specifically, the CPU 202 can start multiple processes or multiple threads. Multiple refers to 2 or more.

[0136] Based on the above multiple processes or multiple threads, the CPU 202 can control the multiple processes or multiple threads to execute the encoding tasks of different video frames in the video to be encoded in a pipelined task execution mode, so as to encode different video frames, thereby realizing spatial parallelism of encoding tasks. That is, the encoding tasks of the multi-level encoding units are controlled to be performed simultaneously, and the encoding tasks of different video frames are processed separately, which can reduce the time that the CPU process or thread is in an idle waiting state, and improve the video encoding efficiency.

[0137] Based on the above pipeline task execution method, when the number of video frames that have not yet been encoded in the video to be encoded is greater than or equal to the number of encoding units, CPU 202 can encode R video frames in parallel at the same time, which helps to improve encoding efficiency.

[0138] In the process of executing the encoding task of different video frames in the to-be-encoded video in a pipelined task execution mode, the CPU 202 can drive the hardware processing unit 201 to act after executing the encoding task of the Kth video frame. The hardware processing unit 201 can use the DMA engine 201b to use the DMA mode to read the (K+1)th video frame from the memory of the virtual machine and write it into the memory 203. The CPU 202 can execute the encoding task of the (K+1)th video frame.

[0139] For the description of CPU 202 returning the video encoding data of the encoded video frame to the virtual machine 30, and recycling the encoding resources corresponding to the virtual machine 30 after the encoded video encoding is completed, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0140] In some embodiments, the CPU 202 may also drive the network card of the data processing module 20 to provide the video encoding data of the encoded video frame to the destination end requesting the video, thereby realizing network unloading on the host side.

[0141] In addition to the above-mentioned system embodiment, the embodiment of the present application also provides a video encoding method. The video encoding method provided by the embodiment of the present application is exemplarily described below.

[0142] Figure 7a The flowchart of the video encoding method provided by the embodiment of the present application is shown in FIG. The method is applicable to a data processing module. The data processing module can be connected to a host for communication. The data processing module is connected to the host; the data processing module includes: a hardware processing unit and a CPU; the hardware processing unit includes: a DMA engine and a control plane transmission channel; and a hardware encoder is added to the hardware processing unit. Figure 7a As shown, the method mainly includes:

[0143] 701. Obtain a video encoding request sent by a host.

[0144] 702. Transmit the video encoding request to the CPU in the data processing module through the control plane transmission channel.

[0145] 703. The CPU runs the firmware of the hardware encoder and drives the hardware encoder to operate based on the video encoding request.

[0146] 704. Driven by the CPU, the hardware encoder uses the DMA engine to obtain the to-be-encoded video corresponding to the video encoding request from the memory of the host.

[0147] 705. The hardware encoder encodes the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

[0148] In an embodiment of the present application, in order to realize the unloading of the video encoding function, the video encoding can be unloaded to the data processing module. In this embodiment, when the host has a video encoding requirement, the video to be encoded can be loaded into the memory of the host. Further, the host can send a video encoding request to the data processing module. In step 701, the hardware processing unit in the data processing module receives the video encoding request, and in step 702, the video encoding request is transparently transmitted to the CPU in the data processing module through the data channel.

[0149] Accordingly, the CPU of the data processing module can run the firmware of the hardware encoder in step 703, and drive the hardware encoder to act based on the video encoding request. Accordingly, in step 704, the hardware encoder can use the DMA engine under the drive of the CPU to obtain the video to be encoded from the memory of the host; and in step 705, encode the video to be encoded to obtain video encoding data corresponding to the video to be encoded, such as video code stream data.

[0150] In this embodiment, the video encoding function is unloaded from the host CPU to the data processing module, which reduces the computing pressure of the host CPU and reduces the probability of the host CPU reaching a performance bottleneck. On the other hand, the hardware encoder in the hardware processing unit uses the DMA engine to access the host memory, which can increase the speed at which the hardware encoder obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency. Moreover, unloading the video encoding to the hardware encoder for processing can achieve hardware acceleration of video encoding, which helps to improve the efficiency of video encoding.

[0151] In addition, due to the flexibility of CPU software update, in this embodiment, the control plane program is executed by the CPU in the data processing module, which facilitates the update of the control plane program.

[0152] For an embodiment in which a data processing module is pre-integrated with a DMA engine and a control plane transmission channel, a hardware encoder added to the hardware processing unit can directly reuse the DMA engine and the control plane transmission channel of the data processing module, thereby reducing the development cost of the hardware encoder.

[0153] In some embodiments of the present application, a host is deployed with a virtual machine. The number of virtual machines may be 1 or more. More means 2 or more. The virtual machine may send a video encoding request to the data processing module according to the video encoding requirement.

[0154] The user of the virtual machine can purchase the corresponding encoding capacity according to actual needs. The scheduling node corresponding to the host can allocate encoding capacity to the virtual machine according to the encoding capacity requirements of the user of the virtual machine. Based on this, the CPU drives the hardware encoder action based on the video encoding request to be implemented as follows: according to the encoding capacity corresponding to the virtual machine (i.e., the encoding capacity allocated to the virtual machine) and the video encoding request, determine whether to accept the video encoding request.

[0155] Specifically, the encoding capacity requirement information of the virtual machine may be obtained based on the video encoding request. For example, the resolution and frame rate of the video frame of the video to be encoded may be obtained from the video encoding request; and the encoding capacity requirement information of the virtual machine may be calculated according to the resolution and frame rate of the video frame of the video to be encoded. For example, the product of the resolution and frame rate of the video frame of the encoded video may be calculated to obtain the encoding capacity requirement information of the virtual machine.

[0156] In actual applications, a single virtual machine may send a single video encoding request, or may send multiple video encoding requests concurrently. Multiple refers to 2 or more. For an embodiment in which a single virtual machine sends multiple video encoding requests concurrently, encoding capacity requirement information corresponding to each video encoding request may be obtained based on the multiple video encoding requests; and the encoding capacity requirement information of the virtual machine may be calculated based on the encoding capacity requirement information corresponding to each of the multiple video encoding requests. For example, the sum of the encoding capacity requirement information corresponding to each of the multiple video encoding requests may be calculated to obtain the encoding capacity requirement information of the virtual machine.

[0157] Further, it is possible to determine whether the remaining coding capacity of the virtual machine meets the requirements of the coding capacity demand information based on the coding capacity corresponding to the virtual machine (i.e., the coding capacity allocated to the virtual machine) and the coding capacity currently consumed by the virtual machine. If the remaining coding capacity of the virtual machine is greater than or equal to the coding capacity demand information, it is determined that the remaining coding capacity of the virtual machine meets the requirements of the coding capacity demand information. If the remaining coding capacity of the virtual machine is less than the coding capacity demand information, it is determined that the remaining coding capacity of the virtual machine does not meet the requirements of the coding capacity demand information.

[0158] Further, if the remaining encoding capacity of the virtual machine meets the requirements of the encoding capacity demand information, it is determined to accept the video encoding request. Further, a request acceptance message may be returned to the virtual machine. The hardware processing unit may transparently transmit the request acceptance message to the virtual machine through the control plane transmission channel. Of course, if the remaining encoding capacity of the virtual machine does not meet the requirements of the encoding capacity demand information, a request rejection message may be returned to the virtual machine. The hardware processing unit may transparently transmit the request rejection message to the virtual machine.

[0159] For an embodiment in which a request acceptance message is returned, the virtual machine may determine a video frame to be encoded from the video to be encoded in response to the request acceptance message, and provide metadata of the video frame to be encoded to the CPU. The metadata of the video frame to be encoded is information describing the attributes of the video frame, and may include: a storage location of the video frame, an identifier of the video frame, and a resolution, etc.

[0160] Accordingly, the CPU can drive the hardware encoder to operate based on the storage location information in the metadata of the video frame to be encoded. Accordingly, the hardware encoder can, under the drive of the CPU, obtain the video frame to be encoded from the memory of the virtual machine using the DMA engine based on the storage location information of the video frame to be encoded, and encode the video frame to be encoded.

[0161] In order to prevent the virtual machine from using encoding resources excessively, the CPU may monitor the actual usage of the encoding capacity of the virtual machine in real time during the encoding of the video to be encoded, after determining to accept the video encoding request of the virtual machine. If it is detected that the actual usage is greater than the encoding capacity allocated to the virtual machine, the frame rate of the video encoding of the hardware processing unit or the CPU may be reduced so that the actual usage of the encoding capacity of the virtual machine does not exceed the encoding capacity allocated to the virtual machine.

[0162] For the above-mentioned embodiment in which a single virtual machine concurrently sends multiple video encoding requests, when the CPU determines to accept multiple video encoding requests concurrently sent by a single virtual machine, the hardware encoder can be driven to execute encoding tasks corresponding to the multiple video encoding requests in parallel in a time-division multiplexing manner. Driven by the CPU, the hardware encoder executes encoding tasks corresponding to the multiple video encoding requests in parallel in a time-division multiplexing manner.

[0163] For an embodiment in which a host is deployed with multiple virtual machines, there may be at least two virtual machines among the multiple virtual machines, which can concurrently send video encoding requests to the data processing module.

[0164] For the embodiment in which at least two virtual machines concurrently send video encoding requests, the CPU may also determine whether to accept the video encoding request of each virtual machine according to the encoding capacity corresponding to each virtual machine and the video encoding request sent by the virtual machine. For any virtual machine A among the at least two virtual machines that concurrently send video encoding requests, it may be determined whether to accept the video encoding request of virtual machine A according to the encoding capacity of virtual machine A and the video encoding request sent by virtual machine A. For the specific implementation method of determining whether to accept the video encoding request of virtual machine A, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0165] In the case of determining to accept the video encoding requests sent concurrently by the at least two virtual machines, the CPU can schedule the encoding tasks of the at least two virtual machines based on the set encoding scheduling strategy; and drive the hardware encoder to execute the encoding task of the currently scheduled virtual machine. Accordingly, the hardware encoder can execute the encoding task of the virtual machine currently scheduled by the CPU under the drive of the CPU.

[0166] The set encoding scheduling strategy refers to a scheduling strategy for concurrent video encoding requests, and the strategy can determine which virtual machine is selected from the concurrent video encoding requests to process the video encoding request.

[0167] In the embodiments of the present application, the specific implementation form of the encoding scheduling strategy is not limited. In some embodiments, the users corresponding to the virtual machines have different priorities, and the encoding scheduling strategy can be implemented as a priority scheduling strategy. Accordingly, the CPU can drive the hardware processing unit to preferentially execute the encoding task of the virtual machine with the highest user priority according to the user priorities corresponding to at least two virtual machines that provide concurrent video encoding requests.

[0168] In other embodiments, the encoding scheduling strategy may be implemented as a time-sharing multiplexing scheduling strategy. Accordingly, the CPU may schedule the encoding tasks of at least two virtual machines that provide concurrent video encoding requests based on the time-sharing multiplexing scheduling strategy, and drive the hardware encoder to execute the encoding task of the currently scheduled target virtual machine. Accordingly, the hardware encoder may execute the encoding task of the currently scheduled target virtual machine under the drive of the CPU, and implement time-sharing multiplexing of the hardware encoder between the encoding tasks of different virtual machines, thereby implementing the parallel execution of the encoding tasks of multiple virtual machines.

[0169] The scheduling method of the encoding task of the virtual machine shown in the above embodiment is only an example and does not constitute a limitation.

[0170] In addition to achieving temporal parallelism, the embodiments of the present application can also achieve spatial parallel processing of video encoding. Specifically, the hardware encoder in the hardware processing unit can be configured as a multi-level encoding unit according to the video encoding process of the video encoding algorithm. Multi-level refers to 2 levels or more. The multi-level encoding units can be connected according to the logical relationship between the video encoding processes of the video encoding algorithm.

[0171] Based on the above multi-level encoding unit, the CPU can drive the multi-level encoding unit to operate. Driven by the CPU, the multi-level encoding unit executes the encoding tasks of different video frames in the video to be encoded in a pipelined task execution mode to encode different video frames, thereby realizing spatial parallelism of encoding tasks, which helps to improve the throughput of the hardware processing unit. That is, the encoding tasks of the multi-level encoding unit are controlled to be performed simultaneously, and the encoding tasks of different video frames are processed separately, which can reduce the time that the hardware encoder is in an idle waiting state and improve the video encoding efficiency.

[0172] After the CPU completes encoding of each video frame, it can also send an encoding success message to the virtual machine. The message may include metadata of the successfully encoded video frame, such as identification information and numbering information of the video frame. The hardware processing unit can transparently transmit the encoding success message to the virtual machine through the control plane transmission channel. Based on the programming success message, the virtual machine can determine which video frame in the video to be encoded is successfully encoded. Accordingly, the virtual machine can determine whether the encoding of the video to be encoded is completed based on the metadata of the video frame in the encoding success message; if the virtual machine receives the encoding success message corresponding to all frames in the video to be encoded, it is determined that the encoding of the video to be encoded is completed.

[0173] Furthermore, the virtual machine may send an encoder destruction request to the data processing module after the programming of the video to be encoded is completed. The hardware processing unit may transparently transmit the encoder destruction request to the CPU through the control plane transmission channel. The CPU may receive the encoder destruction request transparently transmitted by the hardware processing unit, and in response to the encoder destruction request, recycle the encoder resources corresponding to the virtual machine. After reclaiming the encoder resources corresponding to the virtual machine, the CPU may return an encoder destruction success message to the virtual machine, and the hardware processing unit may transparently transmit the encoder destruction success message to the virtual machine so that the virtual machine can be informed of the destruction of its encoder resources.

[0174] The above embodiment only takes the example of unloading video encoding to the hardware encoder to realize hardware acceleration as an example to illustrate the video encoding process. Of course, in some embodiments, the video encoding process of the host can also be unloaded to the CPU of the data processing module to realize the software unloading of video encoding. The software unloading process of video encoding is illustrated below.

[0175] Figure 7bA flow chart of another video encoding method provided by an embodiment of the present application. The method is applicable to a data processing module. The data processing module can be connected to a host for communication. The data processing module includes: a hardware processing unit and a CPU. The hardware processing unit includes: a DMA engine and a control plane transmission channel. Accordingly, Figure 7b As shown, the video encoding method includes:

[0176] 71. Get the video encoding request sent by the host.

[0177] 72. Transmit the video encoding request to the CPU of the data processing module through the control plane transmission channel.

[0178] 73. The CPU runs the firmware of the hardware processing unit and drives the hardware processing unit to act based on the video encoding request.

[0179] 74. The hardware processing unit, under the drive of the CPU, uses the DMA engine to obtain the video to be encoded corresponding to the video encoding request from the memory of the host; and uses the DMA engine to store the video to be encoded in the memory of the CPU of the data processing module.

[0180] 75. The CPU of the data processing module encodes the video to be encoded in the memory to obtain video encoding data corresponding to the video to be encoded.

[0181] In this embodiment, in order to reduce the pressure on the host's CPU and reduce the probability of the CPU reaching a performance bottleneck, the host can map the data processing module to a virtual device of the host through virtualization technology; and through the mapped virtual device, the data processing function is offloaded to the data processing module.

[0182] In an embodiment of the present application, in order to realize the offloading of the video encoding function, a software module, plug-in or container with a video encoding function may be configured in the CPU of the data processing module.

[0183] In this embodiment, when the host has a video encoding requirement, the host can load the video to be encoded into the memory of the host. Further, the host sends a video encoding request to the data processing module. In step 71, the hardware processing unit in the data processing module receives the video encoding request, and in step 72, the video encoding request is transparently transmitted to the CPU of the data processing module through the control plane transmission channel.

[0184] Accordingly, in step 73, the CPU 202 in the data processing module runs the firmware of the hardware processing unit, and drives the hardware processing unit to act based on the video encoding request. Accordingly, in step 74, the hardware processing unit, driven by the CPU 202, uses the DMA engine to obtain the video to be encoded from the memory of the host, and uses the DMA engine 201b to store the video to be encoded in the memory of the CPU of the data processing module. Further, in step 75, the CPU of the data processing module can encode the video to be encoded in the memory of the CPU to obtain video encoding data corresponding to the video to be encoded, such as video code stream data.

[0185] In the embodiment of the present application, the specific algorithm for encoding the video to be encoded by the CPU of the data processing module can be found in the relevant content of the above embodiment, which will not be repeated here.

[0186] In this embodiment, the video encoding function is unloaded from the CPU of the host to the data processing module, which reduces the computing pressure of the host CPU and reduces the probability of the CPU of the host reaching a performance bottleneck. On the other hand, the hardware processing unit uses DMA to access the host memory, which can increase the speed at which the hardware processing unit obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency. In addition, due to the flexibility of CPU software updates, the CPU in the data processing module performs the video encoding operation in this embodiment, which facilitates the update of the program corresponding to the video encoding operation.

[0187] In some embodiments of the present application, a host is deployed with a virtual machine. The number of virtual machines 30 may be one or more. Multiple means two or more. The virtual machine may send a video encoding request to the data processing module according to the video encoding requirement. The hardware processing unit may transparently transmit the video encoding request to the CPU of the data processing module.

[0188] When the CPU of the data processing module drives the hardware processing unit to act based on the video encoding request, it can determine whether to accept the video encoding request according to the encoding capacity corresponding to the virtual machine (i.e., the encoding capacity allocated to the virtual machine) and the video encoding request. For the specific implementation of determining whether to accept the video encoding request, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0189] If it is determined to accept the video encoding request, the CPU of the data processing module may return a request acceptance message to the virtual machine. The hardware processing unit may transparently transmit the request acceptance message to the virtual machine through the control plane transmission channel. Of course, if the remaining encoding capacity of the virtual machine 30 does not meet the requirements of the encoding capacity demand information, a request rejection message may be returned to the virtual machine. The hardware processing unit may transparently transmit the request rejection message to the virtual machine through the control plane transmission channel.

[0190] For an embodiment in which a request acceptance message is returned, the virtual machine may determine a video frame to be encoded from the video to be encoded in response to the request acceptance message, and provide metadata of the video frame to be encoded to the CPU. The metadata of the video frame to be encoded is information describing the attributes of the video frame, and may include: a storage location of the video frame, an identifier of the video frame, and a resolution, etc.

[0191] The CPU of the data processing module can drive the hardware processing unit to act based on the storage location information in the metadata of the video frame to be encoded. Driven by the CPU, the hardware processing unit uses the DMA engine to obtain the video frame to be encoded from the memory of the virtual machine in a DMA manner; and uses the DMA engine to write the video frame to be encoded into the memory corresponding to the CPU of the data processing module in a DMA manner. The CPU of the data processing module can encode the video frame to be encoded in its memory.

[0192] For the above-mentioned embodiment in which a single virtual machine concurrently sends multiple video encoding requests, when the CPU of the data processing module determines to accept multiple video encoding requests concurrently sent by a single virtual machine, the CPU of the data processing module can start multiple processes or multiple threads to execute the encoding tasks corresponding to the multiple video encoding requests in parallel in a time-sharing multiplexing manner.

[0193] In other embodiments, the encoding scheduling strategy may be implemented as a time-division multiplexing scheduling strategy. Accordingly, the CPU of the data processing module may schedule the encoding tasks of at least two virtual machines providing concurrent video encoding requests based on the time-division multiplexing scheduling strategy, thereby implementing time-division multiplexing between encoding tasks of different virtual machines, thereby implementing parallel execution of encoding tasks of multiple virtual machines.

[0194] Specifically, the CPU of the data processing module can receive the metadata of the video frame to be encoded sent by the target virtual machine, and drive the hardware processing unit based on the storage location information in the metadata of the video frame to be encoded. Under the drive of the CPU, the hardware processing unit can use the DMA engine to read the video frame to be encoded from the memory of the target virtual machine based on the storage location information of the video frame to be encoded, and write the video frame to be encoded into the memory of the CPU of the data processing module. The CPU of the data processing module can perform video encoding on the video frame to be encoded in its memory to obtain video encoding data of the video frame to be encoded.

[0195] In addition to achieving temporal parallelism, the embodiments of the present application can also achieve spatial parallel processing of video encoding. Specifically, the CPU of the data processing module can start multiple processes or multiple threads. Multiple refers to 2 or more.

[0196] Based on the above multiple processes or multiple threads, the CPU of the data processing module can control multiple processes or multiple threads to execute the encoding tasks of different video frames in the video to be encoded in a pipelined task execution mode, so as to encode different video frames, thereby realizing spatial parallelism of encoding tasks. That is, the encoding tasks of the multi-level encoding units are controlled to be performed simultaneously, and the encoding tasks of different video frames are processed separately, which can reduce the time that the CPU process or thread is in an idle waiting state and improve the video encoding efficiency.

[0197] Based on the above-mentioned pipeline task execution method, when the number of video frames that have not yet been encoded in the video to be encoded is greater than or equal to the number of encoding units, the CPU of the data processing module can simultaneously encode R video frames in parallel, which helps to improve the encoding efficiency.

[0198] In the process of executing the encoding task of different video frames in the video to be encoded in the pipeline task execution mode, the CPU of the data processing module can drive the hardware processing unit to act after executing the encoding task of the Kth video frame. The hardware processing unit can use the DMA engine to read the (K+1)th video frame from the memory of the virtual machine and write it into the memory of the CPU of the data processing module. The CPU of the data processing module can execute the encoding task of the (K+1)th video frame.

[0199] For a description of the CPU of the data processing module returning the video encoding data of the encoded video frame to the virtual machine, and recycling the encoding resources corresponding to the virtual machine after the encoded video encoding is completed, please refer to the relevant content of the above embodiment, which will not be repeated here.

[0200] In some embodiments, the CPU of the data processing module may also drive the network card of the data processing module to provide the video encoding data of the encoded video frame to the destination end requesting the video, thereby realizing network unloading on the host side.

[0201] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0202] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 701 and 702 can be device A; for another example, the execution subject of step 701 can be device A, and the execution subject of step 702 can be device B; and so on.

[0203] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel, and the sequence numbers of the operations, such as 701, 702, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.

[0204] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the above-mentioned video encoding methods.

[0205] Figure 8a and Figure 8b This is a schematic diagram of the structure of the data processing module provided in the embodiment of the present application. Figure 8a and Figure 8b As shown, the data processing module includes: a hardware processing unit 801 and a CPU 802. The hardware processing unit 801 is communicatively connected with the CPU 802. The hardware processing unit 801 may include: a DMA engine 801b and a control plane transmission channel 801c.

[0206] In some embodiments, Figure 8a As shown, the hardware processing unit 801 is additionally provided with a hardware encoder 801a. When the data processing module is in communication connection with the host, the hardware processing unit 801 can be used to obtain a video encoding request issued by the host; and transparently transmit the video encoding request to the CPU 802 through the control plane transmission channel. The CPU 802 can run the firmware of the hardware encoder 801a, and drive the hardware encoder 801a to act based on the video encoding request.

[0207] Driven by the CPU 802, the hardware encoder 801a uses the DMA engine 801b to obtain the video to be encoded corresponding to the video encoding request from the memory of the host, and encodes the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

[0208] In some embodiments, a host is deployed with a virtual machine; the virtual machine sends a video encoding request to the data processing module. When the CPU 802 drives the hardware encoder to act based on the video encoding request, it is specifically used to: determine whether to accept the video encoding request according to the video encoding request and the encoding capacity allocated to the virtual machine; if the judgment result is yes, return a request acceptance message to the virtual machine; and transparently transmit the request acceptance message to the virtual machine through the control plane transmission channel, so that the virtual machine responds to the request acceptance message, determines the video frame to be encoded from the video to be encoded, and provides the metadata of the video frame to be encoded to the data processing module. Accordingly, the CPU 802 can drive the hardware encoder to act based on the metadata of the video frame to be encoded returned by the virtual machine.

[0209] Accordingly, when the hardware encoder 801a, driven by the CPU 802, uses the DMA engine 801b to obtain the video to be encoded corresponding to the video encoding request from the memory of the host, it is specifically used to: use the DMA engine 801b to obtain the video frame to be encoded from the memory of the host under the drive of the CPU 802. Further, when the hardware encoder 801a encodes the video to be encoded, it can encode the video frame to be encoded to obtain the video encoding data of the video frame to be encoded.

[0210] In some embodiments, when the CPU 802 determines whether to accept a video encoding request based on the encoding capacity allocated to the virtual machine and the video encoding request, it is specifically used to: obtain the encoding capacity requirement information of the virtual machine based on the video encoding request; determine whether the remaining encoding capacity of the virtual machine meets the requirements of the encoding capacity requirement information based on the encoding capacity allocated to the virtual machine and the encoding capacity currently consumed by the virtual machine; if the judgment result is yes, determine to accept the video encoding request.

[0211] In some embodiments of the present application, CPU 802 is also used to: monitor the actual usage of the encoding capacity of the virtual machine when it is determined to accept the video encoding request; and reduce the video encoding frame rate of the hardware encoder when it is monitored that the actual usage is greater than the encoding capacity allocated to the virtual machine, so that the actual usage of the encoding capacity of the virtual machine does not exceed the encoding capacity allocated to the virtual machine.

[0212] In some embodiments, the hardware encoder includes: a multi-level encoding unit; when the hardware encoder encodes the video to be encoded, it is specifically used for: the multi-level encoding unit is driven by the CPU to process the encoding tasks of multiple video frames in the video to be encoded in a pipelined task execution manner.

[0213] In some embodiments, the host deploys multiple virtual machines; at least two of the multiple virtual machines concurrently send video encoding requests. When the CPU 802 drives the hardware encoder to act based on the video encoding request, it is specifically used to: schedule the encoding tasks of at least two virtual machines based on the time-division multiplexing scheduling strategy when the video encoding requests concurrently sent by at least two virtual machines are accepted; and drive the hardware encoder 801a to perform video encoding on the video to be encoded of the currently scheduled virtual machine.

[0214] Accordingly, the hardware encoder 801a, driven by the CPU 802, performs video encoding on the video to be encoded of the virtual machine currently scheduled by the CPU.

[0215] In this embodiment, the data processing module may support SR-IOV. Multiple virtual machines virtualize the data processing module into multiple virtual video devices through SR-IOV and allocate the virtual video devices to the multiple virtual machines.

[0216] In some embodiments, the hardware encoder 801a, driven by the CPU 802, uses the DMA engine 801b to write the video encoding data corresponding to the video frame to be encoded into the memory of the virtual machine.

[0217] In other embodiments, the hardware processing unit 801 may also receive an encoder destruction request sent by the virtual machine; the encoder destruction request is sent by the virtual machine after the encoding of the video to be encoded is completed; and the encoder destruction request is transparently transmitted to the CPU 802 through the control plane transmission channel. In response to the encoder destruction request, the CPU 802 reclaims the encoder resources corresponding to the virtual machine.

[0218] In some embodiments of the present application, Figure 8b As shown, the hardware processing unit 801 has no hardware encoder. Accordingly, the hardware processing unit 801 can obtain the video encoding request sent by the host, and transparently transmit the video encoding request to the CPU 802 through the control plane transmission channel 801c.

[0219] The CPU 802 runs the firmware of the hardware processing unit and drives the hardware processing unit 801 to act based on the video encoding request. Under the drive of the CPU 802, the hardware processing unit 801 uses the DMA engine 801b to obtain the video to be encoded corresponding to the video encoding request from the memory of the host; and uses the DMA engine 801b to store the video to be encoded in the memory 803 of the CPU 802.

[0220] The CPU 802 may encode the video to be encoded in the memory 802 to obtain video encoding data corresponding to the video to be encoded.

[0221] Optionally, when encoding the video to be encoded from the memory, the CPU 802 is specifically configured to: start multiple processes or multiple threads; and process the encoding tasks of multiple video frames in the video to be encoded by using multiple processes or multiple processes in a pipeline task execution manner.

[0222] In some embodiments, a host deploys multiple virtual machines; at least two of the multiple virtual machines concurrently send video encoding requests; when CPU 802 encodes the video to be encoded from the memory, it is specifically used to: when the video encoding requests concurrently sent by at least two virtual machines are accepted, CPU 802 starts multiple processes or multiple threads, and uses multiple processes or multiple processes based on a time-sharing multiplexing scheduling strategy to schedule the encoding tasks of at least two virtual machines; and performs video encoding on the video to be encoded of the currently scheduled virtual machine.

[0223] The data processing module provided in this embodiment can offload the video encoding function from the CPU of the host to the data processing module when communicating with the host, thereby reducing the computing pressure of the CPU of the host and reducing the probability of the CPU of the host reaching a performance bottleneck. On the other hand, the hardware processing unit uses DMA to access the host memory, which can increase the speed at which the hardware processing unit obtains the video to be encoded, thereby helping to improve the subsequent video encoding efficiency.

[0224] It is worth noting that the above computing system may also include: storage, communication components, power components, display components, audio components and other optional components in addition to the memory. In the embodiments of the present application, only some components are schematically shown, which does not mean that the computing system must include all the components shown in the embodiments of the present application, nor does it mean that the computing system can only include the components shown in the embodiments of the present application.

[0225] In an embodiment of the present application, the memory is used to store a computer program and can be configured to store various other data to support operations on the device where it is located. Among them, the processor can execute the computer program stored in the memory to implement the corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random-Access Memory, SRAM), electrically erasable programmable read only memory (Electrically Erasable Programmable Read Only Memory, EEPROM), erasable programmable read only memory (Electrical Programmable Read Only Memory, EPROM), programmable read only memory (Programmable Read Only Memory, PROM), read only memory (Read Only Memory, ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0226] In the embodiment of the present application, the processor can be any hardware processing device that can execute the logic of the above method. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU) or a microcontroller unit (MCU); it can also be a field programmable gate array (FPGA), a programmable array logic device (PAL), a general array logic device (GAL), a complex programmable logic device (CPLD) and other programmable devices; or an application specific integrated circuit (ASIC) chip; or an advanced reduced instruction set (RISC) processor (Advanced RISC Machines, ARM) or a system on chip (SoC), etc., but not limited to this.

[0227] In an embodiment of the present application, the communication component is configured to facilitate wired or wireless communication between the device in which it is located and other devices. The device in which the communication component is located can access a wireless network based on a communication standard, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology or other technologies.

[0228] In an embodiment of the present application, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0229] In an embodiment of the present application, a power supply component is configured to provide power to various components of the device in which it is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0230] In an embodiment of the present application, the audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal may be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal. For example, for a device with a language interaction function, voice interaction with a user can be achieved through an audio component.

[0231] It should be noted that the descriptions such as “first” and “second” in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit “first” and “second” to different types.

[0232] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage, etc.) that contain computer-usable program code.

[0233] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (or systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing module to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing module generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0234] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing module to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0235] These computer program instructions can also be loaded onto a computer or other programmable data processing module so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0236] In a typical configuration, a computing device includes one or more processors (such as a CPU, etc.), an input / output interface, a network interface, and a memory.

[0237] Memory may include non-permanent storage in a computer-readable medium, random-access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0238] The storage medium of a computer is a readable storage medium, which may also be referred to as a readable medium. The readable storage medium includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be a computer-readable instruction, a data structure, a module of a program, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this article, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0239] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the above elements.

[0240] The above contents are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A computing system, characterized in that: include: Host and data processing module; The host is communicatively connected with the data processing module; The data processing module includes: a hardware processing unit and a central processing unit CPU; the hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; the hardware processing unit is additionally provided with a hardware encoder; The host is used to send a video encoding request to the hardware processing unit; The hardware processing unit is used to transparently transmit the video encoding request to the CPU through the control plane transmission channel; The CPU runs the firmware of the hardware encoder and drives the hardware encoder to operate based on the video encoding request; The hardware encoder is used to obtain the video to be encoded corresponding to the video encoding request from the memory of the host using the DMA engine under the drive of the CPU; and encode the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

2. The system according to claim 1, characterized in that The host deploys multiple virtual machines; the data processing module supports single root input and output virtualization SR-IOV; The multiple virtual machines virtualize the data processing module into multiple virtual video devices through SR-IOV technology, and distribute the data processing module to the multiple virtual machines.

3. A computing system, characterized in that: include: Host and data processing module; The host is communicatively connected with the data processing module; The data processing module includes: a hardware processing unit and a central processing unit CPU; the hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; The host is used to send a video encoding request to the hardware processing unit; The hardware processing unit is used to transparently transmit the video encoding request to the CPU through the control plane transmission channel; The CPU, based on the video encoding request, controls the hardware processing unit to use the DMA engine to obtain the video to be encoded from the memory of the host, and uses the DMA engine to store the video to be encoded in the memory of the CPU; reads the video to be encoded from the memory of the CPU, and encodes the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

4. The system according to claim 3, characterized in that The host deploys multiple virtual machines; the data processing module supports single root input and output virtualization SR-IOV; The multiple virtual machines virtualize the data processing module into multiple virtual video devices through SR-IOV technology, and distribute the data processing module to the multiple virtual machines.

5. A video encoding method, characterized in that: Applicable to a data processing module, the data processing module is connected to a host; The data processing module includes: a hardware processing unit and a CPU; the hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; the hardware processing unit is additionally provided with a hardware encoder; The method comprises: Obtaining a video encoding request sent by the host; Transmitting the video encoding request to the CPU through the control plane transmission channel; The CPU runs the firmware of the hardware encoder and drives the hardware encoder to operate based on the video encoding request; The hardware encoder, driven by the CPU, uses the DMA engine to obtain the video to be encoded corresponding to the video encoding request from the memory of the host; and encodes the video to be encoded to obtain video encoding data corresponding to the video to be encoded.

6. The method according to claim 5, characterized in that The host is deployed with a virtual machine; the virtual machine sends the video encoding request to the data processing module; The step of driving the hardware encoder to act based on the video encoding request includes: The CPU determines whether to accept the video encoding request according to the video encoding request and the encoding capacity allocated to the virtual machine; if the determination result is yes, returns a request acceptance message to the virtual machine; transparently transmitting the request acceptance message to the virtual machine through the control plane transmission channel, so that the virtual machine determines the video frame to be encoded from the video to be encoded in response to the request acceptance message, and provides metadata of the video frame to be encoded to the data processing module; The CPU drives the hardware encoder to operate based on the metadata of the to-be-encoded video frame returned by the virtual machine.

7. The method according to claim 6, characterized in that The hardware encoder, under the drive of the CPU, uses the DMA engine to obtain the to-be-encoded video corresponding to the video encoding request from the memory of the host, including: The hardware encoder, driven by the CPU, uses the DMA engine to obtain the to-be-encoded video frame from the memory of the host; The encoding of the video to be encoded includes: the hardware encoder encoding the video frame to be encoded to obtain video encoding data of the video frame to be encoded.

8. The method according to claim 6, characterized in that The CPU determines whether to accept the video encoding request according to the encoding capacity allocated to the virtual machine and the video encoding request, including: Based on the video encoding request, obtaining encoding capacity requirement information of the virtual machine; According to the encoding capacity allocated to the virtual machine and the encoding capacity currently consumed by the virtual machine, determining whether the remaining encoding capacity of the virtual machine meets the requirements of the encoding capacity demand information; If the judgment result is yes, determine to accept the video encoding request.

9. The method according to claim 6, characterized in that Also includes: The CPU monitors actual usage of encoding capacity by the virtual machine when determining to accept the video encoding request; When it is monitored that the actual usage is greater than the encoding capacity allocated to the virtual machine, the frame rate of the video encoding of the hardware encoder is reduced so that the actual usage of the encoding capacity by the virtual machine does not exceed the encoding capacity allocated to the virtual machine.

10. The method according to claim 5, characterized in that The hardware encoder comprises: a multi-level encoding unit; the encoding of the video to be encoded comprises: The multi-level encoding unit, driven by the CPU, processes the encoding tasks of multiple video frames in the video to be encoded in a pipeline task execution manner.

11. The method according to claim 6, characterized in that The host deploys multiple virtual machines; at least two of the multiple virtual machines concurrently send video encoding requests; The step of driving the hardware encoder to act based on the video encoding request includes: When the video encoding requests concurrently sent by the at least two virtual machines are accepted, the CPU schedules the encoding tasks of the at least two virtual machines based on a time-division multiplexing scheduling strategy; and drives the hardware encoder to perform video encoding on the video to be encoded of the currently scheduled virtual machine; The hardware encoder, driven by the CPU, encodes the video to be encoded of the virtual machine currently scheduled by the CPU.

12. The method according to claim 11, characterized in that The data processing module supports SR-IOV; the data processing module is virtualized into multiple virtual video devices by the multiple virtual machines through SR-IOV technology, and is allocated to the multiple virtual machines.

13. The method according to claim 6, characterized in that Also includes: The hardware encoder, driven by the CPU, uses the DMA engine to write the video encoding data corresponding to the video frame to be encoded into the memory of the virtual machine.

14. The method according to claim 6, characterized in that Also includes: The hardware processing unit receives an encoder destruction request sent by the virtual machine; the encoder destruction request is sent by the virtual machine after encoding of the video to be encoded is completed; Transmitting the encoder destruction request to the CPU through the control plane transmission channel; The CPU reclaims encoder resources corresponding to the virtual machine in response to the encoder destruction request.

15. A video encoding method, characterized in that: Applicable to a data processing module, the data processing module is connected to a host; The data processing module includes: a hardware processing unit and a CPU; the hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; the method includes: Obtaining a video encoding request sent by the host; Transmitting the video encoding request to the CPU through the control plane transmission channel; The CPU runs the firmware of the hardware processing unit and drives the hardware processing unit to act based on the video encoding request; The hardware processing unit, under the drive of the CPU, uses the DMA engine to obtain the video to be encoded corresponding to the video encoding request from the memory of the host; and uses the DMA engine to store the video to be encoded in the memory of the CPU; The CPU encodes the video to be encoded in the memory to obtain video encoding data corresponding to the video to be encoded.

16. A data processing module, characterized in that: include: Hardware processing unit and CPU; The hardware processing unit is connected to the CPU for communication; The hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; the hardware processing unit is additionally provided with a hardware encoder; when the data processing module is in communication connection with the host, the CPU is used to execute the steps in the method executed by the CPU in any one of claims 5-14; The hardware encoder is used to execute the steps in the hardware encoder execution method in any one of claims 5-14.

17. A data processing module, characterized in that: include: Hardware processing unit and CPU; The hardware processing unit is connected to the CPU for communication; The hardware processing unit includes: a direct memory access DMA engine and a control plane transmission channel; When the data processing module is in communication connection with the host, the CPU is used to execute the steps in the method executed by the CPU in claim 15; The hardware processing unit is used to execute the steps in the hardware processing unit execution method of claim 15.

Citation Information

Patent Citations

  • Image coding method and system

    CN107277538A

  • Software and hardware hybrid video coding on-card video coding acceleration system

    CN114501027A