Video processing system, method, device, electronic device and medium
By working in a coordinated manner with the first encoder in the heterogeneous video encoding system, the problem of high computing resource occupation when combining the hardware encoder and the software encoder is solved, and more efficient video encoding processing and better system adaptability are achieved.
Patent Information
- Application Number
- CN202411049763.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-08-01
AI Technical Summary
In the existing video encoding technology, the combination of hardware encoder and software encoder has the problem of high computing resource occupation and low encoding efficiency, especially when processing high resolution and high dynamic content, the system processing capability is insufficient.
Using a heterogeneous video encoding system, the hardware encoder works in concert with at least one first encoder. The hardware encoder extracts motion information and frame-level information, and processes it according to the target transmission protocol and sends it to the first encoder. The first encoder uses this information to perform further encoding processing to reduce repeated calculations.
It improves the speed and efficiency of video encoding processing, reduces computing resource consumption, and enhances the flexibility and adaptability of the system, especially when processing high-complexity video content, which significantly improves encoding performance.
Smart Images

Figure CN118870023B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a video processing system, method, device, electronic device and medium. Background Art
[0002] Video encoding is a crucial process in digital media production, which involves converting the original video file into a compressed format suitable for storage, streaming, and broadcasting. This technology is essential for enabling video content to be accessed and efficiently distributed on various platforms. The core of video encoding lies in its ability to balance quality, bitrate, and compatibility, ensuring that the video is small enough to be transmitted over limited bandwidth or stored on a device while maintaining high visual quality.
[0003] There are mainly two types of video encoders: hardware encoders and software encoders. A hardware encoder is a dedicated device or chipset that has the least impact on CPU resources when processing encoding, providing efficiency and speed, and is particularly suitable for applications with high real-time requirements, such as live streaming. A software encoder runs on a general-purpose computer, providing greater flexibility and a larger set of functions, allowing detailed customization of encoding parameters to meet specific quality requirements or implement new standard codecs. The choice between hardware and software encoders usually depends on factors such as cost, required quality, computing resources, and specific use cases. Although some technical solutions comprehensively utilize hardware encoders and software encoders, the computing cost of software encoders remains high, resulting in poor overall system processing capabilities. Summary of the Invention
[0004] The purpose of this application is to provide a video processing system, method, device, electronic device and medium, which can implement a solution to improve video encoding processing capabilities.
[0005] According to the first aspect of the embodiments of this application, a video processing system is provided, including:
[0006] The system includes: a hardware encoder, at least one first encoder; wherein, the hardware encoder is communicatively connected to the first encoder;
[0007] The hardware encoder is configured to obtain a target video, and perform encoding processing on the target video to obtain a first encoded video, motion information, and frame-level information;
[0008] The hardware encoder processes the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data; and sends the transmission data to the first encoder;
[0009] The first encoder is configured to perform encoding processing on the received transmission data and the first encoded video to obtain a second encoded video.
[0010] Optionally, the motion information includes at least one of: a reference frame and a reference frame index, a motion vector and a vector direction, a vector magnitude, a block resolution, a distortion cost, and a prediction unit;
[0011] Allocate at least one corresponding storage interface for the motion information according to at least one resolution, so as to obtain the motion information through the storage interface; wherein, the prediction unit includes multiple different size specifications, so as to provide the required prediction units for different first encoders.
[0012] Optionally, the frame-level information includes at least one of: a session identifier, a motion vector precision, a frame number, a frame type, a block size information, and a reference frame list;
[0013] The hardware encoder is configured to process at least one of the session identifier, the motion vector precision, the frame number, the frame type, the block size information, and the reference frame list according to the transmission protocol to obtain first-level data;
[0014] Process at least one of the reference frame and the reference frame index, the motion vector and the vector direction, the vector magnitude, the resolution, the distortion cost, and the prediction unit according to the transmission protocol to obtain second-level data;
[0015] Based on the first-level data and the second-level data, obtain the transmission data.
[0016] Optionally, when the hardware encoder outputs the reference frame list, it includes at least one of: a forward prediction list, a backward prediction list, and a bidirectional prediction list, and multiple different types of reference frames applicable to multiple first encoders.
[0017] Optionally, the first encoder is further configured to parse the received transmission data to obtain the first-level data and the second-level data;
[0018] Select a required target reference frame list from the first-level data, and select required prediction block information from the second-level data; so as to perform encoding processing on the first encoded video by using the target reference frame list and the prediction block information to obtain a second encoded video.
[0019] Optionally, the first encoder is further configured to determine a target prediction block size and a target prediction block shape corresponding to the first encoder type according to the received first encoded video;
[0020] If the target prediction block size and the target prediction block shape corresponding to the first encoder type are not found, the prediction block size and the prediction block shape are segmented to obtain the target prediction block size and the target prediction block shape corresponding to the first encoder type.
[0021] According to a second aspect of the embodiments of the present application, there is provided a video processing method, including:
[0022] Applied to a main controller or a hardware encoder in a heterogeneous video processing system, the method includes:
[0023] Encoding the target video through a hardware encoder to obtain a first encoded video, motion information, and frame-level information;
[0024] Processing the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data;
[0025] Sending the transmission data to a first encoder so that the first encoder uses the received transmission data and the first encoded video for encoding processing to obtain a second encoded video.
[0026] According to a third aspect of the embodiments of the present application, there is provided a video transmission method, including:
[0027] Applied to a main controller or a hardware encoder in a heterogeneous video system, the method includes:
[0028] Parsing the motion information provided by the hardware encoder to obtain a session, frame metadata, and motion information data;
[0029] Processing the obtained session, frame metadata, and motion information data according to a target transmission protocol matching the first encoder to obtain transmission data including first-level data and second-level data;
[0030] Transmitting the transmission data and the first encoded video encoded by the hardware encoder to the first encoder so that the first encoder uses the transmission data to perform encoding processing on the first encoded video.
[0031] According to a fourth aspect of the embodiments of the present application, there is provided a video processing device, including:
[0032] An encoding module, configured to encode the target video through a hardware encoder to obtain a first encoded video, motion information, and frame-level information;
[0033] A transmission data processing module, configured to process the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data;
[0034] A sending module, configured to send the transmission data to a first encoder, so that the first encoder performs encoding processing on the received transmission data and the first encoded video to obtain a second encoded video.
[0035] According to a fifth aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, where the memory is used to store a computer program executable by the processor; the processor is used to execute the computer program in the memory to implement the above method.
[0036] According to a sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. It is characterized in that when the executable computer program in the storage medium is executed by a processor, the above method can be implemented.
[0037] Compared with the prior art, the beneficial effects of the present application are as follows: In a video processing system, when there is a hardware encoder and at least one first encoder, the working processes of each encoder assist each other. In the system, the hardware encoder is the first encoder to perform encoding processing on the target video. During the processing, since the hardware encoder has sufficient hardware resources, it will not occupy the computing power and storage space of the controller. After the hardware encoder completes the encoding processing of the target video, a first encoded video can be obtained. In addition, the motion information generated by the hardware encoder during the encoding process is also provided externally. When the hardware encoder sends the first encoded video to at least one other first encoder, the motion information will be sent together. That is, during the encoding process, the first encoder can continue to use the motion information provided by the hardware encoder, without the first encoder having to recalculate the motion information, which can effectively reduce the computing power consumption of the first encoder during the encoding process and can effectively improve the encoding processing speed of the first encoder.
[0038] When the hardware encoder sends the motion information to the first encoder, conversion processing needs to be performed, that is, the motion information and frame-level information are converted according to the target transmission protocol, where the target transmission protocol is a protocol matching the first encoder. Thereby enabling the motion information and frame-level information to be successfully sent to the corresponding first encoder and received and processed by the first encoder. Using the target transmission protocol enables the hardware encoder to more smoothly perform data transmission with the first encoders of various reference structures, improving the general effect of the hardware encoder. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of motion estimation provided by an embodiment of the present application.
[0040] Figure 2Schematic diagram of sub-pixels provided by an embodiment of the present application.
[0041] Figure 3 Schematic diagram of coding block partitioning provided by an embodiment of the present application.
[0042] Figure 4 Schematic diagram of reference relationships provided by an embodiment of the present application.
[0043] Figure 5a Schematic diagram of a video processing system provided by an embodiment of the present application;
[0044] Figure 5b Schematic diagram of the structure of a heterogeneous video coding system illustrated by an embodiment of the present application;
[0045] Figure 5c Schematic diagram of an encoder architecture illustrated by an embodiment of the present application;
[0046] Figure 6 Schematic diagram of the process of a video processing method illustrated by an embodiment of the present application;
[0047] Figure 7 Schematic diagram of the process of a video transmission method illustrated by an embodiment of the present application;
[0048] Figure 8 Schematic diagram of the structure of a video processing device proposed by an embodiment of the present application;
[0049] Figure 9 Schematic diagram of the structure of a video transmission device proposed by an embodiment of the present application;
[0050] Figure 10 Block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0051] Unless otherwise defined, technical terms or scientific terms used in this specification and the claims should have the ordinary meanings understood by those of ordinary skill in the technical field to which the present invention belongs. The following will describe the specific implementation manners of the present invention with reference to the accompanying drawings. It should be noted that in the process of the specific description of these implementation manners, for the sake of concise description, this specification may not describe all the features of the actual implementation manners in detail. Without departing from the spirit and scope of the present invention, those skilled in the art can modify and replace the implementation manners of the present invention, and the obtained implementation manners are also within the protection scope of the present invention.
[0052] In modern video compression standards, inter-frame prediction coding is usually adopted, that is, by utilizing the correlation between frame images at different times, video compression is achieved through prediction coding. ME is the basis of inter-frame prediction. Inter-frame prediction finds the best match by performing complex ME operations in the reference frame, that is, finding an image block in the reference image that is most similar to the current coding block in the current frame, as follows Figure 1 is a schematic diagram of motion estimation provided by an embodiment of the present application.
[0053] ME performs motion search based on blocks in the reference frame, that is, within a certain search range, it moves to different positions to obtain corresponding reference blocks. By comparing the pixel differences between the reference block and the current coding block, a position with the smallest difference is found as the best position. The block corresponding to this position is used as the best reference block, and the offset corresponding to this position is the motion vector, as Figure 1 shown by the solid arrow in. Usually, the best position is not at the integer pixel position of the original image. Most of the time, it is at a certain position between two integer pixels. Therefore, before motion search, it is necessary to divide the blank area between integer positions into equally spaced fractional pixel positions and perform interpolation calculations on the fractional pixel positions so that the fractional pixel positions obtain corresponding pixel values, and then perform motion search based on this, which can greatly improve the search accuracy. As Figure 2 is a schematic diagram of fractional pixels provided by an embodiment of the present application. As Figure 2 shown, the pixel points represented by "o" are the pixel points at integer pixel positions, and the pixels represented by "x" are the fractional pixel positions.
[0054] The processing object of a video encoder is a sequence of video images. A video sequence is a set of multiple video images organized in chronological order. The encoder encodes according to frames, and after various divisions, decisions, and selections for each frame, it finally starts encoding with each coding unit CU. As Figure 3 is a schematic diagram of coding block division provided by an embodiment of the present application. For common current video encoders, the basic processing flow is as Figure 3 shown,
[0055] (1) In the first step, each frame image is divided into the largest coding blocks (LCBs) with a fixed size. Each LCB is a coding entry. There are many LCB coding blocks in a reference frame. Generally, the number of LCB coding blocks in a reference frame is the same as the number of integer pixels, that is, one LCB coding block is divided for each integer pixel.
[0056] (2) In the second step, each LCB is divided into coding units (CUs) of various sizes in a quadtree manner. An LCB will have multiple combinations of CU sizes, forming a coding tree unit (CTU).
[0057] (3) In the third step, each CU may be further divided into prediction units (PUs) of different shapes.
[0058] In Figure 3 the coding block division shown, except for the first step, the remaining two steps require rate-distortion optimization (RDO) to decide how to divide, the size of the CU to be divided into, and which PU mode to choose for each CU. When coding a CU / PU, image blocks in adjacent reference frames are needed as references. According to similarity, data compression is achieved through inter-frame prediction. There may be multiple reference relationships here. For each reference relationship, the RDO decision usually traverses and compares various possible divisions, selects the one with the minimum cost as the final coding mode, and motion estimation (ME) and sub-pixel interpolation therein are used. Moreover, different reference frames correspond to different reference relationships and different sub-pixel interpolations. The selection of the reference frame also needs to be decided by RDO. As Figure 4 is a schematic diagram of the reference relationship provided by an embodiment of the present application. As Figure 4 shown (the arrows in the figure represent a possible reference relationship during CU / PU coding). For the sub-pixel interpolation in ME, the common practice of the current encoder is that every time the RDO needs to use sub-pixels as references, the encoder will perform a sub-pixel interpolation. Although this is relatively flexible, it brings a large amount of redundant calculations and affects the speed of the encoder.
[0059] As Figure 5aSchematic diagram of a video processing system provided by an embodiment of the present application. The system includes a hardware encoder and at least one first encoder; wherein, the hardware encoder is communicatively connected to the first encoder. The hardware encoder is configured to obtain a target video and perform encoding processing on the target video to obtain a first encoded video, motion information, and frame-level information. The hardware encoder processes the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data; and sends the transmission data to the first encoder. The first encoder is configured to perform encoding processing on the received transmission data and the first encoded video to obtain a second encoded video.
[0060] Among them, the first encoder can be a hardware encoder or a software encoder. In a heterogeneous video encoder system, the number of first encoders can be one or more. And, the hardware encoder is communicatively connected to the first encoder. The frame-level information mentioned here includes at least one of: session identifier, motion vector precision, frame number, frame type, block size information, and reference frame list.
[0061] Hardware encoders are generally faster than software encoders because they are dedicated chips designed specifically for encoding. They can process video in real-time or faster without burdening the CPU. Due to using dedicated hardware encoding, the impact on CPU resources is minimal, which is crucial for multitasking or real-time streaming. They are generally more energy-efficient, which is particularly beneficial for mobile or embedded devices. And now power consumption is also a very concerned issue in data centers. Hardware encoders generally provide consistent performance and quality because they are less dependent on system resources and software configuration. At the same time, due to the fixed configuration of the hardware encoder, there may be a lack of flexibility to update or adjust advanced settings. Compared with software encoders, the upfront cost may be higher, especially in cases where external hardware or dedicated devices are required. Moreover, the development investment and cycle are also much higher. Most hardware encoders may not be able to provide the same quality as software.
[0062] On the other hand, software encoders are very flexible, allowing updates, customization, and adjustment of codec parameters to optimize quality or meet specific requirements. In short, the choice between hardware and software encoders usually depends on specific needs and environments. Heterogeneous video encoder design refers to the combination of different processing units such as using a CPU and a dedicated hardware encoder, and even can include other types of processors such as GPUs to optimize the video encoding process. This design utilizes the advantages of each processor to achieve higher efficiency, better performance, and improved power consumption characteristics.
[0063] By assigning different tasks to the most suitable processing units, heterogeneous encoding can achieve higher performance levels than single-processor systems. This setup allows for real-time encoding of high-definition videos without significant latency. Each component in the heterogeneous encoder can operate independently at its optimal power level. For example, less demanding tasks can be handled by the CPU, while more intensive tasks can be transferred to dedicated hardware, thus minimizing overall power consumption. This design allows the encoder to be customized according to specific requirements and adjusted for different scenarios. Although the initial setup may be more complex and costly, the long-term benefits of reduced operating costs and the ability to handle more streams, higher resolutions, and even newer video coding standards without the need to upgrade the entire system may outweigh these initial expenses.
[0064] In this system architecture, the hardware encoder communicates with at least one first encoder (which can be another hardware encoder or a software encoder) to optimize the video encoding process. The specific working process of the system is as follows:
[0065] The system is designed as a heterogeneous video encoding system, which includes at least one hardware encoder and one or more first encoders (which can be hardware encoders or software encoders). The hardware encoder and the first encoder are connected through communication and jointly participate in the video encoding process.
[0066] The hardware encoder is responsible for receiving the target video stream and performing preliminary encoding processing on it. During this process, the hardware encoder not only generates the first encoded video but also extracts key motion information and frame-level information. Motion information refers to data related to the movement of objects or regions between video frames, such as motion vectors, distortion costs, etc. Frame-level information contains a wider range of metadata, such as session identifiers, motion vector precision, frame numbers, frame types, block size information, and at least one in the reference frame list, which is used to provide context and control information.
[0067] The hardware encoder processes the motion information and frame-level information according to the target transmission protocol that matches the first encoder. The transmitted data here actually includes detailed motion data including motion information, as well as frame-level information used to synchronize and guide the first encoder on how to use this data. Through the hierarchical data transmission protocol, the motion information (level 2 data) and frame-level information (level 1 data) are efficiently packed and compressed to reduce bandwidth requirements and improve transmission efficiency. In addition, to protect sensitive content, the transmitted data may also be encrypted.
[0068] After the first encoder receives the transmission data, it will use the motion information and the first encoded video therein for further encoding processing to generate an optimized second encoded video. The first encoder can utilize the motion information provided by the hardware encoder to accelerate its own motion estimation process, reduce the search area, optimize prediction and mode decision, thereby improving the encoding efficiency and video quality.
[0069] In a heterogeneous system, multiple first encoders can work in parallel, and each encoder selectively utilizes the motion information and frame-level information provided by the hardware encoder according to its own requirements and capabilities. This design improves the flexibility and efficiency of the system. Especially when dealing with high-resolution or high-dynamic content, it can significantly reduce the consumption of computing resources and improve the encoding speed.
[0070] The first encoder utilizes the motion information of the hardware encoder, reduces the computational requirements of motion estimation, and improves the encoding efficiency. The system supports multiple first encoders and can dynamically adjust the reference list and block size according to requirements to adapt to different encoding strategies and video frame complexities. Through the pre-computed motion information, the first encoder can focus more on mode decision and other encoding tasks, improving the video compression efficiency.
[0071] Through encryption and error detection mechanisms, it ensures the secure transmission of sensitive content, and at the same time guarantees the integrity and reliability of data transmission through a hierarchical protocol.
[0072] Through the above details, this system architecture can not only effectively improve the video encoding efficiency, but also enhance the flexibility and adaptability of video processing, especially showing significant advantages when dealing with high-complexity video content. The system architecture and functions are more comprehensively described, demonstrating the potential of the efficient cooperation between the hardware encoder and the first encoder, as well as the possibility of achieving performance optimization in the field of video encoding.
[0073] For the sake of easy understanding, a schematic diagram of a heterogeneous video encoding system will be illustrated by specific embodiments. As Figure 5b is the structural schematic diagram of the heterogeneous video encoding system illustrated by the embodiments of this application. As can be seen from Figure 5b it,
[0074] After the hardware encoder (HW Encoder) performs encoding processing to obtain the first encoded video, motion information, and frame-level information, it sends the corresponding transmission data to the relevant host CPU via the system bus (system BUS). The motion information, frame-level information, etc. received are then processed by the host CPU to generate transmission data (motion information) suitable for multiple first encoders. After receiving the transmission data, the first encoder can directly use the received transmission data (including motion information, residuals, frame-level information, etc.) to perform subsequent encoding operations, which can significantly improve the encoding efficiency and reduce unnecessary calculations of the software encoder.
[0075] In one or more embodiments of the present application, the motion information includes at least one of: a reference frame and a reference frame index, a motion vector and a vector direction, a vector magnitude, a block resolution, a distortion cost, and a prediction unit. Different storage interfaces are allocated for the motion information according to different resolutions so as to obtain the motion information through the storage interfaces; wherein, the prediction unit includes various different size specifications in order to provide the required prediction unit for different first encoders.
[0076] As mentioned above, the motion information includes a lot of content, and the relevant content will be specifically explained.
[0077] A reference frame and a reference frame index. The reference frame is a previous frame used to predict the content of the current frame in video coding. The reference frame index is the position identifier of a specific frame in the reference frame list, enabling the encoder to clearly call the correct frame as the basis for prediction.
[0078] The motion vector describes the displacement of a block in the current frame relative to the corresponding block in the reference frame. The vector direction represents the direction of motion, and the vector magnitude represents the distance of displacement. These information are crucial for predicting the motion of the block.
[0079] The block resolution refers to the size of the video block to which the motion vector is applied, such as different sizes like 8x8, 16x16, 32x32, etc., which directly affects the accuracy of motion estimation and the encoding efficiency.
[0080] The distortion cost is an index for evaluating the prediction error after using a specific motion vector to predict a block, and it is crucial for optimizing encoding decisions and bitrate control.
[0081] The prediction unit (PU) is the basic unit for performing inter-frame prediction in video coding. It can have different size specifications, such as 2Nx2N, NxN, 2NxN, Nx2N, etc., and even include asymmetrically divided PUs, such as 2NxnU, 2NxnD, nLx2N, nRx2N, etc., to adapt to different prediction requirements.
[0082] To meet the different requirements of different first encoders for motion information, a dedicated storage interface is adopted in the solution of this application, and storage space is allocated for motion information according to different resolution specifications. This is because different first encoders (whether hardware encoders or software encoders) may require different specifications of motion information based on their processing capabilities, codec types, and encoding strategies. For example, one encoder may prefer to use higher-resolution motion vectors to obtain more refined predictions, while another encoder may prefer lower-resolution information due to performance limitations or encoding strategy choices.
[0083] By allocating a dedicated storage interface for motion information of different resolutions, the system can ensure that the first encoder can quickly and accurately access the data that matches their processing capabilities. This not only improves the data transmission efficiency but also simplifies the processing flow of the first encoder for motion information, avoiding unnecessary data conversion and adaptation, thereby enhancing the performance and response speed of the entire encoding system.
[0084] By supporting prediction units of multiple sizes and motion information of different resolutions, the solution of this application can adapt to a wide range of requirements of the first encoder. Whether they are based on hardware or software, they can effectively utilize this information. Moreover, the design of the storage interface ensures fast access and efficient transmission of data, reducing the latency of data processing, and thus improving the efficiency and overall performance of video encoding. By providing a dedicated storage interface for motion information of different resolutions, the system can more reasonably allocate and utilize storage resources, avoid resource waste, and at the same time reduce the processing burden on the first encoder. This design allows the system to easily integrate more first encoders. Whether adding hardware encoders or software encoders, they can be smoothly connected to the system without causing bottlenecks in system performance. In summary, the design of motion information processing and storage interface in the solution of this application greatly enhances the flexibility, efficiency, and compatibility in the field of video encoding.
[0085] Such as Figure 5c is a schematic diagram of the encoder architecture illustrated by an example of an embodiment of this application. From Figure 5cAs can be seen, in this architecture, after the hardware encoder receives the original video, it will perform intra-frame encoding (Intra) and motion estimation processing. After completing intra-frame encoding, mode selection / decision (Mode Decision, MD) can be performed to determine the optimal encoding mode when encoding a video block. In video encoding, each video frame is divided into multiple blocks, and the size of each block can be fixed (such as 16x16 pixels) or variable (such as the macroblock partitioning mode in H.264). For each block, the encoder needs to decide which encoding mode to use to achieve the best compression efficiency and quality. After selecting the encoding mode, further video compression processing will be carried out. VLE (Variable Length Encoding) generally refers to the entropy coding technology in the field of video encoding or data compression. Further, the video can be further processed using an in-loop filter (In-Loop Filtering, ILF) to improve the quality of the encoded video. Next, a decoded image buffer will be used to store the decoded video frames, and these frames can be used as reference frames for the predictive encoding of subsequent frames. In video encoding, especially when using inter prediction technology, the role of the DPB is particularly crucial. After ME / MC, a new storage interface connection needs to be added to the system to write out the motion data and its codec-independent related metadata for each video codec module at different resolutions.
[0086] In one or more embodiments of the present application, the frame-level information includes at least one of a session identifier, motion vector precision, frame number, frame type, block size information, and reference frame list;
[0087] The hardware encoder is configured to obtain first-level data by processing at least one of the session identifier, motion vector precision, frame number, frame type, block size information, and reference frame list according to the transmission protocol;
[0088] Obtain second-level data by processing at least one of the reference frame and reference frame index, motion vector and vector direction, vector magnitude, resolution, distortion cost, and prediction unit according to the transmission protocol;
[0089] Based on the first-level data and the second-level data, obtain the transmission data.
[0090] Among them, in the frame-level information, the session identifier: is used to distinguish and manage different encoding sessions to ensure the correct classification and tracking of data.
[0091] Motion Vector Precision: Defines the resolution of the motion vector, indicating the number of bits for each component (horizontal and vertical) of the motion vector, such as 16 bits.
[0092] Frame Number: Assigns a unique identifier to each frame for easy tracking and reference.
[0093] Frame Type: Indicates whether the frame is an I-frame, P-frame, or B-frame, affecting the correlation and structure of the motion information data.
[0094] Block Size Information: Outlines the block sizes used in the current frame, as the motion information data may vary with the block size.
[0095] Reference Frame List: Provides information about the number of reference lists and their identifiers, helping to interpret the motion information related to past or future frames.
[0096] In motion information, the reference frame and reference frame index: Indicate the reference frame used to predict the current frame and its position in the reference list.
[0097] Motion Vector and Vector Direction, Vector Magnitude: Details the direction, horizontal and vertical motion vector components, and the magnitude of each motion vector in each direction.
[0098] Resolution: Specific data for each block size to ensure that the motion vector is correctly associated with the appropriate block in the frame.
[0099] Distortion Cost: Contains the final distortion metric value of the best motion vector after motion estimation.
[0100] Prediction Unit: Covers PUs of different size specifications to meet the prediction requirements of different encoders.
[0101] In the solution of this application, the transmission and processing of motion information are designed into two levels, aiming to improve the efficiency, flexibility, and accuracy of video coding. The generation process of each level of data and why a hierarchical architecture is adopted are elaborated in detail below. When generating the first-level data (session and frame metadata), the communication between the hardware encoder and the host is initialized, and a session identifier is created at this time. Subsequently, during the execution of Motion Estimation (ME) and Motion Compensation (MC) by the hardware encoder, a series of motion vectors and other related parameters are generated. These parameters include the precision of the motion vector (in bits), the frame number, and information about the current frame type (I, P, or B frame), block size, codec type, and reference frame list. All these metadata constitute the first-level data, which is a high-level description of the video coding session and provides context for the subsequent transmission of motion information.
[0102] Second-level data (motion information data): At the ME / MC stage of the hardware encoder, the motion vector information of each block is extracted, including the direction, magnitude, resolution / block size, reference frame index, and distortion cost of the vector. These specific motion vector information constitute the second-level data, which are the core of the motion prediction and compensation process in video coding. The data at this level describes in detail the motion of each block and is the key to video compression.
[0103] After the hardware encoder generates the first-level and second-level data, these data are packed according to the OpenME protocol. The packing process includes adding a clear header, a payload part, and a checksum for error checking to ensure the integrity and easy parsing of the data. After the data packing is completed, the final transmission data is formed and ready to be sent to the host through DMA (Direct Memory Access) or other high-speed data transmission mechanisms.
[0104] Based on the above scheme, the main advantages of the hierarchical data transmission architecture using the OpenME protocol are as follows: The low-bandwidth characteristic of the first-level data enables it to be transmitted quickly and provide the necessary context information, while the high-bandwidth characteristic of the second-level data details the motion information and reduces the transmission bandwidth requirement through compression technology.
[0105] The hierarchical structure allows the protocol to adapt to different codecs, block sizes, and reference list numbers, ensuring the efficient transmission of information while maintaining the accuracy and validity of the data. By implementing checksum mechanisms such as CRC in the first-level data, the reliability of data transmission is ensured, and the final CRC check of the second-level data guarantees the accurate transmission of all motion information.
[0106] In summary, the multi-level data processing and transmission scheme proposed in this application, through a carefully designed hierarchical architecture, not only optimizes the efficiency and reliability of data transmission, but also improves the overall flexibility and adaptability of the video coding system, providing a solid technical foundation for the efficient cooperation of heterogeneous video encoders.
[0107] In one or more embodiments of this application, when the hardware encoder outputs the reference frame list, it includes at least one of the forward prediction list, the backward prediction list, and the bidirectional prediction list, as well as multiple different types of reference frames applicable to multiple of the first encoders.
[0108] In video coding, Reference Picture Lists (RPLs) play a central role, especially in inter-frame prediction. Inter-frame prediction improves compression efficiency by leveraging previous or future frames to predict the content of the current frame. Reference frame lists typically consist of two main lists: List 0 (L0) and List 1 (L1), which are used for backward prediction and forward prediction respectively, and in some cases for bidirectional prediction.
[0109] Among them, the forward prediction list (L1) contains reference frames used to predict the current frame based on future frames. These frames are decoded after the current frame but are available during the decoding process. The L1 list is typically used for forward prediction but can also be used in combination with the L0 list for bidirectional prediction.
[0110] The backward prediction list (L0) contains reference frames used to predict the current frame based on past frames. These frames are previously decoded frames that the codec can use to estimate motion and predict the current frame. The L0 list is mainly used for backward prediction.
[0111] The bidirectional prediction list (B-frame) When a combination of the L0 and L1 lists is used to predict the current block, bidirectional prediction or B-frames are produced. This helps achieve better compression efficiency because the codec can select the best reference frames from past and future frames to reduce prediction errors.
[0112] Using only the L0 list This is the most common prediction method and is suitable for most video coding scenarios, especially when the motion in the video sequence is not very complex. It exploits the temporal correlation of the video to predict the current frame by referring to past frames.
[0113] Using only the L1 list Although less common, in some special cases, such as fast-forward or rewind playback of a video sequence, future frames may be needed as references to predict the current frame.
[0114] Bidirectional prediction using the L0 and L1 lists This is one of the most efficient prediction methods in video coding because it combines past and future reference frames to predict the current frame, thereby reducing prediction errors and improving compression efficiency. Bidirectional prediction is especially suitable for video sequences with complex motion, where a single reference frame may not be sufficient to accurately predict the changes in the current frame.
[0115] As mentioned above, in a heterogeneous video coding environment, the first encoder can refer to any hardware or software encoder involved in the video coding process. Hardware encoders are typically responsible for fast encoding, while software encoders provide more refined control and higher coding quality. Providing multiple types of reference frame lists provides the necessary information for the first encoder for inter-frame prediction, thereby improving the coding efficiency and video compression ratio.
[0116] Different types of the first encoder may have different performance metrics and coding parameter requirements. For example, they may support different codec standards (such as H.264, H.265 / HEVC, etc.), and these standards have slight differences in the use of reference frames. By providing a list containing multiple types of reference frames, the encoder can select the reference frame that is most suitable for its current coding task. For example, one encoder may prefer to use bidirectional prediction to improve compression efficiency, while another encoder may focus more on real-time performance and thus only use backward prediction.
[0117] For example, assume there is a video sequence that contains fast-moving objects and a complex background. In this case, using bidirectional prediction (B-frames) would be the best choice because it can select the best reference frames from past and future frames to predict the motion of the current frame, thereby reducing prediction errors and improving compression efficiency. The hardware encoder can generate motion information containing L0 and L1 lists, and then the software encoder can utilize the information in these lists for more accurate motion estimation and prediction, thus optimizing the coding quality and efficiency.
[0118] In this way, even in complex video content, the encoder can effectively utilize the reference frame list to achieve the established coding goals, whether it is to improve real-time performance, compression efficiency, or video quality.
[0119] In one or more embodiments of the present application, the first encoder is further configured to parse the received transmission data to obtain the first-level data and the second-level data;
[0120] Select the required target reference frame list from the first-level data, and select the required prediction block information from the second-level data; so as to perform coding processing on the first encoded video by using the target reference frame list and the prediction block information to obtain a second encoded video.
[0121] In a heterogeneous video coding system, the efficient cooperation between the first encoder and the hardware encoder depends on the fine management of data by the OpenME protocol and the intelligent utilization of data by the first encoder.
[0122] The OpenME protocol encapsulates data into a standard format for easy transmission and parsing. The data packet may include header information (such as session ID, frame type, block size, etc.) and payload information (motion vectors, distortion costs, etc.).
[0123] To ensure accurate data transmission, the OpenME protocol may include CRC checks or other error detection methods, as well as a retransmission mechanism to handle loss or corruption during data transmission.
[0124] Considering bandwidth limitations, the OpenME protocol may compress motion information. After receiving the data, the first encoder needs to perform decompression to restore the original motion information.
[0125] The first encoder may use machine learning algorithms to analyze video content and automatically select the optimal reference frame to minimize encoding distortion and maintain encoding efficiency. Based on real-time analysis of video content, the first encoder can dynamically adjust encoding parameters such as quantization parameter QP, resolution, frame rate, etc., to optimize the balance between encoding quality and bandwidth. To accelerate the encoding process, the first encoder may utilize the parallel processing capabilities of a multi-core CPU to split video frames and perform multiple encoding tasks simultaneously. Although the hardware encoder already provides motion information, the first encoder may implement a fast motion estimation algorithm to fine-tune the motion vectors and further improve prediction accuracy.
[0126] The first encoder may adopt an adaptive quantization strategy to dynamically adjust the quantization parameter according to the complexity of the video content to find the best balance between encoding quality and file size. Utilize efficient entropy coding techniques such as arithmetic coding or context-adaptive binary arithmetic coding (CABAC) to reduce the bit rate without sacrificing video quality.
[0127] Through the comprehensive application of the above technical implementation methods, the first encoder can effectively utilize the data provided by the OpenME protocol to achieve high-quality and high-efficiency video encoding while maintaining the flexibility, scalability, and security of the system.
[0128] In one or more embodiments of the present application, the first encoder is further configured to determine a target prediction block size and a target prediction block shape corresponding to the first encoder type according to the received first encoded video; if the target prediction block size and the target prediction block shape corresponding to the software encoder type are not found, then split the prediction block size and the prediction block shape to obtain the target prediction block size and the target prediction block shape corresponding to the software encoder type.
[0129] In a heterogeneous video coding environment, after receiving the motion information provided by the hardware encoder through the OpenME protocol, the first encoder (such as a software encoder) needs to perform a series of processes to ensure the compatibility of the motion information with the prediction block size and shape of the software encoder. Specifically:
[0130] Initialization and configuration: The first encoder first establishes an encoding session with the hardware encoder and sets the hardware encoder parameters to match the use of the software encoder, including reference structure, frame rate, image size, etc. Set the hardware encoder through session-level information to use the same input as the software encoder to ensure the applicability of the motion information.
[0131] Receive the first - level data, including metadata such as session identifier, motion vector precision, frame number / identifier, codec type, frame type, block - size information, and reference - frame list. Parse this information to understand the structure of the video session and the context of the motion information, preparing for processing the motion information.
[0132] Receive the second - level data, including specific motion - information data, including vector direction and magnitude, resolution / block - size, reference - frame index, and distortion cost. These data will be used in the video - encoding process of the first encoder to provide an initial estimate of the motion field of the frame being encoded.
[0133] The first encoder dynamically adjusts the prediction - block size and shape. Specifically, the first encoder parses the prediction - block size and shape information received from the hardware encoder. Determine whether these prediction blocks meet the requirements of the codec standard of the software encoder and the current encoding strategy.
[0134] In practical applications, if the prediction - block size or shape provided by the hardware encoder does not match the standard of the software encoder, the first encoder needs to perform splitting or merging operations to meet the requirements of the software encoder. The splitting operation can break a larger prediction block into smaller blocks supported by the software encoder, while the merging operation does the opposite, combining multiple small blocks into a large block to meet the prediction - block size requirements of the software encoder.
[0135] In addition, if the first encoder is a software encoder, it can dynamically adjust the selection of its reference - frame list and the processing strategy of the block size to adapt to different encoding requirements, such as motion complexity or scene changes. This dynamic - adjustment ability ensures that even if there are differences in the block size and shape between the hardware encoder and the software encoder, efficient information reuse and encoding performance can be achieved.
[0136] Suppose the hardware encoder provides prediction - block sizes of 32x32, 16x16, and 8x8 pixels, while the encoding standard or strategy of the software encoder requires prediction - block shapes of 16x8 or 8x16. In this case, the first encoder needs to split the 16x16 prediction block of the hardware encoder into two 8x16 or 16x8 prediction blocks to meet the requirements of the software encoder.
[0137] Through the above - mentioned technical implementation process, the first encoder can flexibly process the motion information from the hardware encoder, ensuring that it conforms to the prediction - block size and shape requirements of the software encoder. This dynamic adaptability improves the encoding efficiency, reduces unnecessary consumption of computing resources, and also enhances the flexibility and compatibility of the software encoder in the heterogeneous video - encoding system.
[0138] In addition, the desired target prediction block can be derived using an algorithm model: The first encoder may need to process a prediction block with a size of B and a rectangular or square shape. If B is not the optimal prediction block size or shape, the first encoder will split B according to certain rules or algorithm models. The splitting can be performed along the horizontal or vertical direction, generating two or more sub-blocks. The goal of the splitting is to find smaller prediction blocks whose size and shape are closer to the optimal configuration supported by the first encoder type. For each generated sub-block, the first encoder will check again whether the optimal configuration has been found. If not, continue to split the sub-block until the minimum prediction block size supported by the encoder type is reached or other termination conditions are met (such as the split blocks no longer provide additional rate-distortion gain). During the splitting process, the encoder continuously evaluates the rate-distortion performance of each prediction block. This typically involves calculating the number of encoded bits (rate) for each block and the distortion degree between the reconstructed image and the original image (such as MSE, Mean Squared Error). The goal is to find a balance point, that is, to minimize the number of encoded bits while meeting the distortion requirements. The algorithm model may be based on empirical formulas, machine learning models, or complex mathematical optimization algorithms. For example, an optimization algorithm based on gradient descent can be used to minimize the rate-distortion function and find the optimal prediction block configuration. Machine learning models can also be trained based on a large amount of encoding data to predict the best prediction block size and shape in different scenarios. This selection of prediction block size and shape based on the algorithm model allows the encoder to automatically adjust in different scenarios to achieve the best encoding efficiency.
[0139] Based on the same idea, an embodiment of the present application also proposes a video processing method. This method can be applied to the main controller or hardware encoder in a heterogeneous video processing system. As Figure 6 is a schematic flowchart of a video processing method illustrated by an embodiment of the present application. As can be seen from Figure 6 it, the method specifically includes:
[0140] Step 601: Encode the target video through a hardware encoder to obtain a first encoded video, motion information, and frame-level information.
[0141] Step 602: Process the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data.
[0142] Step 603: Send the transmission data to the first encoder so that the first encoder uses the received transmission data and the first encoded video for encoding processing to obtain a second encoded video.
[0143] The first encoder can be a hardware encoder or a software encoder. Here, it is assumed that the first encoder is a software encoder. The software encoder first obtains motion vector information by establishing an encoding session with the hardware encoder and setting the hardware encoder parameters to match its usage (similar reference structures, frame rates, image sizes, etc.). Whether this information is used by one or more software encoder instances will depend on the final application scenario (i.e., multi-codec or multi-resolution encoding). In the case of multiple software encoders, corresponding storage interfaces need to be allocated for each type of software encoder so that different types of transmitted data (including related data such as motion information and residuals) can be stored separately, and the hardware encoder output can meet the encoding requirements of multiple different types of first encoders simultaneously.
[0144] After exchanging the initial system setting information, the hardware encoder and the software encoder can be matched to have a consistent reference structure. The hardware encoder starts sending motion vector data to the software application. The data is divided into two layers: (i) frame-level information, dozens of bytes per frame; (ii) and motion information (motion vectors and distortion / cost), dozens of MB per frame. The software encoder or other client applications must support the new protocol (OpenME) and API to connect to the hardware encoder in order to effectively use the motion information in its internal processes. The information flow is regulated by the firmware layer at the system level. If the software encoder lags too much behind the hardware encoder, the system control will adjust the hardware encoder processing speed so that its motion information input buffer will not overflow.
[0145] The motion information exported by the hardware in the software encoder can effectively improve the overall encoding effect. For example:
[0146] Use the transmitted data provided by the hardware encoder to accelerate the video encoding process. The motion information from the hardware encoder is used to provide seeds for the motion search and motion search refinement processes within the software encoder. This reduces the number of searches and resources allocated for motion estimation and refinement within the software encoder. In some usage cases, even if the hardware encoder does not generate all the block sizes supported by the software encoder, the information derived from the smaller resolution / block size motion map of the hardware encoder can be used to approximate the motion estimation of the larger block sizes in the software encoder. If the content has high motion, and the software encoder has to invest a large amount of computing resources to cover the frame areas affected by the content motion and provide similar compression performance (in terms of the bits generated per content), then the overall acceleration of the software encoding process can be quite significant for the frame / sequence. Additionally, in most software encoder implementations, the complexity of the motion estimation stage is comparable to that of the actual encoding stage (mode decision process), so saving processing cycles will help improve the density of the software encoder running in a heterogeneous ASIC+CPU system.
[0147] Improve video compression efficiency by using the transmission data provided by the hardware encoder. As a control point for accelerating the video encoding process to obtain better real-time performance / increasing the density of system-level encoder instances in the system CPU, the motion information from the hardware encoder can be used to expand the number of hypotheses tested or refined in the software encoder stage. Compared with the non-heterogeneous case of only CPU software encoding, there are different trade-offs between compression gain and execution speed. In this case, the motion information map can be used to generate a likelihood map of the frame region. Compared with traditional grid-based motion tracking methods, more software search resources can be applied and the block matching quality (reducing the matching distortion cost) is most likely to be improved. This type of application is particularly helpful for improving the compression efficiency of sequences with complex motion and multiple moving objects being occluded. Before performing local refinement search using the CPU, all problems can be alleviated through high-level analysis of the motion map at the global level.
[0148] Use the transmission data provided by the hardware encoder to boost the software video frame rate (Frame Rate UpConversion, FRUC). Another software process that can use the motion information of the hardware encoder (mainly the motion vector map) is the video frame rate amplifier. The frame rate boosting algorithm uses the motion map to help track the moving regions between frames as one of the first stages of its processing pipeline. Track the moving blocks / pixels before applying further algorithm stages to interpolate or predict additional frames to be inserted into the sequence to increase the frame rate (the upscaling process). Using the motion map from the hardware encoder will help accelerate the motion estimation stage in the frame rate upconverter pipeline, thus obtaining better performance in terms of throughput for large frame sizes and helping to test more hypotheses for frame interpolation / prediction. Frame rate upconversion can also use the motion vector map to derive multiple hypotheses of objects / pixels for cross-frame interpolation, helping the algorithm handle occlusions and motion boundary regions that can be identified from the dense motion vectors / information map exported by the hardware encoder.
[0149] Use the transmission data provided by the hardware encoder to facilitate more accurate motion tracking in video analysis applications. Object tracking in surveillance application motion analysis is another application that uses motion information (the motion vector map) as part of its input pipeline. The motion information of the hardware encoder can help provide a relatively dense block-level motion map, which can be used to track moving objects in the scene or mark motion regions, segment the foreground and background, compensate for camera motion, improve the overall object tracking accuracy, and reduce the processing load. In object tracking in heterogeneous systems, the tracking algorithm itself focuses on target mapping feature extraction rather than pixel-level processing inside the CPU.
[0150] Based on the same idea, an embodiment of this application also proposes a video transmission method. This method can be applied to the main controller or the hardware encoder in a heterogeneous video system. As Figure 7The flowchart of a video transmission method illustrated by an embodiment of the present application. From Figure 7 it can be seen that the method specifically includes:
[0151] Step 701: Parse the motion information provided by the hardware encoder to obtain a session, frame metadata, and motion information data.
[0152] Step 702: Process the obtained session, frame metadata, and motion information data according to the target transmission protocol matched by the first encoder to obtain transmission data including first-level data and second-level data.
[0153] Step 703: Transmit the transmission data and the first encoded video encoded by the hardware encoder to the first encoder so that the first encoder uses the transmission data to perform encoding processing on the first encoded video.
[0154] In practical applications, motion information is generated within the hardware encoder as part of the motion compensation and motion estimation encoding stages, where inter-frame mode candidates are derived for the rate-distortion optimization stage of the hardware encoder. Motion estimation is a general process that is effectively performed as block-matching motion estimation in the hardware encoder, where the source block content is compared with candidate blocks at a given target region / range. The matching block with the minimum pixel-based distortion (SSD (sum of squared differences) or SAD (sum of absolute differences)) and the minimum bit rate (estimated based on the length of the motion vector).
[0155] Generally, the final calculation of the best block matching decision involves a combination of distortion and number of bits weighted by a scaling parameter derived from the current quantization factor of the encoder.
[0156] In the OpenME case proposed in the present application, the OpenME protocol can transmit the motion information generated during the execution of the hardware encoder to the required first encoder. Compared with the heterogeneous system of traditional multi-encoders, this solution will bypass the final decision and send the original distortion and motion vector information for other hardware / software encoders to use / adjust, so as to derive / improve the decision-making process of the software encoder itself based on the original data provided by the hardware encoder motion.
[0157] Transferring this raw motion estimation information from the hardware encoder to the software encoder first requires establishing an encoding session between the client (software encoder) and the server (controller firmware for the hardware encoder). Session-level information will be used to configure the hardware encoder to use the same input as the software encoder (i.e., there may be use cases where multiple software encoder instances work on the same input frame, outputting bitstreams for different codecs). The hardware encoder responds to the software encoder with the configuration of the raw motion information it expects for the upcoming frames. During the hardware encoder setup process, which depends on the software encoder implementation, it is necessary to match the reference structures of the software and hardware encoders to effectively use the motion information. These are the responsibilities of the encoder firmware / control layer. When performing motion compensation, the software encoder and the hardware encoder will have different engines, producing different bitstreams and reconstructed images. Therefore, the motion information derived from the hardware encoding will be used as seed information to accelerate the software encoder motion estimation process. Additionally, to reduce the difference between the reconstructed information used for motion compensation by the software encoder and the hardware encoder, the hardware encoder can perform motion estimation using the original source image as a better starting point for the software encoder motion compensation process and increase the system's flexibility, thereby allowing it to be codec-independent. In the software encoder implementation, modified to handle the external motion vector information provided by the hardware encoder, the motion vector information will be effectively utilized to provide an initial estimate of the motion field of the frame being encoded. The modified software encoder can utilize the generally larger search area provided by the hardware encoder to improve its own motion estimation process. This improvement can save CPU cycles by performing reduced motion estimation in software on the frame regions most likely to match a given block well. It can also allow the software encoder itself to prioritize search resources on frames with multiple search centers / locations based on the matching likelihood derived from the hardware encoder data and reduce the search area in different resolution block sizes, where the hardware encoder data effectively covers a larger search area. This exemplary technique will allow the software encoder to improve its compression efficiency, especially for high-motion content. The motion information from the encoder hardware encoder can also be used for analysis in the temporal importance algorithm, providing the software encoder with information on the frequency of reuse / reference of a given block in the original image over time, thus saving the software encoder from performing this operation during its own analysis before actual encoding in the motion estimation process. The main criterion is a similar reference structure between the software and hardware encoders to facilitate information reuse and save CPU processing time. It also assumes that the software and hardware encoders will have a similar or overlapping range of block sizes for motion estimation, so the results can be effectively reused without further approximation.
[0158] Based on the same idea, an embodiment of this application also proposes a video processing device. As Figure 8The figure is a schematic structural diagram of a video processing device proposed in an embodiment of the present application. As can be seen from Figure 8 the figure, the device includes:
[0159] An encoding module 81, configured to perform encoding processing on a target video through a hardware encoder to obtain a first encoded video, motion information, and frame-level information.
[0160] A transmission data processing module 82, configured to process the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data.
[0161] A sending module 83, configured to send the transmission data to the first encoder, so that the first encoder performs encoding processing on the received transmission data and the first encoded video to obtain a second encoded video.
[0162] Based on the same idea, an embodiment of the present application also proposes a video transmission device. As Figure 9 shown in the figure which is a schematic structural diagram of a video transmission device proposed in an embodiment of the present application. As can be seen from Figure 9 the figure, the device includes:
[0163] An analysis module 91, configured to analyze the motion information provided by the hardware encoder to obtain a session, frame metadata, and motion information data.
[0164] A processing module 92, configured to process the obtained session, frame metadata, and motion information data according to a target transmission protocol matching the first encoder to obtain transmission data including first-level data and second-level data.
[0165] A transmission module 93, configured to transmit the transmission data and the first encoded video encoded by the hardware encoder to the first encoder, so that the first encoder performs encoding processing on the first encoded video by using the transmission data.
[0166] An embodiment of the present application also proposes an electronic device, including a processor and a memory; the memory is used to store a computer program executable by the processor; the processor is used to execute the computer program in the memory to implement the video processing method described in any one of the above embodiments.
[0167] An embodiment of the present application also proposes a computer-readable storage medium, when the executable computer program in the storage medium is executed by a processor, it can implement the video processing method described in any one of the above embodiments.
[0168] Regarding the device in the above embodiments, the specific manner in which the processor performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0169] Figure 10 is a block diagram of an electronic device shown in accordance with an exemplary embodiment. For example, the electronic device 900 may be provided as a server. Referring to Figure 10 , the device 900 includes a processing component 922 which further includes one or more processors, and memory resources represented by a memory 932 for storing instructions executable by the processing component 922, such as application programs. The application programs stored in the memory 932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute instructions to perform the above-described method for video processing.
[0170] The device 900 may also include a power component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output (I / O) interface 958. The device 900 may operate based on an operating system stored in the memory 932, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0171] In an exemplary embodiment, there is also provided a non-transitory computer-readable storage medium including instructions, such as the memory 932 including instructions, the above instructions being executable by the processing component 922 of the device 900 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0172] In the present invention, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance. The term "plurality" refers to two or more unless otherwise clearly defined.
[0173] The above description of the embodiments is intended to enable those of ordinary skill in the art to understand and apply the present application. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present application is not limited to the embodiments herein, and all improvements and modifications made by those skilled in the art within the scope and spirit of the present application disclosed are within the scope of the present application.
Claims
1. A video processing system, characterized in that, The system includes: a hardware encoder, and at least one first encoder; wherein, the hardware encoder is communicatively connected to the first encoder; The hardware encoder is configured to obtain a target video and perform encoding processing on the target video to obtain a first encoded video, motion information, and frame-level information; The frame-level information includes: a session identifier, motion vector precision, frame number, frame type, block size information, and a reference frame list; the session identifier is created during communication initialization between the hardware encoder and the host; the hardware encoder processes the frame-level information according to a transmission protocol to obtain first-level data; the first-level data is used for a high-level description of a video encoding session and provides context for subsequent transmission of motion information; The motion information includes: a reference frame and a reference frame index, a motion vector and a vector direction, a vector magnitude, a block resolution, a distortion cost, and a prediction unit; the motion information constitutes second-level data; at least one storage interface corresponding to the motion information is allocated according to at least one resolution so as to obtain the motion information through the storage interface; wherein, the prediction unit includes multiple different size specifications so as to provide the required prediction unit for different first encoders; The hardware encoder processes the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data; and sends the transmission data to the first encoder; The first encoder is configured to perform encoding processing on the received transmission data and the first encoded video to obtain a second encoded video.
2. The system according to claim 1, wherein When the hardware encoder outputs the reference frame list, it includes at least one of a forward prediction list, a backward prediction list, and a bidirectional prediction list, and multiple different types of reference frames applicable to multiple first encoders.
3. The system according to claim 2, wherein The first encoder is further configured to parse the received transmission data to obtain the first-level data and the second-level data; Select a required target reference frame list from the first-level data, and select required prediction block information from the second-level data; so as to perform encoding processing on the target reference frame list, the prediction block information, and the first encoded video to obtain a second encoded video.
4. The system according to claim 3, characterized in that, The first encoder is further configured to determine a target prediction block size and a target prediction block shape corresponding to the first encoder type according to the received first encoded video; If a target prediction block size and a target prediction block shape corresponding to the first encoder type are not found, the prediction block size and the prediction block shape are segmented to obtain a target prediction block size and a target prediction block shape corresponding to the first encoder type.
5. A video processing method, characterized in that, Applied to a main controller or a hardware encoder in a heterogeneous video processing system, the method includes: The target video is encoded by a hardware encoder to obtain a first encoded video, motion information, and frame-level information; wherein, the frame-level information includes: a session identifier, motion vector precision, frame number, frame type, block size information, and reference frame list; the session identifier is created during the communication initialization between the hardware encoder and the host; the hardware encoder processes the frame-level information according to a transmission protocol to obtain first-level data; the first-level data is used for the high-level description of the video encoding session and provides context for the subsequent transmission of motion information; The motion information includes: reference frames and reference frame indices, motion vectors and vector directions, vector magnitudes, block resolutions, distortion costs, and prediction units; the motion information constitutes second-level data; at least one storage interface corresponding to the motion information is allocated according to at least one resolution, so as to obtain the motion information through the storage interface; wherein, the prediction unit includes various different size specifications to provide the required prediction units for different first encoders; The motion information and the frame-level information are processed according to a target transmission protocol matching the first encoder to obtain transmission data; The transmission data is sent to the first encoder so that the first encoder uses the received transmission data and the first encoded video for encoding processing to obtain a second encoded video.
6. A video transmission method, characterized in that, Applied to the main controller or hardware encoder in a heterogeneous video system, the method includes: Parse the motion information provided by the hardware encoder to obtain a session, frame metadata, and motion information data; wherein, the frame metadata is the metadata in the frame-level information, including a session identifier, motion vector precision, frame number, frame type, block size information, and reference frame list; the session identifier is created during the communication initialization between the hardware encoder and the host; the hardware encoder processes the frame-level information according to a transmission protocol to obtain first-level data; the first-level data is used for the high-level description of the video encoding session and provides context for the subsequent transmission of motion information; The motion information includes: reference frames and reference frame indices, motion vectors and vector directions, vector magnitudes, block resolutions, distortion costs, and prediction units; The motion information constitutes second-level data; at least one storage interface corresponding to the motion information is allocated according to at least one resolution, so as to obtain the motion information through the storage interface; wherein, the prediction unit includes various different size specifications to provide the required prediction units for different first encoders; Process the obtained session, frame metadata, and motion information data according to a target transmission protocol matching the first encoder to obtain transmission data containing first-level data and second-level data; Transmit the transmission data and the first encoded video encoded by the hardware encoder to the first encoder so that the first encoder uses the transmission data and the first encoded video for encoding processing.
7. A video processing device, characterized in that, The device includes: An encoding module for encoding a target video through a hardware encoder to obtain a first encoded video, motion information, and frame-level information; The frame-level information includes: a session identifier, motion vector precision, frame number, frame type, block size information, and a reference frame list; the session identifier is created during communication initialization between the hardware encoder and the host; the hardware encoder processes the frame-level information according to a transmission protocol to obtain first-level data; the first-level data is used for a high-level description of a video encoding session and provides context for subsequent transmission of motion information; The motion information includes: a reference frame and a reference frame index, a motion vector and a vector direction, a vector magnitude, block resolution, distortion cost, and a prediction unit; the motion information constitutes second-level data; at least one storage interface corresponding to the motion information is allocated according to at least one resolution so as to obtain the motion information through the storage interface; wherein, the prediction unit includes multiple different size specifications so as to provide the required prediction unit for different first encoders; A transmission data processing module for processing the motion information and the frame-level information according to a target transmission protocol matching the first encoder to obtain transmission data; A sending module for sending the transmission data to the first encoder so that the first encoder uses the received transmission data and the first encoded video for encoding processing to obtain a second encoded video.
8. An electronic device comprising a processor and a memory, wherein at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to claim 5 or 6.
9. A computer-readable medium having stored thereon at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the method according to claim 5 or 6.
Citation Information
Patent Citations
Preencoder assisted video encoding
US20150350686A1