Video coding system and method, storage medium and terminal equipment

By introducing a video coding system with macroblock row distribution units and context transmission units, the complexity problem of the H.264 coding standard in low-latency scenarios is solved, and high coding efficiency and low-latency characteristics are achieved, which is suitable for applications such as real-time video streaming and cloud gaming.

CN120751147APending Publication Date: 2025-10-03ZHUHAI HUGE IC CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510893933.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing H.264 encoding standard is highly complex in low-latency scenarios, resulting in large hardware resource usage. In addition, the wavefront parallel processing mechanism of H.265/HEVC has not been optimized in the H.264 ecosystem, resulting in encoding delay and rate-distortion performance loss.

Method used

The macroblock row distribution unit and context transfer unit are introduced to process video input in parallel through N coding units. A specific scheduling mechanism is adopted to ensure that the coding context transfer is free of data competition and is compatible with the H.264 standard.

Benefits of technology

It achieves low-latency video encoding, improves encoding efficiency, and reduces memory access latency. It is suitable for scenarios such as real-time video streaming and cloud gaming, reduces hardware and operating costs, and improves the quality of social services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751147A_ABST
    Figure CN120751147A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding system and method, a storage medium and terminal equipment, and the system comprises a macro block line distribution unit which divides a video input by taking a macro block line as a unit, and respectively distributes the video input to N coding units; the N coding units execute respective video coding tasks according to the H.264 coding standard and output corresponding code stream segments, and N is greater than or equal to 2; n context transmission units, wherein each context transmission unit is used for executing a coding context transmission task between two adjacent coding units; the coding context comprises intra-frame prediction dependency information and / or motion vector prediction dependency information, dependency information of CAVLC operation during entropy coding and deblocking effect filtering dependency information. According to the invention, a small-capacity context information transmission unit group is adopted to carry out low-delay coding context transmission, so that delay caused by data competition is prevented. In addition, a specific scheduling mechanism is adopted, it is guaranteed that it is avoided that multiple coding units read and write the same memory position at the same moment, and the coding efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video signal processing, and in particular relates to a video encoding system, method, storage medium and terminal equipment. Background Art

[0002] With the widespread application of real-time video transmission technology in fields such as drone image transmission, cloud gaming, and video conferencing, low-latency encoding has become a core challenge in hardware encoder design.

[0003] Patent document CN113853786 A discloses a decoding method, comprising obtaining a block end bit and a byte alignment bit with a first value from a video bitstream, wherein the block end bit and the byte alignment bit with the first value are used to indicate that a current coding tree block (CTB) is the last CTB in a block; obtaining a CTB row end bit and a byte alignment bit with a first value from the video bitstream, wherein the CTB row end bit and the byte alignment bit with the first value are used to indicate that waveform parallel processing (WPP) is enabled and that the current CTB is the last CTB in a CTB row but not the last CTB in the block; and reconstructing multiple CTBs in the block based on the block end bit, the CTB row end bit, and the byte alignment bit with the first value.

[0004] The wavefront parallel processing (WPP) in the above-mentioned patented technology is based on a coding standard that uses CTB (such as H.265). Its algorithm complexity is relatively high and has the following defects: on the one hand, for low-latency scenarios, the complexity of the encoder significantly affects the encoding delay; on the other hand, excessive complexity will cause the encoder hardware implementation to occupy a large amount of resources, which is not conducive to deployment; furthermore, its wavefront parallel processing algorithm needs to specify the CTB line end signal at both ends of the codec to achieve consistent encoding and decoding.

[0005] As can be seen from the above, H.264 / AVC is still the most popular video coding standard in production, and its complexity is relatively low. Therefore, an H.264 encoder with low coding delay and low cost still has important application value. However, the design of the H.264 standard is not optimized for parallel processing, which becomes a limitation for further reducing coding delay. Existing H.264 low-latency coding optimization schemes usually adopt Slice partitioning strategy, but Slice partitioning will destroy the spatial continuity of intra-frame prediction and motion estimation, resulting in rate-distortion performance loss at the Slice boundary. In contrast, the Wavefront Parallel Processing (WPP) mechanism in the H.265 / HEVC standard achieves a good balance between coding efficiency and parallelism in software encoders through CTU (coding tree unit) row-level dependency control, but its implementation has so far been limited to the HEVC ecosystem. Summary of the Invention

[0006] The present invention provides a video coding system, method, storage medium and terminal device, which achieve zero rate distortion performance loss by introducing a wavefront parallel mechanism, are fully compatible with the H.264 coding standard, and aim to significantly improve coding efficiency. The present invention is achieved through the following technical solutions.

[0007] In a first aspect, the present invention provides a video encoding system, comprising:

[0008] A macroblock row distribution unit is used to divide the video input into macroblock row units and distribute them to N coding units respectively;

[0009] N coding units, each performing its own video coding task according to the H.264 coding standard and outputting the corresponding bitstream segments, N ≥ 2;

[0010] N context transfer units, each context transfer unit is used to perform the coding context transfer task between two adjacent coding units; the coding context includes: intra-frame prediction dependency information and / or motion vector prediction dependency information, entropy coding dependency information, and deblocking effect filtering dependency information.

[0011] As a preferred technical solution, the coding context includes: the coding context includes the prediction mode of the bottom pixels of the upper and upper right macroblocks and their bottom 4x4 sub-blocks in intra-frame prediction, the motion vector of the bottom sub-block of the upper and upper right macroblocks in motion vector prediction, the number of non-zero coefficients of the bottom 4x4 sub-block of the upper macroblock in the CAVLC operation during entropy coding, and the QP of the relevant macroblock.

[0012] In a second aspect, the present invention provides a video encoding method, based on the above-mentioned video encoding system, comprising:

[0013] The control macroblock row distribution unit divides the video input into macroblock rows and distributes them to N coding units respectively; controls each context transmission unit to transfer the coding context between two adjacent coding units; controls each coding unit to perform its own video coding task based on the coding context and in accordance with the H.264 coding standard and output the corresponding bitstream segment.

[0014] As a preferred technical solution, the video encoding task of the encoding unit includes four pipeline stages:

[0015] Stage 1: Read original pixels from the macroblock row distribution unit;

[0016] Stage 2: Perform intra prediction and / or Inter-frame prediction operate;

[0017] Stage 3: Perform entropy coding and deblocking filtering operations;

[0018] Stage 4: Writing the reconstructed pixels to external storage.

[0019] As a preferred technical solution, during the intra-frame prediction, the current macroblock relies on the reconstructed pixels at the bottom of the macroblocks above and to the right that have not been past-blocking filtered; during the motion vector prediction, the current macroblock relies on the motion vector of the sub-blocks at the bottom of the macroblocks above and to the right; and when performing the CAVLC operation of entropy coding and the deblocking filtering operation, the current macroblock relies on the number of non-zero coefficients of the sub-block at the bottom of the macroblock above and the reconstructed pixels after deblocking filtering, as well as the QP of the macroblock above.

[0020] As an optimal technical solution, the inter-frame prediction is further divided into three stages: integer pixel motion estimation, sub-pixel motion estimation, and motion compensation; while executing the intra-frame prediction or motion compensation stage, the reconstructed pixels required for deblocking effect filtering are pre-fetched from external storage.

[0021] As a preferred technical solution, the video encoding method adopts the following scheduling control mechanism: the operation of loading the top pixels of the deblocking filter operation of the next row of macroblocks is controlled to be performed at least after the storage operation of the reconstructed pixels of the previous row of macroblocks is completed.

[0022] As a preferred technical solution, the video encoding method adopts the following scheduling control mechanism: the intra-frame prediction and / or motion compensation stage of the next row of macroblocks is controlled to lag behind the reconstructed pixel storage stage of the previous row of macroblocks by at least one macroblock.

[0023] In a third aspect, the present invention further provides a storage medium, wherein the computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0024] In a fourth aspect, the present invention further provides a terminal device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0025] The beneficial effects brought about by the technical solution of the present invention include at least: using a small-capacity context information transmission unit group to perform low-latency coding context transfer, without the need for additional measures (such as software locks or arbiters) to prevent delays caused by data contention. In addition, the specific scheduling mechanism adopted by the present invention ensures that when each macroblock is encoded, the coding context information from the adjacent upper and upper right macroblocks (if any) is already available, and different coding units will not read the coding context information required for the same column of macroblocks at the same time, which can ensure that there will not be multiple coding units / threads reading and writing the same memory location at the same time, thereby achieving data contention-free context transfer between coding units (between macroblock rows) without increasing delay and synchronization overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0027] Figure 1 This is a block diagram of a video encoding system provided by an embodiment of the present invention.

[0028] Figure 2 This is a schematic diagram of the data dependency relationship between adjacent macroblock rows during encoding in the video encoding method provided by an embodiment of the present invention.

[0029] Figure 3 This is a timing diagram of the pipeline stages of the encoding unit in the video encoding method provided by an embodiment of the present invention, taking dual-core as an example.

[0030] Figure 4 This is a coding flow chart of the video coding method provided by an embodiment of the present invention, taking three cores as an example, when the adjacent lower macroblock lags behind the upper macroblock by exactly three pipeline stages.

[0031] Figure 5 It is a structural diagram of a terminal device provided by the present invention. DETAILED DESCRIPTION

[0032] To make the technical solution of the present invention clearer and the technical advantages more apparent, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of the present invention.

[0033] The specific embodiments of the present invention are further described below with reference to the accompanying drawings. For the sake of convenience, the present invention defines the orientations with reference to the accompanying drawings. The definitions of these orientations are only for the purpose of clearly describing the relative positional relationships and are not intended to limit the actual orientations of products or devices during production, use, sales, and the like. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Furthermore, in the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "install", "set", "connect", "fix" and the like should be understood in a broad sense. For example, they can be fixedly connected, detachably connected, or integrated; they can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0034] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0035] In order to achieve wavefront parallel processing compatible with the H.264 coding standard, one of the main challenges is to maintain the coding context data dependency between adjacent macroblock lines (MB Lines). In order to solve this problem, the present invention proposes a video coding system and method based on the layout of coding units (or cores) and context transmission units. In addition, this embodiment also sets up a dedicated scheduling mechanism and uses a single-input single-output (SISO) context transmission unit to transmit the coding context between adjacent macroblock lines. A context transmission unit is read and written by only two pre-defined coding units, respectively. Combined with the scheduling mechanism, low-latency context transmission without data competition is achieved.

[0036] like Figure 1 As shown, a video encoding system includes:

[0037] A macroblock row distribution unit is used to divide the video input into macroblock row units and distribute them to N coding units respectively;

[0038] N coding units, each performing its own video coding task according to the H.264 coding standard and outputting the corresponding bitstream segments, N ≥ 2;

[0039] N context transmission units, each context transmission unit is used to perform a coding context transfer task between two adjacent coding units.

[0040] Among them, the encoding unit can be implemented as a hardware core or as a software program; the context transfer unit can be implemented using (on-chip) BRAM and FIFO in the hardware system, and can be implemented using arrays and circular queues in the software system.

[0041] In addition, the coding context includes: intra-frame prediction dependency information and / or motion vector prediction dependency information, entropy coding dependency information, and deblocking filter dependency information. For example, the coding context may include: the coding context includes the prediction mode of the bottom pixel of the upper macroblock and its bottom 4x4 sub-block in intra-frame prediction, the motion vector of the bottom sub-block of the upper macroblock in motion vector (MV) prediction, the number of non-zero coefficients (NB) of the bottom 4x4 sub-block of the upper macroblock in CAVLC operation during entropy coding, and the QP of the relevant macroblock.

[0042] The video coding system uses a small-capacity context information transmission unit group to perform low-latency coding context transfer without using additional measures (such as software locks or arbiters) to prevent data competition.

[0043] Based on the above video encoding system, this embodiment provides a video encoding method, including:

[0044] Controlling the macroblock row distribution unit to divide the video input into units of macroblock rows and distribute them to N coding units respectively;

[0045] Controlling each context transmission unit to transfer the coding context between two adjacent coding units;

[0046] Control each encoding unit to perform its own video encoding task based on the encoding context and in accordance with the H.264 encoding standard and output a corresponding bitstream segment.

[0047] Combine Figure 2 As shown in the figure, the data dependency between adjacent macroblock rows in H.264 encoding is as follows: for intra prediction and motion vector (MV) prediction, the current macroblock depends on the reconstructed pixels and motion vectors at the bottom of the upper and upper right macroblocks that have not been past-blocking filtered; for CAVLC operation and deblocking filter (DB) of entropy coding, the current macroblock depends on the number of non-zero coefficients (NB) and the reconstructed pixels after deblocking filtering in the bottom 4 rows of the upper macroblock, as well as the QP of the upper macroblock, etc.

[0048] Combine Figure 3 As shown, taking the dual-core as an example, the video encoding task of a single encoding unit includes four pipeline stages:

[0049] Phase 1: Read the original pixels from the macroblock row distribution unit (ie Figure 3 Load);

[0050] Stage 2: Perform intra-frame prediction and / or inter-frame prediction (i.e. Figure 3 Predict & Fetch DB in );

[0051] Stage 3: Perform entropy coding and deblocking filtering operations (i.e. Figure 3 EC&DB in ); Among them, entropy coding includes operations such as motion vector prediction, motion vector residual coding, and CAVLC;

[0052] Stage 4: Writing the reconstructed pixels to external storage (i.e. Figure 3 Store).

[0053] In which, during the intra-frame prediction, the current macroblock depends on the reconstructed pixels at the bottom of the macroblocks above and to the right that have not been past-blocking filtered; during the motion vector prediction, the current macroblock depends on the motion vectors of the sub-blocks at the bottom of the macroblocks above and to the right; in addition, when performing the CAVLC operation of entropy coding and the deblocking filtering operation, the current macroblock depends on the number of non-zero coefficients of the sub-block at the bottom of the macroblock above and the reconstructed pixels after deblocking filtering, as well as the QP of the macroblock above, respectively.

[0054] In addition, inter-frame prediction can be further divided into three stages: integer-pixel motion estimation, sub-pixel motion estimation, and motion compensation. While performing intra-frame prediction or motion compensation, the coding unit will pre-fetch reference pixels required for deblocking filtering from external storage.

[0055] Since the deblocking filter is the last stage in the encoding pipeline that has cross-macroblock row dependencies, as long as the data dependencies during the deblocking filter process can be maintained, the remaining data dependencies can naturally be maintained. In order to keep the data dependencies of the DB unchanged, it is necessary to control the operation of loading the top pixels of the DB in the next row of macroblocks to be performed at least after the reconstructed pixel storage operation of the previous row of macroblocks is completed. In addition, since the operation of pre-fetching the reference pixels required for the deblocking filter from the external storage by the encoding unit is performed in the same pipeline stage as the intra-frame prediction or motion compensation, the intra-frame prediction / motion compensation stage of the next row of macroblocks is required to lag behind the reconstructed pixel storage stage of the previous row of macroblocks by at least one macroblock, that is, the next row of macroblocks on the same column lags behind the previous row of macroblocks by at least three pipeline stages.

[0056] The above timing constraints (scheduling mechanism) ensure that two coding units will not process the same pipeline stage for the same column of macroblocks at the same time. Therefore, two coding units will not access the context information for the same column of macroblocks in the same pipeline stage at the same time. This ensures the implementation of a lock-free / arbiter-free context transfer unit. The context used by each pipeline stage is transferred by a pair of independent context transfer units.

[0057] Combine Figure 4 The figure shows the encoding process for a three-core encoding system, where the adjacent lower macroblock lags exactly three pipeline stages behind the upper macroblock. The left side of the dotted line shows the encoding stage of each macroblock, while the right side shows the storage area read or written by different coding units during the (Intra)Predict stage of the corresponding macroblock. The number pairs within the RAM area box are the macroblock coordinates corresponding to the context stored in that area at the current stage.

[0058] Figure 4 Taking the three-core encoding system as an example, the encoding process is demonstrated when the adjacent lower macroblock lags behind the upper macroblock by exactly three pipeline stages. Each context transfer unit defines an input port and an output port, and the two ports are connected to the corresponding modules in the two encoding units respectively. Taking intra-frame prediction as an example, assuming that the Nth row of macroblocks is processed by encoding unit 1, the N+1th row of macroblocks is processed by encoding unit 2, and the N+2th row of macroblocks is processed by encoding unit 3, Figure 4 The three-stage encoding process shown is as follows:

[0059] Phase 1: Coding unit 1 processes macroblock (N, 2), reads reference pixels from columns 2 and 3 of RAM2, and writes reconstructed pixels to column 1 of RAM0; Coding unit 2 processes macroblock row N-2 (or waits); Coding unit 3 processes macroblock row N-1 (or waits);

[0060] Phase 2: Coding unit 1 processes macroblock (N, 3), reads reference pixels from columns 3 and 4 of RAM2, and writes reconstructed pixels to column 2 of RAM0; Coding unit 2 processes macroblock (N+1, 0), reads reference pixels from columns 0 and 1 of RAM0, and writes reconstructed pixels to column 0 of RAM1; Coding unit 3 processes macroblock row N-1 (or waits);

[0061] Stage 3: Coding unit 1 processes macroblock (N, 6), reads reference pixels from columns 6 and 7 of RAM2, and writes reconstructed pixels to column 3 of RAM0; coding unit 2 processes macroblock (N+1, 3), reads reference pixels from columns 3 and 4 of RAM0, and writes reconstructed pixels to column 3 of RAM1; coding unit 3 processes macroblock (N+1, 0), reads reference pixels from columns 0 and 1 of RAM1, and writes reconstructed pixels to column 0 of RAM2.

[0062] Timing constraints ensure that multiple encoding units will not access the same address of the same context transfer unit at the same time, so the context transfer unit does not need to use additional synchronization mechanisms to prevent unpredictable results such as data inconsistency. In addition, observe Figure 4 It can be seen that the contexts of different macroblock columns follow the first-in-first-out rule in the RAM, so the context transfer unit can be implemented using BRAM or array, or FIFO or circular queue.

[0063] The video encoding method provided in the above embodiment, by constructing a scheduling mechanism for the coding unit, makes the next row of macroblocks lag behind the previous row of macroblocks by at least 3 macroblocks (3 pipeline stages: intra-frame prediction / motion compensation, loop filtering (in parallel with entropy coding), and reconstructing pixels after write-back filtering), ensures that when each macroblock is encoded, the coding context information from the adjacent upper and upper right macroblocks (if any) is already available, and different coding units / cores will not read the coding context information required for the same column of macroblocks at the same time. In addition, by reusing the existing context information to construct a context information transmission unit, context transmission between coding units without data competition is achieved without increasing delay and synchronization overhead. Due to the limitations of the scheduling mechanism, different coding units will not read and write the context information required for the same column of macroblocks at the same time. Therefore, it is only necessary to divide the context information transmission unit according to the macroblock column to ensure that there will not be multiple cores / threads reading and writing the same memory location at the same time, thereby achieving context transmission between coding units (between macroblock rows) without data competition without increasing delay and synchronization overhead. The context information transmission unit can be constructed not only by using a random access storage structure that divides addresses according to macroblock columns, but also by using a first-in-first-out storage structure.

[0064] The video encoding system and method provided in the above embodiments have the following beneficial effects:

[0065] From a technical perspective, the present invention achieves a significant improvement in coding efficiency. Experimental results show that compared with the traditional macroblock serial processing method, the dual-core FPGA implementation of the present invention can reduce the overall video coding delay by up to 49.4%. For example, when processing 1080p video, the dual-core FPGA implementation of the encoder of the present invention can achieve a processing speed of 125fps at a frequency of 125MHz. This performance indicator is much higher than the existing technology, providing strong technical support for low-latency application scenarios such as real-time video streaming and cloud gaming. In addition, the present invention effectively reduces memory access latency and improves real-time performance. These technical optimizations improve coding efficiency.

[0066] From an economic perspective, this invention provides a low-cost upgrade path for H.264 codec design, with significant engineering value. By reducing encoding latency and increasing processing speed, this invention can effectively reduce the operating costs of video streaming service providers. For example, in large-scale video streaming services, the use of this invention can reduce the number of servers and energy consumption, thereby lowering operating costs. Furthermore, in drone and cloud gaming applications, this invention can significantly reduce video encoding latency.

[0067] Socially, the low-latency characteristics of this invention are of great significance for improving the quality and efficiency of social services such as telemedicine, online education, and video conferencing. For example, in telemedicine, low-latency video encoding technology can ensure real-time interaction between doctors and patients, improving the accuracy and timeliness of diagnoses. In online education, low-latency video streaming can enhance the interactive experience between teachers and students, improving teaching effectiveness.

[0068] An embodiment of the present invention further provides a storage medium. The computer storage medium can store multiple instructions. The instructions are suitable for being loaded by a processor and executing the method steps of the video encoding method provided in the above embodiment, which will not be described in detail here.

[0069] The present invention also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the quality inspection data extraction configuration method described in the above embodiments.

[0070] See Figure 5 , provides a terminal device for an embodiment of the present invention. Figure 5 As shown, the terminal device 500 may include: at least one processor 501 , at least one network interface 504 , a user interface 503 , a memory 505 , and at least one communication bus 502 .

[0071] The communication bus 502 is used to implement the connection and communication between these components.

[0072] The user interface 503 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 503 may also include a standard wired interface and a wireless interface.

[0073] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0074] The processor 501 may include one or more processing cores. The processor 501 utilizes various interfaces and circuits to connect various components within the entire terminal device 500. It executes instructions, programs, code sets, or instruction sets stored in the memory 505, and accesses data stored in the memory 505 to perform various functions and process data for the terminal device 500. Optionally, the processor 501 may be implemented using at least one hardware form selected from the group consisting of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 501 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 501 and may be implemented separately on a separate chip.

[0075] Among them, the memory 505 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 505 includes a non-transitory computer-readable storage medium. The memory 505 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 505 may also be optionally at least one storage device located away from the aforementioned processor 501. As Figure 5 As shown, the memory 505 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program.

[0076] exist Figure 5In the terminal device 500 shown, the user interface 503 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 501 can be used to call the application stored in the memory 505 and specifically execute the method steps of the video encoding method provided in the above embodiment, which will not be repeated here.

[0077] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0078] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A video coding system, characterized in that: include: A macroblock row distribution unit is used to divide the video input into macroblock row units and distribute them to N coding units respectively; N coding units, each performing its own video coding task according to the H.264 coding standard and outputting the corresponding bitstream segments, N ≥ 2; N context transmission units, each context transmission unit is used to perform a coding context transfer task between two adjacent coding units; The coding context includes: intra-frame prediction dependency information and / or motion vector prediction dependency information, entropy coding dependency information, and deblocking filter dependency information.

2. The video encoding system according to claim 1, wherein The coding context includes: the coding context includes the prediction mode of the bottom pixel of the upper macroblock and its bottom 4x4 sub-block in intra-frame prediction, the motion vector of the bottom sub-block of the upper macroblock in motion vector prediction, the number of non-zero coefficients of the bottom 4x4 sub-block of the upper macroblock in the CAVLC operation during entropy coding, and the QP of the related macroblock.

3. A video encoding method, based on the video encoding system according to claim 1, characterized in that: include: Controlling the macroblock row distribution unit to divide the video input into units of macroblock rows and distribute them to N coding units respectively; Control each context transmission unit to transfer the coding context between two adjacent coding units; control each coding unit to perform its own video coding task based on the coding context and in accordance with the H.264 coding standard and output corresponding code stream segments.

4. The video encoding method according to claim 3, wherein: The video encoding task of the encoding unit includes four pipeline stages: Stage 1: Read original pixels from the macroblock row distribution unit; Stage 2: Perform intra-frame prediction and / or inter-frame prediction; Stage 3: Perform entropy coding and deblocking filtering operations; Stage 4: Writing the reconstructed pixels to external storage.

5. The video encoding system according to claim 4, wherein: During the intra-frame prediction, the current macroblock depends on the reconstructed pixels at the bottom of the macroblocks above and to the right that have not been past-blocking filtered; during the motion vector prediction, the current macroblock depends on the motion vector of the sub-blocks at the bottom of the macroblocks above and to the right; when performing the CAVLC operation of entropy coding and the deblocking filtering operation, the current macroblock depends on the number of non-zero coefficients of the sub-block at the bottom of the macroblock above and the reconstructed pixels after deblocking filtering, as well as the QP of the macroblock above.

6. The video encoding method according to claim 4, wherein: The inter-frame prediction is further divided into three stages: integer pixel motion estimation, sub-pixel motion estimation, and motion compensation. While performing the intra-frame prediction or motion compensation stage, reconstructed pixels required for deblocking filtering are pre-fetched from external storage.

7. The video encoding method according to claim 4, wherein: The video encoding method adopts the following scheduling control mechanism: the operation of loading the top pixels of the deblocking filtering operation of the next row of macroblocks is controlled to be performed at least after the storage operation of the reconstructed pixels of the previous row of macroblocks is completed.

8. The video encoding method according to claim 7, wherein: The video encoding method adopts the following scheduling control mechanism: The intra-frame prediction and / or motion compensation stage of the next row of macroblocks is controlled to lag behind the reconstructed pixel storage stage of the previous row of macroblocks by at least one macroblock.

9. A storage medium, characterized in that: The computer storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executing the method steps according to any one of claims 2 to 8.

10. A terminal device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 2 to 8.

Citation Information

Patent Citations

  • Wavefront parallel processing for tile, brick, and slice

    CN113853786A

  • Image decoding device and image encoding device

    CN101779466A

  • Device used for using image frames as basis to execute parallel video coding and method thereof

    CN104038766A

  • Parallel high efficiency video coding (HEVC) system and method based on distributed computer system

    CN104780377A

  • Video decoding macro-block-grade parallel scheduling method for perceiving calculation complexity

    CN105491377A