Video encoding apparatus and related video encoding method

By introducing a content activity analyzer and a video encoder into the video encoding device, and optimizing the video encoding process using the results of content activity analysis, the problem of low video compression efficiency in mobile networks is solved, and more efficient low bit rate video compression is achieved.

CN116405691BActive Publication Date: 2026-01-16MEDIATEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211733285.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-22
Filing Date
2022-12-30
Publication Date
2026-01-16
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

Existing technologies for video compression in mobile networks, especially for low bitrate video compression, suffer from problems such as large data volume and low transmission efficiency.

Method used

A content activity analyzer circuit is used to perform content activity analysis on multiple consecutive frames, generating content activity analysis results. The video is then processed by a video encoder circuit to generate a bitstream output. The content activity analysis results are used to optimize the video encoding process.

Benefits of technology

By optimizing the video encoding process, the complexity of video encoding is reduced, video quality is improved, and the bit rate is reduced, achieving more efficient video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405691B_ABST
    Figure CN116405691B_ABST
Patent Text Reader

Abstract

A video encoding device includes a content activity analyzer circuit and a video encoder circuit. The content activity analyzer circuit performs a content activity analysis process on successive frames to generate content activity analysis results. The successive frames are derived from input frames of the video encoding device. The content activity analysis process includes deriving a first content activity analysis result from a first frame and a second frame of the successive frames, where the first content activity analysis result includes a processed frame that is different from the second frame, and deriving a second content activity analysis result from a third frame included in the successive frames and the processed frame. The video encoder circuit performs a video encoding process to generate a bitstream output of the video encoding device, where information derived from the content activity analysis results is referenced by the video encoding process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to video compression, and more particularly, to a video encoding apparatus and its related video encoding method, which performs video compression (e.g., low bit rate video compression) by means of content activity analysis. BACKGROUND

[0002] One of the recent goals of mobile telecommunication is to increase the speed of data transmission to enable multimedia services to be incorporated into mobile networks. One of the key components of multimedia is digital video. Transmission of digital video involves continuous data flow. Generally, digital video requires a large amount of data compared to many other types of media. Therefore, there is a need for an innovative method and apparatus for low bit rate video compression. SUMMARY

[0003] It is one of the objects of the present invention to provide a video encoding apparatus and its related video encoding method, which performs video compression (e.g., low bit rate video compression) by means of content activity analysis.

[0004] According to a first aspect of the present invention, an exemplary video encoding apparatus is disclosed. The exemplary video encoding apparatus includes a content activity analyzer circuit and a video encoder circuit. The content activity analyzer circuit is arranged to apply a content activity analysis process to a plurality of consecutive frames to generate a plurality of content activity analysis results, wherein the plurality of consecutive frames is derived from a plurality of input frames of the video encoding apparatus, the content activity analysis process performed by the content activity analyzer circuit includes: deriving, from a first frame and a second frame included in the plurality of consecutive frames, a first content activity analysis result included in the plurality of content activity analysis results, wherein the first content activity analysis result includes a processed frame, which is different from the second frame; deriving, from a third frame included in the plurality of consecutive frames and the processed frame, a second content activity analysis result included in the plurality of content activity analysis results. The video encoder circuit is arranged to perform a video encoding process to generate a bitstream output of the video encoding apparatus, wherein the video encoding process refers to information derived from the plurality of content activity analysis results.

[0005] According to a second aspect of the present disclosure, an example video encoding method is disclosed. The example video encoding method includes applying a content activity analysis process to a plurality of consecutive frames to generate a plurality of content activity analysis results, and performing a video encoding process to generate a bitstream output. The plurality of consecutive frames is derived from a plurality of input frames. The content activity analysis process includes deriving a first content activity analysis result included in the plurality of content activity analysis results from a first frame and a second frame included in the plurality of consecutive frames, wherein the first content activity analysis result includes a processed frame that is different from the second frame, and deriving a second content activity analysis result included in the plurality of content activity analysis results from a third frame included in the plurality of consecutive frames and the processed frame. Information derived from the plurality of content activity analysis results is referenced by the video encoding process.

[0006] These and other objects of the present application will no doubt become apparent to those of ordinary skill in the art after reading the following detailed description of the preferred embodiments that is presented in connection being the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 FIG. 1 shows a first video encoding apparatus illustrating an embodiment of the present application.

[0008] Figure 2 FIG. 2 shows a content activity analysis process according to an embodiment of the present application.

[0009] Figure 3 FIG. 3 shows a second video encoding apparatus according to an embodiment of the present application.

[0010] Figure 4 FIG. 4 shows a 2-mode activity indication derived from consecutive frames (e.g., two input frames, or one input frame and one processed frame) according to an embodiment of the present application.

[0011] Figure 5 FIG. 5 shows a 3-mode activity indication derived from consecutive frames (e.g., two input frames, or one input frame and one processed frame) according to an embodiment of the present application.

[0012] Figure 6 FIG. 6 shows a third video encoding apparatus according to an embodiment of the present application.

[0013] Figure 7 FIG. 7 shows a fourth video encoding apparatus according to an embodiment of the present application.

[0014] Figure 8 FIG. 8 shows a transformed content activity analysis process according to an embodiment of the present application.

[0015] Figure 9 FIG. 9 shows a fifth video encoding apparatus according to an embodiment of the present application.

[0016] Figure 10 FIG. 2 shows a diagram of 2-mode activity indication derived from consecutive frames (e.g., two transformed frames, or one transformed frame and one processed transformed frame) according to embodiments of the application.

[0017] Figure 11 FIG. 3 shows a diagram of 3-mode activity indication derived from consecutive frames (e.g., transformed frames, or one transformed frame and one processed transformed frame) according to embodiments of the application.

[0018] Figure 12 FIG. 4 shows a diagram of a sixth video encoding device according to embodiments of the application. DETAILED DESCRIPTION

[0019] Certain terms are used throughout the following description and claims to refer to particular elements. As one skilled in the art will appreciate, electronic equipment manufacturers can refer to an element by different names. This document does not intend to distinguish between terms, components, items, or modules that differ in name but not function. In the following description and in the claims, the terms "include" and "comprise" are used in an open-ended fashion, and thus should be interpreted as "including, but not limited to." Also, the term "couple" is intended to mean either an indirect or direct electrical connection. Accordingly, if one device is coupled to another device, that connection can be through a direct electrical connection, or through an indirect electrical connection via other devices and connections.

[0020] Figure 1 FIG. 1 shows a diagram of a first video encoding device according to embodiments of the application. The video encoding device 100 includes a content activity analyzer circuit (labeled "content activity analyzer") 102 and a video encoder circuit (labeled "video encoder") 104. The content activity analyzer circuit 102 is configured to apply a content activity analysis process to consecutive frames to generate content activity analysis results. In this embodiment, the consecutive frames received by the content activity analyzer circuit 102 are input frames 101 of the video encoding device 100, and the content activity analysis results generated by the content activity analyzer circuit 102 are processed frames 103. Note that, depending on the content activity analysis process proposed, a previous processed frame 103 generated from a previous input frame 101 can be referenced by the content activity analyzer circuit 102 for content activity analysis of a current input frame 101.

[0021] Figure 2A diagram showing content activity analysis processing according to an embodiment of the present application. Input frame 101 includes consecutive frames, e.g., frames Fl, F2, and F3. Content activity analyzer circuit 102 derives processed frame F2' from input frames Fl and F2. Specifically, content activity analyzer circuit 102 performs content activity analysis on pixel data of input frames Fl and F2 to identify static pixel data in input frame F2, where the static pixel data represents no motion activity between the current frame (i.e., input frame F2) and the previous frame (input frame Fl). In addition, content activity analyzer circuit 102 derives processed pixel data 202, and generates processed frame F2' by replacing the static pixel data identified in input frame F2 with processed pixel data 202. For example, processed pixel data 202 is static pixel data in input frame Fl. As another example, processed static pixel data 202 is generated by applying an arithmetic operation to pixel data in input frames Fl and F2. With the processed pixel data 202 in processed frame F2', the complexity of encoding processed frame F2' can be reduced, resulting in a low bit rate bitstream.

[0022] Processed frame F2' is different from input frame F2, and can replace input frame F2 for subsequent content activity analysis. Compared to content activity analysis of pixel data of input frames F3 and F2, content activity analysis of pixel data of input frame F3 and processed frame F2' can result in more accurate static pixel data detection. As shown, content activity analyzer circuit 102 derives processed frame F3' from input frame F3 and processed frame F2'. Specifically, content activity analyzer circuit 102 performs content activity analysis on pixel data of input frame F3 and processed frame F2' to identify static pixel data in input frame F3, where the static pixel data represents no motion activity between the current frame (i.e., input frame F3) and the previous frame (i.e., processed frame F2'). In addition, content activity analyzer circuit 102 derives processed pixel data 204, and generates processed frame F3' by replacing the static pixel data identified in input frame F3 with processed pixel data 204. For example, processed pixel data 204 is static pixel data in processed frame F2'. As another example, processed static pixel data 204 is generated by applying an arithmetic operation to pixel data in input frame F3 and processed frame F2'. With the processed pixel data 204 in processed frame F3', the complexity of encoding processed frame F3' can be reduced, resulting in a low bit rate bitstream. Similarly, processed frame F3' is different from input frame F3, and can replace input frame F3 for subsequent content activity analysis. For brevity, similar description is omitted here. Figure 2

[0023] ​The video encoder circuit 104 is arranged to perform a video encoding process to generate a bitstream output of the video encoding apparatus 100, wherein information derived from the content activity analysis results (e.g. the processed frames 103) is referenced by the video encoding process. In the present embodiment, the video encoder circuit 104 encodes the input frame Fl to generate a first frame bitstream which is included in the bitstream output, encodes the processed frame F2' to generate a second frame bitstream which is included in the bitstream output, encodes the processed frame F3' to generate a third frame bitstream which is included in the bitstream output, and so on. It should be noted that the video encoder circuit 104 can be implemented by any suitable encoder architecture. That is, the present application does not limit the encoder architecture employed by the video encoder circuit 104.

[0024] Figure 3 A diagram showing a second video encoding apparatus according to an embodiment of the present application is shown. The video encoding apparatus 300 comprises a content activity analyzer circuit (labeled "content activity analyzer") 302 and a video encoder circuit (labeled "video encoder") 304. Like the content activity analyzer circuit 102, the content activity analyzer circuit 302 is arranged to apply a content activity analysis process to consecutive frames to generate content activity analysis results. In the present embodiment, the consecutive frames received by the content activity analyzer circuit 302 are the input frames 101 of the video encoding apparatus 300. The content activity analyzer circuit 302 differs from the content activity analyzer circuit 102 in that the content activity analysis results generated by the content activity analyzer circuit 302 comprise the processed frames 103 and the activity indications 301. Since the principles of generating the processed frames 103 are described above with reference to the content activity analyzer circuit 102, they will not be repeated here for brevity. In some embodiments of the present application, the activity indications 301 can be generated after the content activity analysis of the input frames 101 to derive the processed frames 103. The generation of the activity indications 301 depends on the content activity analysis applied to the input frames 101. In other words, the activity indications 301 can be byproducts of the content activity analysis process performed by the content activity analyzer circuit 302. For example, the activity indications 301 can comprise a plurality of activity indication maps, each activity indication map recording one activity indication for each of a plurality of pixel blocks. Figure 2

[0025] Figure 4 A diagram showing a 2-mode activity indication derived from consecutive frames (e.g. two input frames, or one input frame and one processed frame) according to an embodiment of the present application is shown. As Figure 4 ​, the content activity analyzer circuit 302 can perform content activity analysis on the pixel data in the input frames F1 and F2 to generate an activity map MAP12, which is a 2-mode activity map, including static pixel data indications 402 and non-static pixel data indications 404. Where each static pixel data indication 402 represents no motion activity between blocks located at the same position in the input frames F1 and F2, and each non-static pixel data indication 404 represents motion activity between blocks located at the same position in the input frames F1 and F2. Alternatively, Figure 1 The illustrated input frames F1 and F2 can be replaced with an input frame (e.g., F3) and a processed frame (e.g., F2') such that the 2-mode activity map is derived from content activity analysis of pixel data in the input frame and the processed frame. Since the block-based video encoding process is employed by the video encoder circuit 304, the activity indications 301 generated from the content activity analyzer circuit 302 can be block-based activity indications. For example, each static pixel data indication 402 represented by one block in Figure 4 Each static pixel data indication 402 represented by one block in Figure 4 Each non-static pixel data indication 404 represented by one block in

[0026] Figure 5 A diagram showing a 3-mode activity map derived from consecutive frames (e.g., two input frames, or one input frame and one processed frame) according to an embodiment of the present application is shown. As Figure 5 The content activity analyzer circuit 302 can perform content activity analysis on the pixel data in the input frames F1 and F2 to generate an activity map MAP12, which is a 3-mode activity map, including static pixel data indications 502, non-static pixel data indications 504, and motion (or static) pixel data contour indications 506, where the static pixel data indications 502 represent no motion activity between blocks located at the same position in the input frames F1 and F2, the non-static pixel data indications 504 represent motion activity between blocks located at the same position in the input frames F1 and F2, and the motion (or static) pixel data contour indications 506 represent a contour of the motion activity between the input frames F1 and F2, as shown. It should be noted that the motion (or static) pixel data contour indications 506 can be regarded as a guard ring between the motion pixel data (non-static pixel data) and the static pixel data. Therefore, the terms "motion pixel data contour indication" and "static pixel data contour indication" can be interchangeable.

[0027] Optionally, Figure 5The input frames F1 and F2 in the content activity analyzer circuit 302 can be replaced with input frames (e.g. F3) and processed frames (e.g. F2') such that the 3-mode activity map is derived from content activity analysis of pixel data in the input frames and the processed frames. Since the video encoder circuit 304 employs block-based video encoding processing, the activity indication 301 generated from the content activity analyzer circuit 302 can be a block-based activity indication. For example, each static pixel data indication 502 represented by a block in the activity indication 301 can indicate static activity of a pixel block, each non-static pixel data indication 504 represented by a block in the activity indication 301 can indicate motion (non-static) activity of a pixel block, and each motion (or static) pixel data profile indication 506 represented by a block in the activity indication 301 can correspond to a pixel block at a boundary between motion pixel data (non-static pixel data) and static pixel data. Figure 5 Figure 5 Each static pixel data indication 502 represented by a block in the activity indication 301 can indicate static activity of a pixel block, each non-static pixel data indication 504 represented by a block in the activity indication 301 can indicate motion (non-static) activity of a pixel block, and each motion (or static) pixel data profile indication 506 represented by a block in the activity indication 301 can correspond to a pixel block at a boundary between motion pixel data (non-static pixel data) and static pixel data.

[0028] The video encoder circuit 304 is arranged to perform video encoding processing to generate a bitstream output of the video encoding apparatus 300, wherein information derived from the content activity analysis results (e.g. the processed frames 103 and the activity indication 301) is referenced by the video encoding processing. In the present embodiment, the video encoder circuit 304 encodes the input frame F1 to generate a first frame bitstream included in the bitstream output, encodes the processed frame F2' according to the activity indication 301 (specifically, the activity indication map derived from the input frames F2 and F1) to generate a second frame bitstream included in the bitstream output, encodes the processed frame F3' according to the activity indication 301 (specifically, the activity indication map derived from the input frame F3 and the processed frame F2') to generate a third frame bitstream included in the bitstream output, and so on.

[0029] ​With respect to 2-mode activity indication, it will give two different instructions to the video encoder circuit 304. For example, a plurality of 2-mode activity indication maps are respectively associated with the processing frames 103. That is, the video encoder 304 refers to the 2-mode activity indication map associated with the current processing frame to be encoded by the video encoder 304 to determine how to encode each coding unit (coding block) within the current processing frame. When a coding unit (coding block) in the current processing frame to be encoded is found to be associated with the non-static pixel data indication 404 recorded in the 2-mode activity indication map, the video encoder circuit 304 can encode the coding unit (coding block) in the typical manner specified by the coding standard. When a coding unit (coding block) in the current processing frame to be encoded is found to be associated with the static pixel data indication 402 recorded in the 2-mode activity indication map, the video encoder circuit 304 can force the zeroing of the coded motion vector of the coding unit, or can encode the coding unit in the skip mode. However, these are for illustrative purposes only and are not meant to limit the present application.

[0030] With respect to 3-mode activity indication, it will give three different instructions to the video encoder circuit 304. For example, a plurality of 3-mode activity indication maps are respectively associated with the processing frames 103. That is, the video encoder 304 refers to the 3-mode activity indication map to determine how to encode each coding unit (coding block) within the current processing frame, the 3-mode activity indication map being associated with the current processing frame to be encoded by the video encoder 304. When a coding unit (coding block) in the current processing frame to be encoded is found to be associated with the non-static pixel data indication 504 recorded in the 3-mode activity indication map, the video encoder circuit 304 can encode the coding unit (coding block) in the typical manner specified by the coding standard. When a coding unit (coding block) in the current processing frame to be encoded is found to be associated with the static pixel data indication 502 recorded in the 3-mode activity indication map, the video encoder circuit 304 can encode the coding unit in the skip mode. When a coding unit (coding block) in the current processing frame to be encoded is found to be associated with the motion (or static) pixel data contour indication 506, the video encoder circuit 304 can force the zeroing of the coded motion vector of the coding unit, or can encode the coding unit without residual information, or can encode the coding unit in the skip mode. However, these are for illustrative purposes only and are not meant to limit the present application.

[0031] Compared with the video encoder circuit 104 that encodes each coding unit (coding block) of the processed frames in a typical manner prescribed by a coding standard, the video encoder circuit 304 that encodes each coding unit (coding block) of the processed frames with reference to the activity indication 301 can enable the reconstructed frames (decoded frames) at the decoding end to have better image quality. It should be noted that the video encoder circuit 304 can be implemented by any suitable encoder architecture. That is, the present application does not limit the encoder architecture employed by the video encoder circuit 304.

[0032] Figure 6 A diagram showing a third video encoding apparatus according to an embodiment of the present application is shown. The video encoding apparatus 600 comprises a video encoder circuit (denoted as "video encoder") 604 and the aforementioned content activity analyzer circuit (denoted as "content activity analyzer") 302. The video encoder circuit 604 is configured to perform video encoding processing to generate a bitstream output of the video encoding apparatus 600, wherein information derived from the content activity analysis results (e.g. the activity indication 301) is referenced by the video encoding processing. It should be noted that the processed frames 103 generated by the content activity analyzer circuit 302 are only for content activity analysis, and are not encoded by the video encoder circuit 604 into the output bitstream.

[0033] In this embodiment, the input frames 101 are encoded with the aid of the activity indication 301. For example, the activity indication 301 can comprise a plurality of activity indication maps respectively associated with all the input frames 101 except the first input frame Fl. In the case where a 2-mode activity indication is employed, the 2-mode activity indication will give two different instructions to the video encoder circuit 604. In another case where a 3-mode activity indication is employed, the 3-mode activity indication will give three different instructions to the video encoder circuit 604. When a coding unit (coding block) in the current input frame to be encoded is found to be associated with an activity indication recorded in one of the 2-mode activity indication maps (or 3-mode activity indication maps), the video encoder circuit 604 can encode the coding unit (coding block) in the manner indicated by the activity indication. Thus, the video encoder circuit 604 encodes the input frame Fl to generate a first frame bitstream included in the bitstream output, encodes the input frame F2 according to the activity indication 301 (specifically, the activity indication map derived from the input frames F2 and Fl) to generate a second frame bitstream included in the bitstream output, encodes the input frame F3 according to the activity indication 301 (specifically, the activity indication map derived from the input frames F3 and the processed frame F2') to generate a third frame bitstream included in the bitstream output, and so on. It should be noted that the video encoder circuit 604 can be implemented by any suitable encoder architecture. That is, the present application does not limit the encoder architecture employed by the video encoder circuit 604.

[0034] In the above embodiments, each content activity analyzer circuit 102 and 302 performs the content activity analysis process at the image resolution of the input frame 101. For example, the image resolution of each input frame can be 3840x2160. In order to obtain better video quality and lower bit rate, a pre-processing circuit can be introduced into the video encoding apparatus.

[0035] Figure 7 A diagram of a fourth video encoding apparatus according to an embodiment of the present application is shown. The video encoding apparatus 700 comprises an image transformer circuit (labeled as "image transformer") 702, a content activity analyzer circuit (labeled as "content activity analyzer") 704 and a video encoder circuit (labeled as "video encoder") 706. The image transformer circuit 702 serves as a pre-processing circuit for performing an image transformation process on the input frame 101 of the video encoding apparatus 700 to generate a transformed frame 703 as a consecutive frame for content activity analysis process at the content activity analyzer circuit 704. The image transformation process can include a down-sampling operation such that the image resolution (e.g. 960x540) of one transformed frame output from the image transformer circuit 702 is lower than the image resolution (e.g. 3840x2160) of one input frame received by the image transformer circuit 702. The down-sampling operation can reduce the complexity of the content activity analyzer circuit 704. In addition, the down-sampling operation can reduce the noise level which makes the content activity analyzer circuit 704 more robust to noise.

[0036] The content activity analyzer circuit 704 is arranged to apply the content activity analysis process to the consecutive frame to generate a content activity analysis result. In the present embodiment, the consecutive frame received by the content activity analyzer circuit 704 is the transformed frame 703, and the content activity analysis result generated by the content activity analyzer circuit 704 comprises a processed transformed frame 705 and a processed frame 707. The processed transformed frame 705 and the transformed frame 703 can have the same image resolution (e.g. 960x540). The processed frame 707 and the input frame 101 can have the same image resolution (e.g. 3840x2160).

[0037] It should be noted that according to the proposed content activity analysis process, a previous processed transformed frame 705 generated from a previous transformed frame 703 can be referenced by the content activity analyzer circuit 704 for content activity analysis of a current transformed frame, and a current processed frame can be derived from a current input frame and a previous input frame (or a previous processed frame) according to information given from content activity analysis of the current transformed frame and the previous transformed frame (or the previous processed transformed frame).

[0038] Figure 8A diagram illustrating a post-transformation content activity analysis process according to an embodiment of the present application is shown. The transformed frames 703 include consecutive frames, e.g., frames TF1, TF2 and TF3 derived from input frames F1, F2 and F3, respectively, by image transformation. The content activity analyzer circuit 704 derives a processed transformed frame TF2' from the transformed frames TF1 and TF2. Specifically, the content activity analyzer circuit 704 performs content activity analysis on pixel data of the transformed frames TF1 and TF2 to identify static pixel data in the transformed frame TF2, where the static pixel data represents no motion activity between a current frame (e.g., the transformed frame TF2) and a previous frame (e.g., the transformed frame TF1). In addition, the content activity analyzer circuit 704 derives processed pixel data 802, and generates the processed transformed frame TF2' by replacing the static pixel data identified in the transformed frame TF2 with the processed pixel data 802. For example, the processed pixel data 802 is static pixel data in the transformed frame TF1. As another example, the processed static pixel data 802 is generated by applying an arithmetic operation to pixel data in the transformed frames TF1 and TF2.

[0039] The processed transformed frame TF2' is different from the transformed frame TF2, and can be used as a replacement for the transformed frame TF2 for subsequent content activity analysis. Compared to content activity analysis of pixel data of the transformed frames TF3 and TF2, content activity analysis of pixel data of the transformed frame TF3 and the processed transformed frame TF2' can result in more accurate static pixel data detection. Figure 8 As shown, the content activity analyzer circuit 704 derives a processed transformed frame TF3' from the transformed frame TF3 and the processed transformed frame TF2. Specifically, the content activity analyzer circuit 704 performs content activity analysis on pixel data of the transformed frame TF3 and the processed transformed frame TF2' to identify static pixel data in the transformed frame TF3, where the static pixel data represents no motion activity between a current frame (e.g., the transformed frame TF3) and a previous frame (e.g., the processed transformed frame TF2'). In addition, the content activity analyzer circuit 704 derives processed pixel data 804, and generates the processed transformed frame TF3' by replacing the static pixel data identified in the transformed frame TF3 with the processed pixel data 804. For example, the processed pixel data 804 is static pixel data in the processed transformed frame TF2'. As another example, the processed static pixel data 804 is generated by applying an arithmetic operation to pixel data in the transformed frame TF3 and the processed transformed frame TF2'. Similarly, the processed transformed frame TF3' is different from the transformed frame TF3, and can be used as a replacement for the transformed frame TF3 for subsequent content activity analysis. Similar descriptions are omitted here for brevity.

[0040] As mentioned above, the image resolution of the processed frame 707 is higher than the image resolutions of the transformed frame 703 and the processed transformed frame 705. By appropriate scaling and mapping, the locations of static pixel data in the input frame F2 can be predicted based on the locations of the static pixel data identified in the transformed frame TF2. Accordingly, the content activity analyzer circuit 704 derives processed pixel data, and generates a processed frame F2' by replacing the predicted static pixel data in the input frame F2 with the processed pixel data. For example, the processed pixel data is the static pixel data in the input frame F1. As another example, the processed static pixel data is generated by applying an arithmetic operation to the pixel data in the input frames F1 and F2.

[0041] Similarly, by appropriate scaling and mapping, the locations of static pixel data in the input frame F3 can be predicted based on the locations of the static pixel data identified in the transformed frame TF3. Accordingly, the content activity analyzer circuit 704 derives processed pixel data, and generates a processed frame F3' by replacing the predicted static pixel data in the input frame F3 with the processed pixel data. For example, the processed pixel data is the static pixel data in the processed frame F2'. As another example, the processed static pixel data is generated by applying an arithmetic operation to the pixel data in the input frame F3 and the processed frame F2'.

[0042] The video encoder circuit 706 is arranged to perform a video encoding process to generate a bitstream output of the video encoding apparatus 700, wherein information derived from the content activity analysis results (e.g., the processed frame 707) is referenced by the video encoding process. In the present embodiment, the video encoder circuit 706 encodes the input frame F1 to generate a first frame bitstream included in the bitstream output, encodes the processed frame F2' to generate a second frame bitstream included in the bitstream output, encodes the processed frame F3' to generate a third frame bitstream included in the bitstream output, and so on. It is noted that the video encoder circuit 706 can be implemented by any suitable encoder architecture. That is, the present application does not limit the encoder architecture employed by the video encoder circuit 706.

[0043] Figure 9A diagram showing a fifth video encoding apparatus according to an embodiment of the present application is shown. The video encoding apparatus 900 comprises a content activity analyzer circuit (labeled "content activity analyzer") 904, a video encoder circuit (labeled "video encoder") 906 and the image transformer circuit (labeled "image transformer") 702 described above. Like the content activity analyzer circuit 704, the content activity analyzer circuit 904 is arranged to apply a content activity analysis process to successive frames to generate content activity analysis results. In this embodiment, the successive frames received by the content activity analyzer circuit 904 are the transformed frames 703. The content activity analyzer circuit 904 differs from the content activity analyzer circuit 704 in that the content activity analysis results generated by the content activity analyzer circuit 904 comprise the processed transformed frame 707, the processed transformed frame 705 and the activity indication 901. Since the principles of generating the processed transformed frame 705 and the processed frame 707 are described in the paragraphs above, they are not described further here for brevity. In some embodiments of the present application, the activity indication 901 can be generated after the content activity analysis of the processed transformed frame 705 from the content activity analysis of the transformed frame 703. The generation of the activity indication 901 depends on the content activity analysis applied to the transformed frame 703. In other words, the activity indication 901 can be a by-product of the content activity analysis process performed by the content activity analyzer circuit 904. For example, the activity indication 901 can comprise a plurality of activity indication maps, each activity indication map recording one activity indication for each of a plurality of pixel blocks.

[0044] Figure 10 A diagram showing a 2-mode activity indication derived from successive frames (e.g. two transformed frames, or one transformed frame and one processed transformed frame) according to an embodiment of the present application is shown. Figure 4 The difference between the 2-mode activity indication calculation in Figure 10 The difference between the 2-mode activity indication calculation in Figure 10 As shown, the content activity analyzer circuit 904 can perform content activity analysis on the pixel data in the transformed frames TF1 and TF2 to generate an activity indication map TF_MAP12, which is a 2-mode activity indication map comprising static pixel data indications 1002 and non-static pixel data indications 1004, where the static pixel data indications 1002 represent no motion activity between blocks located at the same position in the transformed frames TF1 and TF2, and the non-static pixel data indications 1004 represent motion activity between blocks located at the same position in the transformed frames TF1 and TF2. Alternatively, Figure 10 As shown, the transformed frames TF1 and TF2 can be replaced with a transformed frame (e.g. TF3) and a processed transformed frame (e.g. TF2'), such that the 2-mode activity indication map is derived from content activity analysis of the pixel data in the transformed frame and the processed transformed frame.

[0045] Figure 11A diagram showing 3-mode activity indication derived from consecutive frames (e.g. transformed frames, or one transformed frame and one processed transformed frame) according to embodiments of the application. Figure 5 The difference between the shown 3-mode activity indication calculations and Figure 11 The difference between the shown 3-mode activity indication calculations and Figure 11 As shown, the content activity analyzer circuit 904 can perform content activity analysis on the pixel data in the transformed frames TF1 and TF2 to generate an activity indication map TF_MAP12, which is a 3-mode activity indication map including a static pixel data indication 1102, a non-static pixel data indication 1104, and a motion (or static) pixel data profile indication 1106, where the static pixel data indication 1102 represents no motion activity between blocks located at the same location in the transformed frames TF1 and TF2, the non-static pixel data indication 1104 represents motion activity between blocks located at the same location in the transformed frames TF1 and TF2, and the motion (or static) pixel data profile indication 1106 represents a profile of motion activity between the transformed frames TF1 and TF2. Optionally, as shown, Figure 11 The shown transformed frames TF1 and TF2 can be replaced with a transformed frame (e.g. TF3) and a processed transformed frame (e.g. TF2') as shown, such that the 3-mode activity indication map is derived from content activity analysis of pixel data in the transformed frame and the processed transformed frame.

[0046] The video encoder circuit 906 is arranged to perform video encoding processing to generate a bitstream output of the video encoding apparatus 900, where information derived from the content activity analysis results (e.g. the processed frames 707 and the activity indication 901) is referenced by the video encoding processing. In this embodiment, the video encoder circuit 906 encodes the input frame F1 to produce a first frame bitstream included in the bitstream output. Further, by appropriate scaling and mapping of the activity indication 901, the video encoder circuit 906 encodes the processed frame F2' according to the activity indication 901 (specifically, the activity indication map derived from the transformed frames TF2 and TF1) to generate a second frame bitstream included in the bitstream output, encodes the processed frame F3' according to the activity indication 901 (specifically, the activity indication map derived from the transformed frame TF3 and the processed transformed frame TF2') to generate a third frame bitstream included in the bitstream output, and so on.

[0047] With respect to the 2-mode activity indication, it will give two different instructions to the video encoder circuit 906. For example, a plurality of 2-mode activity indication maps are respectively associated with the processed frames 707. That is, the video encoder 906 refers to the 2-mode activity indication map associated with the current processed frame to be encoded by the video encoder 906 to determine how to encode each coding unit (coding block) within the current processed frame. When a coding unit (coding block) in the current processed frame to be encoded is found to be associated with the non-static pixel data indication 1004 through appropriate scaling and mapping, the video encoder circuit 906 can encode the coding unit (coding block) in a typical manner as prescribed by the coding standard. When a coding unit (coding block) in the current processed frame to be encoded is found to be associated with the static pixel data indication 1002 through appropriate scaling and mapping, the video encoder circuit 906 can force the coded motion vector of the coding unit to be zero, or can encode the coding unit in a skip mode. However, these are for illustrative purposes only and are not meant to limit the present application.

[0048] With respect to the 3-mode activity indication, it will give three different instructions to the video encoder circuit 906. For example, a plurality of 3-mode activity indication maps are respectively associated with the processed frames 707. That is, the video encoder 906 refers to the 3-mode activity indication map associated with the current processed frame to be encoded by the video encoder 906 to determine how to encode each coding unit (coding block) within the current processed frame. When a coding unit (coding block) in the current processed frame to be encoded is found to be associated with the non-static pixel data indication 1104 through appropriate scaling and mapping, the video encoder circuit 906 can encode the coding unit (coding block) in a typical manner as prescribed by the coding standard. When a coding unit (coding block) in the current processed frame to be encoded is found to be associated with the static pixel data indication 1102 through appropriate scaling and mapping, the video encoder circuit 906 can encode the coding unit in a skip mode. When a coding unit (coding block) in the current processed frame to be encoded is found to be associated with the motion (or static) pixel data contour indication 1106 through appropriate scaling and mapping, the video encoder circuit 906 can force the coded motion vector of the coding unit to be zero, or can encode the coding unit without residual information, or can encode the coding unit in a skip mode. However, these are for illustrative purposes only and are not meant to limit the present application.

[0049] In contrast to video encoder circuit 706 that encodes each coding unit (coding block) of a processed frame in a typical manner specified by a coding standard, video encoder circuit 906 refers to activity indication 901 to encode each coding unit (coding block) of a processed frame so that the reconstructed frame (decoded frame) at the decoding end has better image quality. It should be noted that video encoder circuit 906 can be implemented by any suitable encoder architecture. That is, the present disclosure does not limit the encoder architecture employed by video encoder circuit 906.

[0050] Figure 12 A diagram showing a sixth video encoding apparatus according to an embodiment of the present disclosure is shown. Video encoding apparatus 1200 comprises video encoder circuit (labeled "video encoder") 1206 and aforementioned image transformer circuit (labeled "image transformer") 702 and content activity analyzer circuit (labeled "content activity analyzer") 904. Video encoder circuit 1206 is arranged to perform a video encoding process to generate a bitstream output of video encoding apparatus 1200, wherein information derived from content activity analysis results (e.g. activity indication 901) is referenced by the video encoding process. Activity indication 901 can comprise a plurality of activity maps respectively associated with input frames 101. In this embodiment, video encoder circuit 1206 encodes input frame Fl to produce a first frame bitstream included in the bitstream output. Furthermore, with appropriate scaling and mapping of activity indication 901, video encoder circuit 1206 can encode the rest of input frames 101 with the help of activity indication 901. Video encoder circuit 1206 encodes input frame F2 according to activity indication 901 (specifically, the activity map derived from transformed frame TF2 and TFl) to generate a second frame bitstream included in the bitstream output, encodes input frame F3 according to activity indication 901 (specifically, the activity map derived from transformed frame TF3 and processed transformed frame TF2') to generate a third frame bitstream included in the bitstream output, and so on. It should be noted that video encoder circuit 1206 can be implemented by any suitable encoder architecture. That is, the present disclosure does not limit the encoder architecture employed by video encoder circuit 1206.

[0051] Those skilled in the art will readily observe that numerous modifications and alterations of the device and method can be made while remaining within the teachings of the present disclosure. Accordingly, the above disclosure is intended to be illustrative only and not limiting. The scope of the present disclosure should be limited solely by the appended claims.

Claims

1. A video coding device comprising: a content activity analyzer circuit configured to apply a content activity analysis process to a plurality of consecutive frames to generate a plurality of content activity analysis results, wherein the plurality of consecutive frames are derived from a plurality of input frames of the video coding device, and the content activity analysis process performed by the content activity analyzer circuit comprises: deriving a first content activity analysis result included in the plurality of content activity analysis results from a first frame and a second frame included in the plurality of consecutive frames, wherein the first content activity analysis result comprises a processed frame that is different from the first frame or the second frame; and deriving a second content activity analysis result included in the plurality of content activity analysis results from a third frame and the processed frame, wherein the third frame is included in the plurality of consecutive frames; and a video encoder circuit configured to perform a video coding process to generate a bitstream output of the video coding device, wherein information derived from the plurality of content activity analysis results is referenced by the video coding process, wherein the second frame is positioned between the first frame and the third frame; the processed frame is generated by replacing static pixel data identified by the content activity analysis process in the second frame with processed pixel data, wherein the processed pixel data is static pixel data in the first frame, or the processed data is generated by applying an arithmetic operation on pixel data in the first frame and the second frame; the processed frame is used as a substitute for the second frame to derive the second content activity analysis result, and the content activity analyzer circuit derives the second content activity analysis result by performing content activity analysis on pixel data of the third frame and the processed frame to identify static pixel data in the third frame.

2. The video coding device of claim 1, wherein, The plurality of consecutive frames received by the content activity analyzer circuit are the plurality of input frames of the video coding device.

3. The video coding device of claim 2, wherein, The video coding process performed by the video encoder circuit comprises: encoding the first frame to generate a first frame bitstream included in the bitstream output; and encoding the processed frame to generate a second frame bitstream included in the bitstream output.

4. The video coding device of claim 2, wherein, The first content activity analysis result further comprises an activity indication, the activity indication comprises a static pixel data indication and a non-static pixel data indication.

5. The video coding device of claim 4, wherein, In response to a coding unit associated with the static pixel data indication, the video encoder circuit is arranged to: force a coded motion vector of the coding unit to be zero; or encode the coding unit using a skip mode.

6. The video coding device of claim 4, wherein, The activity indication further comprises a motion pixel data contour indication.

7. The video encoding apparatus of claim 6, wherein, In response to a coding unit associated with the motion pixel data contour indication, the video encoder circuit is arranged to: force a coded motion vector of the coding unit to be zero; or encode the coding unit without residual information; or encode the coding unit using a skip mode.

8. The video coding device of claim 4, wherein, The video coding process performed by the video encoder circuit comprises: encoding the first frame to generate a first frame bitstream included in the bitstream output; and encoding the second frame according to the activity indication to generate a second frame bitstream included in the bitstream output.

9. The video encoding apparatus of claim 4, wherein, The video encoding process performed by the video encoder circuitry comprises: encoding the first frame to generate a first frame bitstream included in the bitstream output; and encoding the second frame according to the activity indication to generate a second frame bitstream included in the bitstream output.

10. The video encoding apparatus of claim 1, further comprising: image transformer circuitry configured to perform an image transformation process on a plurality of input frames of the video encoding apparatus to generate a plurality of transformed frames as a plurality of consecutive frames; wherein the first frame, the second frame and the third frame included in the plurality of consecutive frames are derived from a first input frame, a second input frame and a third input frame included in a plurality of input frames, respectively.

11. The video encoding apparatus of claim 10, wherein, The first content activity analysis result further comprises an activity indication, the activity indication comprising a static pixel data indication and a non-static pixel data indication. In response to a coding unit having the static pixel data indication, the video encoder circuitry is arranged to encode the coding unit in a skip mode. The activity indication further comprises a motion pixel data contour indication.

12. The video encoding apparatus of claim 10, wherein, 15. The video encoding apparatus of claim 14, in response to a coding unit having a motion pixel data contour indication, the video encoder circuitry is arranged to:

13. The video encoding apparatus of claim 12, wherein, force a coded motion vector of the coding unit to be zero; or 14. The video encoding apparatus of claim 12, wherein, encode the coding unit without residual information; or encode the coding unit using a skip mode. The first content activity analysis result further comprises an activity indication, the activity indication comprising a static pixel data indication and a non-static pixel data indication. In response to a coding unit having the static pixel data indication, the video encoder circuitry is arranged to encode the coding unit in a skip mode. The activity indication further comprises a motion pixel data contour indication.

16. The video encoding apparatus of claim 12, wherein, 15. The video encoding apparatus of claim 14, in response to a coding unit having a motion pixel data contour indication, the video encoder circuitry is arranged to: force a coded motion vector of the coding unit to be zero; or encode the coding unit without residual information; or 17. The video encoding apparatus of claim 12, wherein, encode the coding unit using a skip mode. The first content activity analysis result further comprises an activity indication, the activity indication comprising a static pixel data indication and a non-static pixel data indication. In response to a coding unit having the static pixel data indication, the video encoder circuitry is arranged to encode the coding unit in a skip mode. The activity indication further comprises a motion pixel data contour indication.

15. The video encoding apparatus of claim 14, in response to a coding unit having a motion pixel data contour indication, the video encoder circuitry is arranged to: force a coded motion vector of the coding unit to be zero; or encode the coding unit without residual information; or encode the coding unit using a skip mode. The first content activity analysis result further comprises an activity indication, the activity indication comprising a static pixel data indication and a non-static pixel data indication. In response to a coding unit having the static pixel data indication, the video encoder circuitry is arranged to encode the coding unit in a skip mode. The activity indication further comprises a motion pixel data contour indication.

15. The video encoding apparatus of claim 14, in response to a coding unit having a motion pixel data contour indication, the video encoder circuitry is arranged to: force a coded motion vector of the coding unit to be zero; or encode the coding unit without residual information; or encode the coding unit using a skip mode.

18. A video encoding method comprising: applying content activity analysis processing to a plurality of consecutive frames to generate a plurality of content activity analysis results, wherein the plurality of consecutive frames are derived from a plurality of input frames, and the content activity analysis processing comprises: deriving a first content activity analysis result included in the plurality of content activity analysis results from a first frame and a second frame included in the plurality of consecutive frames, wherein the first content activity analysis result comprises a processed frame that is different from the first frame or the second frame; and deriving a second content activity analysis result included in the plurality of content activity analysis results from a third frame and the processed frame, wherein the third frame is included in the plurality of consecutive frames; and performing video encoding processing to generate a bitstream output, wherein information derived from the plurality of content activity analysis results is referenced by the video encoding processing, wherein the second frame is located between the first frame and the third frame; the processed frame is generated by replacing static pixel data identified by the content activity analysis processing in the second frame with processed pixel data, wherein the processed pixel data is static pixel data in the first frame, or the processed data is generated by applying an arithmetic operation to pixel data in the first frame and the second frame; the processed frame is used as a substitute for the second frame to derive the second content activity analysis result, and the content activity analyzer circuit derives the second content activity analysis result by performing content activity analysis on pixel data of the third frame and the processed frame to identify static pixel data in the third frame.

Citation Information

Patent Citations

  • Identifying Content of Interest

    US20160110877A1

  • Video quality by controlling inter frame encoding according to frame position in GOP

    US8774272B1