Block-Based Picture Fusion for Contextual Segmentation and Processing
By dividing video frames into 4x4 blocks and optimizing quantization based on semantic information, the method enhances encoding efficiency and bitrate performance in video compression.
Patent Information
- Application Number
- JP2023025036
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-27
- Filing Date
- 2023-02-21
- Publication Date
- 2025-09-24
- Estimated Expiration
- 2039-11-27
AI Technical Summary
Existing video compression methods partition pictures into large blocks without considering underlying video information, leading to inefficient encoding and poor bitrate performance.
The method involves dividing video frames into smaller 4x4 blocks, determining regions based on semantic information, calculating average measures of information for these regions, and controlling quantization parameters based on these measures to optimize encoding.
This approach allows for finer granularity in encoding, improving encoding efficiency and bitrate performance by aligning with standard transform block sizes and using region-based block merging for contextual and semantic picture analysis.
Smart Images

Figure 0007743089000008 
Figure 0007743089000009 
Figure 0007743089000010
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 771,907, filed November 27, 2018, and entitled "BLOCK-BASED PICTURE FUSION FOR CONTEXTUAL SEGMENTATION AND PROCESSING," which is incorporated herein by reference in its entirety.
[0002] FIELD OF THE INVENTION The present invention relates generally to the field of video compression. Specifically, the present invention is directed to block-based picture fusion for contextual segmentation and processing. [Background technology]
[0003] A video codec may include electronic circuitry or software that compresses or decompresses digital video. It can convert uncompressed video into a compressed format, and vice versa. In the context of video compression, a device that compresses video (and / or performs some of the functions thereof) may typically be called an encoder, and a device that decompresses video (and / or performs some of the functions thereof) may be called a decoder.
[0004] The format of the compressed data can conform to standard video compression specifications. The compression can be lossy, in that the compressed video lacks some information present in the original video. Consequences of this can include the decompressed video having lower quality than the original uncompressed video, because insufficient information exists to accurately reconstruct the original video.
[0005] There can be a complex relationship between video quality, the amount of data used to represent the video (e.g., determined by bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, end-to-end delay (e.g., latency), and the like.
[0006] During encoding, a picture (e.g., a video frame) is partitioned (e.g., divided) into relatively large blocks, such as 128×128, and such structure is fixed. However, by partitioning a picture into large blocks for compression without considering the underlying video information (e.g., video content), the large blocks may not divide the picture in a manner that allows for efficient encoding, thereby resulting in poor bitrate performance. Summary of the Invention [Means for solving the problem]
[0007] In one aspect, an encoder includes circuitry configured to receive a video frame, divide the video frame into blocks, determine a first region within the video frame comprising a first group of a first subset of the blocks, determine a first average measure of information for the first region, and encode the video frame, including controlling a quantization parameter based on the first average measure of information for the first region.
[0008] In another aspect, a method includes, by an encoder, receiving a video frame and dividing the video frame into blocks. The method includes determining a first region within the video frame comprising a first group of a first subset of the blocks. The method includes determining a first average measure of information for the first region. The method includes encoding the video frame, wherein the encoding includes controlling a quantization parameter based on the first average measure of information for the first region.
[0009] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. The present invention provides, for example: (Item 1) 1. An encoder, the encoder comprising a circuit, the circuit comprising: receiving a video frame; Dividing the video frame into blocks; determining a first region within the video frame that includes a first group of the first subset of blocks; determining a first average measure of information in the first region; encoding the video frame, the encoding including controlling a quantization parameter using a first average measure of the information in the first region; an encoder configured to: (Item 2) Item 2. The encoder according to item 1, wherein the block size of the block is 4x4. (Item 3) Item 14. The encoder of item 1, further configured to receive semantic information, wherein the first group is determined based on the received semantic information. (Item 4) Item 5. The encoder of item 4, wherein the semantic information includes data characterizing face detection. (Item 5) Item 1. The encoder of item 1, wherein the first average measure of information is determined by calculating a sum of a plurality of information measures for the plurality of blocks. (Item 6) Item 6. The encoder of item 5, wherein the first average measure of information is further determined by multiplying the sum by a significance coefficient. (Item 7) Item 7. The encoder of item 6, wherein the significant coefficient is determined based on characteristics of the first region. (Item 8) Item 10. The encoder of item 1, wherein the controlling includes determining a first quantization size based on a first measure of the information. (Item 9) a transform and quantization processor; an inverse quantization and inverse transform processor; an in-loop filter; a decoder picture buffer; a motion estimation and compensation processor; an intra-prediction processor; Item 1. The encoder of item 1, further comprising: (Item 10) determining a second region within the video frame comprising a second group of a second subset of the blocks; determining a second average measure of information in the second region, wherein said controlling is further based on the second average measure of said information in the second region; Item 1. The encoder of item 1, further configured to: (Item 11) 1. A method, comprising: receiving, by an encoder, video frames; Dividing the video frame into blocks; determining a first region within the video frame that includes a first group of the first subset of blocks; determining a first average measure of information in the first region; encoding the video frame, the encoding including controlling a quantization parameter using a first average measure of the information in the first region; A method comprising: (Item 12) Item 12. The method according to item 11, wherein the block size of the block is 4x4. (Item 13) Item 12. The method of item 11, further comprising receiving semantic information, wherein the first group is determined based on the received semantic information. (Item 14) Item 14. The method of item 13, wherein the semantic information includes data characterizing face detection. (Item 15) Item 12. The method of item 11, wherein the first average measure of information is determined by calculating a sum of a plurality of information measures for the plurality of blocks. (Item 16) Item 17. The method of item 16, wherein the first average measure of information is further determined by multiplying the sum by a significance coefficient. (Item 17) Item 18. The method of item 17, wherein the significance coefficient is determined based on characteristics of the first region. (Item 18) Item 12. The method of item 11, wherein the controlling includes determining a first quantization size based on a first measure of the information. (Item 19) The encoder comprises: a transform and quantization processor; an inverse quantization and inverse transform processor; an in-loop filter; a decoder picture buffer; a motion estimation and compensation processor; an intra-prediction processor; Item 12. The method of item 11, comprising: (Item 20) determining a second region within the video frame comprising a second group of a second subset of the blocks; determining a second average measure of information in the second region, wherein said controlling is further based on the second average measure of said information in the second region; Item 12. The method of item 11, further comprising: [Brief explanation of the drawings]
[0010] For the purpose of illustrating the invention, the drawings show aspects of one or more embodiments of the invention. It being understood, however, that this invention is not limited to the precise arrangements and instrumentalities shown in the drawings.
[0011] [Figure 1] FIG. 1 is a process flow diagram illustrating an example process for encoding video that may utilize 4x4 blocks as a base partitioning size, which may allow for finer granularity by the encoder, and such sizes align with established transform block sizes.
[0012] [Figure 2] FIG. 2 is an illustrative example of the segmentation and fusion process for a picture with a human face.
[0013] [Figure 3] FIG. 3 is a series of images illustrating another example of the segmentation and fusion process according to some implementations of the present subject matter.
[0014] [Figure 4] FIG. 4 is a system block diagram illustrating an example video encoder capable of block-based picture merging and contextual segmentation and processing.
[0015] [Figure 5] FIG. 5 is a block diagram of a computing system that may be used to implement any one or more of the methods and any one or more portions thereof disclosed herein.
[0016] The drawings are not necessarily to scale and may be illustrated by phantom lines, schematic representations, and partial views. In some instances, details that are not necessary for an understanding of the embodiments or that make other details difficult to perceive may be omitted. Like reference symbols in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION
[0017] Some implementations of the present subject matter are directed to approaches to encoding video that perform picture partitioning using sample blocks as basic units. The sample blocks may have a uniform sample block size, which may be the side length in square pixels of pixels; for example, without limitation, the embodiments disclosed herein use 4x4 sample blocks as the basic unit, and thus may align with some typical standard sizes of transforms in video and image coding. By utilizing 4x4 blocks as the basic partitioning size, some implementations of the present subject matter may allow for finer granularity by the encoder, such sizes being consistent with established transform block sizes and being standard and defined. This allows for the use of transformation matrices, which can improve encoding efficiency. Furthermore, such an approach may contrast with some existing approaches to encoding that use fixed block structures of relatively large size. Those skilled in the art will understand, after reviewing this disclosure in its entirety, that while for brevity, 4x4 sample blocks are described in many of the following examples, in general, sample blocks of any size or shape, according to any method of measurement, may be used for picture division and / or partitioning.
[0018] In some implementations, the present subject matter involves using region-based block merging to enable contextual and semantic picture analysis and processing.
[0019] 1 is a process flow diagram illustrating an example process for encoding video that may utilize 4x4 blocks as a basic partitioning size, which may allow for finer granularity by an encoder implementing and / or performing the process, and such sizes may align with established transform block sizes. In step 105, video frames are received by the encoder. This may be accomplished in any manner suitable for receiving video in the form of a stream and / or file from any device and / or input port. Receiving video frames may include reading from the encoder's memory and / or a computing device in communication with, incorporating, and / or embedded within the encoder. Receiving may include receiving from a remote device via a network. Receiving video frames may include receiving multiple video frames that combine to comprise one or more videos.
[0020] 1, at step 110, the encoder partitions the video frame into blocks, for example, by dividing the video frame into blocks including, but not limited to, blocks having a size of 4 pixels by 4 pixels (4x4). The 4x4 size may be compatible with many standard video resolutions that can be divided into an integer number of 4x4 blocks.
[0021] Continuing with reference to FIG. 1 , block merging is performed at step 115. Block merging may include determining a first region within the video frame that includes a first group of a first subset of blocks. In block merging, each block may be assigned to a region. The assignment logic, such as, but not limited to, semantic information, may be obtained from an external source, and the semantic information may include, as a non-limiting example, information provided from a face detector, such that the semantic information includes data characterizing face detection. Thus, the first group may be determined based on the received semantic information. In some implementations, the assignment logic may be predefined, for example, according to some clustering or grouping algorithm. The encoder may further be configured to determine a second region within the video frame that includes a second group of a second subset of blocks.
[0022] 2 is an illustrative example of a segmentation and fusion process for a picture containing a human face. Any block having at least one pixel belonging to an object of interest, such as, but not limited to, a face, as identified, for example, via received semantic information, may be assigned to the shaded region (A2) and fused with other blocks within that region, for example, according to the received semantic information.
[0023] 1, in step 120, a first average measure of information for a first region may be determined. The measure of information may include, for example, the level of detail of the region. For example, a smooth region or a highly textured region may contain different amounts of information.
[0024] Continuing to refer to FIG. 1, a first average measure of information may be determined, by way of non-limiting example, according to the sum of the information measures relating to the individual blocks within the first region, where the sum of the information measures may be weighted and / or multiplied, for example, by a significance factor as shown in the summation below:
number
number
[0025] In some implementations, integer approximations of the transform matrix may be utilized, which may be used for efficient hardware and software implementations. For example, if the blocks as described above are 4x4 blocks of pixels, the generalized discrete cosine transform matrix may include a generalized discrete cosine transform II matrix of the form:
number
[0026] Continuing with reference to FIG. 1, if the encoder is further configured to determine a second region within the video frame (e.g., including, but not limited to, a second group of a second subset of blocks) as described above with reference to FIG. 1, the encoder may be configured to determine a second average measure of information for the second region, and determining the second average measure of information may be performed as described above for determining the first average measure of information.
[0027] Continuing to refer to Figure 1, the significant coefficient S N may be provided by an external expert and / or calculated based on the characteristics of the first region (e.g., the fused block). A "characteristic" of a region, as used herein, is a measurable attribute of the region that is determined based on the content of the region, and the characteristic may be expressed numerically using the output of one or more operations performed on the first region. The one or more operations may include any analysis of any signal represented by the first region. One non-limiting example is in quality modeling applications, where a higher S is found for regions with a smooth background. N and assigns a lower S for regions with less smooth background. NAs a non-limiting example, smoothness may be determined using Canny edge detection to determine the number of edges, with lower numbers indicating a higher degree of smoothness. A further example of automatic smoothness detection may include the use of a Fast Fourier Transform (FFT) over a signal in spatial variables over a region, in which the signal may be analyzed over any two-dimensional coordinate system and over channels representing red-green-blue color values or equivalent, where a higher relative dominance in the frequency domain of lower frequency components as calculated using the FFT may indicate a higher degree of smoothness, while a higher relative dominance of higher frequencies may indicate more frequent and rapid transitions in color and / or shade values over background regions, which may result in a lower smoothness score, and semantically significant objects may be identified by user input. Alternatively, or in addition, semantic significance may be detected according to edge configuration and / or texture patterns. The background may be identified (without limitation) by receiving and / or detecting portions of a region that represent significant or "foreground" objects, such as faces or other items, including, but not limited to, semantically significant objects. Another example is a higher S for regions that contain semantically significant objects, such as human faces. N may include assigning
[0028] FIG. 3 illustrates another example of the partitioning and fusion process according to some implementations of the present subject matter. A sequence of images. The input image (a) may be partitioned into 4x4 blocks as shown in (b). In (c), the blocks may be merged using predefined logic into a larger set of regions, which may include, but are not limited to, the same 16x16 region set as shown in (d). In (e), segmentation may be performed using semantic information, and the corresponding regions (A1, A2, and A3) of the segmented flower field, clouds, and remaining clear sky are shown in (f). In this example, region A3 may have the lowest significance coefficient, and region A1 may have the highest significance coefficient (smoothest background).
[0029] Referring again to FIG. 1 , at 125, a video frame may be encoded. Encoding may include controlling a quantization parameter based on a first average measure of information in a first region, where the quantization parameter may include, be equal to, proportional to, and / or linearly related to a measure of a quantization size and / or a quantization level. As used in this disclosure, a “quantization level” and / or “quantization size” is a quantity indicating the amount of information to be discarded in the compression of a video frame, where a quantization level may include a numerical value, such as but not limited to, an integer, by which one or more coefficients, including but not limited to, transform coefficients, are divided and / or reduced to reduce the information content of an encoded and subsequently decoded frame. Controlling may include determining a first quantization size based on the first measure of information, where the quantization level may represent a direct or indirect measure of memory storage required to capture information describing luma and / or chroma data of pixels in a block, where a higher number of bits may be required to store information with a higher degree of variance as determined by the first measure of information. The quantization size may be based on a first measure of information as described above, where the quantization size may be larger for a higher first measure of information and smaller for a lower first measure of information, and the quantization size may be proportional and / or linearly related to the first measure of information. Generally, more information content may result in a larger quantization size. By controlling the quantization size, information about the fused block area may be used for rate-distortion optimization for encoding. The controlling may further be based on a second average measure of information of the second region.
[0030] In some implementations, a first average measure of the information may be used for the quality calculation.
[0031] 4 is a system block diagram illustrating an example video encoder 400 capable of block-based picture fusion and contextual partitioning and processing. The example video encoder 400 receives an input video 404, which may first be partitioned or divided into 4x4 blocks for further processing.
[0032] 4, the exemplary video encoder 400 includes an intra-prediction processor 408, a motion estimation / compensation processor 412 (also referred to as an inter-prediction processor), a transform / quantization processor 416, an inverse quantization / inverse transform processor 420, an in-loop filter 424, a decoded picture buffer 428, and an entropy coding processor 432. Bitstream parameters may be input to the entropy coding processor 432 for inclusion in an output bitstream 436.
[0033] Continuing with reference to FIG. 4, the transform / quantize processor 416 may be capable of performing block merging and calculating a measure of information for each region.
[0034] Continuing with reference to FIG. 4, in operation, a decision may be made for each block of a frame of input video 404 whether to process the block via intra-picture prediction or using motion estimation / compensation. The block may be processed by either the intra-picture prediction processor 408 or the motion estimation / compensation processor 409. The block may be provided to a motion estimation / compensation processor 412. If the block is to be processed via intra prediction, the intra prediction processor 408 may perform the processing and output a predictor. If the block is to be processed via motion estimation / compensation, the motion estimation / compensation processor 412 may perform the processing.
[0035] Continuing with reference to FIG. 4, a residual may be formed by subtracting the predictor from the input video. The residual may be received by a transform / quantization processor 416, which may perform a transform operation (e.g., a discrete cosine transform (DCT)) to produce coefficients, which may be quantized. The quantized coefficients and any associated signaling information may be provided to an entropy coding processor 432 for entropy encoding and inclusion in an output bitstream 436. Additionally, the quantized coefficients may be provided to an inverse quantization / inverse transform processor 420, which may reconstruct pixels, which may be combined with the predictor and processed by an in-loop filter 424, the output of which is stored in a decoded picture buffer 428 for use by the motion estimation / compensation processor 412.
[0036] It should be noted that any one or more of the aspects and embodiments described herein may be conveniently implemented using digital electronic circuitry, integrated circuits, specially designed application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof, embodied in and / or implemented on one or more machines programmed according to the teachings herein (e.g., one or more computing devices utilized as user computing devices for electronic documents, one or more server devices such as document servers, etc.), as would be apparent to those skilled in the computer arts. These various aspects or features may include implementation in one or more computer programs and / or software executable and / or readable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device. Appropriate software coding may be readily prepared by skilled programmers based on the teachings of the present disclosure, as would be apparent to those skilled in the software arts. Aspects and implementations discussed above that employ software and / or software modules may also include appropriate hardware to assist in implementing the machine-executable instructions of the software and / or software modules.
[0037] Such software may be a computer program product employing a machine-readable storage medium. A machine-readable storage medium may be any medium capable of storing and / or encoding sequences of instructions for execution by a machine (e.g., a computing device) and causing the machine to perform any one of the methods and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random-access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROMs, EEPROMs, programmable logic devices (PLDs), and / or any combination thereof. As used herein, machine-readable medium is intended to include a single medium as well as a collection of physically separate media, such as, for example, a compact disc or a collection of one or more hard disk drives combined with computer memory. As used herein, machine-readable storage medium does not include transitory forms of signal transmission.
[0038] Such software may also include information (e.g., data) carried in a data signal on a data carrier, such as a carrier wave. For example, machine-executable information may be included as a data carrier signal embodied in a data carrier, which signal encodes a sequence of instructions, or portions thereof, for execution by a machine (e.g., a computing device), and any associated information (e.g., data structures and data) that causes the machine to perform any one of the methods and / or embodiments described herein.
[0039] Examples of computing devices include, but are not limited to, e-book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smartphones, etc.), web appliances, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions that define actions to be taken by the machine, and any combination thereof. In one example, a computing device may include and / or be included within a kiosk.
[0040] 5 shows a diagrammatic representation of one embodiment of a computing device as an exemplary form of computer system 500 upon which a set of instructions for causing a control system to implement any one or more of the aspects and / or methodologies of the present disclosure may be executed. It is also contemplated that multiple computing devices may be utilized to implement a set of instructions specifically configured to cause one or more of the devices to implement any one or more of the aspects and / or methods of the present disclosure. Computer system 500 includes a processor 504 and a memory 508, which communicate with each other and with other components via a bus 512. Bus 512 may include any of several types of bus structures, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures.
[0041] Memory 508 may include a variety of components (e.g., machine-readable media), including, but not limited to, random-access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 516 (BIOS), containing the basic routines that help to transfer information between elements within computer system 500, such as during start-up, may be stored in memory 508. Memory 508 may also include (e.g., stored on one or more machine-readable media) instructions (e.g., software) 520 that embody any one or more of the aspects and / or methods of the present disclosure. In another example, memory 508 may further include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0042] Computer system 500 may also include a storage device 524. Examples of a storage device (e.g., storage device 524) include, but are not limited to, a hard disk drive, a magnetic disk drive, an optical disk drive combined with optical media, a solid-state memory device, and any combination thereof. Storage device 524 may be connected to bus 512 by an appropriate interface (not shown). Exemplary interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE®), and any combination thereof. In one example, storage device 524 (or one or more of its components) may be removably interfaced with computer system 500 (e.g., via an external port connector (not shown)). In particular, storage device 524 and associated machine-readable medium 528 may be connected to bus 512 by an appropriate interface (not shown). The software 520 may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 500. In one example, the software 520 may reside, completely or partially, within the machine-readable media 528. In another example, the software 520 may reside, completely or partially, within the processor 504.
[0043] Computer system 500 may also include input devices 532. In one example, a user of computer system 500 may type commands and / or other information into computer system 500 via input devices 532. Examples of input devices 532 include, but are not limited to, an alphanumeric input device (e.g., a keyboard), a pointing device, a joystick, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), a cursor control device (e.g., a mouse), a touchpad, an optical scanner, a video capture device (e.g., a still camera, a video camera), a touchscreen, and any combination thereof. Input devices 532 may interface to bus 512 via any of a variety of interfaces (not shown), including, but not limited to, a serial interface, a parallel interface, a gameport, a USB interface, a FIREWIRE® interface, an interface directly to bus 512, and any combination thereof. Input devices 532 may include a touchscreen interface, which may be part of or separate from display 536, discussed further below. The input device 532 may be utilized as a user selection device for selecting one or more graphical representations in a graphical interface as described above.
[0044] A user may also input commands and / or other information into computer system 500 via storage device 524 (e.g., a removable disk drive, flash drive, etc.) and / or network interface device 540. A network interface device, such as network interface device 540, may be utilized to connect computer system 500 to one or more of various networks, such as network 544, and one or more remote devices 548 connected thereto. Examples of network interface devices include, but are not limited to, a network interface card (e.g., a mobile network interface card, a LAN card), a modem, and any combination thereof. Examples of networks include, but are not limited to, a wide area network (e.g., the Internet, an enterprise network), a local area network (e.g., a network associated with an office, building, campus, or other relatively small geographic space), a telephone network, a data network associated with a telephone / voice provider (e.g., a mobile communications provider's data and / or voice network), a direct connection between two computing devices, and any combination thereof. A network, such as network 544, may employ wired and / or wireless modes of communication. In general, any network topology may be used. Information (eg, data, software 520 , etc.) may be communicated to and / or from computer system 500 via network interface device 540 .
[0045] Computer system 500 may further include a video display adapter 552 for communicating images displayable on a display device, such as display device 536. Examples of display devices include, but are not limited to, a liquid crystal display (LCD), a cathode ray tube (CRT), a plasma display, a light emitting diode (LED) display, and any combination thereof. Display adapter 552 and display device 536 communicate with processor 504 to provide graphical representations of aspects of the present disclosure. In addition to a display device, computer system 500 may include one or more other peripheral output devices, including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices may be connected to bus 512 via a peripheral interface 556. Examples of peripheral interfaces include, but are not limited to, a serial port, a USB connection, a FIREWIRE® connection, a parallel connection, and any combination thereof.
[0046] The foregoing is a detailed description of illustrative embodiments of the present invention. Various modifications and additions may be made without departing from the spirit and scope of the present invention. Features of each of the various embodiments described above may be combined with features of other described embodiments, as appropriate, to provide a combination of features in related new embodiments. Moreover, while the foregoing describes several separate embodiments, what has been described herein is merely illustrative of the application of the principles of the present invention. In addition, while certain methods herein may be illustrated and / or described as being performed in a particular order, the order can be varied considerably within ordinary skill in order to achieve the embodiments as disclosed herein. Therefore, this description is intended to be taken by way of example only and is not intended to otherwise limit the scope of the present invention.
[0047] In the above description and in the claims, phrases such as "at least one of" or "one or more of" may appear and may be followed by a conjunctive listing of elements or features. The term "and / or" may also occur within a listing of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which such phrase is used, this is intended to mean any of the listed elements or features individually or in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" are each intended to mean "A only, B only, or both A and B." A similar interpretation is intended with respect to listings containing more than two items. For example, the phrases "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, and / or C" are intended to mean "A only, B only, C only, both A and B, both A and C, both B and C, or both A, B, and C," respectively. Additionally, use of the term "based on" above and in the claims is intended to mean "based at least on," such that unrecited features or elements are also allowed.
[0048] The subject matter described herein can be embodied as systems, devices, methods, and / or articles, depending on the desired configuration. The implementations described in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. While some variations have been described in detail above, other modifications or additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features disclosed above. In addition, the logic flow depicted in the accompanying figures and / or described herein does not necessarily require the particular order or sequential order shown to achieve desirable results. Other implementations may be within the scope of the following claims.
Claims
1. 1. A video signal processor, comprising: an inverse quantizer; an inverse transform processor; an in-loop filter; Decoded Picture Buffer and Equipped with the video signal processor is configured to receive a video signal including a picture with pixels quantized using the inverse quantizer; The picture is a first region comprising a first plurality of blocks and having a first quantization parameter controlled by an encoder based on a first average measure of information of the first plurality of blocks, the first average measure of information being determined according to a sum of information measures for individual blocks within the first region weighted by a significance factor; [Equation 1] a first region, where N identifies said first region, S N is a significance coefficient, k is a subscript corresponding to a block of said first plurality of blocks comprising said first region, n is the number of blocks comprising said first region, B k is a measure of spatial activity determined using a discrete cosine transform of block k, and A N is a first average measure of said information; a second region comprising a second plurality of blocks and having a second quantization parameter controlled by the encoder based on a second average measure of information of the second plurality of blocks; Including, the inverse quantizer is configured to inverse quantize pixels of the first plurality of blocks using the first quantization parameter and to inverse quantize pixels of the second plurality of blocks using the second quantization parameter.
2. The processor of claim 1 , wherein the picture comprises at least one 128×128 coding unit.
3. The processor of claim 1 , wherein the significance coefficient is determined based on characteristics of the first region.
4. The processor of claim 1 , wherein the first region has a quantization size determined based on a first measure of the information.
5. The processor of claim 1 , wherein the second average measure of information is a second average measure of spatial activity.
6. The processor of claim 5 , wherein the second average measure of spatial activity is determined using a discrete cosine transform of blocks in the second plurality of blocks.
7. 1. A method, comprising: receiving a video signal including a picture with quantized pixels by a video signal processor including an inverse quantizer, an inverse transform processor, an in-loop filter, and a decoded picture buffer; The picture is a first region comprising a first plurality of blocks and having a first quantization parameter controlled by an encoder based on a first average measure of information of the first plurality of blocks, the first average measure of information being determined according to a sum of information measures for individual blocks within the first region weighted by a significance factor; [Equation 2] a first region, where N identifies said first region, S N is a significance coefficient, k is a subscript corresponding to a block of said first plurality of blocks comprising said first region, n is the number of blocks comprising said first region, B k is a measure of spatial activity determined using a discrete cosine transform of block k, and A N is a first average measure of said information; a second region comprising a second plurality of blocks and having a second quantization parameter controlled by the encoder based on a second average measure of information of the second plurality of blocks; and dequantizing pixels of the first plurality of blocks using the first quantization parameter and dequantizing pixels of the second plurality of blocks using the second quantization parameter; A method comprising:
8. The method of claim 7 , wherein the picture comprises at least one 128×128 coding unit.
9. The method of claim 7 , wherein the significance coefficient is determined based on characteristics of the first region.
10. The method of claim 7 , wherein the first region has a quantization size determined based on a first measure of the information.
11. The method of claim 7 , wherein the second mean measure of information is a second mean measure of spatial activity.
12. The method of claim 11 , wherein the second average measure of spatial activity is determined using a discrete cosine transform of blocks in the second plurality of blocks.
Citation Information
Patent Citations
Method and device for coding image signal and signal-recording medium
JP1998164581A
Image coder, its method and recording medium recording program
JP1999122622A
Motion picture coding device, motion picture coding method, motion picture coding program, and computer readable recording medium recording the program
JP2005269484A
Image encoding apparatus and method
JP2010016467A
Video Quality Measurement
JP2011527544A