Improved stepwise decoder refresh method and apparatus for video encoding and decoding
By dynamically adjusting the refresh area and intra-frame coding within the GDR range, the problem of GDR technology being unable to adapt to video content is solved, improving coding efficiency and bit rate stability, making it suitable for video encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-03-27
AI Technical Summary
Existing Gradual Decoder Refresh (GDR) technology, by uniformly limiting the refresh direction to horizontal or vertical, cannot adapt to video content, resulting in reduced coding efficiency, especially in Ultra Reliable Low Latency Communication (URLLC) where bit rate changes drastically.
By setting the frame length of the GDR interval, the refresh area is specified, and intra-frame encoding is performed on all spatial regions within the GDR interval. The refresh area is dynamically adjusted using functions and transformation patterns, reducing the need for additional data insertion.
While maintaining an appropriate bit rate, it provides effective random access and subjective image quality performance, thus improving encoding efficiency.
Smart Images

Figure CN121753324A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to video compression technology, and more specifically to an improved application method and utilization of the gradual decoder refresh technique, which contributes to bit rate stability, in video encoders and decoders.
[0002] This invention may be in the same technical field as at least one known digital video compression technology standard (e.g., MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG), or may be in the technical field of improving the inherent efficiency of such standards, or may be in the technical field of improving or replacing such standards. Background Technology
[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission via communication networks, video calls, video conversations, video chats, video content recorded and played using optical discs (including VCDs (video compact discs) / DVDs (digital versatile discs) / Blu-ray discs), the entire process of video content production, editing, collection, and distribution, and video shooting devices and cameras used in video shooting and recording activities (including personal, commercial, industrial, and security monitoring purposes) all rely on video encoding and decoding technologies.
[0004] Therefore, physical devices that can be called digital video encoders and decoders can constitute a part of a wide range of devices related to the generation, recording, and playback of digital video, including digital television, digital broadcasting systems, wireless broadcasting systems, laptops, desktop computers, tablets, e-book readers, digital cameras, digital video recording devices, digital multimedia playback devices, video game devices / terminals / control consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and others.
[0005] The aforementioned digital video encoder and decoder can be implemented using digital video compression standards widely used and understood by those skilled in the art. These digital video compression standards may include at least one known compression standard, such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0006] Video encoders and decoders can encode or decode digital video information more efficiently by conforming to the aforementioned standards, or by improving or modifying them. Attempts to modify these standards may also lead to new standards. A well-known example is the Enhanced Compression Model (ECM), developed by the Joint Video Experts Team (JVET), a joint international standards organization comprised of ISO, IEC, and ITU-T, aimed at improving upon and replacing the existing H.266 / VVC standard.
[0007] Among the various techniques applied to existing video encoders and decoders, there exists a technique collectively known as Gradual Decoder Refresh (GDR). GDR was proposed to address the following problem: Because the amount of information in randomly accessed frames, traditionally called "intra-frames," is much higher than that in "inter-frames," which are highly compressed based on time prediction, it causes a sharp change in bit rate when transmitted over ultra-high-speed communication networks, thus hindering the performance of ultra-reliable low-latency communication (URLLC). Summary of the Invention
[0008] The problem that the invention aims to solve
[0009] Current GDR (Generic Refresh Direction) technology applications uniformly limit the refresh direction to either horizontal or vertical, resulting in an inability to adapt the actual refresh direction to the video content and significantly reducing encoding efficiency. However, with the increasing prevalence of URLLC (URL Communication over Linux), the necessity for flexible application of GDR technology is growing. Therefore, a method is needed that can improve existing uniform GDR technology applications while minimizing the extra data inserted due to information indicating the refresh method.
[0010] Solution for solving the problem
[0011] To address the aforementioned technical problems, a video encoding method including a gradual decoder refresh (GDR) function according to an embodiment of the present invention is characterized by comprising the steps of: setting the frame length of the GDR interval for performing the gradual decoder refresh; specifying a refresh region including at least one image coding unit in at least one frame belonging to the GDR interval; performing intra-frame encoding on the at least one image coding unit included in the refresh region; and recording the encoding result into a bitstream, wherein, within the GDR interval, all spatial regions within the video resolution are intra-frame encoded at least once.
[0012] The steps of specifying the refresh area may include specifying the refresh area from the spatial area derived from a function with the frame number as a variable.
[0013] The step of specifying the refresh area may include: specifying the refresh area from the area derived from the function result of inputting the nth frame number, excluding the area derived from the function result of inputting the (n-1)th frame number.
[0014] The above function can be executed repeatedly for each image coding unit, and each time it is executed, the calculation is performed by combining the spatial location of the image coding unit, the number of the current frame, and the length of the GDR interval.
[0015] The above function can be defined by normalization based on the lower limit and the upper limit, and in order to derive the refresh area from the above normalized function, the above function can be projected onto the above frame as the encoding object.
[0016] The lower limit is 0, the upper limit is 1, and the projection method may include at least one of the following methods: projection that maintains aspect ratio, projection that does not maintain aspect ratio, and repeated projection.
[0017] In the encoding method, the above function can be a function that can derive the area of a line through integration, or a function that can calculate the area of a closed curve.
[0018] The above function can be applied after being transformed by at least one of the following operators: the operator for reversing the region to be determined as an area, the rotate operator, the skew operator, the scalaing operator, and the affine transform operator.
[0019] The above encoding method may further include the step of recording information related to the type of the above function into the above encoded bitstream.
[0020] Information related to the type of the above functions may be information indicating the index of the above function selected from the function list.
[0021] Information related to the type of the above functions may include at least one of the following: reference point, angle, magnification, tilt, inner / outer transformation, and affine transformation parameters for applying the above functions.
[0022] Information related to the type of the above functions may be information representing the mathematical expression of the above functions.
[0023] Information related to the type of the above functions can be entropy encoded.
[0024] The step of specifying the refresh area may include: specifying the refresh area from a spatial area specified by a sequentially determined transformation pattern that changes as the frame progresses.
[0025] The aforementioned transformation pattern can be defined by at least one bitmap image.
[0026] The aforementioned transformation pattern can be defined by at least one vector graphic.
[0027] The aforementioned transformation pattern can have a fixed resolution, and can be upscaled or downscaled from the fixed resolution to the resolution of the video frame in the actual application.
[0028] The shape that needs to be transformed in each frame of the above transformation pattern can be defined independently in chronological order.
[0029] The pattern shape that needs to be transformed in each frame of the above transformation pattern can be defined as the residual relative to the pattern that comes first in time.
[0030] The above transformation pattern can be applied after being transformed by at least one of the following operators: the operator for reversing the region to be determined as an area, the rotate operator, the skew operator, the scaling operator, and the affine transform operator.
[0031] The above encoding method may further include the step of recording information related to the type of the above transformation pattern into the above encoded bit stream.
[0032] Information related to the type of the aforementioned transformation pattern may be information indicating the index of the aforementioned transformation pattern selected from the list of transformation patterns.
[0033] Information related to the type of the aforementioned transformation pattern may include at least one of the following: reference point, angle, magnification, tilt, inner / outer transformation, and affine transformation parameters for applying the aforementioned transformation pattern.
[0034] Information related to the type of the aforementioned transformation pattern may be information representing at least one image constituting the aforementioned transformation pattern.
[0035] Information related to the type of the aforementioned transformation pattern can be entropy encoded.
[0036] The step of specifying the refresh area may include: specifying a spatial region such that all spatial regions within the video resolution within the GDR range are intra-coded at least once.
[0037] In the step of specifying the refresh area, when the GDR interval has a length of m frames, at least 1 / m of the spatial region in each frame is specified to be intra-frame encoded.
[0038] The above encoding method may further include the step of storing in the encoder cumulative refresh region information representing a spatial region that is intra-coded at least once within the above GDR interval.
[0039] In the step of specifying the refresh area, if it is the last frame of the GDR interval, then all spatial regions in the GDR interval that are not included in the cumulative refresh area information are specified as refresh areas.
[0040] The step of specifying the refresh area may include: specifying the step of changing at least one inter-frame coding object image coding unit to an intra-frame coding object image coding unit for encoding.
[0041] The aforementioned modification and specification steps may be operated based on the coding efficiency of the aforementioned inter-frame coding object image coding unit, and configured to be modified and specified when the coding efficiency of the aforementioned inter-frame coding object image coding unit is low.
[0042] The above coding efficiency can be predicted based on at least one of the following: the segmentation level of the coding tree unit (CTU), the distribution of DCT coefficient values, coding information including the coding mode set for the coding units of the surrounding image, and the coding result of the coded region.
[0043] Image coding units belonging to the refresh area mentioned above can be encoded without referencing image coding units outside the refresh area.
[0044] To address the aforementioned technical problems, a video decoding method including a gradual decoder refresh (GDR) function according to an embodiment of the present invention may include: receiving an encoded bitstream; reading at least one frame constituting a GDR interval from the bitstream; performing intra-frame decoding on at least one image coding unit contained in the at least one frame belonging to the at least one GDR interval; and recording the result of the intra-frame decoding into an output buffer, wherein, within the GDR interval, all spatial regions within the video resolution are intra-frame decoded at least once.
[0045] In each frame within the aforementioned GDR interval, the image region for which the aforementioned intra-frame decoding is performed can be contained within a functional region of spatial area derived from a function that takes the frame number as a variable.
[0046] The above decoding method can be executed based on at least one function pre-stored in the decoder.
[0047] The above decoding method may further include the step of reading function region information about at least one frame belonging to the above GDR interval from the bit stream.
[0048] The function region information mentioned above may include information indicating the index of the function selected from the function list.
[0049] The aforementioned function region information may include at least one of the following: reference point, angle, magnification, tilt, inner / outer transformation, and affine transformation parameters for applying the aforementioned function.
[0050] The aforementioned function region information may include the mathematical expression of the aforementioned function.
[0051] The aforementioned function region information can be entropy decoded.
[0052] In each frame within the aforementioned GDR interval, the image region for which the aforementioned intra-frame decoding is performed can be contained within a transform pattern region corresponding to a spatial area specified by a transform pattern that changes sequentially as the frames progress.
[0053] The above decoding method can be performed based on at least one transformation pattern pre-stored in the decoder.
[0054] The above decoding method may further include the step of reading transform pattern region information about at least one frame belonging to the above GDR interval from the bit stream.
[0055] The aforementioned transformation pattern area information may include information representing the index of the transformation pattern selected from the list of transformation patterns.
[0056] The aforementioned transformation pattern area information may include at least one of the following: reference point, angle, magnification, tilt, inner and outer transformation, and affine transformation parameters for applying the aforementioned transformation pattern.
[0057] The aforementioned transformation pattern region information may include information representing at least one image constituting the aforementioned transformation pattern.
[0058] The aforementioned transformed pattern region information can be entropy decoded.
[0059] In each frame within the aforementioned GDR interval, when the aforementioned GDR interval has a length of m frames, each frame of the image region for which the aforementioned intra-frame decoding is performed contains at least 1 / m of a spatial region.
[0060] The image coding unit performing the above intra-frame decoding can be decoded without referencing image coding units that do not belong to the same frame or have not been intra-frame coded.
[0061] Invention Effects
[0062] According to the present invention, by applying the improved GDR technology with minimal data consumption, it is expected that, compared with similar technologies in the field of video encoding and decoding, it can provide beneficial effects such as effective random access and subjective image quality while maintaining an appropriate bit rate. Attached Figure Description
[0063] Figure 1 This is a conceptual diagram of a video communication system according to an embodiment of the present invention.
[0064] Figure 2 This is a conceptual diagram of the arrangement of encoders and decoders in a real-time video streaming environment according to an embodiment of the present invention.
[0065] Figure 3 This is a conceptual diagram of the functional unit of a video decoder according to an embodiment of the present invention.
[0066] Figure 5 This is a conceptual diagram of a frame type according to an embodiment of the present invention.
[0067] Figure 6 This is a conceptual diagram showing the structure of a video encoder according to the H.266 / VVC standard.
[0068] Figure 7 This is a conceptual diagram illustrating the application of GDR technology.
[0069] Figure 8 This is a conceptual diagram illustrating the expected effects of applying GDR technology in terms of bit rate.
[0070] Figure 9This is a conceptual diagram illustrating the impact of applying GDR technology on the encoding / decoding process.
[0071] Figure 10 This is a schematic diagram illustrating the provision of frame information in the GDR interval according to an embodiment of the present invention.
[0072] Figure 11 This is a conceptual diagram of regions distinguished by functions according to an embodiment of the present invention.
[0073] Figure 12 This is another conceptual diagram of distinguishing regions by function according to an embodiment of the present invention.
[0074] Figure 13 This is a schematic diagram of a region differentiation method using functions according to a partial embodiment of the present invention.
[0075] Figure 14 This is a conceptual diagram illustrating the use of patterns to distinguish regions according to an embodiment of the present invention.
[0076] Figure 15 This is a schematic diagram illustrating various forms of a pattern according to an embodiment of the present invention.
[0077] Figure 16 This is a schematic diagram illustrating a video projection method for patterns according to various embodiments of the present invention.
[0078] Figure 17 This is a schematic diagram illustrating GDR implementation through forced intra-frame coding according to an embodiment of the present invention.
[0079] Figure 18 This is a flowchart illustrating a forced intra-frame coding method according to an embodiment of the present invention. Detailed Implementation
[0080] This invention can be modified and has various embodiments, with specific embodiments illustrated in the accompanying drawings and described in detail. However, this is not intended to limit the invention to specific implementations, but rather to encompass all modifications, equivalents, and substitutions that fall within the scope of the invention's ideas and techniques.
[0081] While terms such as "first" and "second" may be used to describe various components, the components described above should not be limited by these terms. The purpose of these terms is to distinguish one component from another. For example, a first component may be named a second component without departing from the scope of the invention, and similarly, a second component may be named a first component. The term "and / or" includes a combination of or one of several related described items and is not exclusive unless otherwise stated. The items listed in this application are merely illustrative descriptions intended to facilitate understanding of the concepts and possible implementations of the invention, and are therefore not intended to limit the scope of the embodiments of the invention.
[0082] In this specification, "A or B" can mean "A only", "B only", or "A and B". In other words, "A or B" in this specification can also be interpreted as "A and / or B". For example, "A, B or C" in this specification can mean "A only", "B only", "C only", or "any combination of A, B and C".
[0083] The forward slash ( / ) or comma used in this specification can mean "and / or". For example, "A / B" can mean "A and / or B". Accordingly, "A / B" can mean "A only", "B only", or "A and B". For example, "A, B, C" can mean "A, B, or C".
[0084] In this specification, "at least one of A and B" can mean "A only", "B only", or "A and B". Furthermore, the expressions "at least one of A or B" or "at least one of A and / or B" in this specification can be interpreted as "at least one of A and B".
[0085] Additionally, in this specification, "at least one of A, B and C" can mean "A only", "B only", "C only" or "any combination of all A, B and C". Furthermore, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C".
[0086] When a component is described as "connected" or "accessed" to another component, it should be understood that it may be directly connected to or accessed by another component, or there may be other components in between. Conversely, when a component is described as "directly connected" or "directly accessed" to another component, it should be understood that there are no other components in between.
[0087] The terminology used in this application is merely illustrative of specific embodiments and not limiting of the invention. Where there is no explicit indication of singularity in the context, the singular designation includes the meaning of plural. In this application, terms such as "comprising" or "possessing" indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, rather than precluding the presence or additional possibility of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0088] Unless otherwise specified, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms that are generally used identically to those defined in dictionaries have the same meaning as in the context of the relevant art, and unless explicitly defined, do not have an ideal or excessive meaning in this application.
[0089] In describing the present invention in this application, embodiments may be described or illustrated in the form of unit blocks that perform the described functions or functionalities. In this application, the aforementioned blocks may represent one or more devices, units, modules, components, etc. The aforementioned blocks may be implemented in hardware using one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware methods including but not limited to these. Alternatively, the aforementioned blocks may be implemented in software using application software, operating system software, firmware, or information processing software methods including but not limited to these. A block may be divided into multiple blocks performing the same function, and conversely, the function of simultaneously implementing multiple blocks may be performed by a single block. The aforementioned blocks may also be implemented as physically separated or combined according to any standard. The aforementioned blocks may also be implemented to operate in an environment where their physical location is not specified and they are isolated from each other via communication networks, the Internet, cloud services, or communication methods including but not limited to these. All of the above implementation methods fall within the scope of various embodiments that can be adopted by those skilled in the art of information and communication technology to achieve the same technical concept. Therefore, any specific implementation method should be interpreted as falling within the scope of the inventive technical concept of this application.
[0090] The preferred embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. To facilitate a comprehensive understanding of the invention, the same components will be denoted by the same reference numerals, and repeated descriptions of the same components will be omitted. Furthermore, the various embodiments are not mutually exclusive, and some embodiments can be combined with one or more other embodiments to form new embodiments.
[0091] Digital video codec
[0092] Figure 1 This is a conceptual diagram of a video communication system according to an embodiment of the present invention. The video communication system 100 described above can be configured to include at least two terminals 110 and 120 interconnected via a network 105.
[0093] In one embodiment of the present invention, the above Figure 1 This can be represented as a block diagram for configuring a one-way video communication network. To transmit video data 111 via network 105, the first terminal 110 can encode the video data. The second terminal 120 can be configured to receive the encoded video data 121 via the network, decode it, and display it.
[0094] In another embodiment of the present invention, the above Figure 1This can be represented as a block diagram for configuring a two-way video communication network. To achieve the aforementioned two-way video communication, each of the terminals 110 and 120 can be configured to encode its own acquired video data to transmit the video 112 and 122 to other terminals via the network. Each terminal can also be configured to receive video data sent via the network by other terminals 113 and 123, decode it, and display the decoded video data.
[0095] Figure 1 The terminals 110 and 120 shown above, according to embodiments, can be devices such as server computers, personal computers, portable computers, and smartphones, but are not limited thereto; any commonly used computing device can be used. For example, according to embodiments of the present invention, the terminals 110 and 120 can refer to desktop computers, laptop computers, tablet PCs, mobile phones, smartphones, personal digital assistants (PDAs), workstations, electronic calculators, server computers, cloud computers, virtual computers, quantum computers, or any other electronic, electrical, or quantum computing device implemented in a portable or non-portable form. In particular, the above-mentioned devices can also be interpreted as any device designed to operate as a terminal device according to an embodiment of the present invention, or capable of operating as a terminal device according to an embodiment of the present invention, or capable of installing and / or running computer programs to perform corresponding methods.
[0096] The aforementioned terminals 110 and 120 can be implemented by multiple functional units, which are interconnected in various forms, such as buses, circuits, or relationships between routines and subroutines, and configured to exchange information within the terminals 110 and 120. Furthermore, through these interconnections, these terminals can be configured to include a processor 130 with computing capabilities, and a memory 140 connected to the processor, to execute or support the operation of functional units that primarily require computation between the aforementioned functional units.
[0097] The processor 130 described in this specification may refer to one or more general-purpose or special-purpose computers, such as processors, controllers, arithmetic logic units (ALUs), digital signal processors, microcomputers, field-programmable gate arrays (FPAs), programmable logic units (PLUs), microprocessors, or any other device capable of executing and responding to instructions.
[0098] For ease of understanding, even though the processor 130 is referred to only in the singular, those skilled in the art will understand that the processor 130 may include multiple processing elements and / or various types of processing elements. For example, an apparatus according to an embodiment of the present invention may include multiple processors, or a processor and a controller, as the processor 130. Furthermore, the processor 130 may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.
[0099] The processor 130 can be configured to execute an operating system (OS) and one or more software programs running on the operating system. Furthermore, the processor can access, store, manipulate, process, and generate data in response to the execution of the software. The software may include computer programs, code, instructions, or a combination thereof, and can be configured as needed, or provide independent or collective instruction control over the processing device. The software can be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for parsing by the processor 130, or for providing instructions or data to the processor. The software can be distributed among multiple computer systems (e.g., terminals 110 and 120) connected to the network 105 and stored or executed in a distributed manner.
[0100] The aforementioned software can also be implemented in the form of program instructions executable by various computer means, and recorded or stored in the aforementioned memory 140. The aforementioned memory 140 can be a computer-readable recording medium, in which program commands, data files, and data structures can be recorded individually or in combination. The program commands stored in the aforementioned memory 140 can be implemented based on a command system specifically designed and configured for embodiments of the present invention, or can follow command systems known and available to those skilled in the art of computer software, such as command systems for assembly language, C, C++, Java, Python, etc. It should be understood that the aforementioned command system and the program commands it generates include not only machine language code generated by a compiler, but also high-level language code that can be executed by an interpreter or other means of apparatus according to an embodiment of the present invention and / or the aforementioned processor 130.
[0101] The computer-readable recording medium constituting an embodiment of the apparatus of the present invention includes the memory 140 described herein, and may include: temporary or volatile recording media such as processor cache, random access memory (RAM), flash memory, etc., which are maintained only during the operation of the processor; or relatively non-volatile or long-term recording media such as hard disk, floppy disk, magnetic tape, etc.; optical recording media such as CD-ROM, DVD, etc.; magneto-optical media such as floppy disk; or solid state memory; or read-only recording media such as a hardware-installed read-only memory (ROM). Furthermore, since the circuit wiring is configured to execute a series of program instructions and equivalent actions through a hard-wired structure, it can be considered that the various steps involved in executing the actions described in the embodiments of the present invention can be regarded as being recorded through the connection and arrangement of the aforementioned hardware components. Therefore, it is obvious to those skilled in the art that this connection and arrangement method can be regarded as equivalent to the memory 140 described above.
[0102] The above embodiments regarding the processor 130 and the memory 140 are not exclusive and can be selected or combined as needed. For example, a hardware device can be configured to execute the actions of an embodiment of the present invention as a module composed of more than one of the above-described software, and vice versa. As another example, in this specification, all or part of the actions assigned to a certain functional unit can be implemented by more than one of the above-described software stored in a device according to an embodiment of the present invention (preferably stored in a recording medium belonging to the scope of the above-described memory), and configured to be executed by the above-described processor. In this case, such a functional unit can be referred to as a functional unit "included" in the above-described processor.
[0103] This invention is applicable to any environment in which a one-way or two-way video communication network is established, and it should be understood that the network 105 described above can be configured to transmit encoded video data between the terminals 110 and 120 by any means.
[0104] In one embodiment of the present invention, the network 105 may refer to a wired or wireless communication network. According to this embodiment, the network can be configured to communicate information using any communication standard, including packet-based communication. Packet communication can be understood to include, for example, TCP or UDP packets. The wired communication method of the network 105 may be a method of connecting to an external communication network via RJ-11 standard telephone lines, various types of Ethernet cables conforming to the RJ-45 standard, other coaxial cables, metallic cables, optical cables, and other various wired media. According to the embodiments, the wireless communication method of the network 105 described above may include short-range wireless communication methods, such as Bluetooth, Wi-Fi, Zigbee, and near field communication (NFC); or may include long-range wireless communication methods, such as Wibro, WiMax, Global Systems for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long-Term Evolution (LTE), New Radio (NR), and other wireless communication technology names collectively referred to as 2G, 3G, 4G, 5G, and 6G communication standard generations, as well as international standard wireless communication standards such as IMT-2000, IMT-Advanced, IMT-2020, and IMT-2030. Of course, apart from this, any existing or newly developed wired or wireless communication means, methods, standards, and protocols applied to the implementation of the network 105, as long as they are configured for transmission and reception in the information communication devices such as terminals 110 and 120, will not affect the achievement of the purpose of this invention. Furthermore, it is obvious that a network 105 can be configured using a combination of more than one wired and / or wireless standard.
[0105] However, in another embodiment of the present invention, the network 105 described above can also be understood to include a process of transmitting information using a computer-readable recording medium. In this case, the configuration of the network is not limited to communication media, but should be understood to include a process of temporary storage and physical transmission in computer-readable storage and / or recording media. The computer-readable recording medium used for the above information transmission can be understood as a relatively non-volatile or long-term recordable recording medium, which is mainly used as a means of transmitting data between computing devices, especially magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as floppy disks; or solid-state memory.
[0106] Furthermore, regardless of the information communication or transmission means used, as long as its structure supports the transmission and decoding of video data in an encoded state, it can be considered to fall within the scope of the embodiments of this invention. Therefore, apart from the examples listed above, all known or future information communication or transmission means can be included within the scope of this invention.
[0107] Figure 2 This is a conceptual diagram of the arrangement of encoders and decoders in a real-time video streaming environment according to an embodiment of the present invention. Figure 2 The streaming media system 200 shown can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, even when information transmission is involved via recording media, the same or similar technical structures as described above can be applied.
[0108] According to one embodiment of the present invention, the streaming media system described above may include a video source 210 that generates the video stream. The video source may include a digital video capture unit 212 for capturing uncompressed raw video; for example, the digital video capture unit may be configured as a digital camera or other device. The raw video stream 215 may have a large capacity and can therefore be compressed by combining it with or connecting it to a video encoder 217 connected to the video source.
[0109] The encoder 217 described above can be configured as a unit including hardware, software, or a combination of both, for implementing a video encoding method and / or its implementation method according to an embodiment of the present invention.
[0110] The encoder 217 described above can output an encoded bitstream 219 with a reduced capacity compared to the original video stream. The bitstream 219 can be provided in real time for communication via a relay device (e.g., streaming media server 220) and / or stored on the recording medium 225 of the streaming media server 220 for later use.
[0111] The aforementioned streaming media system 200 may include at least one streaming media client 230, 240 that accesses the aforementioned streaming media server 220 to receive in real time or acquire later the aforementioned encoded bitstream 229. The aforementioned streaming media client may include a video decoder 232, which acquires the aforementioned encoded bitstream 229 (which may also be regarded as a copy of the bitstream 219 received by the aforementioned streaming media server), decodes the aforementioned bitstream 229, and outputs the resulting video data in a form that can be displayed on a display 235 or other visual, auditory, or other sensory display means.
[0112] As mentioned above, the encoding and decoding functions of video data are collectively referred to as a coder-decoder system, or video codec.
[0113] Figure 3 This is a conceptual diagram of the functional units of a video decoder according to an embodiment of the present invention. Figure 3 As shown, the receiving unit 310 can receive at least one encoded video data to be decoded by the decoder 305. In one embodiment of the present invention, each received encoded video data may be independent, and the decoding process of each independent video data may be independent of the decoding process of other video data. The encoded video data can be received to the receiving unit 310 via a hardware or software connection 315 with a storage device. As mentioned above, the storage device may refer to a streaming media server located at the other end of a communication network, or it may be a physical storage medium, but is not limited thereto.
[0114] The receiving unit 310 can receive the encoded video data and other accompanying data, such as encoded audio data or other auxiliary data, and each of the above data can be separated from the video data and provided to other appropriate processing units 312 other than the video decoder.
[0115] When receiving the aforementioned video data via a communication network, a buffer memory 320 can be connected between the receiving unit 310 and the decoder 305 to minimize latency and interruptions caused by the network environment. The buffer memory 320 can be a computer-readable recording medium used to temporarily store the received video data and stably provide it to the parser 330 corresponding to the input of the decoder 305. However, if the bandwidth of the communication network is sufficient, or if the video data is read from a local and not physically isolated recording medium, or if communication latency is not expected in other environments, the buffer memory may not be necessary.
[0116] The video decoder 305 may include the parser 330 as its input for parsing the encoded video data. The parser can separate (parsing) various types of information stored in bitstream form within the encoded video data according to predetermined rules, and, when necessary, perform entropy decoding 335 on the entropy-coded video data, thereby reconstructing symbols 338 as segments of video encoded information. The symbols 338 may include any information for controlling the operation of the decoder 305, and / or may further include information for controlling devices attached to the decoder 305 that perform operations (such as display devices). Control information for controlling the display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).
[0117] As described above, the parser 330 can be configured to perform entropy decoding 335 on the encoded video data. The entropy encoding method for the encoded video data may vary depending on the encoding standard, and the decoding operation should be performed accordingly. Representative examples of the entropy encoding standards may include variable length coding, Huffman coding, and arithmetic coding. These various encoding methods may be context-adaptive or context-sensitive methods according to the standard, or they may be based on principles well known to those skilled in the art.
[0118] The parser 330 described above can be configured to extract at least one local image from the encoded video data described above. The definition of the local image may vary depending on the encoding standard described above. Depending on the standard, one of the examples listed below may apply, or multiple examples may be applied simultaneously. For example, the local image may be defined as the following units: group of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), prediction units (PUs), etc.
[0119] The parser 330 can be configured to extract encoded information from the encoded video data, such as transform coefficients, quantization parameters (QPs), and / or motion vectors. The parser 330 can be configured to perform entropy decoding 335 and parsing operations on the video data received from the buffer memory, and selectively decode symbols 338 representing the encoded information. Furthermore, the parser 330 can also be configured to selectively provide specific symbols 338 to specific decoding functional units within the decoder 305, such as inverse quantization and inverse transform units 340, intra-frame prediction units 350, inter-frame prediction units 355, or loop filter units 360. The control over the provision of such information can be determined by the arrangement of information contained in the encoded video and may vary depending on the encoding standard, and is not limited to the scope of this embodiment, nor is it described in detail in this conceptual diagram.
[0120] The decoder 305 described above may be configured with multiple conceptual functional units for receiving and processing the encoded information from the parser 330 described above. It should be understood that these conceptual functional units can be combined or further subdivided according to implementation needs. For example, they can be further separated for ease of implementation; they can be integrated together to improve operational efficiency. In either case, the individual functional units can be configured to interact closely with each other. Despite the possibility of such integration or separation, the following description will still focus on the combination of conceptual functional units to clarify the decoding process of video data applied as an embodiment of the present invention.
[0121] The decoder described above may include an inverse quantization and inverse transform unit 340. The inverse quantization and inverse transform unit 340 may be configured to receive encoded information from the parser 330, the encoded information including a method for numerical transformation, block size, quantization coefficients for recovering quantized information, and distinguishing information for simplifying and representing the quantization coefficients, and may be configured to output block value 341 that can be input to the aggregator 370 as a result of processing the encoded information.
[0122] In one embodiment of the present invention, the output value of the inverse quantization and inverse transform unit 340 may include intra-frame prediction encoded block values. These intra-frame prediction block values refer to values obtained by decoding using prediction information from the currently being decoded local image (e.g., within the current frame) without using previously decoded local images (e.g., prediction information from the previous frame).
[0123] The prediction information in the current local image can be provided by the intra-prediction unit 350. According to an embodiment of the invention, the intra-prediction unit 350 uses image information of spatially adjacent regions extracted from the currently being decoded and partially decoded local image to generate block values in the same block format as the being decoded as prediction information. The local image information can be provided 381 by the current image buffer (i.e., the so-called line buffer 380). According to an embodiment, the merging unit 370 can be configured to merge the prediction information 351 generated by the intra-prediction unit 350 with the block values 341 provided by the inverse quantization and inverse transform unit 340.
[0124] In another embodiment, the output value of the inverse quantization and inverse transform unit 340 may be a block value encoded for inter-frame prediction. In some cases, the output value may include a block value that has undergone motion compensation. In this case, the inter-frame prediction unit 355 may extract and use sample information 386 for motion-based prediction from the reference picture buffer 385. The information 356 obtained by performing motion compensation on the sample information based on the symbols 338 contained in the block value as the output value may be configured to be merged by the merging unit 370 with the block value 341 provided by the inverse quantization and inverse transform unit 340. In this case, the block value 341 may be referred to as a differential value or a residual value.
[0125] The positional information in the memory, used by the inter-frame prediction unit 355 to extract the sample information from the reference image, can be determined, for example, by a combination of X, Y, and other symbols 338 used to indicate a specific position in the reference image, and by a motion vector transmitted to the inter-frame prediction unit 355. The inter-frame prediction unit 355 may also include the following functions: when providing motion vectors that support "subsampling," it can process the sample values using interpolation; and it can predict and enhance the motion vector values.
[0126] The output value 371 of the merging unit 370 can be provided to the loop filter unit 360 and processed by various loop filtering methods. The loop filter unit 360 can also be configured to receive not only the block unit output 371 of the merging unit 370, but also symbols 338 from the parser 330 to control its operation. The output of the loop filter unit 360 can be transmitted to an external display device such as the aforementioned display device via the output connection 390. However, for subsequent prediction and parsing of intra-frame or inter-frame coded block values, 361 can be stored in the line buffer 380, and then stored in the reference image buffer 385.
[0127] After a specific local image (e.g., a frame) is decoded, it can be used as a reference image for predictive decoding in subsequent decoding processes. A local image (e.g., a frame) can be progressively accumulated into line buffer 380 and decoded, and when a frame is decoded, the contents of the line buffer 380 can be transferred 383 to the reference image buffer 385, and a new line buffer 380 is allocated for decoding the new frame.
[0128] The aforementioned video decoder 305 can be configured to perform decoding operations according to a predetermined video compression technology, which can be documented based on various international standards or commercial standards. These standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sector (ITU-T). Those skilled in the art should understand that these recommendations are equivalent to international standards jointly developed by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may conform to a specific bitstream syntax defined by the relevant standards. This requirement is explicitly specified in video compression specification documents and standard documents, and is specifically defined and constrained through the profiles and levels specified therein. To ensure compliance with the aforementioned profiles and levels, the complexity of the encoded video data also needs to be limited to a certain level. For example, certain profiles or levels may configure limitations such as maximum image size, maximum decoding rate, and maximum reference image size. In some embodiments, the aforementioned limitations can be further limited by the hypothetical reference decoder (HRD) and the HRD buffer management metadata signals included in the encoded video data.
[0129] According to one embodiment of the present invention, the receiving unit 310 may receive additional redundant data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information for use by the decoder 305 to correctly decode the data or more accurately reconstruct an image close to the unencoded image. The additional data may be provided in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, and forward error correction codes.
[0130] Figure 4 This is a conceptual diagram of the functional units of a video encoder according to an embodiment of the present invention. The encoder 405 described above can be configured to receive raw video information 402 from a video source 401 and perform encoding.
[0131] The aforementioned raw video information 402 can have any suitable bit depth, such as 8-bit, 10-bit, 12-bit, etc. Furthermore, the aforementioned raw video information 402 can have any suitable color space, such as R / G / B, Y / U / V, Y / Cb / Cr, etc. Additionally, the aforementioned raw video source can have any suitable sampling structure corresponding to the aforementioned color space, such as Y / Cb / Cr4:2:0, Y / Cb / Cr 4:4:4, etc. A raw video source with this predetermined format can be provided to the aforementioned encoder in the form of a digital video stream.
[0132] In a one-way video communication network, the aforementioned raw video information 402 can be obtained from a recording medium storing pre-prepared original video footage. In a two-way video communication network, the aforementioned raw video information 402 can be obtained from an image acquisition device (e.g., a camera) included in the two-way video communication that generates at least one video transmission stream.
[0133] The video data, including the aforementioned raw video information 402, can be configured to be played sequentially over time to simulate multiple local images of motion. These local images can be represented using concepts such as pictures or frames. Depending on the sampling structure, color space, and other types used, each local image may include more than one sample. Those skilled in the art will understand that the term "sample" is closely related to "pixel" in digital images. The working principle of the encoder will be explained below in relation to such samples.
[0134] According to one embodiment of the present invention, encoder 405 may be configured to encode and compress (partial) images constituting the original video information 402 in real time (or according to other time requirements required by the implementation method) into the form of encoded video information.
[0135] In the encoder 405 described above, the control unit 450 can be a functional unit for appropriately controlling the encoding speed. As described below, the control unit 450 can be configured to control other functional units and combine functions with them. Parameters set by the control unit 450 may include parameters related to bitrate control, such as image skipping, quantizer values, and variable values for applying image quality optimization techniques. Furthermore, it may include values such as image size, group of pictures (GOP) structure, and the maximum search range of motion vectors. Those skilled in the art will understand that the control unit 450 may have various other functions, which can be added or removed depending on the video encoder design optimized for individual system designs.
[0136] According to one embodiment of the present invention, the encoder 405 can be configured to operate in a structure known to those skilled in the art, such as a "coding loop". For simplicity, the encoding loop can, exemplarily, consist of an internal encoder (also called a "source encoder") 410 and an internal decoder 420. The internal encoder is responsible for receiving the image to be encoded and generating symbols based on at least one previously encoded reference image. The internal decoder is configured to connect to the internal encoder. The internal decoder 420 can be configured to, by receiving the output of the internal encoder 410, perform a reproduction operation of sample data generated after a remote decoder 490, which is actually located remotely, receives encoded video information from the encoder 405.
[0137] The video data, composed of sample data reconstructed by the internal decoder 420, can be configured to be input into the reference image buffer of the encoder 405. As described above, since the internal decoder 420 is implemented to reproduce the output of the encoder 405 and the result of decoding by a remote decoder, the video data recorded in the reference image buffer can be identical in bit units to the information in the reference image buffer of the remote decoder. That is, the prediction function unit included in the encoder 405 can read the same value from the encoder 405 reference image buffer as the previous frame sample value that the subsequent decoder will reference during the decoding process.
[0138] As described above, the principle of achieving reference image buffer consistency between encoder 405 and decoder 490 through the internal decoder 420 at encoder 405 is well known to those skilled in the art. Furthermore, those skilled in the art are also familiar with methods for dealing with environments where such guarantees cannot be made (e.g., information loss due to communication failures).
[0139] Referenced Figure 3 An embodiment of the operation method of the aforementioned internal decoder 420 has been described in detail. Figure 3 The decoder in the code can be considered as the "remote" decoder 490 described above. The internal decoder 420 described above may not include lossless encoding and decoding parts, such as the parser 330 or entropy decoder 335. This is because the implementation of the internal encoder 405 is merely to reproduce the actions of the remote decoder, thus allowing direct decoding of symbols without the need for symbol compression and decompression. Therefore, as... Figure 3 As shown, the parser and entropy decoder may not have the previous functional units, or at least partially implement the previous functional units.
[0140] As described above, according to a preferred embodiment of the present invention, any decoder functional unit (excluding the parser and entropy decoder) present in the decoder can naturally exist as substantially the same functional unit in the corresponding encoder 405.
[0141] The operation of the encoding functional units that may be included in the encoder 405 described above can be viewed as the inverse process of the decoder functional units described above. Therefore, this embodiment can generally be explained by reversing the operation of the decoder functional units described above. For example, quantization and transform functional units corresponding to inverse quantization and inverse transform units can be provided, as well as inter-frame prediction coding units corresponding to inter-frame prediction units can be provided. In addition, some further explanation is provided below.
[0142] The aforementioned internal encoder 410 can be configured to perform encoding of input image information (e.g., input frame) for at least one reference image information, such as a local image (e.g., a pin) that has been previously encoded in time sequence at least a distance from video data designated as a reference frame, according to a predictive coding method performed by the predictive coding unit 440 operating via the reference image buffer 430. In this case, the aforementioned encoder 405 can be configured to encode the difference between the sample blocks constituting the input image and the sample blocks constituting the reference image.
[0143] The internal decoder 420 can decode video data that can be designated as the reference image from the symbols generated by the internal encoder 410. As described above, since this video data is the same as the decoding action performed by the remote decoder, the video data used as the reference image can be provided to the encoder 405 in a lossy compressed and somewhat corrupted form, which is intended to ensure consistency with the decoder's action.
[0144] The predictive coding unit 440 can be configured to perform a predictive search operation within the encoder 405. This predictive search operation can refer to the inter-frame prediction or intra-frame prediction operation described in the decoder description. For input image information that is planned to be newly encoded, in order to obtain the motion vectors, block shapes, and metadata that may include these, as well as the actual reference sample blocks, of the reference image's positional information that can provide suitable predictive reference information for the new image information, as well as the sample blocks that actually provide this information, the predictive unit can access the reference image buffer 430 and retrieve the information. The predictive coding unit 440 can operate based on a so-called "sample block by pixel block" reference to obtain suitable predictive reference information. According to one embodiment of the invention, as can be determined based on the search results obtained by the predictive coding unit 440, at least one predictive reference can be specified for the input image to point to at least one reference image information stored in the reference image buffer 430.
[0145] In one embodiment of the present invention, the control unit 450 may be configured to include setting parameters for video data encoding to manage the overall encoding operation of the internal encoder 410.
[0146] The outputs of all the aforementioned functional units can be processed by entropy coding 460 to achieve the final output. For the symbols generated by each of the aforementioned functional units, the entropy coding 460 employs various entropy coding techniques described above (which may include, for example, variable-length coding, Huffman coding, and arithmetic coding). Depending on different standards, these coding methods can be context-adaptive or context-sensitive, and can also be implemented based on principles well known to those skilled in the art. This entropy coding 460 typically achieves lossless compression; therefore, at least one symbol generated by the aforementioned functional units can be converted into encoded video data.
[0147] In controlling the operation of the encoder 405, the control unit 450 can apply a specific encoding type to each local image (e.g., a frame or image) within the encoding range. The encoding method of the local image is also affected by the type. According to an embodiment, the type may include a "frame type" as distinguished below.
[0148] Figure 5 This is a conceptual diagram of a frame type according to an embodiment of the present invention. The following references... Figure 5 Please provide an explanation.
[0149] Intra ("I") images 510 can refer to images that are encoded or decoded using only their own information without requiring predictive coding of other scene information in the reference video data. According to video coding standards, the aforementioned "I" images can refer to key frames, independent / instantaneous decoder refresh (IDR) frames, clean random-access (CRA) frames, etc. As mentioned above, the various names for "I" images can have various variations and applications within the limits allowed by each standard, and there can be some differences between them. In addition to the methods listed above, various application methods for implementing "I" images can be methods known to those skilled in the art or newly provided methods.
[0150] The prediction ("P") image 520 can refer to an image used to predict sample values constituting a block of the aforementioned image. It can be at least one prediction information specifying at least one reference image, and / or an image encoded or decoded based on motion vectors and via intra-frame or inter-frame prediction. According to video coding standards, the aforementioned "P" image can be configured to reference only one reference frame, or refer to more than one reference frame. When referencing more than one reference frame, sample information and / or associated metadata derived from multiple reference images can be used to reconstruct a single block. However, generally, an image designated as a "P" image can be understood as limited to referencing images that are temporally earlier.
[0151] A bidirectional prediction (B) image 530 may refer to at least one prediction information from at least two reference images specified for predicting sample values of blocks constituting the image, and / or an image encoded or decoded based on motion vectors and through intra-frame or inter-frame prediction. Typically, an image designated as a "B" image is distinct from an image designated as a "P" image and can be understood as an image that performs the reference, but is not limited to temporally previous images.
[0152] During encoding and decoding, video data can be spatially divided according to multiple sample blocks and encoded according to the aforementioned block units. It is well known that these block units may include, for example, sizes in horizontal / vertical pixels such as 4x4, 8x8, 4x8, or 16x16, but are not limited to these. These blocks can be encoded using predictive coding methods, based on their respective local images and, depending on the specified type (allowed and / or restricted), by referencing any other (already encoded) blocks. For example, blocks in "I" image 510 may not use predictive coding methods, or may be encoded by referencing already encoded blocks in the same local image. That is, only so-called intra-frame prediction methods may be used. In contrast, "P" image 520 may further refer to at least one reference image encoded in a previous time unit, and therefore may also be encoded using inter-frame prediction along with intra-frame prediction. "B" image 530 may also refer to reference images previously encoded in the encoding sequence but executed later in time unit. However, it is well known that even within "P" or "B" images, there may be blocks encoded without relying on predictive coding. However, it is well known that there may be blocks in a “P” or “B” image that are not encoded using predictive coding.
[0153] The video encoder 405 described above can be configured to perform encoding operations according to predetermined video compression techniques documented according to various international or commercial standards. Examples of such standards may include all the standards described in the decoder described above.
[0154] According to one embodiment of the present invention, the transmitting unit 470 may buffer the encoded video data generated by the entropy encoding described above, so as to provide / send the video data (ultimately transmitted to the remote decoder 490) to the device storing the encoded video data via a hardware or software connection 495. According to the embodiment, when the transmitting unit 470 receives the encoded video data provided / sent from the video encoder 405, it may receive and merge other data (e.g., encoded audio data or other auxiliary data) accompanying the encoded video data from a separate source 480.
[0155] According to one embodiment of the present invention, the transmitting unit 470 may also be configured to transmit additional data along with the encoded video. This additional data can be considered as part of the encoded video data. The additional data may include information for the decoder, used to correctly decode the data or to more accurately reconstruct information closer to the pre-encoded image. Examples of the additional data may include all the examples described above with respect to the decoder receiving unit 310.
[0156] As described above, the present invention can be implemented using digital video compression standards widely used and understood by those skilled in the art. These digital video compression standards may include at least one known compression standard, such as MPEG-2, MPEG-4 video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, MotionJPEG, etc.
[0157] Figure 6 This is a conceptual diagram showing the structure of a video encoder using the H.266 / VVC standard. Figure 6 The structure shown is equivalent to well-known standard encodings such as ITU-T H.266 and ISO / IEC 23090-3, as well as the general structure of video encoders known as MPEG-I Part 3 or Versatile Video Coding (VVC).
[0158] according to Figure 6 The video encoder 605 can be configured to receive an encoded bitstream 602 as input and output uncompressed or unencoded raw video data 601. During intra-frame encoding, the video data 601 can be directly provided to the luma mapping unit 610a, or provided to the luma mapping unit 610b via the inter-frame prediction unit 620, which includes motion vector extraction. In the case of intra-frame encoding, the mapped luma signal can be provided not only separately to the output combiner 606, but also by selecting either the intra-frame prediction encoded signal output by the intra-frame prediction unit 625 (608) or the inter-frame prediction encoded signal output from the luma mapping unit 610b via the inter-frame prediction unit 620. The result of the output combiner can be applied to the chroma scaling unit 615 (the operations of the luma mapping unit 610 and the chroma scaling unit 615 are collectively referred to as the luma mapping / chroma scaling (LMCS) process). The scaled color difference signal can be provided to the transform unit 630, which in particular can perform an adaptive color transform on the color difference signal. The coefficients obtained from the transform are applied to the quantization unit 640 for quantization. Thus, lossy compression can be achieved, and the result of the lossy compression can be output as a bitstream 602 using a lossless compression method based on a multi-hypothesis CABAC unit 650.
[0159] On the other hand, in order to generate the coding loop, the result of the aforementioned lossy compression can undergo inverse quantization 645, inverse transform 635, and luma signal amplification 617 before entering the actual decoding process. The result of the aforementioned luma signal amplification can be provided to the internal combiner 607 together with the result of the selection 608 of at least one previously generated intra-frame predictive coding signal or inter-frame predictive coding signal. After undergoing inverse luma mapping 617, the result of the aforementioned internal combiner can be processed by a deblocking filter 660, a sample adaptive offset (SAO) 670, an adaptive loop filter (ALF), etc., used to reproduce the image quality improvement process in the decoder. As described above, the result of the reproduction action in the decoder can be applied to the reference image buffer 690 and can be reused for predictive coding by the aforementioned inter-frame prediction unit 620.
[0160] This invention can also be implemented or combined with the enhanced compression model (ECM), a next-generation video codec developed by the Joint Video Experts Team (JVET) to improve H.266 / VVC. According to the aforementioned JVET standardization progress document, document number ISO / IEC JTC 1 / SC 29 / WG 5 N 190 (also document number JVET AC2025-v1), the enhanced compression model may include improved intra-frame predictive coding methods, improved inter-frame predictive coding methods, improved transform and transform coefficient coding methods, improved adaptive loop filtering methods, bilateral filtering methods, a novel sample adaptive offset (SAO) method for image quality improvement, extended entropy coding methods, and improved gradual decoding refresh (GDR) technology. These technologies are included as background technology in this invention.
[0161] Overview of GDR Application Methods
[0162] This invention provides an application method and apparatus for an improved gradual decoder refresh (GDR) technique applicable to video encoding and decoding, including the embodiments described above.
[0163] Figure 7This is a conceptual diagram illustrating the application of GDR technology. According to conventional standards, video information consists of multiple still images, or frames. Such a collection of frames can be called a video sequence. As mentioned above, the frame types included in the sequence can include intra-frame frames and inter-frame frames, and may also include frame types that perform functions shown or not shown in this invention.
[0164] refer to Figure 7 In the prior art, refresh information about the entire image frame is provided by an intraframe 10 appearing in the middle of the sequence, while the remaining interframe frames 11 are only used to reference the refresh information. However, according to the GDR method, refresh information is provided by GDR frame intervals 20 defined within the sequence and spanning multiple frames. The GDR frame intervals 20 can be composed of a GDR frame 21, a recovery point frame 23, and intermediate frames 22 of arbitrary length that can be added between them, depending on the length of the GDR frame intervals 20.
[0165] The aforementioned GDR frame interval 20 has a length of at least 2 frames. In this case, it can consist of one GDR frame 21 and one recovery point frame 22. When it has a length of 3 frames or more, it can consist of one GDR frame 21, one recovery point frame 22, and at least one intermediate frame 22. When the length of the GDR frame interval is 1 frame, it is no different from a regular intra-frame.
[0166] Regarding the refresh direction of the aforementioned GDR frame interval 20, structures for performing horizontal progressive refresh 21 or vertical progressive refresh 22 have been disclosed. Furthermore, according to JVET-AE0145, a method for specifying the refresh direction as bidirectional (horizontal and / or vertical) in a GDR frame has been disclosed in a technical proposal submitted to JVET.
[0167] Figure 8 This is a conceptual diagram illustrating the expected effects of applying GDR technology in terms of bit rate. Based on the aforementioned applications of GDR technology, such as... Figure 8 As shown in the schematic diagram 30, unlike the sharp bit rate gap 32 that occurs between existing intra-frame and inter-frame frames, this bit rate burden is distributed among progressive GDR frames, thus the bit rate gap can be expected to be stabilized 34.
[0168] Figure 9This is a conceptual diagram illustrating the impact of applying GDR technology on the encoding / decoding process. When encoding frames within a GDR region, the current frame 900 can be divided into at least two regions. The first region 910 can be a region collectively referred to as a refreshed region or a clean region. Preferably, the first region can be interpreted as a refreshed region 911 and a clean region 911, respectively. Additionally, the second region 920 can be a region collectively referred to as an unrefreshed region or a dirty region.
[0169] The refreshed area 911 mentioned above can refer to the area encoded by decoding only the current GDR interval frame 900. Therefore, for video coding units contained in the refreshed area 911, such as coding units that can be equivalent to at least one of macroblock, block, subblock, coding tree unit (CTU), coding unit, prediction unit, transform unit, partitioned area, and other coding units, they can be encoded using the constrained intra coding method.
[0170] The aforementioned clean region 912 can refer to a region that can be encoded based on information that can be deduced from the refreshed region 911 appearing in the preceding frame 950 (which may be a GDR frame or an intermediate frame). For the aforementioned clean region 912, it is expected that at least complete image information about the clean region 912 can be deduced when decoding begins from the start of the GDR interval. Therefore, for the video coding unit contained in the aforementioned clean region 912, intra-frame coding can be performed, but it can also be configured to perform inter-frame coding with reference to the clean region 912 of the preceding frame. Therefore, the conclusion is that encoding can be performed using a constrained intra / inter coding method.
[0171] The aforementioned dirty region 920 corresponds to a region that can be encoded without referencing information provided by the GDR frame or intermediate frames. From the decoder's perspective, successfully decoding the information in the aforementioned dirty region 920 may require information that can only be obtained by decoding the encoded image bitstream prior to this GDR interval. For video coding units contained in the aforementioned dirty region 920, their encoding can be performed using unrestricted intra-frame or inter-frame methods.
[0172] The aforementioned restricted intra-frame coding can refer to a method that allows complete image information about the coding unit to be derived using only the information contained within the refreshed region 911. Therefore, the aforementioned restricted intra-frame coding can include, for example, completely prohibiting inter-prediction during the coding process, or prohibiting intra-prediction or limiting its spatial range.
[0173] The aforementioned restricted intra / inter-frame coding can refer to a method that allows complete image information about the coding unit to be derived from information in the first interval 910, which includes the refreshed region 911 and the clean region 912, as well as information provided in the preceding GDR interval frame 950. Therefore, based on the features of the aforementioned restricted intra-frame coding, the aforementioned restricted intra / inter-frame coding may further include, for example, the selection of a reference picture for inter-prediction and the temporal or spatial range limitation of the motion vector.
[0174] When performing the aforementioned restricted intra-frame coding and restricted intra / inter-frame coding methods, according to embodiments of the present invention, even if a prohibited or restricted coding method is performed, the information used can be replaced by temporarily generated information. For example, it can be treated as an unavailable value and not reflected in the calculation, or it can be used as a reproducible value from an available value through methods such as padding, filtering, or prediction.
[0175] In one embodiment of the present invention, the coding unit belonging to the refreshed region 911 can also perform intra-frame prediction coding from the information appearing in the coding unit belonging to the refreshed region 911 (e.g., such as...). Figure 9 (As shown in 931). This is because the predictive encoding performed within the refreshed area 911 may not affect the decoder's acquisition of complete image information about the refreshed area 911.
[0176] In one embodiment of the present invention, the coding unit belonging to the refreshed region 911 can perform intra-frame prediction coding (e.g., as shown in the image) from the information appearing in the coding units contained in the clean region 912. Figure 9(As shown in 932), but in other embodiments, such predictive coding may be prohibited. When the decoder is configured to decode the occurrence time of the GDR frame as a random access point, allowing predictive coding of the clean region 912 is advantageous for improving coding efficiency. Conversely, prohibiting predictive coding of the clean region 912 is advantageous when it is desirable to provide complete random access to intermediate frames in the GDR interval, even if it is necessary to tolerate only partial image confirmation for a certain period.
[0177] In a preferred embodiment of the present invention, coding units belonging to the refreshed region 911 may be prohibited from intra-frame predictive coding of information appearing in coding units contained in the dirty region 920 (e.g., ...). Figure 9 (The situation shown in 933 is prohibited). Since, from the decoder's perspective, the information contained in the aforementioned dirty region 920 may not have been acquired, or may be a region encoded based on stale information, it is preferable to prohibit reference to that region.
[0178] Referring to the second edition of the H.266 / VVC standard specification, in order to distinguish between the first interval 910 and the second interval 920, and to smoothly support the aforementioned restricted intra-frame and intra / inter-frame coding methods, a bitstream syntax specifying virtual boundaries can be used. Taking the second edition of the H.266 / VVC standard specification as an example, the above syntax can include syntax signals such as "ph_virtual_boundaries_present_flag", "ph_num_ver_virtual_boundaries", "ph_virtual_boundary_pos_x_minus1", "ph_num_hor_virtual_boundaries", and "ph_num_hor_virtual_boundaries" contained in the picture header structure syntax in Section 7.3.2.8. Schematic, the aforementioned “ph_virtual_boundaries_present_flag” value may be a flag signal indicating that the value of the virtual boundary appears in the aforementioned bitstream; the aforementioned “ph_num_ver_virtual_boundaries” and “ph_virtual_boundary_pos_x_minus1” signals, as well as the aforementioned “ph_num_hor_virtual_boundaries” and “ph_num_hor_virtual_boundaries” signals, may be signal values specified to represent the number and position of the vertical virtual boundaries and the number and position of the horizontal virtual boundaries, respectively.
[0179] However, in embodiments of the present invention, the value of the aforementioned “ph_virtual_boundaries_present_flag” signal can be kept at 0, thereby not providing information about the aforementioned virtual boundaries. Instead, in several embodiments of the present invention described below, an improved method of implementing GDR technology can be provided that does not require the aforementioned signal, or operates in a signaling manner that replaces the aforementioned signal.
[0180] Figure 10 This is a schematic diagram illustrating the provision of frame information in a GDR interval according to an embodiment of the present invention. The presence of GDR frames, the information of GDR frames, and at least a portion of the length of the GDR interval can be included in the encoded video bitstream.
[0181] In one embodiment of the invention, the encoded bitstream may include a flag indicating the presence of GDR interval 1010. Taking the second version of the H.266 / VVC standard specification as an example, this flag could be something like "sps_gdr_enabled_flag" included in the sequence parameter RBSP syntax in section 7.3.2.4, preferably encoded as a 1-bit signal. However, it is obvious that the signal described above can also be provided at any other location in the bitstream with an information capacity of at least 1 bit (or less than 1 bit in entropy encoding).
[0182] In one embodiment of the invention, the encoded bitstream may contain information representing the sequence number of the frame. Referring to the second edition of the H.266 / VVC standard specification, this information may be a variable-length signal such as "ph_pic_order_cnt_lsb" contained in the picture header structure syntax in section 7.3.2.8. The variable length of the aforementioned "ph_pic_order_cnt_lsb" signal may be derived from a signal preferably encoded as 4 bits, such as "sps_log2_max_pic_order_cnt_lsb_minus4" contained in the sequence parameter RBSP syntax in the same specification section 7.3.2.4 prior to it. For example, the value of "sps_log2_max_pic_order_cnt_lsb_minus4" plus the value of 4 could be the length of that bit. However, it is obvious that at any other location in the bitstream, it is also possible to provide the signal described above with an information capacity of at least 1 bit (or less than 1 bit in entropy encoding). Refer to Figure 10 Example 1030 of “ph_pic_order_cnt_lsb” is shown.
[0183] The value of "ph_pic_order_cnt_lsb" can refer to the sequence number representing the display order of the current frame. Its range can be 0 or higher, but less than the "MaxPicOrderCntLsb" value in the H.266 / VVC standard specification. Specifically, the value of "MaxPicOrderCntLsb" is 2 raised to the power of ("sps_log2_max_pic_order_cnt_lsb_minus4" plus the value of 4). Taking the second version of the H.266 / VVC specification as an example, the value of "sps_log2_max_pic_order_cnt_lsb_minus4" is specified to be between 0 and 12. Therefore, the value of "MaxPicOrderCntLsb" can represent a value between 16 and 65536. When the actual sequence number of the above frame is greater than or equal to the value of "MaxPicOrderCntLsb", the value of "ph_pic_order_cnt_lsb" can be moduloed and encoded as the value of "MaxPicOrderCntLsb".
[0184] In one embodiment of the invention, the encoded bitstream may include information representing the frame distance in time to the end point of the GDR frame, i.e., to the aforementioned recovery point frame. Referring to the second edition of the H.266 / VVC standard specification, this information may be a signal such as "ph_recovery_poc_cnt" contained in the picture header structure syntax in Section 7.3.2.8. For example, the value of the "ph_recovery_poc_cnt" signal may be used to determine in which future frame the aforementioned recovery point frame will appear, where the recovery point frame may represent a value specified as 0. See [reference]. Figure 10 Example 1035 of the above-mentioned “ph_recovery_poc_cnt” signal is shown.
[0185] In one embodiment of the invention, the overall length of the GDR interval 1020 can be derived from information in the encoded bitstream. For example, the encoder specifies that a certain GDR interval 1020 must begin when the value of the “ph_pic_order_cnt_lsb” signal is 0, thereby enabling the encoder and / or decoder to calculate the frame distance from the start point of the GDR interval 1020 and the frame distance to the end point of the GDR interval using the current value of the “ph_pic_order_cnt_lsb” signal and the value of the “ph_recovery_poc_cnt” signal. As another example, the decoder is configured to identify the start point of the GDR interval 1020 by storing the value of the “ph_pic_order_cnt_lsb” signal of the current frame in the start frame 1011 of the GDR interval 1020, and to calculate the frame distance to the end point of the GDR interval 1020 using the value of the “ph_recovery_poc_cnt” signal. For example, the encoder can be configured to provide the decoder with the aforementioned length-related information by adding an arbitrary signal specifying the overall length of the aforementioned GDR interval 1020. This information can be applied differently to each GDR interval 1020, or it can be applied equally to each GDR interval 1020 but re-included in each GDR interval 1020, or it can be applied singly to the entire video sequence. Preferably, the aforementioned arbitrary signal can be newly included in the picture header structure syntax of Section 7.3.2.8 in the H.266 / VVC standard specification syntax; however, it is obvious that the signal described above can also be provided at any other location in the bitstream with an information capacity of at least 1 bit (or less than 1 bit in entropy coding).
[0186] First embodiment and its encoding and decoding
[0187] As described above, according to existing GDR technologies, such as the implementation method proposed in the second edition of the H.266 / VVC standard specification, in order to distinguish between "clean" and "dirty" regions refreshed by GDR technology, the existence and location-related syntax signals of virtual boundaries can be used. Therefore, in the first embodiment of the present invention, a method for deriving virtual boundaries by specifying functions is shown, instead of specifying virtual boundaries as described above, thereby providing an improved GDR method and apparatus.
[0188] According to a first embodiment of the present invention, in GDR, the distinction between "clean" and "dirty" regions can be determined by a predefined function with frame number as a variable. This function can be any function that forms boundary lines for spatially dividing regions within a video frame. For example, the function can be defined as a function that draws any straight line, curve, or closed curve. According to one embodiment of the present invention, the function is a normalization function that can be scaled and applied according to the resolution of the frame to which the application is applied. For example, some functions representing the boundary lines of GDR are configured to derive real numbers between 0 and 1 as result values based on variations in frame number n. These result values are multiplied by the horizontal resolution of the frame, and the result can specify the position of the boundary lines on the frame based on these real numbers.
[0189] Figure 11 This is a conceptual diagram illustrating the differentiation of regions using functions according to an embodiment of the present invention. Figure 11 In the example, the method of specifying virtual boundaries in the existing vertical GDR is reproduced through a function. Figure 11 The functions used in the examples can be represented as follows.
[0190] [Formula 1]
[0191]
[0192] Formula 1 above can be applied to all pixels of a frame. In Formula 1 above, R p This can refer to a discrete (Boolcan) value used to indicate whether a pixel is contained within a specific region. The aforementioned R... p This can be determined by the function f(x,y). X can be the value of the pixel's horizontal coordinate in the video, normalized by the resolution. For example, when the horizontal resolution of the video is 1280, 0 can be normalized to 0, 640 to 0.5, and 1280 to 1. Y can be the value of the pixel's vertical coordinate in the video, normalized by the resolution. For example, when the vertical resolution of the video is 720, 0 can be normalized to 0, 360 to 0.5, and 720 to 1. gdr The sequence number of the above frame can be represented within the GDR range. gdr It can represent the overall length of the GDR interval.
[0193] That is, Formula 1 above can be understood as, as the sequence number n of the current frame within the GDR interval... gdr The increase of is a function that advances the points of a specific pixel contained in the "clean" region from 0 to 1 (1151, 1152, 1153, 1154) in the vertical direction. (See reference...) Figure 11 It can be seen that, with Figure 9The "clean" regions 1120, 1121, 1122, and 1123 corresponding to the first interval 910 gradually expand in the vertical direction according to the frame order (1110, 1111, 1112, and 1113).
[0194] An example is given below, referring to Formula 1 above. Figure 11 For example, the GDR interval is 4 frames long, therefore l gdr Set to 4. In the first frame 1110, n gdr It can be determined to be 1. When all the pixels of the first frame 1110 above are substituted into Formula 1 above, the vertical position y is n times the vertical resolution. gdr / l gdr Pixels below a multiple of 1 / 4 (1151) are included in the "clean" region 1120, and other pixels are classified as the "dirty" region 1140. If the vertical resolution of the first frame 1110 is 720 pixels, and the video coordinate axis starts from the top left corner of the video, then the above formula 1 can produce the same result as setting a virtual boundary at a vertical position of 180 pixels in the first frame.
[0195] Similarly, it can be seen that in the second frame 1111, pixels below 2 / 4 (1152), i.e., pixels below 360 in the vertical position 1121; pixels below 3 / 4 (1153) in the third frame 1112, i.e., pixels below 540 in the vertical position 1122; and pixels below 4 / 4 (1154) in the fourth frame 1113, i.e., pixels below 720 in the vertical position 1123, i.e., the entire frame, are included in the "clean" area.
[0196] The advantage of the above embodiments is that the same operation can be performed even if the length of the GDR interval is arbitrarily shortened or extended. For example, in Figure 1 In this context, the function described in Formula 1 specifies the virtual boundary at the y-position {180, 360, 540, 720}. However, when the length l of the GDR interval is increased... gdr When increased to 5 frames, the operation is that the virtual boundary is specified at y position {144, 288, 432, 576, 720} in the first to fifth frames (not shown), when l gdr When increased to 6 frames, the operation is that the virtual boundary is specified at y positions {120, 240, 360, 480, 600, 720} in frames 1 through 6 (not shown). Furthermore, l... gdrWhen set to 7 frames, the virtual boundary of the first frame is specified to 102.857142… pixels. However, the function in Formula 1 above indicates that the vertical position y of the pixel is below the specified boundary line. Therefore, the same effect as specifying it to 102 pixels can be achieved. Similarly, the same effect as specifying the virtual boundary of the second frame to 205 pixels below 205.714285… can be achieved.
[0197] The advantage of the above embodiments is that the same operation can be performed even if the video resolution is arbitrarily changed. This can be understood as the input of the above function being normalized to the interval between 0 and 1. Of course, for those skilled in the art, it is obvious that in order to apply the above function to a changed resolution, in addition to normalizing the input value and the interval between 0 and 1, various mathematical and algorithmic methods can be used, and it is also obvious that examples of such modified implementations fall under the spirit of this invention.
[0198] Figure 12 This is another conceptual diagram illustrating the differentiation of regions using functions according to an embodiment of the present invention. Figure 12 The example reproduces the existing method of specifying virtual boundaries in the vertical GDR through a function, and shows the re-distinguishing of "clean" areas, classifying them as... Figure 9 The case of refresh region 911 and clean region 911. As mentioned above, this embodiment is more conducive to enhancing random access performance by allowing only a limited intra-frame coding method for the aforementioned refresh region, and requires setting two virtual boundary lines.
[0199] exist Figure 12 The example also shows the use of the same function as Equation 1 above. However, to generate two virtual boundary lines, Equation 1 can be applied twice per frame. That is, n can be executed for each frame. gdr The first application of time and n gdr The second application at -1, and the exclusive interval contained in the first application but not in the second application is regarded as the refresh area, the interval contained in both the first and second applications is regarded as the clean area, and the interval not contained in both the first and second applications is regarded as the dirty area.
[0200] Using the above method, when all pixels of the first frame 1210 are substituted into Formula 1, the vertical position y is n times the vertical resolution. gdr / l gdr That is, pixels less than a multiple of 1 / 4 are included in the first application, and the vertical position y is n of the vertical resolution. gdr -1 / l gdrPixels smaller than 0 / 4 are included in the second application. Therefore, if the vertical resolution of the first frame 1210 is 720 pixels, and the video coordinate axis starts from the top left corner of the video, the same result can be obtained as setting virtual boundaries at vertical positions 0 pixel 1250 and 180 pixel 1251 in the first frame, and designating the area between them as the refresh region 1230. Therefore, the aforementioned refresh region 1230 can be defined as the interval from vertical position 0 pixel to 180 pixel.
[0201] Similarly, it can be seen that in the second frame 1211, the pixels in the range from multiples of 1 / 4 (1251) to multiples of 2 / 4 (1252), i.e., pixels 1231 with a vertical position greater than 180 and less than 360; in the third frame, the pixels in the range from multiples of 2 / 4 (1252) to multiples of 3 / 4 (1253), i.e., pixels 1232 with a vertical position greater than 360 and less than 540; and in the fourth frame, the pixels in the range from multiples of 3 / 4 (1153) to multiples of 4 / 4 (1154), i.e., pixels 1233 with a vertical position greater than 540 and less than 720; are designated as refresh areas.
[0202] Figure 13 This is a schematic diagram of a region differentiation method using functions according to a partial embodiment of the present invention. Figure 13 Categories (a) to (g) disclose various embodiments of a method for distinguishing regions using functions, but are not limited to those shown. Those skilled in the art can... Figure 13 It is understood that various implementation methods can be applied based on the shapes of functions and graphs not shown in this specification, and it is also understood that these applied implementation methods are all within the scope of the present invention.
[0203] Figure 13 The function used in example (a) can be represented as follows.
[0204] [Formula 2]
[0205]
[0206] Based on Formula 2 above, it can be understood as follows: Figure 13 As shown in (a), GDR refresh along the diagonal direction can be easily implemented.
[0207] Reference Figure 13(b) It can be understood that, through a variation of a similar function, a GDR refresh area and / or clean area that diffuses outward from the center can be easily implemented. The implementation of the embodiment described above yields similar results to the method in the JVET-AE0145 document, a technical proposal submitted to the JVET, which specifies the refresh direction of GDR frames bidirectionally in the horizontal and / or vertical directions. However, unlike the aforementioned proposal which requires specifying multiple virtual boundary lines for each frame, it has the advantage of being implemented with less information by specifying a single neighborhood function.
[0208] Reference Figure 13 The (c) parameter allows you to specify any curve, such as a rectangular shape, and specify a region function by increasing its size. This function transformation makes it easy to diversify the shape of GDR refresh regions and / or clean regions.
[0209] The aforementioned closed-curve region function can change its reference point. For example, instead of the upper left corner, which is usually the starting point of the video coordinate system, the left center, lower left corner, upper center corner, center, lower center corner, upper right corner, right center corner, lower right corner, and any other point that can be specified as a normalized or unnormalized position value according to any x and y coordinate system can be used as the reference point for function expansion. Furthermore, when the aforementioned function is applied to the aforementioned frame, it can be applied with any aspect ratio. For example, the image information of the frame has a 16:9 aspect ratio, but the aforementioned function can be applied with a 1:1 aspect ratio. According to embodiments, the encoder and / or decoder of the present invention can project the aforementioned function onto the GDR frame and use it through various methods, such as stretching the aforementioned function to the same resolution as the frame and projecting it, or stretching (so-called filing) and projecting it while maintaining the aspect ratio of the original function, or repeatedly applying area calculations based on the function for tiling within the resolution of the video, etc.
[0210] Reference Figure 13 (d) shows that the drawing will be as follows Figure 13 The example shown in (c) illustrates a rectangular region function expansion where the reference point is changed to the exact center of the frame, and the applied aspect ratio is changed to 1:1. This function deformation easily diversifies the shape of GDR refresh and / or clean regions.
[0211] The above function can be applied in reverse. That is, it can be applied by reversing the inside and outside of the boundary line, or by reversing the inside and outside of the closed curve. In this case, the input value of the function can also be from n. gdr Change to l gdr -n gdr Therefore, the order in which they are applied will also be reversed.
[0212] Reference Figure 13 (e) shows the use with Figure 13 (d) is an example of the same expansion method for the region function, but with the application of the inner / outer side and the order of application reversed. Through this function transformation, not only can a GDR be expanded from the starting point to the other side, or from the inner side to the outer side, but also various forms of GDR refresh regions and / or clean regions can be easily achieved by expanding from the other side to the starting point, or from the outer side to the inner side.
[0213] In addition to the embodiments described above, any mathematical expression and / or conditional formula capable of specifying boundary lines or regions can also be used as the functions described above within the scope of the present invention. For example, refer to Figure 13 (f) shows an example using elliptic functions. For example, see [reference 1]. Figure 13 Example (g) illustrates obtaining an amorphous refresh region and / or clean region through the intersection of any two straight lines. Furthermore, the above function can also be applied, depending on the embodiment, after rotation, skew, scaling, and other affine transforms.
[0214] By applying the method according to the first embodiment of the present invention described above, it can be found that existing GDR technology is no longer needed. For example, in the implementation method of the technology proposed in the second version of the H.266 / VVC standard specification, there is no longer a need for virtual boundary and its position-related syntax signal to distinguish between "clean" and "dirty" areas refreshed by GDR technology.
[0215] In a certain codec standard specification, if it is determined that the encoder and decoder use a single function, then it is acceptable not to pass the information of the above function through the bitstream.
[0216] In a certain codec standard specification, if the encoder and decoder use a single function and allow rotation, tilt, scaling and other affine transformations of the function, then the encoder based on the above specification can be designed to insert the parameter values of the above affine transformation, such as angle, reference point, magnification and other information, into the video bitstream, and the decoder based on the above specification can read the above parameter values to decode GDR frames.
[0217] In a certain codec standard specification, if it is determined that the encoder and decoder use multiple functions that belong to or do not belong to the above embodiments, it may be necessary to provide the video bitstream with information indicating that one of the multiple functions has been selected. In this case, the encoder based on the above specification can be designed to insert information about the selection of the above functions into the video bitstream, and the decoder based on the above specification can read the type of the above functions to decode GDR frames.
[0218] In a certain codec standard specification, if the mathematical expression of the transfer function between the encoder and decoder is determined, it may be necessary to provide the video bitstream with a binarized value of the mathematical expression of the function used for GDR. In this case, the encoder, based on the aforementioned specification, can insert the binarized mathematical expression of the function into the video bitstream, and the decoder, also based on the aforementioned specification, can read the expression of the function to decode the GDR frame.
[0219] Furthermore, in a certain codec standard specification, when the encoder encodes a video bitstream containing GDR frames using an appropriate method, if the decoder does not need information about the order in which the GDR frames are refreshed, then regardless of the function used, it is not necessary to transmit information about the function from the encoder to the decoder via the bitstream. This is because, for example, as long as the encoder encodes the specified coding units using the prescribed GDR method with a limited set of intra-frame and / or intra-frame / inter-frame coding methods, it is sufficient for the decoder to perform the decoding process corresponding to the GDR without additional signal transmission.
[0220] In all the aforementioned scenarios, when the encoder inserts certain information into the bitstream, various encoding methods can be applied depending on the specifications and quantity of the information. These encoding methods can include fixed-length encoding and variable-length encoding. It is well known that variable-length encoding can employ Huffman coding, exponential Golomb coding, context-adaptive variable-length coding, context-adaptive arithmetic coding, and other encoding methods, especially entropy coding. It can be confirmed that by using the aforementioned encoding methods, the inserted information can be represented as at least one bit of relevant information at any position in the bitstream.
[0221] It is obvious that when choosing an encoding method in the above examples or types not shown, the decoder should be designed to decode the information by the reverse (inverse) of the same method.
[0222] The second embodiment and its encoding and decoding
[0223] According to a second embodiment of the present invention, in GDR, the distinction between "clean" and "dirty" regions can be determined by a predetermined pattern that changes as the frame progresses. This pattern can be any shape that forms boundary lines for spatially dividing the internal regions of a video frame.
[0224] Figure 14 This is a conceptual diagram illustrating the use of patterns to distinguish regions according to an embodiment of the present invention. Figure 14 The document provides patterns 1415, 1425, 1435, and 1445 that gradually change over a length of 4 frames. Exemplarily, these patterns can store distinctions between refreshed areas, clean areas, and dirty areas. Figure 14 In the example, the length of the GDR interval is fixed at 4 frames. In each GDR frame 1410, intermediate frame 1420, 1430 and recovery point frame 1440 belonging to the GDR interval, the patterns 1415, 1425, 1435 and 1445 corresponding to their respective sequences are projected in a manner corresponding to their resolution, according to the pattern changes consistent with the length of each GDR interval.
[0225] Based on the aforementioned projection pattern, video coding units, such as macroblocks, blocks, subblocks, coding tree units (CTUs), coding units, prediction units, transform units, partitioned areas, and various other coding units, can determine which area each coding unit belongs to. For example, in the projection pattern, all coding units overlapping with the area designated as the refresh area belong to the refresh area and can be encoded using a limited intra-coding method. However, besides this, there may be cases where the area overlaps by more than 50%, and various transformation implementations may be applied, such as designating it as that area.
[0226] Regarding refresh areas, clean areas, dirty areas, etc., the technical methods and limitations imposed on encoding can be referenced from the content of the first embodiment above.
[0227] The aforementioned pattern can be stored as bitmap pixels or as vector graphics. According to an embodiment, the pattern can be configured to have a fixed resolution and used by up-scaling or down-scaling it to the resolution of the video frame being applied. According to an embodiment, for the aforementioned pattern, the pattern that needs to change for each frame can be stored independently according to temporal order. According to an embodiment, in the aforementioned pattern, besides the pattern applied to the initial frame, the pattern that needs to change for each frame can be stored as a residual relative to a temporally earlier pattern, or stored as a vector movement.
[0228] Figure 15 This is a schematic diagram illustrating various forms of a pattern according to an embodiment of the present invention. The pattern according to a second embodiment of the present invention is not limited to, for example... Figure 14 The simple geometric form shown can have any pattern shape and any frame length. As one embodiment, Figure 15 (a) illustrates a pattern transformation method that, as a 6-frame pattern transformation, provides a dithering effect. For example... Figure 15 Providing signals in the complex structured GDR refresh and / or clean regions shown in (a) may not be feasible in existing video codec specifications, or may require a very large amount of information. However, according to the present invention, encoding and / or decoding are achieved through a predetermined GDR refresh pattern, so the method performed with respect to GDR can reduce the amount of information that must be included in the bit rate, while improving the execution process of GDR.
[0229] According to embodiments, the above patterns can also be applied after rotation, skew, scaling, and other affine transforms. For example, Figure 15 (b) shows that it will be with Figure 15 (a) The same pattern is used after being rotated 90 degrees to the right. It will be apparent to those skilled in the art that any other angle and deformation can be achieved.
[0230] Various methods can be used when projecting the above pattern onto a video and applying it. Figure 16This is a schematic diagram illustrating a video projection method of a pattern according to various embodiments of the present invention. Even given the same pattern 1600, according to embodiments, the encoder and / or decoder of the present invention can project the above-described pattern onto a GDR frame and use it in various ways, for example, stretching the above-described pattern to the same resolution as the frame and projecting it (1610); or stretching (so-called filling) while maintaining the aspect ratio of the original pattern and projecting it (1620); or repeatedly applying the pattern within the resolution of the video and tiling it (1630), etc. Furthermore, it is obvious that various variations are possible, such as being able to use a given pattern to derive every region required for the GDR video.
[0231] By employing the application method according to the second embodiment of the present invention described above, it can be found that existing GDR technologies, such as the virtual boundary and its position-related syntax signals that exist in the implementation method of the technology proposed in the second version of the H.266 / VVC standard specification to distinguish between "clean" and "dirty" regions refreshed by GDR technology, are no longer needed.
[0232] In a certain codec standard specification, if it is determined that the encoder and decoder use a single pattern, then it is acceptable not to transmit information of the above pattern through the bitstream.
[0233] In a certain codec standard specification, if the encoder and decoder use a single pattern and specify the aspect ratio of the pattern, or allow rotation, tilt, scaling and other affine transformations, then the encoder according to the above specification can be designed to insert parameter values related to the aspect ratio and / or affine transformations, such as angle, reference point, magnification, etc., into the video bitstream, and the decoder according to the above specification can be designed to decode GDR frames by reading the above parameter values.
[0234] In a certain codec standard specification, if it is determined that the encoder and decoder use multiple patterns that belong to or do not belong to the above embodiments, it may be necessary to provide the video bitstream with information indicating that one of the multiple patterns has been selected. In this case, the encoder based on the above specification can be designed to insert information about the selection of the aforementioned pattern into the video bitstream, and the decoder based on the above specification can read the type of the aforementioned pattern to decode the GDR frame.
[0235] In a certain codec standard specification, if the shape of the pattern transmitted between the encoder and decoder is determined, it may be necessary to provide the video bitstream with a value that binarizes the shape of the pattern used for GDR. In this case, the encoder, based on the aforementioned specification, can be designed to insert the binarized form of the pattern into the video bitstream, and the decoder, based on the aforementioned specification, can read the form of the pattern to decode the GDR frame. According to an embodiment, the binarized form of the pattern may correspond to a binary bitstream in the form of indicating 1 as a first interval and 0 as a second interval; a binary bitstream consisting of a 2-bit bitmap in the form of indicating 0 as a dirty region, 1 as a clean region, and 2 as a refreshed region; or a binary bitstream in another such form.
[0236] Furthermore, in a certain codec standard specification, when the encoder encodes a video bitstream containing GDR frames using an appropriate method, if the decoder does not need information about the order in which the GDR frames are refreshed, then regardless of the pattern used, it is not necessary to transmit information about the pattern from the encoder to the decoder via the bitstream. This is because, for example, as long as the encoder encodes the predetermined coding units using a limited number of intra-frame and / or intra-frame / inter-frame coding methods according to the specified GDR method, it is sufficient for the decoder to perform the decoding process corresponding to the GDR without additional signal transmission.
[0237] In all the aforementioned cases, when the encoder inserts certain information into the bitstream, various encoding methods can be applied depending on the specifications and quantity of the information. These encoding methods can include fixed-length encoding and variable-length encoding. It is well known that variable-length encoding can employ Huffman coding, exponential Golomb coding, context-adaptive variable-length coding, context-adaptive arithmetic coding, and other encoding methods, especially entropy coding. It can be confirmed that by using the aforementioned encoding methods, the inserted information can be represented as at least one bit of relevant information at any position in the bitstream.
[0238] It is obvious that when choosing an encoding method in the above examples or types not shown, the decoder should be designed to decode the information by the reverse (inverse) of the same method.
[0239] Third embodiment and its encoding and decoding
[0240] According to a third embodiment of the present invention, when performing GDR, "clean" and "dirty" regions are no longer distinguished separately. However, it can be configured such that during GDR execution, the actual refreshed regions are inserted into the bitstream in a manner that is displayed through a majority of frames by forcing the number of intra-coded blocks encoded by conventional methods. The third embodiment of the present invention allows implementation in a progressive intra-recovery (PIR) manner as a special implementation of GDR.
[0241] Figure 17 This is a schematic diagram illustrating GDR implementation through forced intra-frame coding according to an embodiment of the present invention. Figure 17 The diagram shows four frames 1720, 1721, 1722, and 1723 contained within a GDR interval 1710 of length 4 frames. In the third embodiment of the invention, the GDR interval can be considered as an interval in which all spatial regions of the video within the aforementioned interval are intra-coded at least once. This method can use existing conventional encoders, but is implemented as follows: among the image coding units belonging to the same frame, for example, those image coding units that can correspond to at least one of network abstraction layer unit (NAL unit), macroblock, block, subblock, coding tree unit (CTU), coding unit, prediction unit, transform unit, partitioned area, and other various coding units, a number of coding units equal to the refresh threshold must be forced to perform intra-frame coding. For example, as Figure 17 As shown in the example, it can be easily deduced that if the GDR interval 1710 occupies 4 frames in length, the refresh threshold (representing the spatial region that must be intra-coded in each frame) can be specified as 25%. That is, it can be confirmed that when a GDR interval has a length of m frames, at least 1 / m of the spatial region in each frame needs to be intra-coded.
[0242] In each frame belonging to the GDR frame and intermediate frames, information about spatial regions that have been intra-coded at least once can be stored in the encoder. When calculating the refresh threshold mentioned above, regions that have been refreshed more than once may be excluded. That is, in the intermediate frames 1721 and 1722 belonging to GDR interval 1710, all frames can be encoded by assigning and according to the following objective: intra-coding at least 25% of the video area in spatial regions that have not been intra-coded in the previous frames belonging to the same GDR interval 1710. Furthermore, in the recovery frame indicating the end of the GDR interval, intra-coding is forced on all previously unrefreshed spatial regions, thereby ensuring that the entire video area within the entire GDR interval can be refreshed by performing intra-coding at least once.
[0243] To enforce intra-frame coding within a regular coding flow, a method may be needed to instruct intra-frame coding to be enforced for blocks that would otherwise be inter-frame coded using existing encoders. According to one embodiment of the invention, a method can be used to instruct blocks planned for inter-frame coding to be arranged in descending order of coding efficiency and to be encoded by converting the inter-frame coded blocks into intra-frame coded data blocks.
[0244] Figure 18 This is a flowchart illustrating a forced intra-frame coding method according to an embodiment of the present invention. Before the description, Figure 18 The flowchart shown is merely one embodiment illustrating the idea of the invention and used to describe a representative implementation method. It will be apparent to those skilled in the art that improvements can be made. Figure 18 All instances where the ideas of the invention are implemented through the steps shown in the flowchart or through other steps shall fall within the scope of the claims of the invention.
[0245] refer to Figure 18A coding loop for all image coding units belonging to the same frame is disclosed (S1810). For each coding unit, the encoder predicts the coding efficiency for intra-frame coding (S1820) and the coding efficiency for inter-frame coding (S1825). This prediction of coding efficiency can be used in existing encoder implementations or implemented by any newly provided method. As a most basic embodiment, a method can be used that evaluates the bit rate after actually performing intra-frame coding and inter-frame coding processes separately for the same image coding unit. However, methods can also be used to predict this efficiency without performing overall coding using various metrics. For example, prediction can be made by the partitioning level of the coding tree unit (CTU), or by the distribution of DCT coefficient values, or by coding information including the coding mode set for surrounding image coding units, or by the coding results of already coded regions, or other methods.
[0246] However, when evaluating the efficiency of the aforementioned intra-coding, it should be noted that the intra-coding is used to implement GDR and must be able to decode completely using the internal information given in each frame. Therefore, the aforementioned intra-coding may need to be evaluated using the aforementioned limited intra-coding methods. As mentioned above, the aforementioned limited intra-coding may include prohibition of intra prediction or spatial range constraints in the intra-coding process.
[0247] Refer again Figure 18 After predicting the efficiency of each coding method, the image coding unit can be encoded according to predetermined conditions and through either intra-frame coding or inter-frame coding, and its content can be recorded in bitstream syntax.
[0248] As a first condition, it can be confirmed whether the current frame is a recovery point frame and whether the image coding unit contains a spatial region within the current GDR interval that has never undergone intra-frame coding (i.e., an unrefreshed unit) (S1830). If the above conditions are met, the coding unit must perform intra-frame coding to achieve the purpose of the GDR interval, and therefore intra-frame coding can be performed (S1860).
[0249] As a second condition, it can be confirmed whether intra-frame coding efficiency is superior to inter-frame coding efficiency, or whether there is contention within a certain threshold range (S1840). This confirmation process is widely used in conventional encoders, so any method already applied or newly applied to such encoders can be used. If the above conditions are met, intra-frame coding can be performed on the coding unit (S1860).
[0250] As a third condition, it can be confirmed whether the inter-frame coding efficiency is within the end threshold ranking within the frame (S1850). As described in the above embodiment, this condition can be applied to situations where a method is used to instruct the blocks planned for inter-frame coding to be arranged in descending order of coding efficiency, and to encode them by changing the inter-frame coding blocks to intra-frame coding blocks. The determination of whether it is within the end threshold ranking can be made after actually performing coding on the entire frame once, but it can also be predicted using various indicators without performing overall coding, as described in the above content on predicting coding efficiency. If it is determined that the inter-frame coding unit belongs to the end threshold ranking, as an auxiliary condition, it can be determined whether the threshold area that must be intra-frame coded within the GDR frame has been reached (S1855). In one embodiment, if the threshold area has been reached, it can be set to no longer perform mandatory intra-frame coding to ensure that the stable bit rate provided by the GDR function is maintained. In another implementation, this judgment can also be determined by ranking the coding efficiencies. Therefore, it can also be achieved by selecting the inter-frame coding target unit required to reach the threshold area in descending order of inter-frame coding efficiency. If the above conditions are met, intra-frame coding can be performed on the coding unit (S1860).
[0251] If none of the above conditions are met, then the coding unit can be inter-frame coded (S1865).
[0252] If intra-frame coding is performed (S1860), the encoder can be configured to store information indicating that the spatial region on the frame where intra-frame coding was performed is the region where information was refreshed through intra-frame coding (S1870). The stored information can be used as a criterion in the above-mentioned judgment conditions, such as the first condition S1830 or the auxiliary condition S1855 of the third condition.
[0253] If the method described above according to the third embodiment of the present invention is adopted, it can be confirmed that, according to existing GDR technology, such as the implementation method of the technology proposed in the second edition of the H.266 / VVC standard specification, the existence and location-related syntax signals for distinguishing the "clean" and "dirty" regions refreshed by GDR technology may be unnecessary. This may be because conventional intra-frame coding syntax is used in the same way to achieve the GDR effect. Furthermore, if the method according to the third embodiment of the present invention is adopted, depending on the implementation method, separate signals identifying GDR intervals or indicating frames associated with GDR may not be provided, and the corresponding frames may also be considered as inter-frame frames with conventional unidirectional or bidirectional references.
[0254] That is, according to a preferred embodiment of the present invention, for a certain codec standard specification, when the encoder encodes a video bitstream containing GDR intervals according to the third embodiment of the present invention using an appropriate method, the decoder can be configured to perform decoding through a conventional process without requiring separate GDR-related information for decoding such frames. In this case, the encoder can be configured to internally use the information required for forced intra-frame coding during the encoding process, but without recording it in the bitstream and erasing it; this information is not transmitted to the decoder, and the decoder can be configured to decode the received information using a conventional method. By applying the above-described limited intra-frame coding, even if the decoder interprets the bitstream syntax in a conventional manner, only limited intra-frame prediction is performed during the decoding of image information designated as GDR intervals, thus enabling progressive recovery of image information in the same manner as implementing GDR.
[0255] Of course, additional information can be inserted into the bitstream as needed to facilitate the implementation of the third embodiment described above. In this case, when the encoder inserts certain information into the bitstream, various encoding methods can be applied according to the specifications and quantity of the information. These encoding methods can include fixed-length encoding and variable-length encoding. It is well known that the variable-length encoding can employ Huffman coding, exponential Golomb coding, context-adaptive variable-length coding, context-adaptive arithmetic coding, and other encoding methods, especially entropy coding. It can be confirmed that by using the above encoding methods, the inserted information can be represented as at least one bit of relevant information at any position in the bitstream.
[0256] It is obvious that when choosing an encoding method in the above examples or types not shown, the decoder should be designed to decode the information by the reverse (inverse) of the same method.
Claims
1. An encoding method, which is a video encoding method including a progressive decoder refresh GDR function, characterized in that, include: The steps for setting the frame length of the GDR interval used to perform stepwise decoder refresh; In at least one frame belonging to the GDR interval, the step of specifying a refresh region including at least one image coding unit is performed. The step of performing intra-frame coding on at least one image coding unit contained in the refresh area; as well as The step of recording the encoded result into a bitstream. Within the GDR range, all spatial regions within the video resolution are intra-coded at least once.
2. The encoding method according to claim 1, characterized in that, The steps for specifying the refresh area include: The step of specifying the refresh area from the spatial area derived from a function with frame number as the variable.
3. The encoding method according to claim 2, characterized in that, The steps for specifying the refresh area include: The step of specifying the refresh area is to exclude the area derived from the function result of the input frame number (n-1) from the area derived from the area derived from the function result of the input frame number (n-1).
4. The encoding method according to claim 2, characterized in that, The function is either a line whose area can be derived through integration, or a closed curve whose area can be calculated.
5. The encoding method according to claim 2, characterized in that, The function is applied after being transformed by at least one of the following operators: the operator for reversing the region to be determined as an area, the rotation operator, the tilt operator, the scaling operator, and the affine transformation operator.
6. The encoding method according to claim 2, characterized in that, Also includes: The step of recording information related to the type of the function into the encoded bitstream.
7. The encoding method according to claim 1, characterized in that, The steps for specifying the refresh area include: The step of specifying the refresh area from the spatial area specified by a sequentially determined transformation pattern that changes as the frame progresses.
8. The encoding method according to claim 7, characterized in that, The pattern shape that needs to be transformed in each frame of the transformed pattern is defined as the residual relative to the pattern that comes first in time.
9. The encoding method according to claim 7, characterized in that, The transformed pattern is applied after being transformed by at least one of the following operators: an operator for reversing the region to be determined as an area, a rotation operator, a tilt operator, a scaling operator, and an affine transformation operator.
10. The encoding method according to claim 7, characterized in that, Also includes: The step of recording information related to the type of the transformation pattern into the encoded bitstream.
11. The encoding method according to claim 1, characterized in that, The steps for specifying the refresh area include: Specify a spatial region such that within the GDR interval, all spatial regions within the video resolution are intra-coded at least once.
12. The encoding method according to claim 11, characterized in that, The steps for specifying the refresh area include: When the GDR interval has a length of m frames, at least 1 / m of the spatial region in each frame is specified to be intra-frame encoded.
13. The encoding method according to claim 11, characterized in that, It also includes the step of storing in the encoder cumulative refresh region information representing a spatial region that is intra-coded at least once within the GDR interval.
14. The encoding method according to claim 11, characterized in that, The steps for specifying the refresh area include: The steps are specified to change at least one inter-frame coding object image coding unit to an intra-frame coding object image coding unit for encoding.
15. The encoding method according to claim 14, characterized in that, The steps of changing and specifying are operated based on the coding efficiency of the inter-frame coding object image coding unit, and are configured to change and specify when the coding efficiency of the inter-frame coding object image coding unit is low.
16. A decoding method, which is a video decoding method including a step-by-step decoder refresh GDR function, characterized in that, include: The steps for receiving an encoded bit stream; The step of reading at least one frame constituting the GDR interval from the bit stream; The step of intra-frame decoding of at least one image coding unit contained in at least one frame belonging to the GDR interval; as well as The step of recording the result of the intra-frame decoding into the output buffer. Within the GDR range, all image regions within the video resolution are decoded at least once per frame.
17. The decoding method according to claim 16, characterized in that, In each frame within the GDR interval, the image region for which the intra-frame decoding is performed is contained within a functional region of spatial area derived from a function that takes the frame number as a variable.
18. The decoding method according to claim 16, characterized in that, In each frame within the GDR interval, the image region for intra-frame decoding is contained within a transform pattern region corresponding to a spatial area specified by a transform pattern that changes sequentially as the frames progress.
19. The decoding method according to claim 16, characterized in that, In each frame within the GDR interval, when the GDR interval has a length of m frames, each frame of the image region for intra-frame decoding contains at least 1 / m of a spatial region.
20. The decoding method according to claim 16, characterized in that, The image coding unit performing the intra-frame decoding decodes in a manner that does not refer to image coding units that do not belong to the same frame or have not been intra-frame encoded.