Method and device for video encoding and decoding based on radial motion vector
Radial global motion vectors in video encoding and decoding improve compression efficiency by minimizing resource and time consumption, addressing the challenges of high compression ratios and image quality degradation in existing technologies.
Patent Information
- Application Number
- PCT/KR2025/007753
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-06-05
- Publication Date
- 2025-12-11
AI Technical Summary
Existing digital video encoding and decoding technologies face challenges in achieving high compression ratios while minimizing image quality degradation, particularly in terms of computational resources and time consumption, which can interfere with the production, recording, and distribution of digital videos.
The use of radial global motion vectors for encoding and decoding, characterized by entropy-encoded bit strings, allows for improved video compression efficiency by replacing multiple individual motion vectors, and includes steps of reading and combining radial and local motion vectors to generate decoded images.
This approach enhances video compression efficiency by reducing computational resources and time consumption, thereby improving image quality and encoding/decoding performance.
Smart Images

Figure KR2025007753_11122025_PF_FP_ABST
Abstract
Description
Method and device for video encoding and decoding based on radial motion vectors
[0001] The present invention relates to the field of encoding and decoding of digital video, and relates to a method for encoding and decoding digital video, a method for recording such data, and components, devices, and systems for realizing such a method.
[0002] The present invention may be in the same technical field as at least one of the digital video compression technology standards known by the names of standards such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, or may be in the technical field for improving the inherent efficiency of the standards, or may be in the technical field for improving or replacing the standards.
[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission over communication networks, video calls / video conversations / video chats, recording and providing video content using optical media including video compact discs (VCDs), digital versatile discs (DVDs), and Blu-Rays, all processes for producing, editing, collecting, and distributing video content, and devices such as video recording devices and camcorders for various purposes, including personal, commercial, industrial, and security purposes, all depend on video encoding and decoding technologies.
[0004] Accordingly, implementations that may be referred to as digital video encoders and decoders may form part of a wide range of devices related to the generation, recording, and provision of digital video, including digital televisions, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and other devices.
[0005] The above digital video encoders and decoders can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art. The digital video compression standard may include at least one of compression standards known by a standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0006] Video encoders and decoders can be implemented to more efficiently encode or decode digital video information while complying with the above standards, or by improving or modifying the above standards. Attempts to modify the above standards can also lead to the development of new standards. A well-known example is the so-called enhanced compression model (ECM), an attempt to improve and replace the existing H.266 / VVC standard, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.
[0007] Digital video is known to require a significant amount of information to describe its content in its uncompressed state. Therefore, recording or transmitting this information in its original form can be inefficient. Therefore, digital video is compressed using various methods prior to recording or transmission. These compression methods include lossy and lossless encoding. Lossy encoding sacrifices some image quality to achieve high compression performance, while lossless encoding sacrifices some compression performance to prevent image quality degradation. Regardless of the encoding method, there is a need to implement a technology that achieves a high compression ratio while minimizing image quality degradation to meet the growing demand for high-quality digital video within the constraints of limited memory storage capacity and communication transmission bandwidth.
[0008] As described above, the encoding process for compression requires various operations, such as spatial segmentation of digital video, segmentation and / or processing in color channels, removal of spatial redundancy, removal of temporal redundancy, tracking of motion vectors within the video, encoding of differential images, quantization, coefficient scan, run-length coding, entropy coding, and loop filtering. These encoding operations generally consume computing resources and take a certain amount of time to complete. Similarly, the decoding operation for the encoding operation also requires certain computing resources and a certain amount of time. The main goal of video encoding and decoding technology is to ensure that the above resource consumption and time consumption do not interfere with the production, recording, distribution, and viewing of digital videos.
[0009] Accordingly, the present invention provides a new technology that can contribute to at least one of the following: improvement of encoding efficiency, improvement of decoding efficiency, improvement of video quality, reduction of computational amount, reduction of software size, reduction of hardware size, and improvement of other performances related to encoding and decoding in the technical tasks in the field of video encoding and decoding.
[0010] A method for decoding an encoded video bitstream according to an embodiment of the present invention for solving the above-described technical problem may include a step of reading information related to a radial global motion vector from the bitstream, a step of executing prediction decoding based on the radial global motion vector to generate a predicted image, a step of reading information related to a residual image from the bitstream, and a step of combining the predicted image and the residual image to generate a decoded image.
[0011] The information related to the above radial global motion vector may be characterized by including at least one of first information related to the magnitude of the radial global motion vector, second information related to the aspect ratio of the radial global motion vector, and third information related to the centric point of the radial global motion vector.
[0012] The method may be characterized in that at least one of the first information, the second information, and the third information is read from at least one entropy-encoded bit string included in the bitstream.
[0013] The above entropy encoded bit string may be characterized in that it is interpreted by an entropy decoding algorithm corresponding to context-adaptive binary arithmetic coding (CABAC).
[0014] The above first information may be characterized by indicating the size of a basis vector constituting the radial global motion vector.
[0015] The size of the above basic vector may be characterized in that it is defined in proportion to at least one of the horizontal and vertical resolutions of the predicted image.
[0016] The above second information may be characterized by indicating an angle of a basic vector constituting the above radial global motion vector.
[0017] The above second information may be characterized in that it is defined as a ratio of one direction to the other direction of the horizontal or vertical direction of the radial global motion vector.
[0018] The third information may be characterized by indicating a distance at which the center point of the radial global motion vector is separated from the reference point of the predicted image.
[0019] The third information may be characterized by including at least one of information about a first center point determined through translation of the radial global motion vector, and information about a second center point determined through eccentricity of the radial global motion vector.
[0020] The method may be characterized in that the first information, the second information, and the third information are simultaneously defined by one vector capable of determining the size, angle, and position.
[0021] The above method may be characterized in that the first information, the second information, and the third information are simultaneously defined by a single rectangle capable of determining a horizontal length, a vertical length, and a position.
[0022] The method may be characterized in that the first information, the second information, and the third information are simultaneously defined by one ellipse capable of determining a diameter, ellipticity, and position.
[0023] The step of generating the above predicted image may be characterized in that it is performed without referring to any motion vector other than the above radial global motion vector.
[0024] The method may further include a step of reading information related to at least one local motion vector from the bitstream, and the step of generating the predicted image may be performed with reference to the radial global motion vector and the local motion vector together.
[0025] The step of generating the above predicted image may be characterized in that it is performed by inter-frame prediction.
[0026] A method for encoding a video according to an embodiment of the present invention for solving the above-described technical problem may include a step of calculating a radial global motion vector for an image included in the video, a step of performing prediction coding for the image based on the radial global motion vector to generate a residual image, and a step of encoding information related to the radial global motion vector and information related to the residual image in a bitstream.
[0027] The step of calculating the above radial global motion vector may be characterized in that it is performed with reference to information related to the direction and intensity of acceleration obtained from an accelerometer associated with the source of the video.
[0028] According to an embodiment of the present invention for solving the above-described technical problem, a decoder device configured to decode a bitstream includes a processor, a memory, a bitstream parser, and a prediction decoding unit, wherein the bitstream parser is configured to read information related to a radial global motion vector and information related to a residual image from the bitstream, and the prediction decoding unit may be configured to execute prediction decoding based on the radial global motion vector to generate a predicted image, and to combine the predicted image and the residual image to generate a decoded image.
[0029] According to an embodiment of the present invention for solving the above-described technical problem, a video encoding device configured to output a bitstream includes a processor, a memory, a global motion calculation unit, a prediction encoding unit, and a bitstream generation unit, wherein the global motion calculation unit is configured to calculate a radial global motion vector for an image included in the video, the prediction encoding unit is configured to perform prediction encoding on the image based on the radial global motion vector to generate a residual image, and the bitstream generation unit can be configured to encode information related to the radial global motion vector and information related to the residual image in an output bitstream.
[0030] According to the present invention, by using a radial motion vector, the effect of replacing a plurality of individual motion vectors can be obtained, thereby improving video compression efficiency.
[0031] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention;
[0032] FIG. 2 is a conceptual diagram of the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention.
[0033] Figure 3 is a functional unit conceptual diagram of a video decoder according to one embodiment of the present invention;
[0034] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention;
[0035] Figure 5 is a conceptual diagram of a frame type according to one embodiment of the present invention;
[0036] Figure 6 is a conceptual diagram showing the structure of a video encoder using VVC.
[0037] Figure 7 is a flowchart of a video encoding process according to one embodiment of the present invention.
[0038] FIG. 8 is an example diagram of a radial global motion vector according to one embodiment of the present invention;
[0039] FIG. 9 is a conceptual diagram of a method for representing first information related to the intensity of a radial global motion vector according to one embodiment of the present invention;
[0040] FIG. 10 is a conceptual diagram of a method for representing second information related to the aspect ratio of a radial global motion vector according to one embodiment of the present invention;
[0041] FIG. 11 is a conceptual diagram of a method for representing third information related to the center point of a radial global motion vector according to one embodiment of the present invention;
[0042] FIG. 12 is a conceptual diagram of a first center point reference defining the center point of a radial global motion vector according to one embodiment of the present invention;
[0043] FIG. 13 is a conceptual diagram of a second center point reference defining the center point of a radial global motion vector according to one embodiment of the present invention; and
[0044] FIG. 14 is a flowchart of a decoding process of an encoded video bitstream according to one embodiment of the present invention.
[0045] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.
[0046] Although terms such as “first,” “second,” etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term “and / or” includes any combination of multiple related listed items or any of multiple related listed items, and is non-exclusive unless otherwise indicated. The listing of items in this specification is merely an exemplary description to easily explain the spirit and possible implementation methods of the invention herein, and therefore is not intended to limit the scope of embodiments of the present invention.
[0047] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0048] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0049] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0050] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0051] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0052] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0053] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0054] In describing the invention herein, embodiments may be described or illustrated in terms of unit blocks that perform the described function or functions. The blocks may be expressed herein as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware implementation methods, but not limited thereto. Alternatively, the blocks may be implemented in software by application software, operating system software, firmware, or information processing software implementation methods, but not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may be implemented to operate in an environment where their physical locations are not specified and are separated from each other by a communication network, the Internet, a cloud service, or a communication method, but not limited thereto. All of the above implementation methods are within the scope of various embodiments that can be taken by a person skilled in the field of information and communication technology to implement the same technical idea, and therefore, any detailed implementation method should be interpreted as being included within the scope of the technical idea of the invention in this specification.
[0055] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. To facilitate a comprehensive understanding of the present invention, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted. Furthermore, the multiple embodiments are not mutually exclusive, and it is assumed that some embodiments may be combined with one or more other embodiments to form new embodiments.
[0056]
[0057] digital video codec
[0058] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other via a network (105).
[0059] In one embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a unidirectional video communication network. A first terminal (110) among the terminals may encode video data in order to transmit (111) the video data via a network (105). A second terminal (120) among the terminals may be configured to receive (121) the encoded video data via a network and decode and display the same.
[0060] In another embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a two-way video communication network. For the two-way video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal via the network. Each terminal may also be configured to receive (113, 123) video data transmitted by another terminal via the network, decode the same, and display the decoded video data.
[0061] The terminals (110, 120) shown in Fig. 1 may be exemplified as devices such as server computers, personal computers, portable computers, and smartphones, depending on the embodiment, but are not limited thereto. The present invention is applicable to all environments for establishing a one-way or two-way video communication network, and it should be understood that the network (105) may be established by any means for transporting encoded video data between the terminals (110, 120).
[0062] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. Depending on the embodiment, the network may be configured to communicate information using any communication standard, which may include packet-based communication. The packet communication may be understood to include packets, for example, known as TCP or UDP.
[0063] However, in another embodiment of the present invention, the network (105) may be understood to include a process of information transmission using a recording medium. In this case, the configuration of the network is not limited to a communication medium, and should be understood to include a process of temporarily storing and physically transporting information on a hard disk, a solid state disk (SSD), a flash memory, a compact disc (CD), a digital versatile disc (DVD), a Blu-ray disc, and other mechanical, electronic, or optical recording media.
[0064] Any other means of information communication or transport, regardless of the method employed, can be considered within the scope of the present invention as long as it has a structure that supports the transmission and decoding of video data in an encoded state. Therefore, in addition to the examples listed above, any means of information communication or transport, whether known in the past or newly available, can fall within the scope of application of the present invention.
[0065] FIG. 2 is a conceptual diagram illustrating the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention. The streaming system (200) illustrated in FIG. 2 can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, it should be noted that technical structures identical or similar to the streaming system can be equally applied even when information is transmitted via a recording medium, as described above.
[0066] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212), which may be, for example, a digital camera or other device, for acquiring uncompressed raw video. The raw video stream (215) may have a large capacity and may therefore be compressed by a video encoder (217) coupled or connected to the video source.
[0067] The above encoder (217) may be configured as a means including hardware, software, or a combination of the two, configured to implement an image encoding method and / or an implementation method thereof according to one embodiment of the present invention.
[0068] Through the encoder (217), an encoded bitstream (219) having a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real time via a relay device, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.
[0069] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit string (229) in real time or obtain it later. The streaming client may include a video decoder (232) that obtains the encoded bit string (229) (which may also be considered as a copy of the bit string (219) received by the streaming server), decodes the bit string (229), and outputs the resulting video data as video data in a form that can be displayed by a display (235) or other visual, auditory, or other sensory display means.
[0070] As described above, the functions for encoding and decoding video data are collectively called a coder-and-decoder, or video codec.
[0071] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiving unit (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiving unit (310) through a hardware or software connection (315) to a device storing the same, and as described above, the storing device may be a type of streaming server located at the other end of a communication network, or may mean a physical recording medium, but is not limited thereto.
[0072] The above-described receiving unit (310) can receive the encoded video data together with other data accompanying it, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing function unit (312) other than the video decoder.
[0073] When the video data is provided through a communication network, a buffer memory (320) may be coupled between the receiving unit (310) and the decoder (305) to minimize delay and disconnection according to the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and stably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, if the bandwidth of the communication network is sufficient, if video data is read from a recording medium in a local location that is not physically separated, or if the possibility of communication delay is not predicted in other environments, the buffer memory may be unnecessary.
[0074] The video decoder (305) may include the parser (330) as its input terminal to interpret the encoded video data. The parser may perform a function of separating (parsing) a plurality of pieces of information stored in the form of a bit string in the encoded video data according to a predetermined rule, and, if necessary, performing an entropy decoding (335) of entropy-coded video data, thereby performing a function of reconstructing symbols (338), which are paragraphs of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that is attached to and operable with the decoder (305), such as a display device. Control information for controlling the above display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).
[0075] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The entropy encoding method of the encoded video data may vary depending on the encoding standard, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be a context-adaptive or context-sensitive method depending on the standard, and may also be based on principles widely known to those skilled in the art.
[0076] The parser (330) may be configured to extract at least one picture from the encoded video data. The definition of the picture may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may correspond simultaneously and overlappingly. The picture may be grouped, defined, and / or divided into encoding / decoding units such as, for example, groups of pictures (GOPs), pictures / frames, tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs).
[0077] The parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. In addition, the parser (330) may be configured to selectively supply a specific symbol (338) to a specific decoding function unit within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). Control of such information supply can be determined by the information sequence contained in the encoded video, and may vary depending on the encoding standard, and is not limited within the scope of the embodiments of the present invention, and is not described in detail in this conceptual diagram.
[0078] The decoder (305) may be comprised of a number of conceptual functional units that receive and process the encoded information provided by the parser (330). It should be readily apparent that these conceptual functional units may be combined or further subdivided, depending on implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with each other. However, despite the possibility of such integration or separation, the following description will be given as a combination of conceptual functional units to illustrate the video data decoding procedure applied as an embodiment of the present invention.
[0079] The decoder may include an inverse quantization and inverse transformation unit (340). The inverse quantization and inverse transformation unit (340) may be configured to receive encoding information including a method to be used for numerical transformation (transform), a block size, quantization coefficients for recovering quantized information, and distinction information of a quantization matrix that simplifies and represents the quantized coefficients from the parser (330), and may be configured to output block values (341) that can be input to an aggregator (370) as a result of processing the encoding information.
[0080] In one embodiment of the present invention, the output values of the inverse quantization and inverse transformation unit (340) may include intra-prediction encoded block values. The intra-predicted block values may refer to values that can be decoded using prediction information within a picture currently being decoded, for example, the current frame, but not using prediction information from a previously decoded picture, for example, the previous frame.
[0081] The prediction information within the current picture may be provided by the intra prediction unit (350). According to an embodiment of the present invention, the intra prediction unit (350) generates a block value of the same form as the block being decoded as the prediction information by using picture information of a spatially adjacent area derived from a picture currently being decoded and of which decoding has been partially completed. The picture information may be provided (381) from a buffer for the current picture, a so-called line buffer (380). The merging unit (370), according to an embodiment, may be configured to merge the prediction information (351) generated by the intra prediction unit (350) with the block values (341) provided by the inverse quantization and inverse transformation unit (340).
[0082] In another embodiment, the output values of the inverse quantization and inverse transformation unit (340) may include block values subjected to motion compensation as inter-prediction encoded block values, and in some cases, block values subjected to motion compensation. In this case, the inter-prediction unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values as the output values may be configured to be merged with the block values (341) provided by the inverse quantization and inverse transformation unit (340) by the merger unit (370). In this case, the block values (341) may be referred to as so-called differential or residual values.
[0083] The position information within the memory used by the inter prediction unit (355) to extract the sample information from the reference picture may be determined by a motion vector provided to the inter prediction unit (355) and composed of a combination of symbols (338) for indicating, for example, X, Y, and other specific points of the reference picture. The inter prediction unit (355) may also include a function for interpolating and using the sample values when a so-called 'subsampling'-capable motion vector is provided, and may further include a function for predicting and reinforcing the value of the motion vector.
[0084] The output values (371) of the above merging unit (370) may be provided to the loop filter unit (360) and processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided from the parser (330) to control its operation. The output of the loop filter unit (360) may be output to an external display means such as the display device through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction for interpreting a subsequent intra or inter coded block value, and may also be stored in a reference picture buffer (385) through this.
[0085] Certain pictures, such as frames, after their decoding is completed, can be utilized as reference pictures for performing predictive decoding in a subsequent decoding process. A picture (or frame) can be gradually accumulated in a line buffer (380) and decoded. When the decoding of a frame is completed, the contents of the line buffer (380) are transferred (383) to the reference picture buffer (385), and a new line buffer (380) can be allocated for decoding the new frame.
[0086] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technique that may be documented by various international standards or commercial standards. The standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sub-Division (ITU-T). Those skilled in the art will understand that each of the above recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the relevant standards, as defined and required by the video compression standard documents and standard documents, and specifically by the profiles and levels specified within such documents. In addition, the complexity of the encoded video data may be limited to a certain level to comply with the profiles and levels. For example, a profile or level may be configured to limit a maximum picture size, a maximum decoding speed, and a maximum reference picture size. These limitations may, in some embodiments, also be further limited by metadata signals for a hypothetical reference decoder (HRD) and HRD buffer management included in the encoded video data.
[0087] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data along with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that may be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that approximates the original image. The additional data may be provided in the form of, for example, layers for temporal, spatial, or signal-to-noise ratio (SNR) enhancement, redundant slices, redundant pictures, and forward error correction codes.
[0088] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.
[0089] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. In addition, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. In addition, the original video information (402) may have any suitable sampling structure corresponding to the color space, for example, Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4, etc. The original video information (402) having such a predetermined format may be provided to the encoder in the form of a digital video stream.
[0090] In a one-way video communication network, the original video information (402) can be obtained from a recording medium storing a previously prepared video source. In a two-way video communication network, the original video information (402) can be obtained from an image acquisition device, such as a camera, that generates at least one video transmission stream included in the two-way video communication.
[0091] The video data including the above original video information (402) may be configured as a plurality of pictures configured to simulate motion by playing them in chronological order. In addition to pictures, the pictures may also be expressed in concepts such as frames. The pictures may include one or more samples depending on the type of sampling structure, color space, etc. being used. Those skilled in the art will understand that the samples are closely related terms to pixels in digital images. The operation of the encoder will be described below with reference to such samples.
[0092] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress pictures (and / or grouped or segmented information) constituting the original video information (402) in real time (or according to other temporal requirements required according to the implementation method) into the form of encoded video information.
[0093] In the encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and be functionally coupled to the following functional units as described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as picture skip, quantizer, and variable values for applying image quality optimization techniques, and may also include values such as picture size, group of pictures (GOP) structure, and maximum search range of motion vectors. A person skilled in the art will be able to understand various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of a video encoder optimized for individual system design.
[0094] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" well known to those skilled in the art. To simplify the description by way of example, the coding loop may be configured with an internal encoder (so-called "source coder") (410) which is responsible for receiving a picture to be encoded and generating symbols based on at least one reference picture that has been encoded in the past, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that will receive encoded video information from the encoder (405) by receiving an output of the internal encoder (410).
[0095] Video data composed of sample data reconstructed by the internal decoder (420) may be configured to be input to the reference picture buffer of the encoder (405). As described above, the internal decoder (420) is implemented to reproduce the result output by the encoder (405) and to be decoded by a remote decoder, so the video data recorded in the reference picture buffer may also be identical in bit unit to the information of the reference picture buffer of the remote decoder. That is, the prediction function unit that may be included in the encoder (405) may read the same values as the sample values of the previous frame that the decoder will later refer to in the decoding process from the reference picture buffer of the encoder (405).
[0096] As described above, the principle of achieving matching of the reference picture buffers between the encoder (405) and the decoder (490) by means of the internal decoder (420) on the encoder (405) side is well known to those skilled in the art, and a method of responding to an environment in which such an environment is not guaranteed (e.g., information loss due to communication failure, etc.) can also follow what is known to those skilled in the art.
[0097] An embodiment of the operation method of the internal decoder (420) has been described in detail above with reference to FIG. 3. The decoder of FIG. 3 may be regarded as the aforementioned "remote" decoder (490). The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335). This is because the internal encoder (405) is implemented to simply reproduce the operation of a decoder located at a remote location, and thus may directly decode symbols without requiring a process of compressing and then decompressing symbols. Accordingly, the functional units preceding the parser and entropy decoder as shown in FIG. 3 may not be provided or may be implemented at least partially.
[0098] As described above, according to a preferred embodiment of the present invention, any decoder function (excluding a parser and an entropy decoder) present in the decoder can naturally exist as a substantially identical function in the corresponding encoder (405).
[0099] The operation of the encoding function unit that may be included in the above encoder (405) can be considered as the reverse operation of the decoder function unit. Therefore, the embodiment can be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter prediction encoding unit corresponding to the inter prediction unit may be provided. In addition, some additional explanations will be added.
[0100] The internal encoder (410) may be configured to perform encoding on input picture information, for example, an input frame, by a prediction encoding method executed by a prediction encoding unit (440) that operates by referencing at least one reference picture information, for example, at least one temporally previous encoded picture (or frame) from a reference picture buffer (430) from video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input picture and blocks of samples constituting the reference picture.
[0101] The internal decoder (420) can decode video data that can be designated as the reference picture from symbols generated by the internal encoder (410). As described above, since the video data is identical to the decoding operation performed by the remote decoder, the video data used as the reference picture may be provided to the encoder (405) in a form that has undergone lossy compression and has been partially damaged, and this operation may be intended to ensure operational consistency with the decoder.
[0102] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter prediction or intra prediction described in the description of the decoder. For picture information that is input and scheduled to be newly encoded, the prediction unit may access the reference picture buffer (430) to retrieve information, such as a motion vector, a block shape, and metadata that may include the same, which are information indicating a point of a reference picture that can function as prediction reference information suitable for the new picture information, and a sample block to be actually referenced. The prediction encoding unit (440) may operate on the basis of a so-called "sample block by pixel block" to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information designating at least one reference picture information stored in the reference picture buffer (430) may be designated for the input picture, as determined based on the search results obtained by the prediction encoding unit (440).
[0103] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including setting parameters used to encode video data.
[0104] All outputs of the above-described functional units may be subjected to entropy encoding (460) in order to be finally output. The entropy encoding (460) may include various entropy coding techniques, such as variable length coding, Huffman coding, and arithmetic coding, for the symbols generated by the various functional units as described above, and each encoding method may be a context-adaptive or context-sensitive method according to the standard, or may be based on principles widely known to those skilled in the art. Such entropy encoding (460) can typically achieve lossless compression, and thus can be configured to convert at least one symbol generated by the functional units into encoded video data.
[0105] The above control unit (450) may, when controlling the operation of the encoder (405), apply to each picture (or frame) the type of encoding in which a specific picture is encoded during the encoding period. Depending on the type, the method by which the picture is encoded may be affected. Depending on the embodiment, the type may include what is categorized as the following "frame type."
[0106] Fig. 5 is a conceptual diagram of a frame type according to one embodiment of the present invention. The following description will be made with reference to Fig. 5.
[0107] An intra (“I”) picture (510) may refer to a picture that can be encoded and decoded using only its own information without referring to other picture information in video data through predictive encoding. The “I” picture may be designated by names such as a key frame, an independent / instantaneous decoder referh (IDR) frame, and a clean random-access (CRA) frame, depending on the video encoding standard, and the “I” pictures designated by the various names as described above may have various modifications and application methods as permitted by each standard and may be partially different from each other. In addition to those listed above, various application methods for implementing the “I” picture may be by various methods that are already known to those skilled in the art or may be newly provided.
[0108] A prediction ("P") picture (520) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one prediction information and / or a motion vector that designates at least one reference picture to predict sample values of a block constituting the picture. The "P" picture may be configured to refer to only one reference frame, or may be configured to refer to one or more reference frames, according to a video encoding standard. When referring to more than one reference frame, sample information and / or associated metadata derived from multiple reference pictures may be used to reconstruct a single block. However, in common cases, a picture designated as a "P" picture may be understood as a picture that performs reference only to a temporally preceding picture.
[0109] A bidirectional prediction ("B") picture (530) may refer to a picture that can be encoded and decoded through intra or inter prediction based on at least one piece of prediction information and / or a motion vector that designates at least two reference pictures in order to predict sample values of blocks constituting the picture. In a common case, a picture designated as the "B" picture is distinguished from a picture designated as the "P" picture, and may be understood as a picture that performs a reference without being limited to a temporally preceding picture.
[0110] Video data may be spatially divided into a plurality of sample blocks during the encoding and decoding process, and encoding may be performed in units of the blocks. The block units may include, but are not limited to, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known. The block may be encoded using a predictive encoding method with reference to any other (already encoded) blocks, as permitted and / or restricted by the type specified for each picture in which the block is included. For example, the blocks of the "I" picture (510) may be encoded without using a predictive encoding method, or with reference to blocks that have already been encoded within the same partial picture. That is, only the so-called intra prediction method may be used. In contrast, the "P" picture (520) may further reference a reference picture encoded in at least one previous time unit, and thus, inter prediction may also be used for encoding along with intra prediction. In the case of a "B" picture (530), reference can be made not only to a previously encoded picture in the encoding order but also to a later reference picture in terms of time unit. However, it is widely known that there may be blocks within a "P" picture or a "B" picture that are encoded without relying on predictive encoding.
[0111] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technique that may be documented by various international standards or commercial standards. Examples of the above standards may include all of those described in the above decoder.
[0112] According to one embodiment of the present invention, the transmitter (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a remote decoder (490)) to a device storing the encoded video data via a hardware or software connection (495). According to an embodiment, when providing / transmitting the encoded video data from the video encoder (405), the transmitter (470) may receive and merge other data accompanying the encoded video data, for example, encoded audio data or other auxiliary data, from a separate source (480).
[0113] According to one embodiment of the present invention, the transmitter (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data or to more accurately reconstruct an image that approximates the original image. Examples of the additional data may include all of the examples previously presented with respect to the receiver (310) of the decoder.
[0114] The present invention can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art as described above. The digital video compression standard may include at least one of compression standards known by the standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.
[0115] Fig. 6 is a conceptual diagram illustrating the structure of a video encoder according to another embodiment of the present invention. What is depicted in Fig. 6 may be a rough structure of a video encoder widely known as a standard code such as ITU-T H.266 and ISO / IEC 23090-3, and also known as MPEG-I Part 3 or versatile video coding (VVC).
[0116] According to FIG. 6, a video encoder (605) may be configured to receive uncompressed and unencoded original video data (601) as input and output an encoded bit stream (602). The video data (601) may be supplied directly to a luma mapping unit (610a) when intra-encoded, or may be supplied to a luma mapping unit (610b) via an inter-prediction unit (620) including motion vector extraction. In the intra-encoded case, the mapped luma signal may be supplied to an output merger (606) by selecting (608) at least one of an intra-prediction encoded signal via an intra-prediction unit (625) or an inter-prediction encoded signal output from the luma mapping unit (610b) via the inter-prediction unit (620). The result of the above output merger can be applied to a chroma scaling unit (615). (The operation of the luminance signal mapping unit (610) and the operation of the chroma scaling unit (615) are collectively referred to as a luma mapping / chroma scalaing (LMCS) process.) The reduced chroma signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the chroma signal. The coefficients derived as a result of the transform are applied to a quantization unit (640) and quantized. As a result, lossy compression is achieved, and the result of the lossy compression can be output as a bit string (602) through a multi-hypothesis CABAC (650), which is a lossless compression method.
[0117] Meanwhile, the result of the lossy compression may actually enter the decoding process by going through the processes of inverse quantization (645), inverse transform (635), and luminance signal expansion (617) to generate a coding loop. The result of the luminance signal expansion may be supplied to the internal merger (607) together with the result of selecting (608) at least one of the previously generated intra prediction encoding signal or inter prediction encoding signal. The result of the internal merger may go through inverse luma mapping (618), and then may go through processing such as a deblocking filter (660), sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference picture buffer (690) and can be reused for prediction encoding by the inter prediction unit (620).
[0118] The present invention can also be utilized by or incorporated into an enhanced compression model (ECM), which is an implementation of a next-generation video codec currently being developed by the Joint Video Experts Team (JVET), an international standardization expert organization. The enhanced compression model can include an enhanced intra prediction coding method, an enhanced inter prediction coding method, an enhanced transform and transform coefficient coding method, an enhanced adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for improving image quality, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique.
[0119]
[0120] Composition of the present invention
[0121] A global motion vector is one of the technical means that can be used for the above-described inter prediction and / or intra prediction. The global motion vector may, for example, mean a motion vector that represents the overall motion that appears throughout a single picture. For example, when a camera that captures a video moves to the left, most objects captured in the video move to the right, and in this case, it can be seen that the pictures acquired while the camera moves to the left have a global motion vector that points to the right. The global motion vector is defined in a wide range, such as a picture unit, so that it can offset the amount of information of individual motion vectors, i.e., local motion vectors, for smaller encoding / decoding units (such as blocks), and as a result, it can lead to improved encoding efficiency.
[0122] The present invention proposes a method for improving encoding efficiency for special global motion occurring in a video by configuring the global motion vector radially, and also for describing such radial global motion vector in an encoding bitstream.
[0123] FIG. 7 is a flowchart illustrating a video encoding process according to an embodiment of the present invention. The video encoding method according to an embodiment of the present invention may include a step of calculating a radial global motion vector for a picture (image) included in the video to be encoded (S710), a step of calculating a local motion vector based on the radial global motion vector (S720), a step of generating a residual image by performing prediction coding on the picture based on the calculated motion vectors (S730), and a step of encoding information related to the radial global motion vector, information related to the local motion vector, and information related to the residual image in an encoding bitstream (S740).
[0124] According to one embodiment of the present invention, the step (S710) of calculating a radial global motion vector for a picture (image) included in the video to be encoded may mean a step of calculating a global motion vector by comparing at least one previous picture with a current picture. The method for calculating the radial global motion vector is not limited, but may be calculated by, for example, a method using SAD (sum of absolute difference), a method using optical flow, etc. In addition, it is obvious that various conventional or newly provided calculation methods that can be used for calculating a motion vector for a person skilled in the art can be applied to the practice of the present invention. In the present invention, the calculation method may be configured to calculate a radial global motion vector in particular. Fig. 8 is an exemplary diagram of a radial global motion vector according to one embodiment of the present invention. As illustrated in FIG. 8, a radial global motion vector may refer to a vector field (820) configured radially around a predetermined center point for a specific picture (810). The vector field may be expressed as a plurality of independent vectors, or may refer to a field that can derive a vector corresponding to a position at any position on the picture by being defined by a mathematical method. According to a general understanding, the radial motion vector may be understood to indicate a change in the picture corresponding to zooming out when it is inward, and a change in the picture corresponding to zooming in when it is outward, but is not necessarily limited thereto. Although FIG. 8 illustrates a case where the radial global motion vector is inward, the radial global motion vector may also be defined as outward, regardless of the illustrated content.
[0125] According to one embodiment of the present invention, the step of calculating the radial global motion vector may be performed with reference to information related to the direction and intensity of acceleration acquired from an accelerometer associated with the source of the video. For example, if there is an accelerometer installed in connection with the movement of a photographing device that captures the video to be encoded, the calculation of the radial global motion vector may be performed based on the acceleration measured by the accelerometer. For example, if the forward movement of the photographing device is measured through the accelerometer, the movement of an object substantially corresponding to zoom-in may be recorded in the video, and thus the corresponding radial global motion vector may be inferred from the measured acceleration. Various features of the radial global motion vector may be derived depending on the direction and intensity of the acceleration, and the features may include size, direction, angle, center point, ellipticity, eccentricity, etc.
[0126] According to one embodiment of the present invention, the step (S720) of calculating a local motion vector based on the radial global motion vector may mean a step of calculating a local motion vector by comparing at least one previous picture with a current picture. The calculation of the local motion vector may be similar to or identical to the calculation process of the motion vector used in the predictive encoding process of the general digital video codec described above. However, the local motion vector may be calculated as a value that cancels out a vector represented by the global motion vector. In addition, depending on the embodiment and the local content of the picture to be encoded, some or all of the local motion vectors may be completely canceled and / or replaced by the global motion vector and thus omitted.
[0127] According to one embodiment of the present invention, the step (S730) of generating a residual image by performing prediction coding on the picture based on the calculated motion vectors may be similar to or identical to the prediction coding process of the general digital video codec described above. That is, various algorithms may be utilized to generate a residual image by subtracting a value of an image of an actual picture from a value of a predicted image predicted by a motion vector. According to one embodiment of the present invention, the prediction coding may be performed by inter prediction (prediction between screens). However, it is not necessarily limited thereto, and even if the prediction coding includes intra prediction (prediction within a screen), there is no problem in implementing the present invention.
[0128] According to one embodiment of the present invention, the step (S740) of encoding information related to the radial global motion vector, information related to the local motion vector, and information related to the residual image into the encoded bitstream may be included in the step of creating a bitstream in the encoder of the general digital video codec described above. A detailed method of encoding information related to the radial global motion vector into the bitstream will be described later. The present invention can be equally implemented using any method of a digital video encoder that is known in the art or can be newly provided for encoding information related to the local motion vector and information related to the residual image into the bitstream.
[0129] According to an embodiment of the present invention, the step (S720) of calculating the local motion vector may be omitted. In this case, the step (S730) of generating the residual image may be configured to be performed without referring to any motion vector other than the radial global motion vector. In this case, the encoding step (S740) may be configured to encode the bitstream while omitting information related to the local motion vector. In addition, in this case, a signal indicating that information related to the local motion vector has been omitted may be encoded in the bitstream.
[0130] According to one embodiment of the present invention, the radial global motion vector may be defined by at least one of first information related to its magnitude, second information related to its aspect ratio, and third information related to its centric point.
[0131] FIG. 9 is a conceptual diagram of a method for representing first information related to the intensity of a radial global motion vector according to one embodiment of the present invention.
[0132] According to one embodiment of the present invention, the first information may be defined in a manner indicating the magnitude (935) of the basis vector (930) constituting the radial global motion vector (920). According to an embodiment of the present invention, the radial global motion vector (920) may be defined as a vector field obtained by repeatedly rearranging the basis vector (930) radially according to a predetermined mathematical rule. In this case, when the magnitude of the basis vector (930) is defined, by changing the magnitude of the entire radial global motion vector (920), the magnitude of the motion vector on the prediction encoding calculated at all positions at all angles by the radial global motion vector (920) may be changed.
[0133] According to one embodiment of the present invention, the size (935) of the basic vector (930) may be defined by an absolute unit system. For example, if the picture (910) for which the radial global motion vector (920) is calculated is an image composed of pixel units, the size (935) of the basic vector (930) may be defined in units of a predetermined pixel.
[0134] According to another embodiment of the present invention, the size (935) of the base vector (930) may be defined by a relative unit system. For example, it may be defined in proportion to at least one of the horizontal and vertical resolutions of the picture (910). For another example, if the base size of the base vector (930) is previously specified by the encoder, it may be defined as a fixed multiple of the base size.
[0135] FIG. 10 is a conceptual diagram of a method for representing second information related to the aspect ratio of a radial global motion vector according to one embodiment of the present invention.
[0136] According to one embodiment of the present invention, the second information may be defined in a manner indicating an angle (1040) of a base vector (1030) constituting the radial global motion vector (1020). According to an embodiment of the present invention, when the reference angle of the base vector (1030) is determined, the overall aspect ratio of the radial global motion vector (1020) may be derived according to the degree to which the angle of the base vector changes from the reference angle. For example, when the reference angle is 45 degrees, it may be calculated that the ratio in the horizontal direction becomes longer as it approaches 0 degrees, and the ratio in the vertical direction becomes longer as it approaches 90 degrees.
[0137] According to one embodiment of the present invention, the second information may be defined as a ratio of one direction of the horizontal (1050) or vertical (1060) direction of the radial global motion vector (1020) to the other direction. Depending on the embodiment, the ratio may be defined by specifying the horizontal (1050) length and the vertical (1060) length, respectively, or may be defined as a proportional value (e.g., 1:1, 2:1, 1:2, 3:4, etc.), or may be defined as a ratio based on one of the horizontal or vertical directions (e.g., 1.0, 0.5, 2.0, etc.), and various other methods for indicating a predetermined aspect ratio may be applied and used.
[0138] FIG. 11 is a conceptual diagram of a method for representing third information related to the center point of a radial global motion vector according to one embodiment of the present invention.
[0139] Referring to FIG. 11a, according to one embodiment of the present invention, the third information may include information about the center point (1130) of the radial global motion vector (1120). According to one embodiment, the information about the center point (1130) may be defined as the distance that the center point (1130) is spaced apart from the reference point of the corresponding picture (1110). Referring to FIG. 11b, in one embodiment of the present invention, the reference point may be a coordinate system reference point of the picture (1110), for example, the upper left corner (1140), and in this case, the center point (1130) may be defined by the distance (1145) spaced apart from the upper left corner (1140). In another embodiment of the present invention, the reference point may be the center point (1150) of the picture (1110), in which case the center point (1130) of the radial global motion vector may be defined by a distance (1155) from the center point (1150) of the picture.
[0140] FIG. 12 is a conceptual diagram of a first center point reference defining the center point of a radial global motion vector according to one embodiment of the present invention.
[0141] According to one embodiment of the present invention, the center point (1230) of the radial global motion vector may refer to a first center point determined through translation (1260) of the radial global motion vector. In this case, the movement of the center point (1230) may be interpreted as a movement of the entire vector in which the overall shape of the radial global motion vector remains the same but only its center moves, and thus may be reflected in the calculation and / or prediction encoding process of the corresponding local motion vector.
[0142] FIG. 13 is a conceptual diagram of a second center point reference defining the center point of a radial global motion vector according to one embodiment of the present invention.
[0143] According to one embodiment of the present invention, the center point (1330) of the radial global motion vector may refer to a second center point determined through the eccentricity (1360) of the radial global motion vector. In this case, the movement of the center point (1230) may be interpreted as a movement of the position of the vanishing point where the radial global motion vector converges, and accordingly, may be reflected in the calculation and / or prediction encoding process of the corresponding local motion vector.
[0144] According to one embodiment of the present invention, the center point of the first center point reference and the center point of the second center point reference may be defined in parallel. That is, a single radial global motion vector may be defined to have both a center point of movement and a center point based on eccentricity. In this case, the third information may include information regarding the center points of both references.
[0145] According to one embodiment of the present invention, the first information, the second information, and the third information can be simultaneously defined by a single vector that can determine a size, an angle, and a position. Referring to FIGS. 9, 10, 11A, and 11B collectively, if a single basic vector (e.g., 930, 1030) is defined to have a magnitude (935), an angle (1040), and a distance (1145, 1155) from a reference point (e.g., 1140, 1150) of the basic vector, it can be seen that information corresponding to the first information, the second information, and the third information are all derived from the definition of the basic vector, thereby allowing the radial global motion vector to be derived.
[0146] According to one embodiment of the present invention, the first information, the second information, and the third information can be simultaneously defined by a single rectangle capable of determining a horizontal length, a vertical length, and a position. Referring again to FIGS. 9, 10, 11A, and 11B, the radial global motion vector can be expressed in the form of a rectangle having a predetermined horizontal length (1050) and a vertical length (1060) and having a position spaced a predetermined distance (1145, 1155) from a reference point (e.g., 1140, 1150). Therefore, it can be seen that information corresponding to the first information, the second information, and the third information can all be derived from the definition of a single rectangle, thereby deriving the radial global motion vector.
[0147] According to one embodiment of the present invention, the first information, the second information, and the third information may be simultaneously defined by a single ellipse whose diameter, ellipticity, and position can be determined. Similar to the definition as a rectangle, the radial global motion vector may be expressed in the form of an ellipse with a constant diameter, an ellipticity defined according to an aspect ratio, and a center point spaced a constant distance from a reference point. Therefore, it can be seen that information corresponding to the first information, the second information, and the third information can all be derived from the definition of a single ellipse, thereby deriving the radial global motion vector. In this case, the definition of the ellipse may be defined by the diameter, ellipticity, and position as described above, or may be defined in the form of coefficients of an ellipse formula that derives an ellipse with a predetermined diameter, ellipticity, and position.
[0148] At least one of the first information, the second information, and the third information may be recorded in an encoded bitstream generated by an encoder. According to various embodiments of the present invention, the at least one piece of information may be located in either a header area or a data area of an encoded video bitstream. Non-limiting examples of the header area include a NAL unit header, a picture (frame) header, and a slice header. Non-limiting examples of the data area include picture (frame) data, slice data, macroblock data, block data, and coding tree unit data.
[0149] According to one embodiment of the present invention, at least one of the first information, the second information, and the third information can be recorded in the encoded bitstream as a fixed-length signal of at least 1 bit.
[0150] According to one embodiment of the present invention, at least one of the first information, the second information, and the third information may be recorded as at least one entropy-encoded bit string in the encoded bit stream. According to an embodiment, the entropy-encoded bit string may be encoded by various entropy coding and / or variable length coding methods including Huffman coding, exponential Golomb coding, run length coding, variable length coding, context-adaptive variable length coding, and context-adaptive binary arithmetic coding (CABAC). Depending on the coding method, the at least one piece of information may be recorded by overlapping one coding symbol. Depending on the coding method, various coding methods including the exemplary coding methods described above may be applied in duplicate.
[0151] As described above, it is self-evident that the radial global motion vector encoded in the bitstream can be decoded by the reverse order of the calculation and encoding process, as is widely known.
[0152] Fig. 14 is a flowchart illustrating a decoding process of an encoded video bitstream according to an embodiment of the present invention. The method for decoding an encoded video bitstream according to an embodiment of the present invention may include a step of decoding the bitstream to read information related to a radial global motion vector from the bitstream (S1410), a step of executing prediction decoding based on the radial global motion vector to generate a predicted image (S1420), a step of decoding the bitstream to read information related to a residual image from the bitstream (S1430), and a step of combining the predicted image and the residual image to generate a decoded image (S1440).
[0153] According to one embodiment of the present invention, the method for decoding the bitstream further includes a step of reading information related to at least one local motion vector from the bitstream, and the step (S1420) of generating the predicted image can be performed with reference to the radial global motion vector and the local motion vector together.
[0154]
[0155] Encoder and decoder
[0156] The encoding method according to the present invention described above can be implemented through an encoder as a device. The encoder as a device can be implemented in a form that maintains the conventional encoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of encoder structure that can function as a video encoder should be considered an encoder established by the present invention as long as it implements the technical idea of the present invention.
[0157] According to one embodiment of the present invention, a video encoding device configured to output a bitstream may be configured to include at least a processor, a memory, a global motion calculation unit, a prediction encoding unit, and a bitstream generation unit. At this time, the global motion calculation unit is configured to calculate a radial global motion vector for an image included in the video, the prediction encoding unit is configured to perform prediction encoding on the image based on the radial global motion vector to generate a residual image, and the bitstream generation unit may be configured to encode information related to the radial global motion vector and information related to the residual image in an output bitstream. It can be easily understood that the operation process of each of the above-described functional units can be implemented with reference to what has been described above as an encoding method according to the present invention.
[0158] The method for decoding an encoding result according to the present invention described above can be implemented through a decoder as a device. The decoder as a device can be implemented in a form that maintains the conventional decoder structure exemplified through FIGS. 1 to 6 above or applies a predetermined change thereto. However, the form of implementation is not necessarily limited to that exemplified, and any form of decoder structure that can function as a video decoder can be considered a decoder established according to the present invention as long as it implements the technical idea of the present invention.
[0159] According to one embodiment of the present invention, a decoder device configured to decode a bitstream may be configured to include a processor, a memory, a bitstream parser, and a predictive decoding unit. At this time, the bitstream parser is configured to read information related to a radial global motion vector and information related to a residual image from the bitstream, and the predictive decoding unit may be configured to perform predictive decoding based on the radial global motion vector to generate a predicted image, and to combine the predicted image and the residual image to generate a decoded image. It can be easily understood that the operation process of each of the above-described functional units can be implemented with reference to a method of decoding a bitstream encoded by the encoding method according to the present invention described above by applying the same or a reverse method.
[0160] The processor that may be included in the encoder and / or decoder described herein may mean one or more general purpose computers or special purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding.
[0161] Even if the processor is expressed singularly for ease of understanding, those skilled in the art will appreciate that the processor may include multiple processing elements and / or multiple types of processing elements. For example, a device according to one embodiment of the present invention may include multiple processors or one processor and one controller as the processor. Furthermore, the processor may be implemented using various processing configurations, such as a parallel processor or a multi-core processor.
[0162] The processor may be configured to execute an operating system (OS) and one or more software programs running on the operating system. Furthermore, the processor may access, store, manipulate, process, and generate data in response to the execution of the software.
[0163] The software may include a computer program, code, instructions, or a combination of one or more of these, and may be configured to control the processor to perform a desired operation and to issue commands to the processor, either independently or collectively. The software may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processor or for providing commands or data to the processor. The software may also be distributed over networked computer systems, and stored or executed in a distributed manner.
[0164] The software may also be implemented in the form of program commands that can be executed through various computer means and recorded or stored in the memory. The memory may be a computer-readable recording medium, and program commands, data files, data structures, etc. may be recorded singly or in combination in the computer-readable recording medium. The program commands stored in the memory may be based on a command system specifically designed and configured for the embodiment of the present invention, or may follow a command system known and available to those skilled in the art of computer software, such as a command system exemplified by the assembly language, C, C++, Java, Python, etc. It should be understood that the command system and the program commands therefrom include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by the device and / or the processor according to an embodiment of the present invention using an interpreter or the like.
[0165] The computer-readable recording medium constituting the device according to one embodiment of the present invention, including the memory described in this specification, may include a temporary or volatile recording medium that is maintained only while the processor is operating, such as a processor cache, a RAM, a flash memory, or a relatively non-volatile or long-term recording medium, such as a magnetic media such as a hard disk, a floppy disk, and a magnetic tape, an optical media such as a CD-ROM, a DVD, a magneto-optical media such as a floptical disk, or a solid state memory, or may include a read-only recording medium, such as a ROM arranged on hardware, and further, the hardware itself configured to perform an operation equivalent to a series of program commands by a hard-wired structure by circuit wiring, and each step for performing the operation for implementing the embodiment of the present invention can be viewed as being recorded by the connection and arrangement of the hardware components, and thus the connection and arrangement method is the memory and It is obvious to a person skilled in the art that they can be considered equivalent.
[0166] The embodiments described above with respect to the processor and the memory are not mutually exclusive and may be selected or combined and implemented as needed. For example, a single hardware device may be configured to operate as a module composed of one or more of the software to perform the operations of an embodiment of the present invention, and vice versa. As another example, in the present specification, all or part of the operations assigned to a certain functional unit may be implemented by one or more of the software stored in a device according to an embodiment of the present invention (preferably, in any one of the recording media belonging to the category of the memory) and configured to be executed by the processor. In such a case, such a functional unit may be referred to as a functional unit "included" in the processor.
[0167]
[0168] Although the present invention has been described with reference to drawings and embodiments, as already mentioned above, it does not mean that the scope of protection of the present invention is limited to the drawings or embodiments presented above, and it will be understood that a person skilled in the relevant technical field can modify and change the present invention in various ways within a scope that does not depart from the spirit and scope of the present invention described in the claims of the present invention patent.
Claims
1. In a method for decoding an encoded video bitstream, A step of reading information related to a radial global motion vector from the above bitstream; A step of generating a predicted image by executing prediction decoding based on the above radial global motion vector; A step of reading information related to a residual image from the bitstream; and A method comprising: a step of generating a decoded image by combining the predicted image and the residual image.
2. In paragraph 1, Information related to the above radial global motion vector is, A method characterized in that it includes at least one of first information related to the magnitude of the radial global motion vector, second information related to the aspect ratio of the radial global motion vector, and third information related to the centric point of the radial global motion vector.
3. In paragraph 2, At least one of the first information, the second information, and the third information is, A method characterized in that it is read from at least one entropy encoded bit string included in the bitstream.
4. In paragraph 3, A method characterized in that the above entropy encoded bit string is interpreted by an entropy decoding algorithm corresponding to context-adaptive binary arithmetic coding (CABAC).
5. In paragraph 2, The above first information is, A method characterized in that it indicates the size of a basis vector constituting the above radial global motion vector.
6. In paragraph 5, The magnitude of the above basic vector is, A method characterized in that it is defined in proportion to at least one of the horizontal and vertical resolutions of the predicted image.
7. In paragraph 2, The above second information is, A method characterized in that it represents the angle of the basic vector constituting the above radial global motion vector.
8. In paragraph 2, The above second information is, A method characterized in that it is defined as a ratio of one direction to the other direction of the horizontal or vertical direction of the above radial global motion vector.
9. In paragraph 2, The above third information is, A method characterized in that the center point of the above radial global motion vector represents a distance from a reference point of the above predicted image.
10. In paragraph 2, The above third information is, A method characterized in that it includes at least one of information about a first center point determined through translation of the radial global motion vector, and information about a second center point determined through eccentricity of the radial global motion vector.
11. In paragraph 2, The above first information, the above second information, and the above third information, A method characterized in that it is simultaneously defined by a single vector capable of determining size, angle, and position.
12. In paragraph 2, The above first information, the above second information, and the above third information, A method characterized in that it is simultaneously defined by a single rectangle capable of determining a length, a width, and a position.
13. In paragraph 2, The above first information, the above second information, and the above third information, A method characterized in that it is simultaneously defined by an ellipse having a diameter, ellipticity, and position that can be determined.
14. In paragraph 1, The step of generating the above predicted image is: A method characterized in that it is performed without reference to any other motion vector other than the above radial global motion vector.
15. In paragraph 1, Further comprising a step of reading information related to at least one local motion vector from the bitstream; A method, characterized in that the step of generating the above predicted image is performed by referring to the radial global motion vector and the local motion vector together.
16. In paragraph 1, The step of generating the above predicted image is: A method characterized in that it is performed by inter-frame prediction.
17. In a method of encoding a video, A step of calculating a radial global motion vector for an image included in the above video; A step of generating a residual image by performing prediction coding on the image based on the radial global motion vector; and A method comprising: encoding information related to the radial global motion vector and information related to the residual image in a bitstream.
18. In paragraph 17, The step of calculating the above radial global motion vector is: A method characterized in that it is performed by referring to information related to the direction and intensity of acceleration obtained from an accelerometer associated with the source of the above video.
19. In a decryption device configured to decrypt a bitstream, processor; memory; bitstream parser; and Includes a predictive decoding unit, The bitstream parser is configured to read information related to a radial global motion vector and information related to a residual image from the bitstream, A device wherein the above prediction decoding unit is configured to generate a prediction image by executing prediction decoding based on the radial global motion vector, and to generate a decoded image by combining the prediction image and the residual image.
20. In a video encoding device configured to output a bitstream, processor; memory; Global motion calculation unit; Predictive encoding unit; and including a bitstream generation unit; The above global motion calculation unit is configured to calculate a radial global motion vector for an image included in the video, The above prediction encoding unit is configured to perform prediction encoding on the image based on the radial global motion vector to generate a residual image, A device wherein the bitstream generation unit is configured to encode information related to the radial global motion vector and information related to the residual image into an output bitstream.
Citation Information
Patent Citations
Method and apparatus for adaptive coding of motion information
JP2020188483A
Mothod of estimating motion vector using global motion vector, apparatus, encoder, decoder and decoding method
KR101356735B1
Method and apparatus for estimating motion homogeneity for video quality assessment
KR1020150052049A
Outdoor terminal of hydrogen stations, methods for providing operational status information of hydrogen stations and program stored in recording medium
KR102613383B1
Global motion modeling for automotive image data
US20240078684A1