Motion refinement using bilateral matching and twisted refinement models
By using bilateral matching and a warped motion model for motion refinement during encoding and decoding, the problem of insufficient efficiency and accuracy in non-translational motion processing in existing technologies is solved, achieving more efficient video code processing.
Patent Information
- Application Number
- CN202480026376.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-26
- Filing Date
- 2024-04-03
- Publication Date
- 2026-01-27
AI Technical Summary
Existing block-based hybrid video codec processing techniques or codecs are inefficient and inaccurate when processing non-translational motion, and the signal transmission is not efficient enough in terms of distorted motion model parameters.
Motion refinement is achieved by employing bilateral matching and a twisted motion model. By using a four-parameter scaling refinement model, a three-parameter scaling refinement model, and a four-parameter rotation refinement model during encoding and decoding, motion vectors are refined without explicit indication, thereby improving the accuracy and efficiency of motion representation.
It improves the efficiency and accuracy of video code processing technology, reduces the signal transmission requirements for motion vector parameters, and optimizes the encoding and decoding process.
Smart Images

Figure CN121420554A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority and benefit to U.S. Provisional Patent Application Serial No. 63 / 456,860, filed April 4, 2023; U.S. Provisional Patent Application Serial No. 63 / 526,099, filed July 11, 2023; and U.S. Provisional Patent Application Serial No. 63 / 529,085, filed July 26, 2023, the entire disclosure of which is incorporated herein by reference. Background Technology
[0002] Digital images and videos can be used, for example, over the Internet for remote business meetings via video conferencing, high-definition video entertainment, video advertising, or the sharing of user-generated content. Due to the large amounts of data involved in the transmission and processing of image and video data, high-performance compression can facilitate transmission and storage. Therefore, providing high-resolution images and videos for transmission over communication channels with limited bandwidth is advantageous. Summary of the Invention
[0003] This application relates to encoding and decoding image data, video stream data, or both, for transmission, storage, or both. This document discloses aspects of systems, methods, and apparatuses for encoding and decoding, including motion thinning using bilateral matching and one or more distortion-based thinning models.
[0004] The changes in these and other aspects will be described in more detail below.
[0005] One aspect is a method for decoding, which includes motion refinement using bilateral matching and a warped motion model. Decoding including motion refinement using bilateral matching and a warped motion model involves generating reconstructed block data by decoding the current block of the current frame from an encoded bitstream. Decoding the current block may include obtaining a refined prediction block for decoding the current block using bilateral matching. Obtaining the refined prediction block may include obtaining a refinement model from available warped refinement models, including four-parameter scaling refinement models, three-parameter scaling refinement models, and four-parameter rotation refinement models. Obtaining the refined prediction block may include obtaining refined motion vectors using the warped refinement model and previously obtained reference frame data when no data explicitly indicating refined motion vectors exists in the encoded bitstream. Obtaining the refined prediction block may include generating refined prediction block data using the refined motion vectors. Decoding the current block may include: generating reconstructed block data using refined prediction block data; including the reconstructed block data in the reconstructed frame data of the current frame; and outputting the reconstructed frame data.
[0006] One aspect is a method for encoding that includes motion thinning using bilateral matching and a warped motion model. Encoding that includes motion thinning using bilateral matching and a warped motion model involves generating encoded block data by encoding the current block of the current frame. Encoding the current block may include obtaining a thinned prediction block for encoding the current block using bilateral matching. Obtaining the thinned prediction block may include obtaining a thinning model from available warped thinning models, including four-parameter scaling thinning models, three-parameter scaling thinning models, and four-parameter rotation thinning models. Obtaining the thinned prediction block may include obtaining thinned motion vectors using the warped thinning model and previously obtained reference frame data. Encoding that includes motion thinning using bilateral matching and a warped motion model omits including data explicitly indicating the thinned motion vectors in the encoded bitstream. Obtaining the thinned prediction block may include generating thinned prediction block data using the thinned motion vectors. The encoding, which includes motion refinement using bilateral matching and a twisted motion model, involves obtaining a residual by subtracting the refined prediction block data from the current block and including the residual data in the encoded bitstream.
[0007] One aspect is a method for decoding, which includes motion refinement using bilateral matching and a warped motion model. Decoding including motion refinement using bilateral matching and a warped motion model involves generating reconstructed block data by decoding the current block of the current frame from an encoded bitstream. Decoding the current block may include obtaining refined prediction blocks for decoding the current block using bilateral matching. Obtaining refined prediction blocks may include: obtaining refined motion vectors using a rotation and scaling refinement model and previously obtained reference frame data when no data explicitly indicating refined motion vectors exists in the encoded bitstream. Obtaining refined prediction blocks may include generating refined prediction block data using the refined motion vectors. Decoding the current block may include: generating reconstructed block data using the refined prediction block data; including the reconstructed block data in the reconstructed frame data of the current frame; and outputting the reconstructed frame data.
[0008] One aspect is a method for encoding that includes motion refinement using bilateral matching and a warped motion model. Encoding that includes motion refinement using bilateral matching and a warped motion model includes generating encoded block data by encoding the current block of the current frame. Encoding the current block may include obtaining a refined prediction block for encoding the current block using bilateral matching. Obtaining the refined prediction block may include obtaining refined motion vectors using a rotation and scaling refinement model and previously obtained reference frame data. Encoding that includes motion refinement using bilateral matching and a warped motion model omits including data explicitly indicating the refined motion vectors in the encoded bitstream. Obtaining the refined prediction block may include generating refined prediction block data using the refined motion vectors. Encoding that includes motion refinement using bilateral matching and a warped motion model includes obtaining a residual by subtracting the refined prediction block data from the current block, and including the residual data in the encoded bitstream.
[0009] One aspect is a method for decoding, which includes motion refinement using bilateral matching and a warped motion model. Decoding including motion refinement using bilateral matching and a warped motion model involves generating reconstructed block data by decoding the current block of the current frame from an encoded bitstream. Decoding the current block may include obtaining a refined prediction block for decoding the current block using bilateral matching. Obtaining the refined prediction block may include obtaining a refinement model from available warped refinement models, including four-parameter scaling refinement models, three-parameter scaling refinement models, four-parameter rotation refinement models, and four-parameter rotation and scaling models. Obtaining the refined prediction block may include, in the absence of data in the encoded bitstream explicitly indicating refined motion vectors, obtaining these refined motion vectors using the warped refinement model and previously obtained reference frame data, wherein obtaining these refined motion vectors includes obtaining a combination of the following as these refined motion vectors: block-based warped motion parameters obtained using the warped refinement model and sub-block-based translational motion parameters. Obtaining a refined prediction block may include generating refined prediction block data using refined motion vectors. Decoding the current block may include: generating reconstructed block data using the refined prediction block data; including the reconstructed block data in the reconstructed frame data of the current frame; and outputting the reconstructed frame data.
[0010] One aspect is a method for encoding that includes motion refinement using bilateral matching and a warped motion model. Encoding that includes motion refinement using bilateral matching and a warped motion model involves generating encoded block data by encoding the current block of the current frame. Encoding the current block may include obtaining a refined prediction block for encoding the current block using bilateral matching. Obtaining the refined prediction block may include obtaining a refinement model from available warped refinement models, including four-parameter scaling refinement models, three-parameter scaling refinement models, four-parameter rotation refinement models, and four-parameter rotation and scaling models. Obtaining the refined prediction block may include obtaining refined motion vectors using the warped refinement model and previously obtained reference frame data. Encoding that includes motion refinement using bilateral matching and a warped motion model omits including data explicitly indicating the refined motion vectors in the encoded bitstream, wherein obtaining these refined motion vectors includes obtaining a combination of the following as these refined motion vectors: block-based warped motion parameters obtained using the warped refinement model and sub-block-based translational motion parameters. Obtaining a refined prediction block may include generating refined prediction block data using refined motion vectors. Encoding for motion refinement, including using bilateral matching and a twisted motion model, includes obtaining a residual by subtracting the refined prediction block data from the current block, and including the residual data in the encoded bitstream. Attached Figure Description
[0011] The description herein is taken with reference to the accompanying drawings, wherein, unless otherwise expressly indicated or clearly understood from the context, the same reference numerals refer to the same parts in all views.
[0012] Figure 1 This is a diagram of a computing device according to an implementation of the present disclosure.
[0013] Figure 2 It is a diagram of a computing and communication system according to an implementation of this disclosure.
[0014] Figure 3 It is a diagram of a video stream used in encoding and decoding according to an implementation of this disclosure.
[0015] Figure 4 This is a block diagram of an encoder according to the implementation of this disclosure.
[0016] Figure 5 This is a block diagram of a decoder implemented according to this disclosure.
[0017] Figure 6 It is a block diagram representing a portion of a frame according to an implementation of this disclosure.
[0018] Figure 7 This is a block diagram illustrating an example of side motion vector refinement for a translational motion decoder based on bilateral matching.
[0019] Figure 8 This is a block diagram illustrating an example of using optical flow for prediction refinement.
[0020] Figure 9 This is a flowchart illustrating an example of a process for obtaining a refined prediction block using bilateral matching with one or more twisted and refined models. Detailed Implementation
[0021] Image and video compression schemes may include breaking down an image or frame into smaller parts, such as blocks, and generating an output bitstream using techniques to minimize the bandwidth consumption of information included for each block in the output. In some implementations, the information included for each block in the output can be limited by reducing spatial redundancy, reducing temporal redundancy, or a combination thereof. For example, temporal or spatial redundancy can be reduced by predicting a frame, or a portion thereof, based on information available from both the encoder and decoder, and by including information in the encoded bitstream representing the difference or residual between the predicted frame and the original frame. The residual information can be further compressed by transforming the residual information into transform coefficients (e.g., energy compression), quantizing the transform coefficients, and entropy-coding the quantized transform coefficients. The encoded bitstream may include other code processing information, such as motion information, which may include transmitting differential information based on predictions of the encoded information, which may be entropy-coded to further reduce the corresponding bandwidth consumption. The encoded bitstream can be decoded to reconstruct blocks and the source image from the limited information. In some implementations, the accuracy, efficiency, or both of using inter-frame prediction or intra-frame prediction for block code processing may be limited.
[0022] Some block-based hybrid video codecs may be limited by using translational motion models to reduce temporal redundancy, which may not efficiently or accurately represent non-translational motion. Some block-based hybrid video codecs may include warped motion video codecs (including warped motion compensation), which, for non-translational motion, can improve efficiency, accuracy, or both, compared to block-based hybrid video codecs limited by using translational motion models to reduce temporal redundancy. For example, some block-based hybrid video codecs may include warped motion video codecs using a global warped motion model, a local warped motion model, or both.
[0023] Some block-based hybrid video codec processing techniques or codecs (including warped motion video codec processing) may inefficiently signal warped motion model parameters. For example, some block-based hybrid video codec processing techniques or codecs (including warped motion video codec processing) may signal warped motion model parameters, such as global affine motion parameters, on a frame-by-frame or frame-by-frame basis. For example, some block-based hybrid video codec processing techniques or codecs (including warped motion video codec processing) may omit signaling warped motion model parameters, such as warped motion model parameters of local warped motion models.
[0024] This article describes encoding and decoding that includes motion refinement using bilateral matching and a twisted motion model to improve video codec processing techniques or codecs by refining twisted motion parameters without explicit signaling.
[0025] Figure 1 This is a diagram of a computing device 100 according to an implementation of the present disclosure. The computing device 100 shown includes a memory 110, a processor 120, a user interface (UI) 130, an electronic communication unit 140, a sensor 150, a power supply 160, and a bus 170. As used herein, the term "computing device" includes any unit or combination of units capable of performing any method disclosed herein or any one or more portions thereof.
[0026] The computing device 100 may be a fixed computing device, such as a personal computer (PC), server, workstation, minicomputer, or mainframe computer; or a mobile computing device, such as a mobile phone, personal digital assistant (PDA), laptop computer, or tablet PC. Although shown as a single unit, any one or more components of the computing device 100 may be integrated into any number of separate physical units. For example, the user interface 130 and processor 120 may be integrated into a first physical unit, and the memory 110 may be integrated into a second physical unit.
[0027] Memory 110 may include any non-transitory computer-usable or computer-readable medium, such as any tangible means that can, for example, contain, store, transmit, or transfer data 112, instructions 114, operating system 116, or any information associated therewith, for use by or in conjunction with other components of computing device 100. Non-transitory computer-usable or computer-readable media may be, for example, solid-state drives, memory cards, removable media, read-only memory (ROM), random access memory (RAM), any type of disk (including hard disks, floppy disks, optical disks, magnetic cards, or optical cards), application-specific integrated circuits (ASICs), or any type of non-transitory medium suitable for storing electronic information, or any combination thereof.
[0028] Although shown as a single unit, memory 110 may include multiple physical units, such as one or more main memory units (e.g., random access memory units), one or more secondary data storage units (e.g., disks), or combinations thereof. For example, data 112, or a portion of that data, instructions 114, or a portion of those instructions, or both, may be stored in secondary storage units and may be processed in combination with the corresponding data 112, the corresponding instructions 114, or both, loaded or otherwise transferred to the main storage units. In some implementations, memory 110, or a portion thereof, may be removable memory.
[0029] Data 112 may include information such as input audio data, encoded audio data, and decoded audio data. Instructions 114 may include instructions, such as code, for performing any method disclosed herein, or any one or more portions thereof. Instructions 114 may be implemented in hardware, software, or any combination thereof. For example, instructions 114 may be implemented as information stored in memory 110, such as a computer program, which may be executed by processor 120 to perform any of the corresponding methods, algorithms, aspects, or combinations thereof as described herein.
[0030] Although shown as being included in memory 110, in some implementations, instructions 114, or portions thereof, may be implemented as a dedicated processor or circuit, which may include dedicated hardware for performing any of the methods, algorithms, aspects, or combinations thereof described herein. Portions of instructions 114 may be distributed across multiple processors on the same or different machines, or across a network such as a local area network, a wide area network, the Internet, or combinations thereof.
[0031] Processor 120 may include any device or system capable of manipulating or processing existing or future-developed digital signals or other electronic information, including optical processors, quantum processors, molecular processors, or combinations thereof. For example, processor 120 may include a dedicated processor, a central processing unit (CPU), a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic array, a programmable logic controller, microcode, firmware, any type of integrated circuit (IC), a state machine, or any combination thereof. As used herein, the term "processor" includes a single processor or multiple processors.
[0032] User interface 130 may include any unit capable of interacting with a user, such as a virtual or physical keypad, touchpad, display, touch display, speaker, microphone, camera, sensor, or any combination thereof. For example, user interface 130 may be an audiovisual display device, and computing device 100 may use user interface 130 to present audio (such as decoded audio) in conjunction with displayed video (such as decoded video). Although shown as a single unit, user interface 130 may include one or more physical units. For example, user interface 130 may include an audio interface for audio communication with the user and a touch display for visual and touch-based communication with the user.
[0033] Electronic communication unit 140 can transmit, receive, or transmit and receive signals via wired or wireless electronic communication medium 180 (such as radio frequency (RF) communication medium, ultraviolet (UV) communication medium, visible light communication medium, fiber optic communication medium, wired communication medium, or combinations thereof). For example, as shown, electronic communication unit 140 is operatively connected to electronic communication interface 142 (such as an antenna) configured to communicate via wireless signals.
[0034] Although the electronic communication interface 142 is in Figure 1 While shown as a wireless antenna, the electronic communication interface 142 can be a wireless antenna as shown, a wired communication port (such as an Ethernet port), an infrared port, a serial port, or any other wired or wireless unit capable of interfacing with wired or wireless electronic communication medium 180. Although Figure 1 A single electronic communication unit 140 and a single electronic communication interface 142 are shown, but any number of electronic communication units and any number of electronic communication interfaces can be used.
[0035] Sensor 150 may include, for example, an audio sensing device, a visible light sensing device, a motion sensing device, or a combination thereof. For example, sensor 150 may include a sound sensing device, such as a microphone, or any other existing or later-developed sound sensing device capable of sensing sounds near computing device 100, such as speech or other utterances made by a user operating computing device 100. In another example, sensor 150 may include a camera, or any other existing or later-developed image sensing device capable of sensing images, such as images of a user operating computing device. Although a single sensor 150 is shown, computing device 100 may include multiple sensors 150. For example, computing device 100 may include a first camera oriented toward the user's field of view towards computing device 100 and a second camera oriented toward the user's field of view away from computing device 100.
[0036] The power supply 160 can be any device suitable for powering the computing device 100. For example, the power supply 160 may include a wired external power interface; one or more dry cell batteries, such as nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel hydride (NiMH), lithium-ion (Li-ion); a solar cell; a fuel cell; or any other device capable of powering the computing device 100. Although Figure 1 A single power source 160 is shown, but the computing device 100 may include multiple power sources 160, such as batteries and wired external power interfaces.
[0037] Although shown as separate units, the electronic communication unit 140, electronic communication interface 142, user interface 130, power supply 160, or portions thereof, can be configured as a combined unit. For example, the electronic communication unit 140, electronic communication interface 142, user interface 130, and power supply 160 can be implemented as a communication port capable of interfacing with an external display device, providing communication, power, or both.
[0038] One or more of the following components—memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, or power supply 160—may be operatively coupled via bus 170. Although Figure 1 A single bus 170 is shown, but computing device 100 may include multiple buses. For example, memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, and bus 170 may receive power from power supply 160 via bus 170. In another example, memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, power supply 160, or combinations thereof, may transmit data, for example, by sending and receiving electronic signals via bus 170.
[0039] although Figure 1 Although not shown separately, one or more of the processor 120, user interface 130, electronic communication unit 140, sensor 150, or power supply 160 may include internal memory, such as internal buffers or registers. For example, processor 120 may include internal memory (not shown) and may read data 112 from memory 110 into internal memory (not shown) for processing.
[0040] Although shown as separate components, memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, power supply 160 and bus 170, or any combination thereof, may be integrated into one or more electronic units, circuits or chips.
[0041] Figure 2This is a diagram of a computing and communication system 200 according to an implementation of this disclosure. The computing and communication system 200 shown includes computing and communication devices 100A, 100B, 100C, access points 210A, 210B, and a network 220. For example, the computing and communication system 200 may be a multiple access system that provides communication (such as voice, audio, data, video, messaging, broadcasting, or combinations thereof) to one or more wired or wireless communication devices (such as computing and communication devices 100A, 100B, 100C). Although for simplicity... Figure 2 Three computing and communication devices 100A, 100B, and 100C, two access points 210A and 210B, and a network 220 are shown, but any number of computing and communication devices, access points, and networks can be used.
[0042] The computing and communication devices 100A, 100B, and 100C can be, for example, computing devices, such as... Figure 1 The computing device 100 is shown in the diagram. For example, computing and communication devices 100A and 100B can be user devices, such as mobile computing devices, laptop computers, thin clients, or smartphones, and computing and communication device 100C can be a server, such as a mainframe or cluster. Although computing and communication devices 100A and 100B are described as user devices, and computing and communication device 100C is described as a server, any computing and communication device can perform some or all of the functions of a server, some or all of the functions of a user device, or some or all of the functions of a server and a user device. For example, server computing and communication device 100C can receive, encode, process, store, transmit, or a combination thereof of audio data, and one or both of computing and communication devices 100A and 100B can receive, decode, process, store, present, or a combination thereof of audio data.
[0043] Each computing and communication device 100A, 100B, 100C (which may include user equipment (UE), mobile station, fixed or mobile subscriber unit, cellular phone, personal computer, tablet computer, server, consumer electronics, or any similar device) may be configured to perform wired or wireless communication, such as via network 220. For example, computing and communication devices 100A, 100B, 100C may be configured to transmit or receive wired or wireless communication signals. Although each computing and communication device 100A, 100B, 100C is shown as a single unit, the computing and communication device may include any number of interconnecting elements.
[0044] Each access point 210A, 210B may be any type of device configured to communicate with computing and communication devices 100A, 100B, 100C, network 220, or both via wired or wireless communication links 180A, 180B, 180C. For example, access points 210A, 210B may include base stations, base transceivers (BTS), node Bs, enhanced node Bs (eNode-Bs), home node Bs (HNode-Bs), wireless routers, wired routers, hubs, repeaters, switches, or any similar wired or wireless devices. Although each access point 210A, 210B is shown as a single unit, the access point may include any number of interconnecting elements.
[0045] Network 220 may be any type of network configured to provide services such as voice, data, applications, Voice Internet Protocol (VoIP), or any other communication protocol or combination of communication protocols via wired or wireless communication links. For example, network 220 may be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a mobile or cellular telephone network, the Internet, or any other electronic communication method. The network may use communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Internet Protocol (IP), Real-Time Transfer Protocol (RTP), Hypertext Transfer Protocol (HTTP), or combinations thereof.
[0046] Computing and communication devices 100A, 100B, and 100C can communicate with each other via network 220 using one or more wired or wireless communication links, or a combination of wired and wireless communication links. For example, as shown, computing and communication devices 100A and 100B can communicate via wireless communication links 180A and 180B, and computing and communication device 100C can communicate via wired communication link 180C. Any one of computing and communication devices 100A, 100B, and 100C can communicate using any one or more wired or wireless communication links. For example, the first computing and communication device 100A can communicate via a first access point 210A using a first type of communication link, the second computing and communication device 100B can communicate via a second access point 210B using a second type of communication link, and the third computing and communication device 100C can communicate via a third access point (not shown) using a third type of communication link. Similarly, access points 210A and 210B can communicate with network 220 via one or more types of wired or wireless communication links 230A and 230B. Although Figure 2The computing and communication devices 100A, 100B, and 100C are shown communicating via network 220, but the computing and communication devices 100A, 100B, and 100C can communicate with each other via any number of communication links (such as direct wired or wireless communication links).
[0047] In some implementations, communication between one or more of the computing and communication devices 100A, 100B, and 100C may omit communication via network 220 and may include data transfer via another medium (not shown) (such as a data storage device). For example, server computing and communication device 100C may store audio data (such as encoded audio data) in a data storage device (such as a portable data storage unit), and one or both of computing and communication devices 100A or 100B may access, read, or retrieve the stored audio data from the data storage unit, such as by physically disconnecting the data storage unit from server computing and communication device 100C and physically connecting the data storage unit to computing and communication device 100A or computing and communication device 100B.
[0048] Other implementations of the computing and communication system 200 are also possible. For example, in one implementation, network 220 may be an ad hoc network, and one or more of access points 210A and 210B may be omitted. The computing and communication system 200 may include Figure 2 Devices, units, or elements not shown. For example, computing and communication system 200 may include more communication devices, networks, and access points.
[0049] Figure 3 This is a diagram of a video stream 300 used in encoding and decoding according to an implementation of this disclosure. The video stream 300 (such as a video stream captured by a camera or a video stream generated by a computing device) may include a video sequence 310. The video sequence 310 may include a sequence of adjacent frames 320. Although three adjacent frames 320 are shown, the video sequence 310 may include any number of adjacent frames 320.
[0050] Each frame 330 from adjacent frames 320 can represent a single image from the video stream. Although Figure 3 Not shown, but frame 330 may include one or more segments, tiles, or planes that can be independently (e.g., in parallel) or otherwise processed. Frame 330 may include one or more tiles 340. Each tile 340 may be a rectangular region of the frame that can be independently processed. Each tile 340 may include a corresponding block 350. Although Figure 3Not shown, but a block can include pixels. For example, a block can include a 16×16 pixel group, an 8×8 pixel group, an 8×16 pixel group, or any other pixel group. Unless otherwise indicated herein, the term 'block' can include a superblock, macroblock, segment, slice, or any other portion of a frame. Frames, blocks, pixels, or combinations thereof can include display information, such as luminance information, chrominance information, or any other information that can be used to store, modify, communicate, or display a video stream or a portion thereof.
[0051] Figure 4 This is a block diagram of an encoder 400 according to an implementation of this disclosure. The encoder 400 can be used in, for example... Figure 1 The computing device 100 shown in the figure or Figure 2 The computing and communication devices 100A, 100B, and 100C shown herein are implemented, for example, stored in a storage medium such as... Figure 1 The computer software program is located in the data storage unit of the memory 110 shown in the figure. The computer software program may include programs that can be generated by, for example, Figure 1 The processor 120 shown executes machine instructions and can cause the device to encode video data as described herein. The encoder 400 can be implemented as dedicated hardware, for example, included in the computing device 100.
[0052] Encoder 400 can handle, for example, Figure 3 The input video stream 402 of the video stream 300 shown is encoded to generate an encoded (compressed) bitstream 404. In some implementations, the encoder 400 may include a forward path for generating the compressed bitstream 404. The forward path may include an intra / inter-frame prediction unit 410, a transform unit 420, a quantization unit 430, an entropy coding unit 440, or any combination thereof. In some implementations, the encoder 400 may include a reconstruction path (indicated by a dotted connection line) for reconstructing frames for encoding further blocks. The reconstruction path may include a dequantization unit 450, an inverse transform unit 460, a reconstruction unit 470, a filtering unit 480, or any combination thereof. Other structural variations of the encoder 400 may be used to encode the video stream 402.
[0053] In order to encode video stream 402, each frame within video stream 402 can be processed in blocks. Therefore, the current block can be identified from the blocks in the frame, and the current block can be encoded.
[0054] In the intra / inter-frame prediction unit 410, the current block can be encoded using either intra-frame prediction (which can occur within a single frame) or inter-frame prediction (which can occur from frame to frame). Intra-frame prediction may include generating a prediction block based on previously encoded and reconstructed samples in the current frame. Inter-frame prediction may include generating a prediction block based on samples from one or more previously constructed reference frames. Generating a prediction block for the current block in the current frame may include performing motion estimation to generate motion vectors indicating appropriate reference portions of the reference frames.
[0055] Intra / inter-frame prediction unit 410 can subtract the predicted block from the current block (original block) to produce a residual block. Transform unit 420 can perform a block-based transform that may include transforming the residual block into transform coefficients in, for example, the frequency domain. Examples of block-based transforms include the Karhunen-Loève transform (KLT), discrete cosine transform (DCT), singular value decomposition transform (SVD), and asymmetric discrete sine transform (ADST). In the example, DCT may include transforming the block to the frequency domain. DCT may include using transform coefficient values based on spatial frequency, with the lowest frequency (i.e., DC) coefficients in the upper left of the matrix and the highest frequency coefficients in the lower right of the matrix.
[0056] Quantization unit 430 can convert transform coefficients into discrete quantum values, which may be referred to as quantized transform coefficients or quantization levels. The quantized transform coefficients can be entropy-encoded by entropy coding unit 440 to produce entropy-encoded coefficients. Entropy coding may include the use of probability distribution indices. The entropy-encoded coefficients and information for the decoded block can be output to a compressed bitstream 404, which may include the prediction type used, motion vectors, and quantizer values. The compressed bitstream 404 can be formatted using various techniques such as run-length encoding (RLE) and zero-run coding.
[0057] The reconstruction path can be used to maintain encoder 400 and such Figure 5The decoder 500 shown in the diagram synchronizes reference frames between corresponding decoders. The reconstruction path can be similar to the decoding process discussed below and can include decoding the encoded frame, or a portion thereof. This can include decoding the encoded block, which may include dequantizing the quantized transform coefficients at dequantization unit 450 and performing an inverse transform on the dequantized transform coefficients at inverse transform unit 460 to produce a derived residual block. Reconstruction unit 470 can add the prediction block generated by intra / inter-frame prediction unit 410 to the derived residual block to create a decoded block. Filtering unit 480 can be applied to the decoded block to generate a reconstructed block that can reduce distortion such as block artifacts. Although Figure 4 A filtering unit 480 is shown, but filtering a decoded block can include loop filtering, deblocking filtering, or other types of filtering, or a combination of multiple filtering types. The reconstructed block can be stored or otherwise made accessible as a reconstructed block, which can be part of a reference frame, used to encode another part of the current frame, another frame, or both, as indicated by the dashed line at 482. Code processing information of the frame (such as the deblocking threshold index value) can be encoded, included in the compressed bitstream 404, or both, as indicated by the dashed line at 484.
[0058] Other variations of encoder 400 can be used to encode the compressed bitstream 404. For example, a non-transform-based encoder 400 can directly quantize the residual block without transform unit 420. In some implementations, quantization unit 430 and dequantization unit 450 can be combined into a single unit.
[0059] Figure 5 This is a block diagram of a decoder 500 according to an implementation of this disclosure. The decoder 500 can be configured in, for example... Figure 1 The computing device 100 shown in the figure or Figure 2 The computing and communication devices 100A, 100B, and 100C shown herein are implemented, for example, stored in a storage medium such as... Figure 1 The computer software program is located in the data storage unit of the memory 110 shown in the figure. The computer software program may include programs that can be generated by, for example, Figure 1 The processor 120 shown executes machine instructions and enables the device to decode video data as described herein. The decoder 500 can be implemented as dedicated hardware, for example, included in the computing device 100.
[0060] Decoder 500 can receive compressed bitstream 502, such as Figure 4The compressed bitstream 404 is shown, and the compressed bitstream 502 can be decoded to generate the output video stream 504. The decoder 500 may include an entropy decoding unit 510, a dequantization unit 520, an inverse transform unit 530, an intra / inter-frame prediction unit 540, a reconstruction unit 550, a filtering unit 560, or any combination thereof. Other structural variations of the decoder 500 can be used to decode the compressed bitstream 502.
[0061] Entropy decoding unit 510 can use, for example, context-adaptive binary arithmetic decoding to decode data elements within the compressed bitstream 502 to produce a quantized set of transform coefficients. Dequantization unit 520 can dequantize the quantized transform coefficients, and inverse transform unit 530 can perform an inverse transform on the dequantized transform coefficients to produce a derived residual block, which can be combined with... Figure 4 The inverse transform unit 460 shown corresponds to the derived residual block generated. Using header information decoded from the compressed bitstream 502, the intra / inter-frame prediction unit 540 can generate a prediction block corresponding to the prediction block created in the encoder 400. At the reconstruction unit 550, the prediction block can be added to the derived residual block to create a decoded block. The filtering unit 560 can be applied to the decoded block to reduce artifacts (such as block artifacts), which may include loop filtering, deblocking filtering, or other types of filtering or a combination of multiple filtering types, and may include generating a reconstructed block that can be output as the output video stream 504.
[0062] Other variations of decoder 500 can be used to decode the compressed bitstream 502. For example, decoder 500 can produce an output video stream 504 without deblocking filtering.
[0063] Figure 6 A frame (such as) according to an implementation of this disclosure. Figure 3The diagram shows a block representation of a portion 600 of frame 330. As shown, portion 600 of the frame includes four 64×64 blocks 610, which are represented in two rows and two columns in a matrix or Cartesian plane. In some implementations, the 64×64 block can be a maximum code processing unit, N=64. Each 64×64 block can include four 32×32 blocks 620. Each 32×32 block can include four 16×16 blocks 630. Each 16×16 block can include four 8×8 blocks 640. Each 8×8 block 640 can include four 4×4 blocks 650. Each 4×4 block 650 can include 16 pixels, which can be represented in four rows and four columns in each corresponding block in the Cartesian plane or matrix. Pixels can include information representing the image captured in the frame, such as brightness information, color information, and position information. In some implementations, a 16×16 pixel block, as shown, may include: a luma block 660, which may include luma pixels 662; and two chroma blocks 670 and 680, such as a U or Cb chroma block 670 and a V or Cr chroma block 680. Chroma blocks 670 and 680 may include chroma pixels 690. For example, luma block 660 may include 16×16 luma pixels 662, and each chroma block 670 and 680 may include 8×8 chroma pixels 690, as shown. Although one arrangement of blocks is shown, any arrangement can be used. Figure 6 An N×N block is shown, but in some implementations, an N×M block can be used. For example, 32×64 blocks, 64×32 blocks, 16×32 blocks, 32×16 blocks, or blocks of any other size can be used. In some implementations, N×2N blocks, 2N×N blocks, or combinations thereof can be used.
[0064] In some implementations, video frame coding may include ordered block-level coding. Ordered block-level coding may involve coding blocks of a frame in an order such as raster scan order, where blocks can be identified and processed starting with the top-left block of the frame or a portion of the frame, and then proceeding row by row from left to right and from top to bottom, thus identifying each block sequentially for processing. For example, the 64×64 block in the top left column of the frame may be the first coded block, and the 64×64 block immediately to the right of the first block may be the second coded block. The second row from the top may be the second coded row, such that the 64×64 block in the left column of the second row can be coded after the 64×64 block in the rightmost column of the first row.
[0065] In some implementations, code processing of blocks may include using quadtree code processing, which may involve code processing smaller block units within the block in raster scan order. For example, quadtree code processing can be used to... Figure 6The 64×64 block shown in the lower left corner of the frame portion is encoded, where the upper left 32×32 block can be encoded, then the upper right 32×32 block, then the lower left 32×32 block, and then the lower right 32×32 block. Quadtree encoding can be used to encode each 32×32 block, where the upper left 16×16 block can be encoded, then the upper right 16×16 block, then the lower left 16×16 block, and then the lower right 16×16 block. Quadtree encoding can be used to encode each 16×16 block, where the upper left 8×8 block can be encoded, then the upper right 8×8 block, then the lower left 8×8 block, and then the lower right 8×8 block. Quadtree coding can be used to code each 8×8 block. This can be done by coding the top-left 4×4 block, then the top-right 4×4 block, then the bottom-left 4×4 block, and finally the bottom-right 4×4 block. In some implementations, the 8×8 blocks can be omitted from the 16×16 block, and quadtree coding can be used to code the 16×16 block. This involves coding the top-left 4×4 block, and then coding the remaining 4×4 blocks in the raster scan order.
[0066] In some implementations, video bitcoding may include compressing information included in the original or input frame by, for example, omitting some information from the original frame from the corresponding encoded frame. For instance, bitcoding may include reducing spectral redundancy, reducing spatial redundancy, reducing temporal redundancy, or a combination thereof.
[0067] In some implementations, reducing spectral redundancy may include using a color model based on a luma component (Y) and two chromaticity components (U and V or Cb and Cr), which may be referred to as the YUV or YCbCr color model or color space. Using the YUV color model may involve using a relatively large amount of information to represent the luma component of a portion of the frame and using a relatively small amount of information to represent each corresponding chromaticity component of that portion of the frame. For example, a portion of the frame may be represented by a high-resolution luma component that may include a 16×16 pixel block and two lower-resolution chromaticity components, where each chromaticity component represents a portion of the frame as an 8×8 pixel block. Pixels may indicate values, for example, values in the range of 0 to 255, and may be stored or transmitted using, for example, eight bits. Although this disclosure is described with reference to the YUV color model, any color model may be used.
[0068] In some implementations, reducing spatial redundancy may include transforming the block into the frequency domain using, for example, the discrete cosine transform (DCT). For example, encoder units (such as...) Figure 4 The transform unit 420 shown can perform DCT using transform coefficient values based on spatial frequency.
[0069] In some implementations, reducing temporal redundancy may include encoding frames with a relatively small amount of data based on the similarity between frames using one or more reference frames, which may be previously encoded, decoded, and reconstructed frames of the video stream. For example, a block or pixel in the current frame may resemble a spatially corresponding block or pixel in a reference frame. In some implementations, a block or pixel in the current frame may resemble a block or pixel in a reference frame at different spatial locations, and reducing temporal redundancy may include generating motion information indicating the spatial difference or translation between the position of a block or pixel in the current frame and its corresponding position in the reference frame.
[0070] In some implementations, reducing temporal redundancy may include identifying a portion of a reference frame corresponding to the current block or pixel of the current frame. For example, a reference frame or a portion thereof, which may be stored in memory, may be searched to identify a portion for generating a prediction used to encode the current block or pixel of the current frame with maximum efficiency. For example, the search may identify a block of the reference frame where the difference in pixel values between its current block and a predicted block generated based on that portion of the reference frame is minimized, and may be referred to as a motion search. In some implementations, the portion of the reference frame being searched may be restricted. For example, the portion of the reference frame being searched may include a restricted number of rows of the reference frame, which may be referred to as a search region. In an example, identifying the portion of the reference frame used to generate the prediction may include calculating a cost function, such as the sum of absolute differences (SAD), between the pixels of each portion of the search region and the pixels of the current block.
[0071] In some implementations, the spatial difference between the position of the portion of the reference frame used to generate the prediction within the reference frame and the position of the current block in the current frame can be represented as a motion vector. The difference in pixel values between the prediction block and the current block can be referred to as differential data, residual data, prediction error, or residual block. In some implementations, generating the motion vector can be referred to as motion estimation, and the pixels of the current block can be indicated using Cartesian coordinates based on their position as f. x, y Similarly, the pixels of the search area in the reference frame can be indicated using Cartesian coordinates based on their position as r. x, y The motion vector (MV) for the current block can be determined based on, for example, the SAD between pixels in the current frame and corresponding pixels in a reference frame.
[0072] Although this document describes frames with reference to matrix or Cartesian representations for clarity, frames can be stored, transmitted, processed, or any combination thereof using any data structure, making it possible to efficiently represent the pixel values of frames or images. For example, frames can be stored, transmitted, processed, or any combination thereof using two-dimensional data structures (such as the matrix shown) or one-dimensional data structures (such as vector arrays). In implementations, the representation of a frame (such as the two-dimensional representation shown) can correspond to its physical location when the frame is rendered as an image. For example, the top-left corner position of a block at the top-left corner of a frame can correspond to the physical location of the top-left corner when the frame is rendered as an image.
[0073] In some implementations, block-based code processing efficiency can be improved by dividing the input block into one or more prediction partitions, which can be rectangular (including square) partitions for predictive code processing. In some implementations, video code processing using prediction partitions can include selecting a prediction partitioning scheme from multiple candidate prediction partitioning schemes. For example, in some implementations, candidate prediction partitioning schemes for a 64×64 code processing unit can include rectangular prediction partitions ranging in size from 4×4 to 64×64, such as 4×4, 4×8, 8×4, 8×8, 8×16, 16×8, 16×16, 16×32, 32×16, 32×32, 32×64, 64×32, or 64×64. In some implementations, video code processing using prediction partitions can include a complete prediction partition search, which can include selecting a prediction partitioning scheme by encoding the code processing unit with each available candidate prediction partitioning scheme and selecting the optimal scheme (such as the one that produces the minimum rate distortion error).
[0074] In some implementations, coding a video frame may include identifying a prediction partitioning scheme for encoding the current block (such as block 610). In some implementations, identifying the prediction partitioning scheme may include determining whether to encode the block as a single prediction partition of the largest code processing unit size (which could be 64×64 as shown), or to divide the block into multiple prediction partitions (which may correspond to sub-blocks, such as 32×32 block 620, 16×16 block 630, or 8×8 block 640 as shown), and may include determining whether to divide it into one or more smaller prediction partitions. For example, a 64×64 block may be divided into four 32×32 prediction partitions. Three of the four 32×32 prediction partitions may be encoded as 32×32 prediction partitions, and the fourth 32×32 prediction partition may be further divided into four 16×16 prediction partitions. Three of the four 16×16 prediction partitions can be encoded as 16×16 prediction partitions, and the fourth 16×16 prediction partition can be further divided into four 8×8 prediction partitions, each of which can be encoded as an 8×8 prediction partition. In some implementations, identifying the prediction partitioning scheme may include using a prediction partitioning decision tree.
[0075] In some implementations, video code processing for the current block can include identifying the optimal predictive code processing (RCC) mode from a variety of candidate RCC modes. This provides flexibility for processing video signals with diverse statistical properties and can improve compression efficiency. For example, the video code processor can evaluate each candidate RCC mode to determine the optimal RCC mode, which could be, for example, the RCC mode that minimizes the error metric (such as rate-distortion cost) of the current block. In some implementations, the complexity of searching for candidate RCC modes can be reduced by limiting the set of available candidate RCC modes based on the similarity between the current block and the corresponding predictive block. In some implementations, the complexity of searching for each candidate RCC mode can be reduced by performing a targeted refinement mode search. For example, metrics can be generated for a finite set of candidate block sizes (such as 16×16, 8×8, and 4×4), the error metrics associated with each block size can be sorted in descending order, and other candidate block sizes (such as 4×8 and 8×4 block sizes) can be evaluated.
[0076] In some implementations, block-based code processing efficiency can be improved by dividing the current residual block into one or more transform partitions, which can be rectangular (including square) partitions for transform code processing. In some implementations, video code processing (such as video code processing using transform partitions) can include selecting a uniform transform partitioning scheme. For example, the current residual block (such as block 610) can be a 64×64 block and can be transformed using a 64×64 transform without partitioning.
[0077] although Figure 6 While not explicitly stated, a unified transform partitioning scheme can be used to transform partition the residual blocks. For example, a unified transform partitioning scheme comprising four 32×32 transform blocks, sixteen 16×16 transform blocks, sixty-four 8×8 transform blocks, or 256 4×4 transform blocks can be used to transform partition the 64×64 residual blocks.
[0078] In some implementations, video bitcoding (such as video bitcoding using transform partitioning) may include using polyform transform partitioning to identify multiple transform block sizes for the residual block. In some implementations, polyform transform partitioning may include recursively determining whether to transform the current block using the current block size transform, or by partitioning the current block and performing polyform transform partitioning on each partition. For example, Figure 6 The lower left block 610 shown may be a 64×64 residual block, and the multiform transform partitioning code processing may include determining whether to code the current 64×64 residual block using a 64×64 transform, or to code the 64×64 residual block by dividing it into partitions (such as four 32×32 blocks 620) and performing multiform transform partitioning code processing on each partition. In some implementations, determining whether to partition the current block may be based on a comparison of the cost of encoding the current block using a current block size transform with the sum of the costs of encoding each partition using a partition size transform.
[0079] Figure 7 This is a block diagram illustrating an example of translational motion decoder-side motion vector refinement 700 based on bilateral matching (BM). The bilateral matching-based translational motion decoder-side motion vector refinement 700 can be implemented by a decoder (such as...) Figure 5 The decoder 500 shown is implemented. The translational motion decoder side motion vector refinement 700 based on bilateral matching can be achieved by the encoder (such as...). Figure 4 The encoder 400 shown is implemented where “decoder side” indicates translational motion based on bilateral matching (BM). Decoder side motion vector refinement 700 includes using information currently available for decoding the encoded video sequence or a portion thereof (such as previously reconstructed reference frame data); omitting or excluding the use of information not available for decoding the encoded video sequence or a portion thereof; and omitting or excluding the inclusion of information not available for decoding the encoded video sequence or a portion thereof in the compressed bitstream unless otherwise described herein or otherwise clear from the context.
[0080] The translational motion decoder based on bilateral matching, implemented by the encoder, refines the side motion vectors 700, including the input video stream (such as... Figure 4 The input video stream 402 shown, or one or more portions thereof, is encoded to generate an encoded (compressed) output bitstream (such as...). Figure 4 The encoded (compressed) bitstream shown is 404.
[0081] The decoder-based translational motion refinement 700, implemented by the decoder, includes the refinement of the encoded video stream (such as...) Figure 5 The compressed bitstream 502 shown, or one or more portions thereof, is decoded to generate an encoded (compressed) output bitstream (such as...). Figure 4 The encoded (compressed) bitstream shown is 404.
[0082] In block-based hybrid video codec processing, in order to reduce or minimize the resource consumption (such as bandwidth consumption) of signaling, storage, or both of compressed or encoded video data, redundant data (such as spatial redundant data, temporal redundant data, or both) is omitted or eliminated from the compressed or encoded data.
[0083] The translational motion decoder side motion vector refinement 700 based on bilateral matching includes encoding or decoding using bidirectional merging candidates, such as in bidirectional prediction (or bi-prediction). Previously reconstructed reference frame data may include backward reference frames, which can be previously reconstructed frames such as those following the current frame in temporal or frame index order. Previously reconstructed reference frame data may also include forward reference frames, which can be previously reconstructed frames such as those preceding the current frame in temporal or frame index order. Reference frames may be indicated in one or more lists of reference images, such as a first list of reference images (L0), a second list of reference images (L1), or both. In some implementations, the first list of reference images (L0) may be a forward prediction list of reference images (forward reference image list), and the second list of reference images (L1) may be a backward prediction list of reference images (backward reference image list).
[0084] Bidirectional prediction involves obtaining refined motion vectors by searching regions of previously reconstructed reference frame data, such as from a forward reference image list (L0), a second reference image list (L1), or both, based on one or more previously obtained motion vector identifiers. Bidirectional matching involves obtaining, determining, or calculating distortions between candidate blocks obtained from previously reconstructed reference frame data, such as the distortion between a first candidate block obtained from previously reconstructed reference frame data from the forward reference image list (L0) and a second candidate block obtained from previously reconstructed reference frame data from the second reference image list (L1).
[0085] Figure 7 The current frame 710, the first reference frame 720 from the forward reference image list (L0), and the second reference frame 730 from the backward reference image list (L1) are shown.
[0086] The current frame 710 includes the current block 712. A first portion of the current frame 710 is shown against a dark dotted background to indicate that the reconstructed frame data corresponding to the first portion of the current frame 710 can be used as (spatial) reference data for decoding or encoding the current block 712. A second portion of the current frame 710 is shown against a white background to indicate that the reconstructed frame data corresponding to the second portion of the current frame 710 cannot be used as reference data for decoding or encoding the current block 712. The current block 712 is shown against a background of diagonally upward-sloping lines (from left to right) to indicate that the bilateral matching translational motion decoder side motion vector refinement 700 includes decoding or encoding the current block 712.
[0087] The first reference frame 720 is shown as including a first block 722 located at a position within the first reference frame 720 indicated by a first previously acquired motion vector 740 (MV0) relative to the position of the current block 712 in the current frame 710. The first reference frame 720 is shown against a light-colored dotted background to indicate that the reconstructed frame data corresponding to the first reference frame 720 can be used as (time) reference data for decoding or encoding the current block 712. The first block 722 is shown against a background of diagonally downward-sloping lines (from left to right) to indicate that the first block 722 corresponds to the first previously acquired motion vector 740 (MV0).
[0088] The second reference frame 730 is shown as including a second block 732 located at a position within the second reference frame 730 indicated by a second previously obtained motion vector 742 (MV1) relative to the position of the current block 712 in the current frame 710. The second reference frame 730 is shown against a light-colored dotted background to indicate that the reconstructed frame data corresponding to the second reference frame 730 can be used as (time) reference data for decoding or encoding the current block 712. The second block 732 is shown against a background of diagonally downward-sloping lines (from left to right) to indicate that the second block 732 corresponds to the second previously obtained motion vector 742 (MV1).
[0089] The time distance (d0) from the first reference frame 720 to the current frame 710 matches the time distance (d1) from the current frame 710 to the second reference frame 730.
[0090] Figure 7 The diagram shows the first refined motion vector 750 (MV). 0' The first direction line and its representation are related to the first refined motion vector 750 (MV). 0' Aligned second refined motion vector 752 (MV) 1' The second direction line indicates that the motion or trajectory from the first reference frame 720 through the current frame 710 to the second reference frame 730 is continuous, such as being constructively continuous. The first refined motion vector 750 (MV 0' ) relative to the current frame 710 and the second refined motion vector 752 (MV 1' It is a mirror image.
[0091] The translational motion decoder-side motion vector refinement 700 based on bilateral matching includes searching for the first reference frame 720 (as indicated by the first motion vector 740 (MV0)) in the region surrounding the first block 722 to obtain the position of the first refined matching block 724, and obtaining the first refined motion vector 750 (MV0) based on that position. 0' The first refined matching block 724 is shown against a background of diagonally upward lines (from left to right) to indicate that the first refined matching block 724 corresponds to the first refined motion vector 750 (MV). 0' The first refined matching block 724 is instructed to be used as a reference for encoding the current block 712. Obtaining the first refined matching block 724 may include generating the first refined matching block 724 based on the reconstructed data in the first reference frame 720, such as by using interpolation.
[0092] Figure 7 The first offset motion vector 760 (MV) is shown. DIFFThe third direction line, the first offset motion vector indicating the position in the first reference frame 720 indicated by the first motion vector 740 (MV0) and the first refined motion vector 750 (MV0) 0' The offset of the position in the first reference frame 720 indicated by )
[0093] The translational motion decoder-side motion vector refinement 700 based on bilateral matching includes searching for a second reference frame 730 (as indicated by the second motion vector 742 (MV1)) in the region surrounding the second block 732 to obtain the position of the second refined matching block 734, and obtaining the second refined motion vector 752 (MV1) based on that position. 1' The second refined matching block 734 is shown against a background of diagonally upward-sloping lines (from left to right) to indicate that the second refined matching block 734 corresponds to the second refined motion vector 752 (MV). 1' The second refined matching block 734 is instructed to be used as a reference for encoding the current block 712. Obtaining the second refined matching block 734 may include generating the second refined matching block 734 based on the reconstructed data in the second reference frame 730, such as by using interpolation.
[0094] Figure 7 The second offset motion vector 762(-MV) is shown. DIFF The fourth direction line, the second offset motion vector indicating the position in the second reference frame 730 indicated by the second motion vector 742 (MV1) and the second refined motion vector 752 (MV2) 1' The offset of the position in the second reference frame 730 indicated by )
[0095] Searching the first reference frame 720 and searching the second reference frame 730 includes evaluating candidate motion vectors or candidate motion vector pairs, wherein the first refined motion vector 750 (MV 0' ) and the second refined motion vector 752 (MV 1' ) are candidate motion vector pairs. Obtaining the first refined matching block 724 and obtaining the second refined matching block 734 includes determining that the error metric (such as the sum of absolute differences) between the first refined matching block 724 and the second refined matching block 734 is the smallest among the candidate blocks corresponding to the respective candidate motion vector pairs obtained from the first reference frame 720 and the second reference frame 730.
[0096] In some implementations, the translational motion decoder side motion vector refinement 700 based on bilateral matching can be used for translational motion vectors with two parameters representing translational motion, and may not be used for twisted motion vectors with more than two parameters or that otherwise represent motion other than translational motion.
[0097] In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 can be used for code processing blocks or code processing units according to the following characteristics: In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on blocks or code processing units that use a code processing unit-level merging mode and have bidirectional predicted motion vectors. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on blocks or code processing units that use backward prediction reference pictures or frames and forward prediction reference pictures or frames for code processing. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on blocks or code processing units where the distance from the reference frame to the current frame (which may be matched with the picture sequence count (POC) difference) is performed. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on block or code processing units where the reference frame is a short-term reference frame. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on block or code processing units where the current block or code processing unit includes sixty-four or more luma (or luminance) samples. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on block or code processing units where the height of the block or code processing unit is greater than or equal to eight luma samples, and the width of the block or code processing unit is greater than or equal to eight luma samples, or both. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on blocks or code processing units where the weights indicated by the bidirectional prediction weighting (BCW) weight index of the code processing unit are equal. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on blocks or code processing units where weighted prediction is disabled for the current block. In some implementations, the bilateral matching-based translational motion decoder-side motion vector refinement 700 includes performing bilateral matching-based translational motion decoder-side motion vector refinement 700 on blocks or code processing units where the use of combined inter-frame combining and intra-frame prediction (CIIP) modes is omitted for the current block.The translational motion decoder side motion vector refinement 700 based on bilateral matching may not be applicable to block or code processing units other than those with the aforementioned characteristics.
[0098] One or more refined motion vectors obtained using the bilateral matching-based translational motion decoder-side motion vector refinement 700 are used to generate inter-frame prediction samples. One or more refined motion vectors obtained using the bilateral matching-based translational motion decoder-side motion vector refinement 700 are used for temporal motion vector prediction in subsequent code processing. The first motion vector, the second motion vector, or both can be used for deblocking. The first motion vector, the second motion vector, or both can be used for spatial motion vector prediction in subsequent block or code processing units.
[0099] Figure 8 This is a block diagram illustrating an example of prediction refinement using optical flow. Prediction refinement using optical flow can be performed by a decoder (such as...) Figure 5 The decoder 500 shown is implemented. Prediction refinement 800 using optical flow can be achieved by an encoder (such as...). Figure 4 The encoder 400 shown is implemented. As used herein, the term “twisting” motion refers to non-translational motion other than or in lieu of translational motion, such as affine motion, homography, or similarity motion.
[0100] Motions other than translational motion (which may not be accurately represented using translational motion vectors) can be represented using warp-torsion models, such as homography warp-torsion models, affine warp-torsion models, similarity warp-torsion models, or other warp-torsion models. In some implementations, a six-parameter warp-torsion can be called a six-parameter affine motion, and the corresponding model can be called a six-parameter affine motion model. In some implementations, a model corresponding to a six-parameter warp-torsion can be called a six-parameter warp-torsion model. In some implementations, a four-parameter warp-torsion can be called a four-parameter affine motion, and the corresponding model can be called a four-parameter affine motion model. In some implementations, a model corresponding to a four-parameter warp-torsion can be called a four-parameter warp-torsion model.
[0101] Figure 8 The current block 810 is shown, which includes the current sub-block 812, which includes pixels (shown as circles), including the current pixel 814. Figure 8 The translational motion vector 820 (V) from the current sub-block 812 to the translation prediction block 830 is shown. SB ). Figure 8The diagram shows the warped motion vector 840 (V(i, j)) from the current pixel 814 in the current sub-block 812 to the corresponding pixel position in the warped prediction block 850. Prediction refinement 800 using optical flow includes refining the translational motion vector 820 (V... SB This is used to obtain the refined motion vector for each pixel.
[0102] Compared to pixel-based motion compensation, sub-block-based twisted or affine motion compensation can reduce memory access bandwidth, computational complexity, or both. However, sub-block-based twisted or affine motion compensation may decrease prediction accuracy compared to pixel-based motion compensation. To achieve finer-grained motion compensation (such as smaller block sizes), predictions after sub-block-based twisted or affine motion compensation are refined using prediction refinement 800 leveraging optical flow, avoiding increased memory access bandwidth for motion compensation. In some implementations, brightness prediction samples are refined after sub-block-based twisted or affine motion compensation.
[0103] Prediction refinement using optical flow 800 may include using sub-block-based twisting or affine motion compensation to obtain sub-block predictions I(i, j).
[0104] Prediction refinement using optical flow 800 can include using filters (such as a three-tap filter [-1, 0, 1]) to obtain the spatial gradient g of the sub-block prediction at each sample location. x (i, j) and g y (i, j). In some implementations, gradient calculations can be performed using bidirectional optical flow (BDOF), where the spatial gradient using the gradient precision control parameter (shift1) can be represented as follows: g x (i, j)=(I(i+1, j)≫shift1)-(I(i-1, j)≫shift1), g y (i, j)=(I(i, j+1)≫shift1)-(I(i, j-1)≫shift1).
[0105] Prediction refinement using optical flow 800 can include sub-block prediction (such as for 4×4 sub-blocks), which can obtain gradients by expanding samples on each side. To avoid increased memory bandwidth usage, interpolation calculations, or both, the nearest integer pixel location in a reference image or frame can be used as the expanded sample.
[0106] Prediction refinement using optical flow 800 may include brightness prediction refinement based on optical flow, which may include using the sample motion vector 840 (V(i, j)) of the sample position (i, j) obtained with a warp or affine model and the translational motion vector 820 (V(i, j)) of the sub-block including the sample (i, j). SB The difference (ΔV(i, j)) between them can be expressed as follows:
[0107] The difference (ΔV(i, j)) is quantized in units of 1 / 32 brightness sample precision, and can be referred to as brightness prediction refinement.
[0108] The distortion or affine model parameters and the sample position relative to the sub-block center can remain constant across different sub-blocks, allowing the difference (ΔV(i,j)) obtained for the first sub-block (such as the top-left sub-block 812) to be used for other sub-blocks in block 810 or the code processing unit. The distance from the sample position (i,j) to the sub-block center (...) can be used... , The horizontal offset (dx(i,j)) and vertical offset (dy(i,j)) are used to obtain the motion vector difference (ΔV(x, y)) relative to the center of the sub-block, which can be represented as follows:
[0109] For accuracy, it can be determined based on the sub-block width ( ) and sub-block height ( To obtain the center of the sub-block ( , ), where the width of the sub-block ( Subtracting one from the result and dividing it by two yields the horizontal component of the center, and the height of the sub-block ( Subtracting one and dividing by two yields the result as the vertical component of the center, which can be represented as follows: .
[0110] The homography-warp motion model includes eight parameters that indicate the displacement between pixels of the current block and pixels of a reference frame (such as in a quadrilateral portion of the reference frame) to generate the prediction block. The homography-warp motion model can represent translation, rotation, scaling, aspect ratio changes, shearing, and other non-parallelogram distortions.
[0111] The affine warp motion model comprises six parameters that indicate the displacement between pixels of the current block and pixels of a reference frame (such as in a parallelogram portion of the reference frame) to generate the prediction block. The affine warp motion model is a linear transformation between coordinates in two spaces, represented by these six parameters. The affine warp motion model can represent translation, rotation, scaling, aspect ratio changes, and shearing. The parameters of the affine warp motion model include a first pair of parameters (h... 13 , h 23 The first pair of parameters represents translational motion (translation parameters), such as horizontal translational motion parameters (h). 13 ) and vertical translational motion parameters (h 23 The parameters of the affine torsion motion model include a second pair of parameters (h). 11 , h 22 The second pair of parameters represents scaling (scaling parameters), such as the horizontal scaling parameter (h). 11 ) and vertical scaling parameter (h) 22 The parameters of the affine torsion motion model include a third pair of parameters (h). 12 , h 21 The third pair of parameters, together with the scaling parameters, represents the angular rotation (rotation parameters). For example, for the current pixel at position (x, y) in the current frame, the corresponding position (x', y') in the reference frame can be indicated using an affine twisted motion model, which may include: a horizontal displacement for encoding the current block (x', y'). The horizontal displacement is the result of multiplying the horizontal scaling parameter by the current horizontal position, multiplying the first rotation parameter by the current vertical position, and adding the horizontal translation motion parameter; and the vertical displacement used to encode the current block ( The vertical displacement is the result of multiplying the vertical scaling parameter by the current horizontal position, multiplying the second rotation parameter by the current vertical position, and adding the vertical translation motion parameter. This can be expressed as follows:
[0112] A six-parameter affine warp motion model (including the motion vector of the upper left control point) , ), upper right control point motion vector ( , ), lower left control point motion vector ( , The width (w) and height (h) of the block or code processing unit can be represented as follows:
[0113] The similarity-through-warp motion model comprises four parameters indicating the displacement between pixels of the current block and pixels of a reference frame (such as within a square portion of the reference frame) to generate the prediction block. The similarity-through-warp motion model is a linear transformation between coordinates in two spaces, represented by these four parameters. For example, these four parameters could be translation along the x-axis, translation along the y-axis, rotation, and scaling. The similarity-through-warp motion model can represent a square-to-square transformation with rotation and scaling. The parameters of the similarity-through-warp motion model include a first pair of parameters (h... 13 , h 23 The first pair of parameters represents translational motion (translation parameters), such as horizontal translational motion parameters (h). 13 ) and vertical translational motion parameters (h 23 The parameters of the similarity-through-distortion motion model include the second parameter (h). 11 The second parameter represents scaling (scaling parameter) (h) 22 = h 11 The parameters of the similarity-through-distortion motion model include the third parameter (h). 21 This third parameter, together with the scaling parameter, represents the angular rotation (rotation parameter) (h). 12 = h 21 For example, for the current pixel at position (x, y) in the current frame, a similarity-distorted motion model can be used to indicate the corresponding position (x', y') in the reference frame. This similarity-distorted motion model can include: a horizontal displacement for encoding the current block ( The horizontal displacement is the result of subtracting the result of multiplying the rotation parameter by the current vertical position from the result of multiplying the horizontal scaling parameter by the current horizontal position, and adding the horizontal translation motion parameter; and the vertical displacement used to encode the current block ( The vertical displacement is the result of multiplying the rotation parameter by the current horizontal position, multiplying the horizontal scaling parameter by the current vertical position, and adding the vertical translation motion parameter. This can be expressed as follows:
[0114] Four-parameter similarity through a twisted motion model (including the motion vector of the top left control point) , ), upper right control point motion vector ( , The width (w) of the block or code processing unit can be represented as follows:
[0115] The parameters of the torsional motion model, except for the translation parameters, are all non-translation parameters.
[0116] The brightness prediction refinement (ΔI(i, j)) is added to the sub-block prediction (I(i, j)) to obtain the final prediction (I'), which can be represented as follows: I'(i, j)= I(i, j)+ΔI(i, j).
[0117] In some implementations, for blocks or sub-blocks that are code-processed using a twisted motion model (where control point motion vectors (CPMVs) (which may be referred to as refined motion vectors) are matched), the prediction refinement using optical flow can be omitted. This matching indicates that the block or code-processing unit has translational motion and other motions are omitted.
[0118] In some implementations, for blocks or sub-blocks that are code-processed using a twisted motion model (where the twisted motion parameters are greater than defined limits), the optical flow prediction refinement can be omitted by using motion compensation based on the block or code processing unit and omitting the use of twisted motion compensation based on the sub-block, in order to avoid relatively large memory access bandwidth usage.
[0119] Although not shown separately, in some implementations, code processing may include multiple passes of decoder-side motion vector refinement. The first pass of multi-pass decoder-side motion vector refinement may include bilateral matching for the code processing block. The second pass of multi-pass decoder-side motion vector refinement may include bilateral matching for individual sub-blocks (such as 16×16 sub-blocks) within the code processing block. The third pass of multi-pass decoder-side motion vector refinement may include refining individual motion vectors within the individual sub-blocks (such as 8×8 sub-blocks) by applying bidirectional optical flow. The refined motion vectors may be stored for spatial motion vector prediction, temporal motion vector prediction, or both.
[0120] The first pass of multi-pass decoder-side motion vector refinement may include block-based bilateral matching motion vector refinement. Refined motion vectors are obtained or derived by applying bilateral matching to the code processing blocks. The refined motion vectors are obtained by searching for portions of reference frames obtained from reference image lists L0 and L1 based on the previously obtained motion vectors (MV0 and MV1). The refined motion vector (MV0) is obtained or derived based on the minimum bilateral matching cost between two reference or prediction blocks in L0 and L1, based on the previously obtained motion vectors (MV0 and MV1). pass1 and MV1 pass1 ).
[0121] Bilateral matching involves a local search to obtain the integer sample precision difference motion vector (intDeltaMV). The local search uses a 3×3 square search pattern to iterate through the defined search area, which is ([–sHor, sHor]) in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined according to the block size, and the maximum value of both sHor and sVer is eight.
[0122] Obtaining (such as calculating) the bilateral matching cost (bilCost) can be expressed as bilCost = mvDistanceCost + sadCost. The block size (cbW * cbH) can be greater than sixty-four, and the mean-eliminated sum of absolute differences (MRSADWha) cost function can be used to eliminate the DC effect of distortion between reference or predictor blocks. The bilateral matching cost (bilCost) at the center point of the 3×3 search pattern can have a minimum cost, in which case the local search for the integer sample precision difference motion vector (intDeltaMV) can be omitted. The bilateral matching cost (bilCost) at the center point of the 3×3 search pattern can have costs other than the minimum cost, and the current minimum cost search point can be used as the center point of the 3×3 search pattern. The search for the minimum cost can continue to the end of the search range.
[0123] Fractional sample refinement can be applied to obtain or derive the difference motion vector (deltaMV). The refined motion vector obtained after the first pass can be represented as follows: MV0pass1 = MV0 + deltaMV, MV1pass1 = MV1 – deltaMV.
[0124] The second pass of multi-pass decoder-side motion vector refinement can include sub-block-based bilateral matching motion vector refinement. In the second pass, refined motion vectors are obtained by applying bilateral matching to 16×16 grid sub-blocks. For each sub-block, in the reference image lists L0 and L1, the refined motion vectors are obtained from the first pass around the motion vector (MV0). pass1 and MV1 pass1 The corresponding refined motion vector is searched. The refined motion vector (MV0) is obtained based on the minimum bilateral matching cost between two reference or predictive sub-blocks in L0 and L1. pass2 (sbIdx2) and MV1 pass2 (sbIdx2)).
[0125] For each sub-block, bilateral matching involves a search to obtain the integer sample precision difference motion vector (intDeltaMV). The search has a search range of [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum value of sHor and sVer can be eight.
[0126] The bilateral matching cost is obtained (e.g., calculated) by applying a cost factor (costFactor) to the sum of absolute Hadamard transform differences (SATD) between two reference or predictive sub-blocks, which can be expressed as bilCost = satdCost * costFactor. The search area (2*sHor + 1) * (2*sVer + 1) can be divided into five diamond-shaped search regions. A cost factor (costFactor) is assigned to each search region, determined by the distance (intDeltaMV) between each search point and the coarse motion vector, and each diamond-shaped search region is processed in order starting from the center of the search area. Within each region, search points are processed in raster scan order, from the upper left corner to the lower right corner. The minimum bilateral matching cost (bilCost) within the current search region can be less than a threshold equal to the result of multiplying the sub-block width by the sub-block height (sbW * sbH), in which case integer pixel searches can be omitted; otherwise, integer pixel searches continue to the next search region until all search points have been checked. In some implementations, the difference between the previous minimum cost and the current minimum cost during iteration can be less than a threshold equal to the block area, in which case the search process can be omitted.
[0127] In some implementations, decoder-side motion vector refinement fractional sample refinement can be used to obtain or derive the difference motion vector deltaMV(sbIdx2). The refined motion vector obtained in the second pass can be represented as follows: MV0 pass2 (sbIdx2) = MV0 pass1 + deltaMV(sbIdx2), MV1 pass2 (sbIdx2) = MV1 pass1 – deltaMV(sbIdx2).
[0128] The third pass of multi-pass decoder-side motion vector refinement may include sub-block-based bidirectional optical flow motion vector refinement. In the third pass, refined motion vectors are obtained by applying bidirectional optical flow to 8×8 grid sub-blocks. For each 8×8 sub-block, bidirectional optical flow refinement is applied starting from the refined motion vectors of the parent-child blocks from the second pass to obtain unclipped scaled Vx and Vy. The obtained bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped to between -32 and 32.
[0129] The refined motion vector (MV0) is obtained from the third pass. pass3 (sbIdx3) and MV1 pass3 (sbIdx3)) can be represented as follows: MV0 pass3 (sbIdx3) = MV0 pass2 (sbIdx2) + bioMv, MV1 pass3 (sbIdx3) = MV0 pass2 (sbIdx2) – bioMv.
[0130] Although not shown separately, in some implementations, code processing may include twisted or affine decoder-side motion vector refinement. Applying twisted or affine decoder-side motion vector refinement to twisted or affine merged code processing blocks improves code processing efficiency. The twisted or affine model (including the motion vector of the upper left control point) , ), upper right control point motion vector ( , ), lower left control point motion vector ( , The width (w) and height (h) of the block or code processing unit can be represented as follows:
[0131] Basic motion vector This represents the translational motion of the affine model. Affine decoder-side motion vector refinement is used, refining the fundamental motion vectors of the affine model in the affine merging mode code processing block by applying the first step of multi-pass decoder-side motion vector refinement. Other steps of multi-pass decoder-side motion vector refinement are omitted. The translational motion vector offset is added to the affine merging list of candidates that satisfy the decoder-side motion vector refinement conditions. , and The motion vector offset is obtained or derived by minimizing the cost of bilateral matching, which is the same approach taken in motion vector refinement on the decoder side.
[0132] Multi-pass decoder-side motion vector refinement is efficient but computationally complex. Decoder-side motion vector refinement is limited by the use of translation models. In affine decoder-side motion vector refinement, the translational portion is refined, while other portions remain unrefined. In some implementations, affine motion estimation first solves the affine equations and then converts the solution into control point motion vectors, which introduces errors during the conversion. Decoder-side motion vector refinement may introduce quantization errors. Decoder-side motion vector refinement may introduce rounding errors.
[0133] Although this paper describes prediction refinement using optical flow, it is also possible to refine using optical flow motion vectors. Optical flow motion vector refinement involves bilateral matching that uses a composite reference to derive local motion vector offsets on a per-sub-block basis (e.g., each 8×8 or 4×4 sub-block). Optical flow motion vector refinement involves obtaining two translation parameters for each sub-block based on optical flow equations (e.g., Equation 1). Optical flow motion vector refinement involves obtaining two composite prediction reference blocks or prediction blocks (P0 and P1) based on forward and backward motion vectors (MV0 and MV1). Optical flow motion vector refinement involves obtaining the x and y spatial gradients (Gx0, Gx1, Gy0, Gy1) of the composite prediction reference blocks or prediction blocks (P0 and P1). Optical flow motion vector refinement involves parameter solving, where d0 and d1 are signed temporal distances (positive for the past, negative for the future). Optical flow motion vector refinement involves obtaining four motion vector offsets based on a two-dimensional inverse autocorrelation problem. Optical flow motion vector refinement involves using refined motion vectors obtained based on motion vector offsets to obtain motion compensation on a per-sub-block basis.
[0134] Figure 9 This is a flowchart illustrating an example of decoding 900 that includes motion refinement using bilateral or composite matching with one or more warped refinement models. Decoding 900, including motion refinement using bilateral or composite matching with one or more warped refinement models, can be performed by a decoder (such as...) Figure 5 The decoder 500 shown is implemented. Although Figure 9 Not explicitly shown, but including motion refinement using bilateral or composite matching with one or more distorted refinement models, is similar to decoding 900 including motion refinement using bilateral or composite matching with one or more distorted refinement models, unless otherwise described herein or clearly understood from the context. Encoding including motion refinement using bilateral or composite matching with one or more distorted refinement models can be performed by an encoder (such as...) Figure 4 The encoder 400 shown is implemented.
[0135] Decoding 900, including motion refinement using bilateral or composite matching with one or more twisted refinement models, includes refining from encoded bitstreams (such as...) Figure 4 The current block of the current frame is decoded from the encoded (compressed) bitstream 404 shown to generate reconstructed block data.
[0136] The decoding 900, which includes motion refinement using bilateral or composite matching with one or more warped refinement models, includes accessing a flag (at 910), determining whether to use warped refinement (at 915), obtaining the refined prediction block (at 920), generating the reconstructed block (at 930), and outputting (at 940). The decoding 900, which includes motion refinement using bilateral or composite matching with one or more warped refinement models, may include other aspects of the decoding, which are omitted in this description for simplicity.
[0137] The decoder accesses or otherwise obtains instructions from the encoded bitstream to decode the value of the current block (at 910) using motion thinning based on bilateral or composite matching with one or more twisted thinning models. For example, this value can be a flag, one or more bits, or one or more symbols, such as those included in a sequence parameter set, picture parameter set, frame header, block header, or another unit of encoded or compressed image or video data. Access flags or values are indicated by dashed borders to indicate that in some implementations, access to this value (at 910) can be omitted, excluded, or skipped.
[0138] Encoding that includes motion thinning using bilateral or composite matching with one or more warped thinning models may include signaling values (such as bits, flags, syntax elements, or symbols) in the encoded bitstream (such as in a sequence parameter set, picture parameter set, frame header, block header, or another element of the encoded or compressed image or video data) indicating whether motion thinning based on bilateral or composite matching with one or more warped thinning models is applied to the current block. In some implementations, signaling values in the encoded bitstream indicating whether motion thinning based on bilateral or composite matching with one or more warped thinning models is applied to the current block may be omitted or skipped.
[0139] Determine whether to use motion refinement utilizing bilateral or composite matching with one or more twisted refinement models (at 915).
[0140] For example, in decoding that includes motion refinement using bilateral or composite matches with one or more distorted refinement models, the decoder determines whether to use motion refinement utilizing bilateral or composite matches with one or more distorted refinement models (at 915). For example, the decoder may determine whether to use motion refinement utilizing bilateral or composite matches with one or more distorted refinement models to decode the current block based on or in response to a value indicated (obtained at 915) that suggests this motion refinement is being used (at 910). In some implementations, access to this value (at 910) may be omitted, and the determination (at 915) of whether to use motion refinement based on bilateral or composite matches with one or more distorted refinement models to decode the current block may be based on one or more rules or configurations. In some implementations, this value or flag may indicate that motion refinement based on bilateral or composite matching with one or more twisted refinement models may be available or enabled for decoding the current block, and may be determined (at 915) based on one or more rules or configurations to determine whether to use motion refinement based on bilateral or composite matching with one or more twisted refinement models to decode the current block.
[0141] In another example, in encoding that includes motion thinning using bilateral or composite matching with one or more twisted thinning models, the encoder determines whether to use motion thinning utilizing bilateral or composite matching with one or more twisted thinning models based on rate-distortion optimization (at 915).
[0142] Decoding 900, which includes motion refinement using bilateral or composite matching with one or more distorted refinement models, includes: such as in response to a value indicating that the current block is decoded using motion refinement based on bilateral or composite matching with one or more distorted refinement models, such as in response to determining (at 915) that the flag (obtained at 910) indicates that the current block is decoded using motion refinement based on bilateral or composite matching with one or more distorted refinement models, obtaining a refined prediction block (at 920).
[0143] Encoding for motion refinement using bilateral or composite matching with one or more twisted refinement models includes obtaining a refined prediction block (at 920).
[0144] Obtaining the refined prediction block (at 920) includes obtaining the twisted refined model (at 950), obtaining the coarse motion vector (at 960), obtaining the coarse prediction block and gradient (at 970), obtaining the block-based refined motion vector (at 980), and obtaining the sub-block-based refined translation motion vector (at 990). In some implementations, obtaining the sub-block-based refined translation motion vector (at 990) can be omitted, as indicated by the dashed border.
[0145] Obtaining the refined prediction block (at 920) includes, for example, obtaining the refined motion vector using a twisted refinement model and previously obtained reference frame data in cases where data explicitly indicating the signaling of refined motion vectors is not available or cannot be obtained in the encoded bitstream.
[0146] Encoding that uses bilateral or composite matching with one or more twisted thinning models for motion thinning can omit or exclude data that explicitly indicates the signaling of thinned motion vectors from the encoded bitstream.
[0147] Obtaining the twisted thinning model (at 950) involves deriving the twisted thinning model from the available twisted thinning models. In some implementations, the available twisted thinning models include a four-parameter scaling thinning model (with at least four parameters), a three-parameter scaling thinning model, and a four-parameter rotation thinning model (with at least four parameters). Other twisted thinning models can also be used.
[0148] Obtaining the twisted refinement model (at 950) involves refining the motion vector using a twisted refinement model or a combination of twisted refinement models (such as in bilateral or composite matching). Refining the motion vector may include using combinations of twisted refinement models, such as sequential combinations. Aspects and elements of code processing for motion refinement using bilateral or composite matching with twisted motion models (such as decoding 900 including motion refinement using bilateral or composite matching with one or more twisted refinement models or encoding including motion refinement using bilateral or composite matching with one or more twisted refinement models (not explicitly shown)) may be performed sequentially, in combination, or both (not explicitly described herein) unless otherwise described herein or otherwise clear from the context.
[0149] In some implementations, obtaining the twisted thinning model (at 950) includes identifying the target twisted motion pattern used to decode the current block and obtaining the twisted thinning model based on the target twisted motion pattern used to decode the current block.
[0150] For example, obtaining the twisted refinement model (at 950) may include identifying a six-parameter twisted motion pattern as the target twisted motion pattern for decoding the current block, and in response to identifying the six-parameter twisted motion pattern as the target twisted motion pattern for decoding the current block, identifying a four-parameter scaling refinement model as the twisted refinement model.
[0151] In another example, obtaining the warped thinning model (at 950) may include identifying a warped motion pattern with four or more parameters (such as a similarity warped motion pattern with four parameters, or another warped motion pattern with four or more parameters) as the target warped motion pattern for decoding the current block, and in response to identifying a warped motion pattern with four or more parameters as the target warped motion pattern for decoding the current block, identifying a three-parameter scaling thinning model as the warped thinning model.
[0152] In another example, obtaining the twisted thinning model (at 950) may include identifying a twisted motion pattern with four or more parameters (such as a similar twisted motion pattern with four parameters, or another twisted motion pattern with four or more parameters) as the target twisted motion pattern for decoding the current block, and in response to identifying a twisted motion pattern with four or more parameters as the target twisted motion pattern for decoding the current block, identifying a four-parameter rotation thinning model as a twisted thinning model.
[0153] Code processing that uses bilateral or composite matching with a twisted motion model for motion refinement (such as decoding 900 that includes motion refinement using bilateral or composite matching with one or more twisted refinement models or encoding that includes motion refinement using bilateral or composite matching with one or more twisted refinement models (not explicitly shown)) may include using a twisted refinement model (V x V y (Such as an affine model), this distorted and refined model indicates scaling in the x or horizontal direction (S x ), y-axis scaling or vertical scaling (S) y ), rotation angle ( ), shear factor (k) and translational motion ( , This can be represented as follows:
[0154] Rotation angle ( The scaling factor (k) and shearing factor (d0) can be linearly scaled with respect to the time distance between the current frame and each reference frame (d0, d1). The scaling factor (S) in the x or horizontal direction... x ) and scaling in the y or vertical direction (S)y It can be proportional to the time distance between the current frame and each reference frame (d0, d1) in an exponential manner.
[0155] The time distance between the reference frame and the current frame can be obtained (such as by calculating) the first time distance (d0) and the second time distance (d1) from the reference frame to the current frame, where d>0 indicates a forward reference (first in display order) and d<0 indicates a backward reference (last in display order).
[0156] Obtain the coarse motion vector (at 960).
[0157] Obtaining the coarse motion vector (at 960) involves obtaining the forward coarse motion vector ((a0,b0) or MV0) of the current block relative to a portion of a forward reference frame that is temporally preceding the current frame in display order, such as the portion obtained from the forward reference picture list (L0) that is temporally preceding the current frame. Figure 7 The current frame before 710 shown Figure 7 The first reference frame 720 is shown. The forward reference frame is a first time distance (d0) from the current frame.
[0158] Obtaining the coarse motion vector (at 960) involves obtaining the backward coarse motion vector ((a1,b1) or MV1) of the current block relative to a portion of a backward reference frame that is temporally located after the current frame in display order, such as the one obtained from the backward reference picture list (L1). Figure 7 The current frame after 710 shown Figure 7 The second reference frame 730 is shown. The backward reference frame is a second time distance (d1) from the current frame. In some implementations, the second time distance (d1) matches the first time distance (d0).
[0159] For simplicity, the forward coarse motion vector (a0, b0) and the backward coarse motion vector (a1, b1) can be collectively referred to as the coarse motion vector pair.
[0160] Based on the coarse motion vector pairs (obtained at 960), a coarse prediction block (at 970) can be obtained, for example, by using motion compensation.
[0161] Obtaining the prediction block (at 970) involves obtaining a forward coarse prediction block (P0(a0,b0) or P0) from a portion of the forward reference frame indicated by the forward coarse motion vector (a0,b0).
[0162] Obtaining the coarse prediction block (at 970) involves obtaining the backward coarse prediction block (P1(a1,b1) or P1) from a portion of the backward reference frame indicated by the backward coarse motion vector (a1,b1).
[0163] The gradient (at 970) is obtained from the coarse motion vector pair (obtained at 960).
[0164] Obtaining (such as generating) gradients (at 970) involves obtaining one or more gradients based on the coarse prediction block (P0, P1).
[0165] For example, obtaining the gradient (at 970) can include obtaining the forward horizontal spatial gradient (G) of the forward coarse prediction block (P0(a0,b0)) in the x or horizontal direction. 0x ), obtain the forward vertical spatial gradient (G) of the forward coarse prediction block (P0(a0,b0)) in the y or vertical direction. 1x ), obtain the backward horizontal spatial gradient (G) of the backward translation prediction block (P1(a1,b1)) in the x or horizontal direction. 1x ), and obtain the backward vertical spatial gradient G of the backward translation prediction block (P1(a1,b1)) in the y or vertical direction. 1y ).
[0166] In some implementations, bicubic interpolation can be used to obtain the gradient.
[0167] In some implementations, the prediction block size may be relatively large, such as greater than 16×16, and the size of the gradient array (Gx0, Gy0, Gx1, Gy1) can be reduced, for example, by using average pooling, which can reduce the complexity of parameter derivation. For example, a prediction block may have a first width (W) and a first height (H), and a gradient (Gx0, Gy0, Gx1, Gy1) with the first width (W) and the first height (H) may be obtained. The gradient (Gx0, Gy0, Gx1, Gy1) may be divided into appropriate sub-regions, which have a second width (w) less than or equal to the first width (w <= W) (where the second width (w) is a factor of the first width (W)) and a second height (h) less than or equal to the first height (h <= H) (where the second height (h) is a factor of the first height (H). The average gradient value within each sub-region may be obtained as the gradient (Gx0, Gy0, Gx1, Gy1) with the second width (w) (such as sixteen) and the second height (h) (such as sixteen).
[0168] Obtain the block-based refined motion vector (at 980).
[0169] In some implementations, obtaining the block-based refined motion vector (at 980) includes obtaining the block-based refined control point motion vector (Mv0) at the top left corner of the current block, the block-based refined control point motion vector (Mv1) at the top right corner of the current block, and the block-based refined control point motion vector (Mv2) at the bottom left corner of the current block.
[0170] Code processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) includes obtaining a refined prediction block (at 980) based on the refined motion vectors of the block.
[0171] In some implementations, the motion trajectory is continuous or constructively continuous, and obtaining the refined prediction block (at 980) includes obtaining the forward refined prediction block (P0(V)). x0 V y0 (P0') and the backward refined prediction block (P1(V)) x1 V y1 ) or P1′), which can be mirror symmetric with respect to the current frame.
[0172] In some implementations, the coarse motion vector pair ((a0,b0) and (a1,b1)) can approximate the motion vector of the upper left control point (V). x0 V y0 ) and the motion vector of the upper right control point (V x1 V y1 For example, within the defined threshold of the motion vectors of these control points, the difference in forward motion vectors (dV) can be obtained relative to the forward coarse motion vector (a0, b0). 0x dV 0y The difference in backward motion vectors (dV) can be obtained relative to the backward coarse motion vector (a1, b1). 1x dV 1y ), and approximately the forward refined prediction block (P0(V) x0 V y0 (or P0′) and backward refined prediction block (P1(V)) x1 V y1 P1) or P1′) can be represented as follows: .
[0173] The difference in motion vectors may be affected by the affine model described in this paper.
[0174] In an implementation where the refined prediction block is mirror-symmetric relative to the current frame, the mirror symmetry indicates the forward motion vector difference (dV). 0x dV 0y ) and backward motion vector difference (dV) 1x dV 1yThe affine model used. Rotation angle at the reference point ( The shearing factor (k) and the shearing factor (k) can be mirror images of zero (0), and the scaling factor (S) at the reference point x , S y They can be reciprocals of each other. An affine model can be achieved by minimizing P0(V). x0 V y0 ) and P1(V x1 V y1 The minimum is obtained by summing the squared errors between the two sides. This minimization can be solved by equations such as the Wiener-Hopf equation.
[0175] In some implementations, the sine, cosine, and reciprocal may be nonlinear, thus making it impossible to obtain a solution to the Wiener-Hopf equation.
[0176] In the example, the scaling factor (S) relative to the forward reference image list (L0) can be the result of adding one to the scaling parameter (s), which can be expressed as S=1+s, while the approximate value of the scaling factor (1 / S) relative to the backward reference image list (L1) can be the result of subtracting the scaling parameter (s) from one, which can be expressed as 1 / S=1-s.
[0177] In some implementations, the twisted thinning model is a three-parameter scaling thinning model, where the rotation angle ( The shearing factor (k) is zero or constructively zero, and the non-scaling twisted motion is constructively omitted, excluded, or ignored, resulting in a refined motion vector (at 980) that includes scaling factors for the x or horizontal and y or vertical directions. The three-parameter scaling refinement model can be represented as follows:
[0178] Forward refined prediction block (P0(V)) x0 V y0 The refined motion vector at the top left corner of (or P0′) can be represented as ( Forward refined prediction block (P0(V)) x0 V y0 The refined motion vector at the upper right corner of (or P0′) can be represented as ( ), backward refined prediction block (P1(V x1 V y1 The refined motion vector at the top left corner of (or P1′) can be represented as ( ), backward refined prediction block (P1(V x1 V y1 The refined motion vector at the upper right corner of (or P1′) can be represented as ( Furthermore, the three-parameter scaling refinement model can be represented as follows:
[0179] The three-parameter scaling refinement model includes three independent parameters ( , and Other values for obtaining, determining, or deriving the refined motion vector can be represented as follows:
[0180] The refined motion vector (at 980) determined using a three-parameter scaling and thinning model can be represented as follows:
[0181] In some implementations, the warped thinning model is a four-parameter scaling thinning model, which is similar to the three-parameter scaling thinning model, unless otherwise described herein or clearly understood from the context. For example, the scaling factor of the four-parameter scaling thinning model differs from that of the three-parameter scaling thinning model.
[0182] The refined motion vector (at 980) determined using a four-parameter scaling and thinning model can be represented as follows:
[0183] The four-parameter scaling and thinning model uses three thinned motion vectors. The four-parameter scaling and thinning model can be achieved through two parameters (…). and This is available in a six-parameter affine model, where the two parameters may be non-independent.
[0184] In some implementations, the twisted thinning model is a four-parameter rotational thinning model, and scaling and shearing can be omitted, excluded, or ignored. The four-parameter rotational thinning model can be represented as follows:
[0185] Obtaining the refined prediction block (at 980) involves determining, obtaining, or generating the refined prediction block (forward and backward) based on the refined motion vector.
[0186] In some implementations, obtaining the refined prediction block (at 920) may include multiple iterations of obtaining a coarse motion vector (at 960), obtaining the coarse prediction block and gradient (at 970), and obtaining the block-based refined motion vector (at 980), as indicated by the dashed direction line (995) from obtaining the block-based refined motion vector (at 980) to obtaining the coarse motion vector (at 960). In iterations following the first iteration, obtaining the coarse motion vector (at 960) involves using the refined motion vector obtained in the immediately preceding iteration as the coarse motion vector.
[0187] The maximum number of iterations, count, or cardinality (such as three) can be defined, either at the encoder and decoder without signaling, or signaled in the bitstream (such as in a sequence parameter set, picture parameter set, picture header, or slice header). The maximum number of iterations, count, or cardinality can be specific to the refinement model.
[0188] In some implementations, iteration may include determining whether to perform a subsequent iteration or the current iteration based on one or more defined conditions. For example, the absolute difference between the refined motion vector and the coarse motion vector in an iteration may be below a minimum difference threshold, and the encoder or decoder may determine to omit or exclude subsequent portions of the current iteration or the next iteration. The minimum difference threshold may be defined, such as at the encoder and decoder without signaling, or it may be signaled in the bitstream.
[0189] In some implementations, multiple candidate coarse motion vector pairs for the current block can be obtained (at 960), and code processing using motion refinement with bilateral or composite matching of twisted motion models (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) can be included to obtain candidate refined motion vectors on a per-candidate coarse motion vector pair. For simplicity, the candidate coarse motion vector pair for which corresponding candidate refined motion vectors have been obtained may be referred to herein as a processed candidate coarse motion vector pair. Code processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding (not explicitly shown) including motion refinement with bilateral or composite matching of one or more twisted refinement models) can include obtaining the best candidate refined motion vector as the refined motion vector from the processed candidate coarse motion vector pairs, such as by minimizing costs (such as the sum of absolute differences (SAD)).
[0190] In some implementations, obtaining the coarse motion vector (at 960) may include determining whether to use a candidate coarse motion vector pair. Determining whether to use a candidate coarse motion vector pair may include determining whether the maximum absolute difference between the candidate coarse motion vector pair and at least one processed candidate coarse motion vector pair is less than or equal to a minimum difference threshold, such as a one-pixel difference between the components of the motion vector. In response to determining that the maximum absolute difference between the candidate coarse motion vector pair and at least one processed candidate coarse motion vector pair is less than or equal to the minimum difference threshold, the candidate coarse motion vector pair is ignored, discarded, or skipped. The minimum difference threshold may be defined, such as at the encoder and decoder without signaling, or it may be signaled in the bitstream.
[0191] In some implementations, obtaining the block-based refined motion vector (at 980) may include determining whether the maximum distance between the refined motion vector and the coarse motion vector pair is greater than a maximum distance threshold, such as two pixels. In response to determining that the maximum distance between the refined motion vector and the coarse motion vector pair is greater than the maximum distance threshold, the refined motion vector may be omitted, skipped, excluded, or avoided. The maximum distance threshold may be defined, such as at the encoder and decoder without signaling, or it may be signaled in the bitstream.
[0192] although Figure 9 While not explicitly stated, in some implementations, code processing for motion refinement using bilateral or composite matching with a distorted motion model (such as decoding 900 including motion refinement using bilateral or composite matching with one or more distorted refinement models or encoding including motion refinement using bilateral or composite matching with one or more distorted refinement models (not explicitly stated)) may include determining whether to perform predictive refinement using optical flow, such as... Figure 8The diagram illustrates prediction refinement using optical flow 800. For example, in response to determining that prediction refinement using optical flow is disabled or unavailable based on code processing using motion refinement with bilateral or composite matching of distorted motion models (such as decoding 900 including motion refinement with bilateral or composite matching of one or more distorted refinement models or encoding including motion refinement with bilateral or composite matching of one or more distorted refinement models (not explicitly shown)), prediction refinement using optical flow can be omitted, skipped, or avoided. It can be determined based on defined rules, signaling in the bitstream, or a combination thereof: that prediction refinement using optical flow is disabled or unavailable based on code processing using motion refinement with bilateral or composite matching of distorted motion models (such as decoding 900 including motion refinement with bilateral or composite matching of one or more distorted refinement models or encoding including motion refinement with bilateral or composite matching of one or more distorted refinement models (not explicitly shown)).
[0193] In some implementations, defined weights (such as (½,½)) can be used as bidirectional prediction weights for code-based processing units corresponding to the refined motion vectors.
[0194] In some implementations, values such as those defined at the encoder and decoder without signaling or values sent by signals (such as false) can be used as flags, bits, or symbols to indicate illumination compensation corresponding to the refined motion vector.
[0195] although Figure 9 Not explicitly shown, but in some implementations, code processing using motion refinement with bilateral or composite matching of the twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) may include determining whether to perform translational motion decoder-side motion vector refinement based on bilateral matching (BM), such as Figure 7The illustrated translational motion decoder-side motion vector refinement 700 based on bilateral matching (BM) is shown. For example, in response to determining that translational motion decoder-side motion vector refinement based on bilateral matching (BM) is disabled or unavailable based on code processing using motion refinement with bilateral or composite matching with a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching with one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching with one or more twisted refinement models (not explicitly shown)), translational motion decoder-side motion vector refinement based on bilateral matching (BM) can be omitted, skipped, or avoided. The determination can be based on defined rules, signaling in the bitstream, or a combination thereof: depending on the code processing that uses motion refinement with bilateral or composite matching with a twisted motion model (such as decoding 900 which includes motion refinement with bilateral or composite matching with one or more twisted refinement models or encoding which includes motion refinement with bilateral or composite matching with one or more twisted refinement models (not explicitly shown)), translational motion decoder-side motion vector refinement based on bilateral matching (BM) is disabled or unavailable. For example, a flag, bit, or symbol can be sent to indicate one of the following: translational motion decoder-side motion vector refinement based on bilateral matching (BM) is used, and code processing using motion refinement with bilateral or composite matching of distorted motion models (such as decoding 900 including motion refinement with bilateral or composite matching of one or more distorted refinement models or encoding including motion refinement with bilateral or composite matching of one or more distorted refinement models (not explicitly shown)) is disabled or unavailable; code processing using motion refinement with bilateral or composite matching of distorted motion models (such as including...) is disabled or unavailable. Decoding 900 for motion refinement or encoding (not explicitly shown) that includes motion refinement using bilateral or composite matching with one or more distorted refinement models is used, and translational motion decoder-side motion vector refinement based on bilateral matching (BM) is disabled or unavailable; or code processing that utilizes motion refinement with bilateral or composite matching with distorted motion models (such as decoding 900 that includes motion refinement using bilateral or composite matching with one or more distorted refinement models or encoding (not explicitly shown) that includes motion refinement using bilateral or composite matching with one or more distorted refinement models) and translational motion decoder-side motion vector refinement based on bilateral matching (BM) are both used.
[0196] In some implementations, translational motion decoder-side motion vector refinement based on bilateral matching (BM) is used to obtain a translationally refined motion vector and a corresponding first matching cost; coding processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or coding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) is used to obtain a refined motion vector and a corresponding second matching cost; the minimum matching cost among the first and second matching costs is determined; and the motion vector corresponding to the minimum matching cost is used as the refined motion vector for the current block. In some implementations, the first matching cost may be matched with the second matching cost, and whether translational motion vector refinement based on bilateral matching (BM) is used at the encoder, decoder, or both, or whether code processing is used to refine motion vectors using bilateral or composite matching with a twisted motion model (such as decoding 900 that includes motion refinement using bilateral or composite matching with one or more twisted refinement models or encoding that includes motion refinement using bilateral or composite matching with one or more twisted refinement models (not explicitly shown)).
[0197] Code processing that utilizes motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 that includes motion refinement with bilateral or composite matching of one or more twisted motion models or encoding that includes motion refinement with bilateral or composite matching of one or more twisted motion models (not explicitly shown)) can be used in combination with one or more code processing modes that include bilateral or composite prediction.
[0198] For example, code processing using motion refinement with bilateral or composite matching of distorted motion models (such as decoding 900 including motion refinement with bilateral or composite matching of one or more distorted refinement models or encoding including motion refinement with bilateral or composite matching of one or more distorted refinement models (not explicitly shown)) can be used to refine motion vectors from merge candidates in which the code processing mode is a merge mode. In merge mode, motion information of the current block is obtained from or based on motion information of adjacent blocks (merge candidates) from one or more previously processed code processes in the absence of explicitly signaled motion information for the current block. Merge candidates can be indexed (e.g., in a merge candidate list), and index values of merge candidates for code processing of the current block can be signaled in the encoded bitstream.
[0199] In some implementations, the code processing mode can be a merging mode, the coarse motion vector obtained from the merging candidate can be a translational motion vector, the index value of the merging candidate (such as a regular merging candidate, a merging candidate of the motion vector difference merging mode, or both) can be signaled in the encoded bitstream, and code processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) can be used to refine the motion vector from the merging candidate of the merging mode corresponding to the merging candidate index value sent by the signal.
[0200] In some implementations, the code processing mode can be a merging mode, the coarse motion vector obtained from the merging candidate can be a translational motion vector, and the code processing using motion refinement with bilateral or composite matching with the twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching with one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching with one or more twisted refinement models (not explicitly shown)) can include obtaining a refined motion vector as a twisted motion vector based on the coarse motion vector obtained from the merging candidate.
[0201] In some implementations, the code processing mode can be a merging mode, the coarse motion vector obtained from the merging candidate can be a translational motion vector, and the code processing using motion refinement with bilateral or composite matching with the twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching with one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching with one or more twisted refinement models (not explicitly shown)) can include obtaining a refined motion vector as a translational motion vector based on the coarse motion vector obtained from the merging candidate, and the merging candidate and the corresponding refined motion vector can be identified as usable for subsequent twisted motion prediction.
[0202] In some implementations, the code processing mode can be a merging mode, the coarse motion vector obtained from the merging candidate can be a translational motion vector, and the code processing using motion refinement with bilateral or composite matching with a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching with one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching with one or more twisted refinement models (not explicitly shown)) can include obtaining a refined motion vector as a translational motion vector based on the coarse motion vector obtained from the merging candidate, the refined motion vector being identified as usable for subsequent motion prediction, and the coarse motion vector obtained from the merging candidate being identified as usable for subsequent motion prediction.
[0203] In some implementations, the code processing mode is a merge mode, where the coarse motion vector obtained from the merge candidates is a translational motion vector, and the refined motion vector is included in the merge candidate list, such as a sub-block merge candidate list or a twisted merge candidate list. In other implementations, the refined motion vector is a translational motion vector and is omitted or excluded from the merge candidate list, such as from a sub-block merge candidate list or a twisted merge candidate list.
[0204] In some implementations, the code processing mode is a merging mode (such as a sub-block merging mode or a twisted merging mode), the coarse motion vector obtained from the merging candidates is a twisted motion vector, and code processing using motion refinement with bilateral or composite matching with the twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching with one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching with one or more twisted refinement models (not explicitly shown)) can be used to refine the twisted coarse motion vector. In the example, the coarse motion vector is a four-parameter twisted motion vector, and the available twisted refinement models include a three-parameter scaling refinement model, a four-parameter rotation refinement model, or both.
[0205] In some implementations, the code processing mode is a merging mode (such as a sub-block merging mode or a twisted merging mode), the coarse motion vector obtained from the merging candidates is a twisted motion vector, and obtaining the twisted refinement model (at 950) involves obtaining the twisted refinement model from available twisted refinement models based on or in response to the twisted motion model corresponding to the coarse motion vector pair. For example, the twisted motion model corresponding to the coarse motion vector pair could be a four-parameter twisted motion model, and a three-parameter scaling refinement model could be used in the refinement. In another example, the twisted motion model corresponding to the coarse motion vector pair could be a six-parameter twisted model, and a four-parameter scaling refinement model could be used.
[0206] In some implementations, obtaining the refined prediction block (at 920) may include using a four-parameter rotation refinement model to obtain translation and rotation parameters. In some implementations, the scaling parameters may not be available when using a four-parameter rotation refinement model without employing a scaling refinement model (such as a three-parameter or four-parameter scaling refinement model).
[0207] In some implementations, obtaining the refined prediction block (at 920) may include obtaining translation and scaling parameters using a three-parameter scaling refinement model or a four-parameter scaling refinement model. Without using a rotation refinement model (such as a four-parameter rotation refinement model), it may be impossible to obtain rotation parameters using a three-parameter scaling refinement model or a four-parameter scaling refinement model.
[0208] In some implementations, obtaining the refined prediction block (at 920) may include obtaining rotation and scaling parameters using a multi-stage process or pipeline. In a multi-stage process or pipeline, obtaining the refined prediction block (at 920) may include obtaining translation and rotation parameters using a four-parameter rotation refinement model, a three-parameter scaling refinement model, or a four-parameter scaling refinement model. Furthermore, obtaining the refined prediction block (at 920) may include combining the translation and rotation parameters obtained using the four-parameter rotation refinement model with the translation and scaling parameters obtained using the three-parameter scaling refinement model or the four-parameter scaling refinement model.
[0209] For example, after obtaining translation and rotation parameters using a four-parameter rotation refinement model, gradients (different from those obtained using the four-parameter rotation refinement model) can be obtained (e.g., computed) based on the translation and rotation parameters obtained using the four-parameter rotation refinement model; translation and scaling parameters can be obtained using gradients (which are obtained based on the translation and rotation parameters obtained using the four-parameter rotation refinement model); translation and scaling parameters obtained using the four-parameter rotation refinement model can be combined with those obtained using the three-parameter scaling refinement model (or four-parameter scaling refinement model), and the combined parameters (predictions) can be further refined.
[0210] In some implementations, to reduce complexity, resource consumption, or both, compared to a multi-stage process or pipeline that combines parameters obtained using a four-parameter rotational refinement model with parameters obtained using a three-parameter scaling refinement model (or a four-parameter scaling refinement model), obtaining a twisted refinement model (at 950) may include obtaining a rotational and scaling refinement model as a twisted refinement model.
[0211] In rotation and scaling refinement models, shearing can be omitted, excluded, or ignored.
[0212] In the rotation and scaling refinement model, rotation parameters, scaling parameters (in the x and y directions), and translation parameters can be obtained.
[0213] In rotation and scaling refinement models, the rotation angle ( ) and the time distance between the corresponding reference frame and the current frame ( It is proportional in a linear manner.
[0214] In rotation and scaling refinement models, translation parameters ( and ) and the time distance between the corresponding reference frame and the current frame ( It is proportional in a linear manner.
[0215] In rotation and scaling thinning models, the scaling factor is related to the time distance between the corresponding reference frame and the current frame. The scaling factor is proportional in an exponential manner. It can be expressed as follows:
[0216] In rotation and scaling refinement models, the scaling parameter ( and Since the scaling factor is relatively small, it can be approximated as follows:
[0217] In the rotation and scaling refinement model, the rotation angle is relatively small, therefore the cosine (cos( It can be approximated as one (1), and its sine (sin( It can be approximated as This avoids non-linear operations.
[0218] In the rotation and scaling refinement model, the rotation parameters are obtained ( ).
[0219] The twisted model (x, y) → (x', y') (indicating the twisted motion from the horizontal component (x) and vertical component (y) of the coarse motion vector to the horizontal component (x') and vertical component (y') of the refined motion vector) can be represented as follows:
[0220] In rotation and scaling refinement models, the scaling parameter ( and ) and rotation parameters ( The second-order terms, which are relatively small and can be omitted, ignored, discarded, or discarded, can be represented as follows:
[0221] Equations 8 and 9 with unknown parameters ( The relationship is linear. The optical flow equation (such as Equation 1) may include Equations 8 and 9, and the parameters can be obtained by solving the optical flow equation using a five-dimensional inverse autocorrelation problem.
[0222] The optical flow equation for intensity (such as (I), such as luminance value) can be expressed as follows:
[0223] In some implementations, rotation parameters ( ) equals or constructively equals the horizontal scaling factor ( The horizontal scaling factor is equal to or constructively equal to the vertical scaling factor. The vertical scaling factor is equal to or constructively equal to zero. This is equivalent to the translational optical flow model.
[0224] In some implementations, rotation parameters ( ) equals or constructively equals zero ( This is equivalent to a four-parameter scaling refinement model.
[0225] In some implementations, rotation parameters ( ) equals or constructively equals zero ( And the horizontal scaling factor ( The scaling factor is equal to or constructively equal to the vertical scaling factor. ()( This is equivalent to a three-parameter scaling refinement model.
[0226] In some implementations, the scaling parameters can be equal or constructively equal. This allows the use of or solution of a four-parameter refinement model to obtain a rotation parameter, a scaling parameter, and two translation parameters.
[0227] To obtain the forward refined prediction block (i=0) and the backward refined prediction block (i=1), the twisting motion can be represented as follows:
[0228] In some implementations, obtaining the refined motion vectors includes obtaining the scaling parameters ( Multiply by the time distance between the reference frame indicated by the reference frame data (i=0 for obtaining the forward refined motion vector, i=1 for obtaining the backward refined motion vector) and the current frame. The result obtained is used as the first value. In some implementations, obtaining the refined motion vector includes obtaining the rotation parameters. Multiply by time distance ( The result obtained is used as the second value. In some implementations, obtaining the refined motion vector involves summing the following items to obtain the horizontal component of the refined motion vector. ): Horizontal translational motion value ( Multiply by time distance ( ) results ( ); and the result of the following operations: from the horizontal position value indicating the horizontal position of the current block in the current frame ( Multiply by the first value ( The result obtained by adding one () Subtract the vertical position value indicating the vertical position of the current block in the current frame from the value in the current block. Multiply by the second value ( The results obtained ( ),Right now( In some implementations, obtaining the refined motion vector involves summing the following items to obtain the vertical component of the refined motion vector (). ): The vertical translation motion value ( Multiply by time distance ( ) results ( ); Set the horizontal position value ( Multiply by the second value ( ) results ( ); and vertical position value ( Multiply by the first value ( The result of adding one The result obtained is .
[0229] In some implementations, the twist matrix is defined for predictions from the past to the future; therefore, the twist parameters can be used for forward reference, while the twist parameters for backward reference are unavailable. To obtain the twist parameters for backward reference, the inverse twist matrix can be obtained.
[0230] For example, the forward twist matrix (A) and its corresponding inverse matrix (A). (Where 1 / (1+αt) is approximated as 1-αt to avoid division operations) can be expressed as follows:
[0231] The mapping for the distortion prediction can be obtained through matrix operations, which can be represented as follows:
[0232] Based on the above closed-form solution, and given that cos(t) ~=1 and sin(t) ~=t, the backward translation parameter in the horizontal (x) direction is: In the vertical (y) direction, This avoids division operations.
[0233] In some implementations, the twisting motion of obtaining the forward refined prediction block (i=0) and the backward refined prediction block (i=1) can be expressed as shown in Equations 10 and 11.
[0234] In some implementations, the twisting motion to obtain the forward refined prediction block (i=0) can be expressed as shown in Equations 10 and 11, and the twisting motion to obtain the backward refined prediction block (i=1) can be expressed as follows:
[0235] In some implementations, obtaining the backward refined motion vector includes obtaining the scaling parameter ( Multiply by the second time distance between the backward reference frame and the current frame, as indicated by the reference frame data. The result obtained is used as the third value. In some implementations, obtaining the back-refined motion vector includes obtaining the rotation parameters. Multiply by the second time distance ( The result obtained is used as the fourth value. In some implementations, obtaining the backward refined motion vector involves obtaining the sum of the following items as the horizontal component of the backward refined motion vector ( The result of multiplying the following items: Second time distance ( Subtract the third value from one. The results obtained, and the horizontal translation motion values ( ) and the second time distance ( ), rotation parameters ( ) and vertical translational motion values ( The result of multiplying the results of the following operations; and the result of multiplying the results of the following operations: from the horizontal position value ( Multiply by the third value ( Subtract the vertical position value from the result obtained by adding one. Multiply by the fourth value ( The result obtained. In some implementations, obtaining the backward refined motion vector includes obtaining the vertical component of the backward refined motion vector by summing the following items ( The result of multiplying the following items: Second time distance ( Subtract the third value from one. The results obtained, and from the second time distance ( ), rotation parameters ( ) and horizontal translational motion values ( Subtract the vertical translational motion value from the result of the multiplication. The result obtained; the horizontal position value ( The result of multiplying by the fourth value; and the vertical position value ( Multiply by the third value ( The result obtained by adding one.
[0236] In some implementations, the twisting motion of obtaining the forward refined prediction block (i=0) can be expressed as shown in Equations 13 and 14, and the twisting motion of obtaining the backward refined prediction block (i=1) can be expressed as shown in Equations 10 and 11.
[0237] For implementations that omit or exclude obtaining (or otherwise use) the sub-block-based refined translation motion vector (at 990), determining the refined prediction block (at 980) includes obtaining a combination (such as the average) of the forward refined prediction block and the backward refined prediction block as the refined prediction block.
[0238] In some implementations, a refined translational motion vector based on sub-blocks is obtained (at 990).
[0239] Optical flow refinement includes bilateral or composite matching that uses two prediction blocks and their respective coarse composite motion vectors, and uses optical flow derivation to obtain fine (or refined) motion vectors, such as on each 8×8 or 4×4 sub-block of the prediction cell or block. In some implementations, optical flow refinement includes two translational motion parameters. In some implementations, optical flow refinement includes three or four warped motion parameters and two translational motion parameters, such as in four-parameter scaling refinement models, three-parameter scaling refinement models, four-parameter rotation refinement models, or rotation and scaling refinement models, as described herein.
[0240] In some implementations, code processing can be performed on a per-prediction-unit basis, where the encoder determines whether to use sub-block-based thinning or block-twist (such as affine) thinning. Determining whether to use sub-block-based thinning or block-twist (such as affine) thinning involves encoder searching, such as by comparing candidates obtained according to sub-block-based thinning with candidates obtained according to block-twist (such as affine) thinning. This has relatively high resource consumption and complexity and utilizes bandwidth to signal data in the encoded bitstream.
[0241] In some implementations, code processing that utilizes motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) (including obtaining sub-block-based refined translational motion vectors) can achieve similar or improved prediction quality, lower complexity, lower resource (processing) consumption, and without increasing bandwidth consumption, compared to code processing that includes determining whether to use sub-block-based refinement or block-based twisted (such as affine) refinement.
[0242] Code processing that utilizes motion refinement with bilateral or composite matching of twisted motion models (such as decoding 900 which includes motion refinement with bilateral or composite matching of one or more twisted refinement models, or encoding (not explicitly shown) which includes motion refinement with bilateral or composite matching of one or more twisted refinement models) includes block-based (rather than sub-block) twisted (or affine) refinement and sub-block-based translation refinement to improve code processing efficiency. In some implementations, as described herein, encoder complexity and resource consumption may be lower compared to encoders that use two or more models for search or evaluation.
[0243] Obtaining the sub-block-based refined translation motion vector (at 990) involves optical flow refinement using forward-refined and backward-refined prediction blocks, obtaining sub-block translation parameters on a per-8×8 or 4×4 sub-block basis. Optical flow refinement can be similar to... Figure 8 The prediction refinement using optical flow is shown in 800 unless otherwise described herein or otherwise clear from the context.
[0244] Code processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) includes obtaining a combination of block-based (or prediction unit) twisted motion parameters and sub-block-based translation parameters for sub-block prediction refinement.
[0245] Equations 10 and 11 use a four-parameter rotation and scaling model to describe the acquisition of the sub-block-based refined translational motion vector (at 990), where the rotation and scaling refinement model is a four-parameter warped model. In some implementations, equations 13 and 14 can be used. In some implementations, other warped motion refinement models can be used.
[0246] Obtaining the sub-block-based refined translational motion vector (at 990) involves using optical flow motion refinement to obtain the optical flow refined translational motion vector for each 4×4 or 8×8 sub-block. ),in Indicator sub-block k. Optical flow motion refinement is similar to... Figure 8 The translational optical flow refinement shown is unless otherwise described herein or otherwise clear from the context.
[0247] For implementations that include obtaining (or otherwise using) a sub-block-based thinned translation motion vector (at 990), obtaining the sub-block-based thinned translation motion vector (at 990) includes using twisted motion parameters of a block-based or prediction unit that incorporates the sub-block-based thinned translation motion vector to obtain a thinned prediction block (thinned prediction pixel) using twisted prediction.
[0248] For sub-blocks ( ), such as Figure 8 The top-left sub-block 812 shown performs warp prediction using rotation, scaling, and translation parameters based on the block or prediction cell (obtained at 980) and sub-block offset (based on the sub-block translation parameters) (obtained at 990). The refined prediction block is obtained by averaging the warp predictions at i=0 (forward) and i=1 (backward).
[0249] In the example, using a four-parameter rotation and scaling model, the sub-blocks ( The distortion prediction is based on the rotation angle ( scaling factor Translation parameters (in the horizontal (x) direction) and ( (in the vertical (y) direction), where i=0 and i=1.
[0250] In some implementations, obtaining (or otherwise using) the refined translational motion vector (at 990) based on sub-blocks involves performing sub-block predictions using twist predictions while omitting translation predictions. Using sub-block predictions leveraging twist predictions includes defining constraints on the twist parameters to perform low-complexity twist predictions using a two-step interpolation filter. In some implementations, the derived twist parameters may be inconsistent with or incompatible with the defined constraints, and translational motions may be obtained (e.g., computed) at the block center of each 4×4 region based on a twist or affine model, with translation predictions used on a per-4×4 sub-block basis.
[0251] In some implementations, obtaining (or otherwise using) the sub-block-based thinned translational motion vector (at 990) involves omitting the sub-block-based thinning of the chroma block, where the chroma block uses the distortion or affine parameters obtained for the corresponding iso-luminance block.
[0252] In some implementations, for luma block sizes larger than 8×8, optical flow refinement is performed on a per-8×8 luma sub-block basis, or for 8×8 luma block sizes, optical flow refinement is performed on a per-4×4 luma sub-block basis.
[0253] In some implementations, for luma block sizes larger than 8×8, refinement based on a combination model is performed on each 8×8 luma sub-block, and for co-located chroma blocks, refinement based on a combination model is performed on each 4×4 chroma sub-block.
[0254] In some implementations, for an 8×8 luma block size, combined (block-based (at 980) and sub-block-based (at 990)) refinement is performed on each 4×4 luma sub-block, and then combined refinement is performed on the corresponding 4×4 chroma block (instead of the sub-block), such as according to the constraint of performing distortion prediction at the 4×4 level, where the average of the refined motion vectors of the four sub-blocks is obtained and the average is combined with the distortion or affine model to obtain distortion prediction.
[0255] In some implementations, obtaining the gradient (at 970) involves expanding the pixel range used to derive the translation parameters. For an 8×8 sub-block, a 10×10 pixel region can be used, where each edge or boundary of the sub-block is expanded by one pixel, and the corresponding gradient values are available therein. For a 4×4 sub-block, a 6×6 pixel region can be used, where each edge or boundary of the sub-block is expanded by one pixel, and the corresponding gradient values are available therein. Using an expanded pixel region can improve code processing gain when combining a warp or affine model with a sub-block-based translation model, compared to using a pixel region based on the sub-block size.
[0256] Code processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding (not explicitly shown) including motion refinement with bilateral or composite matching of one or more twisted refinement models) includes using refined prediction blocks, refined motion vectors, or both to obtain or generate reconstructed block data (at 930) (such as by adding refined prediction blocks to decoded residual block data).
[0257] In encoding that includes motion refinement using bilateral or composite matching with one or more twisted refinement models, the encoder refines the motion from the corresponding coarse blocks ( Figure 9 (Not explicitly shown in the text) Subtract the refined prediction block to obtain or generate residual block data, and use the refined prediction block, refined motion vector, or both to obtain or generate reconstructed block data (at 930) (such as by adding the refined prediction block to the decoded residual block data).
[0258] Code processing using motion refinement with bilateral or composite matching of a twisted motion model (such as decoding 900 including motion refinement with bilateral or composite matching of one or more twisted refinement models or encoding including motion refinement with bilateral or composite matching of one or more twisted refinement models (not explicitly shown)) includes the decoder including the reconstructed block data in the reconstructed frame data of the current frame and outputting the reconstructed frame data (at 940). For example, the decoder may output the reconstructed frame to be presented to the user. In another example, the decoder may store the reconstructed frame data for subsequent code processing of another frame.
[0259] In coding that includes motion thinning using bilateral or composite matching with one or more twisted thinning models, the encoder includes the reconstructed block data in the reconstructed frame data of the current frame and stores the reconstructed frame data for subsequent coding of another frame.
[0260] As used herein, the terms “best,” “optimized,” “optimized,” or other forms thereof are relative to the appropriate context and do not indicate absolute theoretical optimization unless expressly stated herein.
[0261] As used herein, the term “set” refers to a distinguishable collection or grouping of zero or more distinct elements or members, which may be represented as a one-dimensional array or vector, unless explicitly described herein or otherwise clear from the context.
[0262] The terms “example” or “exemplary” are used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “example” or “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, the use of the terms “example” or “exemplary” is intended to present concepts in a specific manner. As used in this application, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise specified or clearly apparent from the context, “X comprises A or B” is intended to mean either of the natural inclusive arrangements. That is, if X comprises A; X comprises B; or X comprises both A and B, then “X comprises A or B” is satisfied under any of the foregoing examples. Additionally, the articles “a” and “an” as used in this application and the appended claims should generally be interpreted as meaning “one or more” unless otherwise specified or clearly apparent from the context. Furthermore, the use of the terms “embodiment” or “an embodiment” or “an implementation” or “an implementation” throughout the document is not intended to represent the same embodiment or implementation unless so described. As used herein, the terms “identify” and “identify”, or any variations thereof, include the use of… Figure 1 One or more of the devices shown may be selected, detected, calculated, searched, received, determined, established, obtained, or otherwise identified or determined in any way.
[0263] Furthermore, for the sake of simplicity, although the figures and descriptions herein may include a series of steps or stages, the elements of the methods disclosed herein may occur in various orders and / or simultaneously. Additionally, the elements of the methods disclosed herein may appear together with other elements not explicitly presented and described herein. Moreover, one or more elements of the methods described herein may be omitted from the implementation of the methods according to the disclosed subject matter.
[0264] The implementation of the transmitting computing and communication device 100A and / or the receiving computing and communication device 100B (as well as the algorithms, methods, instructions, etc. stored on and / or executed by these devices) can be implemented in hardware, software, or any combination thereof. Hardware may include, for example, a computer, intellectual property (IP) core, application-specific integrated circuit (ASIC), programmable logic array, optical processor, programmable logic controller, microcode, microcontroller, server, microprocessor, digital signal processor, or any other suitable circuitry. In the claims, the term "processor" should be understood to cover any of the foregoing hardware individually or in combination. The terms "signal" and "data" are used interchangeably. Furthermore, the parts of the transmitting computing and communication device 100A and the receiving computing and communication device 100B need not be implemented in the same manner.
[0265] Furthermore, in one implementation, for example, the transmitting computing and communication device 100A or the receiving computing and communication device 100B may be implemented using a computer program that, when executed, performs any of the corresponding methods, algorithms, and / or instructions described herein. Alternatively or alternatively, for example, a dedicated computer / processor may be utilized, which may include dedicated hardware for performing any of the methods, algorithms, or instructions described herein.
[0266] The transmitting computing and communication device 100A and the receiving computing and communication device 100B can be implemented, for example, on a computer in a real-time video system. Alternatively, the transmitting computing and communication device 100A can be implemented on a server, and the receiving computing and communication device 100B can be implemented on a device separate from the server (such as a handheld communication device). In this example, the transmitting computing and communication device 100A can use an encoder 400 to encode content into an encoded video signal and transmit the encoded video signal to the communication device. The communication device can then use a decoder 500 to decode the encoded video signal. Alternatively, the communication device can decode content stored locally on the communication device (e.g., content not transmitted by the transmitting computing and communication device 100A). Other suitable implementations of the transmitting computing and communication device 100A and the receiving computing and communication device 100B are available. For example, the receiving computing and communication device 100B can be a generally fixed personal computer instead of a portable communication device, and / or the device including the encoder 400 can also include the decoder 500.
[0267] Furthermore, the implementation may take the form of a computer program product accessible from, for example, a tangible computer-usable or computer-readable medium. A computer-usable or computer-readable medium may be any means that can, for example, tangibly contain, store, transmit, or transport the program for use by or in conjunction with any processor. The medium may be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable media may also be available.
[0268] It should be understood that the aspects can be implemented in any convenient form. For example, the aspects can be implemented by a suitable computer program that can be carried on a suitable carrier medium, which can be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). The aspects can also be implemented using suitable devices, which can take the form of a programmable computer running a computer program arranged to implement the methods and / or techniques disclosed herein. The aspects can be combined such that features described in the context of one aspect can be implemented in another aspect.
[0269] The above implementations have been described for ease of understanding of this application, but are not restrictive. Rather, this application covers various modifications and equivalent arrangements within the scope of the appended claims, which should be interpreted in the broadest possible sense to cover all such modifications and equivalent structures permitted by law.
Claims
1. A method comprising: Reconstructed block data is generated by decoding the current block of the current frame from the encoded bitstream, wherein decoding the current block includes: A refined prediction block for decoding the current block is obtained using bilateral matching, wherein obtaining the refined prediction block includes: The refinement model is obtained from the available twisted refinement models, wherein the available twisted refinement models include a four-parameter scaling refinement model, a three-parameter scaling refinement model, and a four-parameter rotation refinement model; In the absence of data explicitly indicating the refined motion vector in the encoded bitstream, the refined motion vector is obtained using the twisted refinement model and previously obtained reference frame data; and The refined motion vectors are used to generate refined prediction block data; The refined prediction block data is used to generate reconstructed block data; The reconstructed block data is included in the reconstructed frame data of the current frame; and Output the reconstructed frame data.
2. The method as described in claim 1, wherein, Obtaining the refined motion vector includes: Obtain coarse motion vector pairs, the coarse motion vector pairs including: Forward coarse motion vector; and Backward coarse motion vector; A forward coarse prediction block is obtained based on the forward coarse motion vector and forward reference frames from the forward reference image list, wherein the forward reference frames are at a first time distance from the current frame; A backward coarse prediction block is obtained based on the backward coarse motion vector and the backward reference frame from the forward reference image list, wherein the backward reference frame is a second time distance from the current frame; Obtain the gradient of the forward coarse prediction block; Obtain the gradient of the backward coarse prediction block; The refined motion vector is obtained based on the forward coarse prediction block, the backward coarse prediction block, the gradient of the forward coarse prediction block, and the gradient of the backward coarse prediction block.
3. The method of claim 2, wherein: Obtaining the gradient of the forward coarse prediction block includes: Obtain the horizontal spatial gradient of the forward coarse prediction block; and Obtain the vertical spatial gradient of the forward coarse prediction block; and Obtaining the gradient of the backward coarse prediction block includes: Obtain the horizontal spatial gradient of the backward coarse prediction block; and The vertical spatial gradient of the backward coarse prediction block is obtained.
4. The method of claim 2, wherein, The second time distance matches the first time distance.
5. The method of claim 1, wherein, Decoding the current block includes: Accessing the encoded bitstream indicates that the value of the current block is decoded using motion thinning based on bilateral matching and one or more twisted thinning models; Based on the value, determine whether to use motion thinning based on bilateral matching and one or more twisted thinning models to decode the current block; and In response to determining that the current block is decoded using motion refinement based on bilateral matching and one or more twisted refinement models, the refined motion vector is obtained.
6. The method of claim 1, wherein, Obtaining the refined motion vector includes: Obtain the first refined motion vector of the upper left corner of the current block; Obtain the second refined motion vector of the upper right corner of the current block; and Obtain the third refined motion vector of the lower left corner of the current block.
7. The method of claim 1, wherein, Obtaining the twisted and refined model includes: In response to identifying the six-parameter warped motion pattern as the target warped motion pattern for decoding the current block, the four-parameter scaling and thinning model is identified as the warped thinning model.
8. The method of claim 1, wherein, The distorted, refined model is identified by: In response to identifying a warped motion pattern with at least four parameters as the target warped motion pattern for decoding the current block, a four-parameter rotational refinement model is identified as the warped refinement model.
9. The method of claim 1, wherein, The distorted, refined model is identified by: In response to identifying a warped motion pattern with at least four parameters as the target warped motion pattern for decoding the current block, a three-parameter scaling refinement model is identified as the warped refinement model.
10. The method of claim 1, wherein, The available twisted thinning models include rotation and scaling thinning models.
11. The method of claim 1, wherein, Obtaining the refined motion vector includes obtaining the following combination as the refined motion vector: block-based twisted motion parameters and sub-block-based translational motion parameters obtained using the twisted refinement model.
12. A method comprising: Reconstructed block data is generated by decoding the current block of the current frame from the encoded bitstream, wherein decoding the current block includes: A refined prediction block for decoding the current block is obtained using bilateral matching, wherein obtaining the refined prediction block includes: A refined motion vector for decoding the current block is obtained using bilateral matching, wherein obtaining the refined motion vector includes: obtaining the refined motion vector using a rotation and scaling refinement model and previously obtained reference frame data when no data explicitly indicating the refined motion vector exists in the encoded bitstream; and The refined motion vectors are used to generate refined prediction block data; The refined prediction block data is used to generate reconstructed block data; The reconstructed block data is included in the reconstructed frame data of the current frame; and Output the reconstructed frame data.
13. The method of claim 12, wherein: The horizontal scaling factor is equal to the vertical scaling factor; and The rotation and scaling thinning model is a four-parameter twisted model.
14. The method of claim 12, wherein, Obtaining the refined motion vector includes: The first value is obtained by multiplying the scaling parameter by the time distance between the reference frame indicated by the reference frame data and the current frame; The result of multiplying the rotation parameter by the time distance is used as the second value; The following sums are obtained as the horizontal component of the refined motion vector: The result obtained by multiplying the horizontal translational motion value by the time distance; and The result of the following operation is: the result obtained by subtracting the result obtained by multiplying the vertical position value indicating the current block's horizontal position in the current frame by the second value from the result obtained by multiplying the horizontal position value indicating the current block's horizontal position in the current frame by the first value plus one; and The following sums are obtained as the vertical component of the refined motion vector: The result obtained by multiplying the vertical translational motion value by the time distance; The result obtained by multiplying the horizontal position value by the second value; and The result is obtained by multiplying the vertical position value by the result of adding one to the first value.
15. The method of claim 14, wherein, Obtaining the refined motion vector includes: Obtaining the forward-refined motion vector, wherein obtaining the forward-refined motion vector includes using a forward reference frame as the reference frame.
16. The method of claim 15, wherein, Obtaining the refined motion vector includes: Obtaining the backward refined motion vector, wherein obtaining the backward refined motion vector includes using a backward reference frame as the reference frame.
17. The method of claim 15, wherein, Obtaining the refined motion vector includes obtaining a backward refined motion vector, wherein obtaining the backward refined motion vector includes: The third value is obtained by multiplying the scaling parameter by the second time distance between the backward reference frame indicated by the reference frame data and the current frame; The result of multiplying the rotation parameter by the second time distance is used as the fourth value; The following sums are obtained as the horizontal component of the backward refined motion vector: The result of multiplying the following items: the second time distance, the result obtained by subtracting the third value from the first, and the result obtained by adding the horizontal translation value to the result of multiplying the second time distance, the rotation parameter, and the vertical translation value; and The result of the following calculation is: the result obtained by subtracting the vertical position value multiplied by the fourth value from the result obtained by multiplying the horizontal position value by the result obtained by adding one to the third value; and The following sums are obtained as the vertical component of the back-refined motion vector: The result of multiplying the following items: the second time distance, the result obtained by subtracting the third value from the first, and the result obtained by subtracting the vertical translation value from the result of multiplying the second time distance, the rotation parameter, and the horizontal translation value; The result obtained by multiplying the horizontal position value by the fourth value; and The result is obtained by multiplying the vertical position value by the result of adding one to the third value.
18. The method of claim 12, wherein, Obtaining the refined motion vector includes obtaining the following combination as the refined motion vector: block-based twisted motion parameters and sub-block-based translational motion parameters obtained using the rotation and scaling refinement model.
19. A method comprising: Reconstructed block data is generated by decoding the current block of the current frame from the encoded bitstream, wherein decoding the current block includes: A refined prediction block for decoding the current block is obtained using bilateral matching, wherein obtaining the refined prediction block includes: The refinement model is obtained from the available twisted refinement models, wherein the available twisted refinement models include a four-parameter scaling refinement model, a three-parameter scaling refinement model, a four-parameter rotation refinement model, and a four-parameter rotation and scaling model; In the absence of data explicitly indicating a refined motion vector in the encoded bitstream, the refined motion vector is obtained using the twisted refinement model and previously obtained reference frame data. Obtaining the refined motion vector includes obtaining a combination of the following as the refined motion vector: block-based twisted motion parameters and sub-block-based translational motion parameters obtained using the twisted refinement model; and The refined motion vectors are used to generate the refined prediction block data of the refined prediction block; The refined prediction block data is used to generate reconstructed block data; The reconstructed block data is included in the reconstructed frame data of the current frame; and Output the reconstructed frame data.
20. The method of claim 19: wherein, Obtaining the combination includes: The non-translational parameters of the combination are obtained from the block-based twisted motion parameters; and The combined translation parameters are obtained by adding the offset obtained from the translational motion parameters based on the sub-block to the translational motion parameters obtained from the twisted motion parameters based on the block.