Encoding method or apparatus based on indication of camera motion information

By determining the association between camera motion information and coding blocks in video encoding, utilizing the depth map and camera parameters provided by the game engine, and using a dedicated inter-frame prediction tool to encode the camera motion coding blocks, the problems of low coding efficiency and high latency in cloud game compression are solved, achieving a more efficient encoding process.

CN120660347APending Publication Date: 2025-09-16INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093665.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-12-08
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing video encoding methods suffer from low encoding efficiency and high latency when encoding 2D rendered videos for game engines, especially in cloud gaming compression, where intensive computing power is required, leading to increased latency.

Method used

By determining whether camera motion information is associated with the coding block, using the depth map and camera parameters provided by the game engine to perform video encoding and decoding, and using a dedicated inter-frame prediction tool to encode the camera motion coding block, unnecessary computing tool testing is reduced and coding efficiency is improved.

Benefits of technology

It reduces the delay in the encoding process, improves encoding efficiency, optimizes the encoder's decision-making process, and reduces the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660347A_ABST
    Figure CN120660347A_ABST
Patent Text Reader

Abstract

At least one method and apparatus for efficiently encoding or decoding a video are presented. For example, an indication of camera motion information is associated with a block in a current image, where the current image is part of a game engine 2D rendered video. Obtaining first motion information from the game engine, wherein the first motion represents motion information of a sample between the current image and the reference image; obtaining second motion information from a depth map and camera parameters provided by the game engine, wherein the second motion information represents motion information due to camera motion between the current image and the reference image; and associating an indication of camera motion information with the block when the first motion information and the second motion information approach. An indication of camera motion information is used in encoding or decoding of a block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of European Patent Application No. 22306848.7 filed on December 12, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0003] At least one of the present embodiments relates generally to a method or apparatus for video encoding or decoding, and more particularly, to a method or apparatus that includes determining whether an indication of camera motion information is associated with an encoding block. Background Art

[0004] To achieve high compression efficiency, image and video coding schemes typically employ prediction (including motion vector prediction) and transforms to exploit spatial and temporal redundancy in video content. Typically, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations. The difference between the original image and the predicted image (often expressed as a prediction error or prediction residual) is then transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded through the inverse process of entropy coding, quantization, transform, and prediction.

[0005] To achieve encoding gains, modern codec standards define increasingly complex tools and let the codec encoder decide the best tool to use. In the context of cloud gaming compression, minimizing latency is key. While recent encoders require intensive computational power, this introduces latency between the rendering of game content and its encoding.

[0006] Existing methods for encoding and decoding show some limitations in the field of encoding 2D rendered video for game engines. Therefore, there is a need to improve the existing technology. Summary of the Invention

[0007] The shortcomings and deficiencies of the prior art are solved and addressed by the general aspects described herein.

[0008] According to a first aspect, a method is provided. The method includes performing video encoding by the following steps: obtaining a coding block in a current image; determining whether an indication of camera motion information is associated with the coding block; and encoding the coding block based on the determination. For example, the current image is part of a 2D rendered video by a game engine. In a specific variant, for at least one sample in a coding block of a current image to be inter-coded relative to a reference image: obtaining first motion information from a game engine, the first motion information representing motion information of at least one sample between the current image and the reference image; determining second motion information based on a depth map and camera parameters published from the game engine, the second motion information representing motion information caused by camera motion between the current image and the reference image; and determining whether the indication of camera motion information is associated with the coding block based on a comparison between the first motion information and the second motion information.

[0009] According to another aspect, a second method is provided, which includes performing video decoding by: obtaining a coding block in a current image; decoding an indication of whether camera motion information is associated with the coding block; and decoding the coding block based on the indication.

[0010] According to another aspect, an apparatus is provided. The apparatus includes one or more processors, wherein the one or more processors are configured to implement the method for video encoding according to any of its variants. According to another aspect, the apparatus for video encoding includes means for implementing the method for video decoding according to any of its variants.

[0011] According to another aspect, another apparatus is provided. The apparatus includes one or more processors, wherein the one or more processors are configured to implement the method for video decoding according to any of its variants. According to another aspect, the apparatus for video decoding includes means for implementing the method for video decoding according to any of its variants.

[0012] According to another general aspect of at least one embodiment, camera parameters represent the position and characteristics of a game engine virtual camera that captures images of a game engine 2D rendered video.

[0013] According to another general aspect of at least one embodiment, values ​​in a depth map represent depths of samples of an image of a 2D rendered video by a game engine.

[0014] According to another general aspect of at least one embodiment, an indication of camera motion information associated with an encoded block is signaled from an encoder to a decoder.

[0015] According to another general aspect of at least one embodiment, there is provided an apparatus comprising: an apparatus according to any of the decoding embodiments; and at least one of: (i) an antenna configured to receive a signal comprising a video block; (ii) a bandlimiter configured to limit the received signal to a frequency band comprising the video block; or (iii) a display configured to display an output representing the video block.

[0016] According to another general aspect of at least one embodiment, there is provided a non-transitory computer-readable medium containing data content generated according to any of the described encoding embodiments or variations.

[0017] According to another general aspect of at least one embodiment, there is provided a signal comprising video data generated according to any of the described encoding embodiments or variants.

[0018] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.

[0019] According to another general aspect of at least one embodiment, there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to perform any of the described encoding / decoding embodiments or variants.

[0020] These and other aspects, features and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments, which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In the accompanying drawings, examples of several embodiments are illustrated.

[0022] Figure 1 Illustrated is a block diagram of an example apparatus in which various aspects of the embodiments may be implemented.

[0023] Figure 2 Illustrated is a block diagram of an embodiment of a video encoder in which various aspects of the embodiments may be implemented.

[0024] Figure 3 Illustrated is a block diagram of an embodiment of a video decoder in which various aspects of the embodiments may be implemented.

[0025] Figure 4 Illustrated are example texture frames and corresponding depth maps from a video game.

[0026] Figure 5Illustrated is an example architecture of a cloud gaming system.

[0027] Figure 6 A general encoding method according to a general aspect of at least one embodiment is illustrated.

[0028] Figure 7 A general decoding method according to a general aspect of at least one embodiment is illustrated.

[0029] Figure 8 An exemplary camera motion block segmentation example in VTM is illustrated.

[0030] Figure 9 An encoding method according to a general aspect of at least one embodiment is illustrated.

[0031] Figure 10 The diagram illustrates the principle of the pinhole camera model of the virtual camera in the cloud gaming system.

[0032] Figure 11 Illustration of the projection plane of a virtual camera in a cloud gaming system.

[0033] Figure 12 2D to 3D conversion is illustrated according to a general aspect of at least one embodiment. DETAILED DESCRIPTION

[0034] Various embodiments relate to video coding systems, wherein, in at least one embodiment, it is proposed to adapt video coding tools for cloud gaming systems. Different embodiments are proposed below, which introduce some tool modifications to improve coding efficiency and improve codec consistency when processing 2D rendered game engine video. Among others, encoding methods, decoding methods, encoding devices, decoding devices based on this principle are proposed. Although the present embodiments are presented in the context of a cloud gaming system, they can be applied to any system in which 2D video can be associated with camera parameters, such as video captured by a mobile device and information of a sensor that allows the position and characteristics of the camera of the device capturing the video to be determined. Depth information can be obtained from the sensor or other processing.

[0035] Furthermore, although principles related to a specific draft of the VVC (Universal Video Coding) or HEVC (High Efficiency Video Coding) specification or ECM (Enhanced Compression Model) reference software are described, the present invention is not limited to VVC or HEVC or ECM and can be applied to, for example, other standards and recommendations (whether pre-existing or developed in the future) and any such standards and recommendations (including VVC and HEVC and ECM). Unless otherwise indicated or technically excluded, the various aspects described in this application can be used alone or in combination.

[0036] The acronyms used herein reflect the current state of video coding developments and should therefore be considered as examples of nomenclature, which may be renamed at a later stage while still referring to the same technology.

[0037] Figure 1 A block diagram illustrating an example of a system in which various aspects and embodiments can be implemented is illustrated. System 100 can be implemented as a device including the various components described below, and is configured to perform one or more aspects of the various aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances and servers. The elements of system 100 can be implemented individually or in combination in a single integrated circuit, multiple ICs and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is coupled to other systems or to other electronic devices via, for example, a communication bus or by dedicated input ports and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects of the various aspects described in this application.

[0038] The system 100 includes at least one processor 110, which is configured to execute instructions loaded therein for implementing various aspects described in the present application. The processor 110 may include embedded memory, input and output interfaces, and various other circuits as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drive, and / or optical disk drive. As non-limiting examples, the storage device 140 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0039] The system 100 includes an encoder / decoder module 130, which is configured to process data to provide encoded video or decoded video, for example, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module(s) that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. In addition, the encoder / decoder module 130 may be implemented as a separate element of the system 100, or may be incorporated into the processor 110 as a combination of hardware and software as known to those skilled in the art.

[0040] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described herein may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of the various entries during execution of the processes described herein. Such stored entries may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.

[0041] In several embodiments, memory internal to the processor 110 and / or encoder / decoder module 130 is used to store instructions and to provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory (such as RAM) is used as working memory for video encoding and decoding operations, such as for HEVC or VVC.

[0042] Input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to: (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcaster; (ii) a composite input terminal; (iii) a USB input terminal; and / or (iv) an HDMI input terminal.

[0043] In various embodiments, the input device of block 105 has associated corresponding input processing elements as known in the art. For example, the RF part can be associated with elements applicable to the following: (i) selecting the desired frequency (also referred to as selecting a signal, or limiting the signal band to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select the signal band that (for example) can be referred to as a channel in certain embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF part of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF part can include a tuner that performs various functions in these functions, including, for example, down-converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or down-converting to baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and by filtering, down-conversion and filtering to desired frequency band again to perform frequency selection.Various embodiments rearrange the order of (and other) element described above, remove some elements in these elements, and / or add other elements of execution similar or different functions.Adding element can be included in and inserts element between existing element, for example, inserts amplifier and analog to digital converter.In various embodiments, the RF part comprises antenna.

[0044] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It is understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 110, as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 110, as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and the encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.

[0045] The various elements of system 100 may be provided within an integrated housing within which the various elements may interconnect and transfer data therebetween using a suitable connection arrangement 115 (e.g., an internal bus as known in the art, including an I2C bus, wiring, and printed circuit boards).

[0046] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0047] In various embodiments, a Wi-Fi network (such as IEEE 802.11) is used to stream data to the system 100. The Wi-Fi signals of these embodiments are received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to the system 100, which delivers data through the HDMI connection of the input block 105. Still other embodiments use the RF connection of the input block 105 to provide streaming data to the system 100.

[0048] System 100 can provide output signals to various output devices, including a display 165, speakers 175, and other peripherals 185. In various examples of embodiments, other peripherals 185 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are transmitted between system 100 and display 165, speakers 175, or other peripherals 185 using signaling, such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 using communication channel 190 via communication interface 150. In electronic devices (e.g., televisions), display 165 and speakers 175 can be integrated into a single unit with other components of system 100. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0049] For example, if the RF portion of input 105 is part of a separate set-top box, the display 165 and speaker 175 may alternatively be separate from one or more of the other components. In various embodiments where the display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0050] Figure 2 An example video encoder 200 is illustrated, such as a VVC (Versatile Video Coding) encoder. Figure 2 It is also possible to illustrate an encoder in which the VVC standard is improved or an encoder that adopts a technique similar to VVC.

[0051] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "encoding" or "encoded" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstruction" is used at the encoder side, while "decoding" is used at the decoder side.

[0052] Before being encoded, the video sequence may undergo a pre-encoding process (201), for example, applying a color transform to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components to obtain a more resilient to compression signal distribution (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.

[0053] In encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is segmented (202) and processed in units such as CUs. For example, each unit is encoded using intra or inter mode. When a unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) whether to use intra mode or inter mode to encode the unit and indicates the intra / inter decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the predicted block from the original image block.

[0054] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.

[0055] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. A loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).

[0056] Figure 3 FIGURE 3 illustrates a block diagram of an example video decoder 300. In the decoder 300, the bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs the same operations as described above. Figure 2 The decoding process is the reverse of the encoding process described in . The encoder 200 also typically performs video decoding as part of encoding the video data.

[0057] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoding information. Picture segmentation information indicates how the picture is segmented. Thus, the decoder can divide (335) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (355) to reconstruct the image block. The prediction block can be obtained (370) from intra-frame prediction (360) or motion compensated prediction (i.e., inter-frame prediction) (375). A loop filter (365) is applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380).

[0058] The decoded picture may further undergo post-decoding processing (385), such as an inverse color transform (e.g., from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.

[0059] A video encoding system (such as a cloud gaming server or a device with light detection and ranging (LiDAR) capabilities) can receive input video frames (e.g., texture frames) and possibly associated depth information (e.g., depth map) and / or motion information.

[0060] Figure 4The diagram illustrates an example texture frame 402 of a video game that can be extracted (e.g., directly) from a game engine that is rendering a game scene, along with a corresponding depth map 404, horizontal motion data 406, and vertical motion data 408. The depth map can be represented by a grayscale image that can indicate the distance between the camera and real-world objects. The depth map can represent the basic geometry of the captured video scene. The depth map can correspond to a texture picture of the video content and can include a dense monochrome picture of the same resolution as the luma picture. In an example, the depth map and the luma picture can have different resolutions.

[0061] Figure 5 An example architecture of a cloud gaming system is shown, where a game engine can run on a cloud server. The gaming system can render a game scene based on player actions. The rendered game scene can be represented as a 2D video including a set of texture frames. For example, a video encoder can be used to encode the rendered game engine 2D video into a bitstream. The bitstream can be encapsulated by a transport protocol and can be sent to the player's device as a transport stream. The player's device can decapsulate and decode the transport stream and present the decoded 2D video representing the game scene to the player. Figure 5 As illustrated in , additional information such as depth information, motion information, object ID, occlusion mask, camera parameters, etc. can be obtained from the game engine (e.g., as output of the game engine) and made available to the cloud server (e.g., the cloud's encoder) as prior information.

[0062] The information described herein (such as depth information, motion information, camera parameters, or a combination thereof) can be used to segment the rendered game engine 2D video in a video processing device (e.g., the encoder side of a video codec). Such segmentation of the rendered game engine 2D video can be used to simplify encoding decisions while still maintaining coding gain (e.g., compression gain). Thus, a high degree of flexibility in block representation of video in the compressed domain can be achieved, for example, in a manner that may have a limited increase in the rate-distortion optimization search space (e.g., on the encoder side).

[0063] At least some embodiments relate to a method for encoding or decoding a video, wherein two types of coding units CU (or coding blocks CB) can be used to improve motion estimation and motion compensation in inter-frame coding. The first type of CU includes camera motion coding units, where the motion is only due to the movement (or characteristics) of the virtual camera of the game engine. The second type of CU includes coding units, where the motion is due to the intrinsic motion of objects in the scene (or a combination of camera movement and object motion). Advantageously, this classified information can be used by the encoder to make coding decisions and optionally transmitted to the decoder.

[0064] Figure 6 A general encoding method 600 is illustrated according to general aspects of at least one embodiment. Figure 6 The block diagram partially represents the modules of the encoder or encoding method, such as in Figure 2 Modules implemented in an exemplary encoder.

[0065] according to Figure 6 In preliminary steps not shown above, the game engine generates a 2D video, at least one image (texture image) of the rendered game engine 2D video, and side information. According to non-limiting examples, the side information may include motion information relative to the game scene, depth information relative to the game scene, or camera parameters of a virtual camera capturing the game scene. The current image of the rendered game engine 2D video is segmented and processed in blocks (or units), such as coding blocks (CBs) or coding tree units (CTUs) (corresponding to higher-level image segmentation in the codec). Each block is encoded using, for example, intra or inter mode. In inter mode, motion estimation and compensation are performed. The encoding determines which of multiple inter modes to use for the block, and the inter decision is indicated by, for example, signaling motion information to obtain an inter-prediction block during decoding. According to a first step 610, a coded block is obtained from the segmentation process of the current image. According to a second step 620, an indication is determined as to whether the coded block has motion information related to camera motion. The determination step 620 allows the current coding block of the current image to be classified as a camera motion coding block or a non-camera motion coding block using an indication of camera motion information associated with the current coding block (true or false). Advantageously, the indication of camera motion information associated with the current coding block (true or false) is used in the encoding process 670 to assist at least one of the segmentation process or the inter-frame prediction process. The encoding process may be accelerated. According to an embodiment, the determination step 620 is based on a comparison of sample motion information provided by the game engine and sample motion information estimated based on camera parameters of a virtual camera in the game engine and the position of samples of a 2D image in a 3D game scene. For example, first motion information is obtained 630 from the game engine, the first motion information including a motion vector of the sample relative to a reference image (i.e., a game engine MV). Thus, the first motion information represents the motion information of the sample between the current image and the reference image. Second motion information is obtained 640 from the depth map and camera parameters also provided by the game engine. The second motion information includes a motion vector of the sample relative to the reference image (i.e., a camera MV), and the second motion information represents the motion information of the sample due to camera motion between the current image and the reference image. Reference is made below Figure 12A variant embodiment of calculating the second motion information is described. Then, in step 650, the first and second motion information are compared. A difference between the camera MV and the game engine MV can be calculated. If the difference is above a certain level, this means that the camera MV and the game engine MV are very different, and we assume that the motion of the sample is not only due to camera motion within the 3D game scene, but also due to motion of the sample as part of an object. Conversely, if the difference is below or equal to this level, the camera MV and the game engine MV are close, and we assume that the motion of the sample is due to camera motion within the 3D game scene. In that case, inter-frame motion prediction can be enhanced in the encoder using a dedicated inter-frame motion estimation tool or by selecting only a subset of inter-frame motion estimation tools suitable for this type of homogeneous motion. For example, expensive affine motion estimation can be removed from the subset of inter-frame tools tested, as it may be redundant for camera motion estimation that accounts for camera motion (including rotation in 3D space). In one variant, the level of the MV difference is non-zero, thereby accounting for distortion and / or accuracy in the calculation of the game engine MV and / or camera MV. In one variant, the MV difference is an absolute difference, and the level of the MV difference is above zero. According to the following reference Figure 9 Various embodiments are described, where MV differences are processed for one to all samples of a coding block and an indication of camera motion information is derived at the level of the coding block. It should be emphasized that the texture information of the coding block is not used when determining the indication of camera motion information, which is based only on information obtained from the game engine about the defined position and size of the block. Finally, in step 670, the camera motion information associated with the coding block is used to drive coding decisions for the segmentation process or the inter-coding process of the coding method. According to yet another variant, during the entropy decoding process, a CABAC context is derived for each value of the indication of camera motion associated with the coding block. According to another variant, the indication of camera motion information associated with the coding block is encoded and can be used during decoding.

[0066] Figure 7 A general decoding method 700 is illustrated according to general aspects of at least one embodiment. Figure 7 The block diagram partially represents a module of a decoder or decoding method, such as in Figure 3Modules implemented in an exemplary decoder of . According to a first step 710, a coding block to be decoded is obtained in a current image. According to a second step 720, an indication of whether the coding block has motion information related to camera motion is decoded from the bitstream. In step 730, the coding block is decoded based on the indication that the camera motion information is associated with the coding block. For example, an inter-frame prediction process suitable for camera motion content can be enabled and used to generate a prediction for the coding block. For example, a CABAC context is derived by determining the indication of camera motion associated with the coding block.

[0067] Various embodiments of a general encoding or decoding method are described below.

[0068] According to at least one embodiment, the encoder classifies blocks for encoding into camera motion CBs and non-camera motion CBs. Advantageously, the encoder can use this classification to make decisions. For example, it can decide to encode the camera motion CB using a dedicated inter-frame prediction tool, referred to herein as a camera motion tool. Without information about the camera motion CB, the encoder would have to test the performance of all inter-frame tools to decide which one to use. To achieve coding gains, modern codec standards define an increasing number of tools and let the encoder select the best one to use. Making this decision requires intensive computing power, which introduces a delay between the rendering of game content and its encoding. Unlike offline compression, minimizing this delay is key within the scope of cloud gaming compression. The duration between a player's action and its consequences should be minimized.

[0069] Figure 8 An exemplary camera motion block segmentation example in VTM is illustrated. The determination of camera motion blocks has been implemented in VTM (JVET reference software for VVC) on an exemplary 2D video rendered by a game engine. Figure 8 The result of this implementation is presented. The game scene represents the moving character, while the virtual camera is moving backward. The encoder splits the content into encoding blocks. By applying Figure 6 Based on the determination 620, coding blocks 810 for stationary areas of the 3D scene can be segmented, where the only apparent motion is due to the game engine's camera. The coding blocks 820 for the characters have appropriate motion and can be encoded using any of the most advanced inter-frame motion modes (regular, affine, merged, ...). The camera motion CBs 810 can be processed differently from the CBs representing moving characters. For example, a camera motion tool suitable for encoding camera motion can be applied to these CBs by default, thereby improving the inter-frame tool decision process and minimizing latency.

[0070] Figure 9The diagram illustrates an encoding method according to a general aspect of at least one embodiment. A game engine 910 provides a 2D rendered image 918 to the encoder. As described above, the game engine 910 also provides the encoder with a motion vector 916, a depth map 914, and camera parameters 912 for its virtual camera. According to this embodiment, these three types of information are used to determine the camera motion CB. The camera parameters 912 represent the characteristics and position of the game engine's virtual camera. They are provided for a reference image and for the current image to be encoded. The depth information represents the depth of the 3D points of the game content for each sample (or pixel) of the current frame after projection by the game engine's camera. Therefore, the depth information can be referred to as the depth map 914 of the current image. For the current sample, a motion vector 932 is calculated by the "Calculate Motion Vector" processing block 930. Because this vector is calculated using the depth and camera parameters, it represents the displacement of the current sample between the reference frame and the current frame due to camera motion (translation and / or rotation) or changes in camera characteristics (focal length, etc.). The game engine provides the motion vector 916 for the current sample. If the sample represents a 3D point that does not move in the 3D scene, the game engine motion vector is due only to the game engine's camera. In this case, the game engine motion vector is the same as the calculated motion vector. By comparing 940 the calculated motion vector 932 with the game engine's motion vector 916, it is possible to determine whether the vector is due only to the camera or to motion in the 3D scene. In one variant, to compare motion vectors 932 and 916, the difference between the game engine motion vector and the calculated motion vector is calculated for the samples in the block; and if the MV difference is above a certain level, the motion vectors are different, while if the MV difference is below or equal to that level, the motion vectors are considered to be the same. In one variant, the MV difference is an absolute difference that ends in a positive difference value.

[0071] In one variant, some errors due to calculation, rounding or quantization may be taken into account when comparing the calculated motion vectors with the vectors of the game engine. According to this variant, the level is higher than zero.

[0072] By analyzing 950 at least one game engine motion vector of the coded block, a decision is made as to whether the CB is a camera motion CB or not a camera motion CB (referred to as a non-camera motion CB). For the samples of the coded block, the sample-by-sample calculation of the motion vector and the comparison with the game engine motion vector can be processed in parallel or sequentially. Advantageously, once a decision can be made at the block level, the process can be terminated.

[0073] In one variant, all samples of the CB are processed, and if one or more game engine motion vectors in the coded block differ from corresponding calculated motion vectors, then the CB is not a camera motion CB. Thus, in response to determining that, for at least one sample of the block, the difference between the first motion information and the second motion information is above a certain level, an indication that the camera motion information is not associated with the coded block is determined.

[0074] In another variation, a CB may be considered a camera motion CB even if a small number of game engine motion vectors are not due to the game engine's camera. Thus, a number of non-camera motion samples is determined in the coding block, where the difference between the calculated motion vector and the game engine motion vector is above a certain level; and in response to determining that the number of non-camera motion samples is below a certain number of samples, an indication of camera motion information is associated with the coding block. In this variation, all samples in the coding block are processed. In this variation, the prediction of the CB by the dedicated camera motion tool may not be optimal, but may still benefit coding efficiency.

[0075] In yet another variant, the CB may be subsampled. The motion vector calculation and comparison with the motion vectors provided by the game engine are applied to a subset of the subsamples in the CB.

[0076] According to other variant embodiments, the indication of camera motion information is used to speed up the encoder. Instead of testing all inter-frame tools to evaluate their performance on the current camera motion CB, the encoder may decide to stop splitting 920. Thus, in this variant, in response to determining that the indication of camera motion is associated with a coding block, the splitting of the coding block into smaller coding blocks and the inter-frame motion estimation process for the smaller blocks are skipped. Furthermore, in this case, the encoder may decide to encode the camera motion CB using a dedicated camera motion inter-frame tool 964. However, such camera motion interaction tools are outside the scope of this disclosure. In another variant, since the nature of the motion in the CB is known (camera motion), the encoder may select a subset of inter-frame coding tools 962 during the encoding loop to speed up the encoding process. In another variant, the encoder may further split 920 the camera motion CB to find the optimal split, and the encoder may apply only the camera motion inter-frame tools to encode the smaller splits of the camera motion CB or the selected subset of inter-frame coding tools. Advantageously, the lack of competition between inter-frame tools saves encoding time.

[0077] On the other hand, if there are some moving objects in the current coding block, the CB may not be well predicted by the camera motion inter-frame tool. It can be encoded by another state-of-the-art inter-frame tool. According to another variant, the encoder may decide to further split the non-camera motion CB so that the camera motion area is covered by a smaller CB. Therefore, in this variant, in response to determining that the indication of camera motion information is not associated with the coding block, the coding block is split into smaller coding blocks, and it is determined whether the indication of camera motion information is iteratively associated with the smaller coding blocks.

[0078] In the following, at least one embodiment of the calculation of camera motion information is detailed, wherein the camera motion information allows determining an indication of camera motion information associated with a coding block.

[0079] like Figure 4 Neutralization Figure 5 The depth map shown in is a representation of the depth of a point belonging to a 2D projected image. However, the depth value in the depth map does not directly represent the depth of a 3D point in a 3D scene. When a 3D point is projected onto a 2D image, it is projected to the image location (x, y). In fact, mathematically there is a third coordinate, which is however discarded when considering the depth in a 2D image, but is stored in the Z buffer of the game engine. Advantageously, the game engine generates a third coordinate called "zbuff". The third coordinate called "zbuff" is used to calculate the camera motion vector.

[0080] According to at least one embodiment, the video to be encoded is composed of Figure 5 The 3D game engine generated in the cloud gaming system shown. Figure 10 The diagram illustrates the principle of the pinhole camera model of the virtual camera in the cloud gaming system. The 3D engine uses a virtual camera 1010 to project the 3D scene 1020 onto a plane 1030 to generate a 2D image. In the pinhole camera representation, the physical properties of the camera (focal length, sensor size, field of view, ...) can be used to calculate the projection matrix, which is the intrinsic matrix of the camera. This matrix defines the point Pi (x, y) in the 2D image, where the point P (X, Y, Z) in the 3D space is projected. In the following, the matrix is ​​referred to as the camera projection matrix, and the 2D image is referred to as the game engine 2D rendered image.

[0081] Figure 11The diagram illustrates the projection planes of a virtual camera in a cloud gaming system. In fact, unlike a physical camera that projects objects from 0 to infinity, a virtual camera of a game engine projects objects between two projection planes: a near plane 1110 and a far plane 1120. It means that these two planes represent the minimum and maximum depth for rendering: the near plane 1110 is usually mapped to a depth of 0, while the far plane 1120 is mapped to a depth of 1. However, according to a variant, the depth values ​​associated with the far and near planes can be represented oppositely. The camera projection matrix depends on the positions of the planes 1110, 1120. The way to construct this matrix is ​​not described here, but is well known in, for example, OpenGL or DirectX. Furthermore, the camera projection matrix performs its projection relative to its own coordinate system, which is as Figure 10 and Figure 11 As shown in . Since the camera is not placed at the origin of the 3D world coordinate system, another matrix is ​​needed to transform the position of the 3D point from the 3D world coordinate system to the camera coordinate system. The world to camera matrix 1130 is used to represent the rotation and translation of the camera relative to the 3D world coordinate system. The relationship between a 3D point in the 3D world of the game and its 2D position in the 2D projection image is defined by the world to camera projection matrix 1130 and the camera projection matrix. Conversely, the position of a 2D image point can be associated to a 3D world point by the inverse projection matrix and the camera to world matrix. In order to achieve the reconstruction of a 3D world point from a 2D projection image, a third image coordinate Zbuff representing a depth value is used for each sample of the 2D projection image. This is the depth information provided by the game engine, such as Figure 4 and Figure 5 As presented in .

[0082] Figure 12 2D to 3D transformation according to a general aspect of at least one embodiment is illustrated. Figure 12 As shown in , four matrices representing changes in coordinate system or projection / deprojection are used [PM1] -1 , [C1toW], [WtoC0] and [PM0] are used to calculate the camera motion vector. The matrix is ​​4x4. The world to camera matrix [WtoC0] and the projection matrix [PM0] represent the camera and its position relative to the reference image I0. The camera to world matrix [C1toW] and the inverse projection matrix [PM1] -1 Corresponds to the camera in the current image I1. Also uses zbuf1 information representing the depth of the current sample P1.

[0083] Compute the motion vector representing the motion of the current sample P1(x1, y1) in the current image I1 relative to the corresponding sample P0(x0, y0) in the reference image I0. In a first step, use the matrices representing the inverse projection and the camera-to-world matrix of the current image [PM1] respectively. -1, [C1toW], calculates the 3D position P(X, Y, Z) of P1 in the 3D world. In the second step, the 3D point is projected onto the reference image I0 using the matrices [WtoC0] and [PM0] that represent the projection into the reference image's camera C0 and the 2D reference image, respectively. The projection provides the point P0(x0, y0). The difference between the sample position in the current image and the reference image provides the motion vector.

[0084] In a variant embodiment of the first step, the transformation of a 2D image point P1 (x1, y1) of the current image to a point P in the 3D world is as follows. The coordinates of the point P1 (x1, y1) in the current image are expressed in normalized device coordinates (NDC) in the range [-1, 1]. The center of the image is [0, 0]. Additional depth information zbuf1 representing the depth of the current sample P1 in the 3D scene is added as an additional coordinate. The inverse projection matrix C1 [PM1] is then transformed. -1 And the C1 camera to world matrix [C1toW] is applied to the 2D+1 coordinates of the point to obtain coordinates in the 3D world. In order to multiply the projection matrix, the 2D image point P1 must be represented using a 4-dimensional vector in homogeneous coordinates:

[0085] P1=[x1,y1,zbuf1,1] T

[0086] Where x1 is the horizontal coordinate of P1 relative to the center of the image in the range [-1, 1], y1 is the vertical coordinate of P1 relative to the center of the image in the range [-1, 1], and zbuf1 is the depth coordinate of P1 relative to the 3D scene provided by the game engine in the range [-1, 1]. cam1 ,y cam1 , z cam1 , w cam1 ] T The intermediate 3D position of the representation is obtained from the deprojection of P1.

[0087] The inverse projection matrix of camera C1 [PM1] -1 is applied to the vector representing P1:

[0088] [x cam1 ,y cam1 ,z cam1 , w cam1 ] T =[PM1] -1 *[x1, y1, zbuf1, 1] T

[0089] To represent the true 3D Cartesian position, the fourth vector coordinate w cam1 should be equal to 1. Therefore, the four coordinates [x′ can1 , y′cam1 , z′ cam1 , w′ cam1 =1] T By dividing their values ​​by w cam1 To normalize:

[0090]

[0091] The camera C1 to world matrix [C1toW] is then applied to the intermediate normalized 3D Cartesian position to obtain the coordinates of the 3D point [X, Y, Z, W = 1] T :

[0092] [X,Y,Z,W=1] T =[G1toW]*[x′ cam1 y′ cam1 , z′ cam1 , w cam1 =1] T

[0093] In a variant embodiment of the second step, the transformation of a point P in the 3D world to a point P0(x0, y0) of the 2D reference image is obtained as follows. The world-to-camera C0 matrix [WtoC0] is applied to the 4-dimensional vector representing the coordinates of the 3D point p to generate an intermediate position [x cam0 ,y cam0 , z cam0 , w cam0 =1] T .

[0094] [x cam0 ,y cam0 ,z cam0 ,w cam0 =1] T =[WtoC0]*[X,Z,Y,W=1] T

[0095] The projection matrix [PM0] is then applied to the intermediate position to obtain the coordinates in the 2D plane of the reference image as a four-dimensional vector [x′0, y′0, z′0, w′0] T :

[0096] [x′0,y′0,z′0,w′0] T =[PM2]*[x cam0 ,y cam0 , z cam0 ,w cam0 =1] T

[0097] Normalize the fourth coordinate w to obtain the position of the 2D point P0 (x0, y0) in the reference image:

[0098]

[0099] Another variant embodiment calculated based on camera motion information determines a motion vector between P1 and P0. The motion vector MV represents the vector for the position P0 of the 3D point to be applied to the projection in the reference image to obtain the corresponding position P1 of the same 3D point projected in the current image. Thus, MV is obtained by the difference between the coordinates of points P1(x1, y1) and P0(x0, y0) using the following formula:

[0100] MV.x = x0 - x1 and

[0101] MV.y = y0 - y1.

[0102] According to one variant, if the motion vector provided by the game engine (i.e., game engine MV) is the same as the calculated motion vector, it is considered a camera motion vector caused only by changes in the characteristics and / or position of the camera between the reference frame and the current frame. In another variant, some small differences can be accepted as they are due to calculation, rounding, or quantization errors. Depending on the implementation, a threshold defining the acceptable difference between MVs can be set.

[0103] According to another variant, the process is applied to all pixels of the block or after subsampling of the block to save processing time. In one variant, when all motion vectors of the block are camera motion vectors, the block is considered a camera motion block. In another variant, even if some vectors are not camera motion vectors, the block can still be a camera motion block. For example, a simple threshold can define the proportion of non-camera motion vectors that can be accepted. In this case, the coding efficiency using camera motion inter-frame tools will not be optimal.

[0104] However, when encoding a block using dedicated camera motion inter-frame tools, some motion vectors may not be camera motion vectors without introducing strong distortion. On the contrary, when encoding using this tool, some non-camera motion vectors may imply strong distortion. In one variant, a threshold regarding acceptable distortion can be set instead of defining a threshold regarding the number of non-camera motion vectors.

[0105] According to another variant embodiment, an indication as to whether the block is a camera motion block is signaled from the encoder to the decoder. Since the original motion vector from the game engine is only available on the encoder side, it is impossible for the decoder to determine such a classification, and thus an indication as to whether the block is a camera motion block can be signaled from the encoder to the decoder.

[0106] A simple signaling strategy would be to encode a flag at the block level, such as cu_camera_motion_flag in bold in the table below, to indicate whether the block is a camera motion block. For example, when the block is a camera motion block, set the flag.

[0107]

[0108]

[0109]

[0110]

[0111] Table 1: Examples of camera motion signaling

[0112] According to another variant embodiment, if camera_motion_flag is set to 1, some modes are disabled, as depicted in Table 1. For example, camera_motion_flag is encoded first, and then pred_mode_flag is inferred to be inter, and the merge flag is inferred to be equal to 1, which means that all other modes (intra, IBC, amvp, etc.) are disabled. In another variant, a hierarchical signaling approach is adopted, where, for example, a single flag is signaled for a group of blocks that characterize a common characteristic, i.e., when all blocks within a group are camera motion blocks or not. One such example is signaling the flag at the CTU level, where all CU blocks within the CTU are camera motion blocks. In one such example, the flag is set at the CTU level when all blocks within the CTU are camera motion blocks. When the flag is not set, each block within the CTU must encode a flag indicating whether the block is a camera motion block. The same principle can be applied at the picture level, where, in one example, the flag is signaled at the picture level, where all CU blocks in the picture are camera motion blocks.

[0113] According to another variant embodiment, the camera motion flag is decoded and used at the decoder. For example, the camera motion flag can be used at the decoder to derive two different CABAC contexts for encoding the residual data. The different contexts are updated based on the camera motion flag. This can be beneficial in terms of compression efficiency, as the statistics of the residual in regions with camera motion differ from those in regions without camera motion. However, this approach does increase the memory required to store more CABAC tables.

[0114] Additional Examples and Information

[0115] Various methods are described herein, and each method in the method includes one or more steps or actions for realizing the method. Unless the correct operation method requires the steps or actions of a specific order, the order and / or use of specific steps and / or actions can be modified or combined. In addition, in various embodiments, terms such as "first", "second" etc. can be used to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding". Unless otherwise required, using such terms does not mean the sequencing of the operation to modification. Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can occur in, for example, before, during, or in the time period overlapping with the second decoding.

[0116] The various methods and other aspects described in this application can be used to modify modules, e.g. Figure 2 and Figure 3 , and the like. The video encoder 200 and decoder 300 shown in FIG. 1 show the segmentation and inter-frame prediction modules (202, 270, 275, 335, 375) of FIG. 1 . Furthermore, the present aspects are not limited to VVC or HEVC and can be applied, for example, to other standards and recommendations, as well as extensions of any such standards and recommendations. Unless otherwise indicated or technically excluded, the various aspects described in this application can be used alone or in combination.

[0117] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0118] Various implementations involve decoding. As used in this application, "decoding" can encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description and is considered to be well understood by those skilled in the art.

[0119] Various implementations involve encoding.In a similar manner to the discussion above regarding "decoding", "encoding" as used in this application may encompass all or part of a process performed on an input video sequence to produce an encoded bitstream, for example.

[0120] Note that the grammatical elements as used herein are descriptive terms. Therefore, they do not exclude the use of other grammatical element names.

[0121] The implementations and aspects described herein can be implemented as, for example, fragments of various information that can be transmitted or stored, such as, for example, syntax. This information can be packaged or arranged in a variety of ways, including, for example, those commonly found in video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or a slice header), or SEI message. Other approaches are also available, including, for example, those commonly found in system-level or application-level standards, such as placing the information in one or more of the following:

[0122] SDP (Session Description Protocol), a format for describing multimedia communication sessions, used for the purpose of session announcements and session invitations, e.g. as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transport;

[0123] DASH MPD (Media Presentation Description) descriptor, e.g. as used in DASH and transmitted over HTTP, a descriptor is associated with a representation or a set of representations to provide additional characteristics to the content representation;

[0124] RTP header extensions, e.g. as used during RTP streaming;

[0125] ISO Base Media File Format, e.g., as used in OMAF, and uses boxes, which are object-oriented building blocks defined by a unique type identifier and a length, also referred to as "atoms" in some specifications;

[0126] HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can, for example, be associated with a version or set of versions of the content to provide characteristics of the version or set of versions.

[0127] The implementations and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., a device or program). The device can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as computers, cellular phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the transmission of information between end users.

[0128] Reference to "one embodiment" or "an embodiment" or "an implementation" or "an implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations thereof in various places throughout this application are not necessarily all referring to the same embodiment.

[0129] Furthermore, the present application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0130] Furthermore, the present application may involve "accessing" various pieces of information. Accessing information may include, for example, one or more of: receiving information, retrieving information (e.g., retrieving information from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0131] Furthermore, the present application may involve "receiving" various pieces of information. As with "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of: accessing information or retrieving information (e.g., retrieving information from a memory). Furthermore, during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is generally involved in one way or another.

[0132] It is to be understood that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," use of any of the following " / ," "and / or," and "at least one of..." is intended to encompass selecting only the first-listed option (A), or only the second-listed option (B), or both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to encompass selecting only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first-listed option and the second-listed option (A and B), or only the first-listed option and the third-listed option (A and C), or only the second-listed option and the third-listed option (B and C), or all three options (A, B, and C). As will be apparent to one of ordinary skill in this and related arts, this can be extended to as many items as listed.

[0133] Furthermore, as used herein, the term "signaling" refers to, among other things, indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals a quantization matrix for dequantization. Thus, in embodiments, the same parameters are used at both the encoder and decoder sides. Thus, for example, an encoder can transmit (explicitly signal) specific parameters to a decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters along with other parameters, signaling can be used without transmitting (implicitly signaling) them, allowing only the decoder to know and select the specific parameters. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, information is signaled to a corresponding decoder using one or more syntax elements, flags, and the like. Although the verb form of the term "signaling" has been used above, the term "signal" can also be used herein as a noun.

[0134] As will be apparent to one of ordinary skill in the art, implementations can generate a variety of signals formatted to carry information that can, for example, be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.

Claims

1. A method comprising: Get the coding block in the current image; determining whether an indication of camera motion information is associated with the coding block; as well as The coding block is encoded based on the determination.

2. The method of claim 1, wherein the current image is part of a 2D rendered video by a game engine.

3. The method of claim 2 , wherein the determining comprises, for at least one sample in a coding block of a current picture to be inter-coded relative to a reference picture: Obtaining first motion information from a game engine, wherein the first motion information represents motion information of at least one sample between a current image and a reference image; determining second motion information based on the depth map and the camera parameters, wherein the second motion information represents motion information caused by camera motion between the current image and the reference image; as well as In response to determining that the difference between the first motion information and the second motion information is above a certain level, determining that the indication of camera motion information is not associated with the coding block. The method of claim 3 , wherein the level is greater than zero.

5. The method according to any one of claims 3 or 4, further comprising determining a number of non-camera motion samples in the coding block, wherein the non-camera motion samples have a difference between the first motion information and the second motion information above the level; and In response to determining that the number of non-camera motion samples is less than the number of samples, determining that an indication of camera motion information is associated with the coding block.

6. The method according to any one of claims 3 or 4, further comprising Subsampling the coding block; determining a number of non-camera motion samples among samples of the subsampled coding block, wherein the non-camera motion samples have a difference between the first motion information and the second motion information above the level; and In response to determining that the number of non-camera motion samples is below the number of samples, determining that an indication of camera motion information is associated with the coding block.

7. The method according to any one of claims 3 to 6, wherein the camera parameters represent the position and characteristics of a game engine virtual camera that captures images of the game engine 2D rendered video.

8. A method according to any one of claims 3 to 7, wherein the values ​​in the depth map represent the depth of samples of an image of a 2D rendered video by a game engine.

9. The method of any one of claims 1 to 8, further comprising encoding an indication of camera motion information associated with the coding block.

10. The method according to any one of claims 1 to 9, wherein encoding the coding block further comprises: In response to determining that the indication of camera motion is associated with the coding block, testing to split the coding block into smaller coding blocks is stopped.

11. The method according to any one of claims 1 to 9, wherein encoding the coding block further comprises: In response to determining that the indication of camera motion is not associated with the coding block, the coding block is partitioned into smaller coding blocks, and iteratively determining whether the indication of camera motion information is associated with the smaller coding blocks.

12. The method according to any one of claims 1 to 11, wherein encoding the coding block further comprises: In response to determining that the indication of camera motion information is associated with the coding block, a subset of inter-prediction coding tools is selected during the encoding process.

13. The method according to any one of claims 1 to 11, wherein encoding the coding block further comprises: A CABAC context is derived having an indication of camera motion associated with the coding block.

14. The method of claim 3, wherein determining second motion information based on the depth map and camera parameters, wherein the second motion information represents motion information caused by camera motion between the current image and the reference image, further comprises: Get the depth value of the current sample from the depth map; A 3D point position corresponding to the current sample in the current image is determined based on the position of the current sample in the current image, the depth value of the current sample, and based on the 2D to 3D transformation specified by the camera parameters. Determine the 2D point location corresponding to the current sample in the reference image based on the 3D-to-2D transformation specified by the camera parameters; as well as The second motion information of the current sample is determined as a displacement between a position of the current sample in the current image and a position of the current sample in the reference image.

15. A method comprising: Get the coding block in the current image; decoding an indication that camera motion information is associated with a coded block; as well as The coded block is decoded based on the indication.

16. The method of claim 15, wherein decoding the coded block further comprises: A CABAC context is derived having an indication of camera motion associated with the coding block.

17. An apparatus comprising a memory and one or more processors, wherein the one or more processors are configured to: Get the coding block in the current image; determining whether an indication of camera motion information is associated with the coding block; and The coding block is encoded based on the determination. The apparatus of claim 17 , wherein the current image is part of a 2D rendered video by a game engine.

19. The apparatus of claim 18, wherein the one or more processors are further configured to: for at least one sample in a coding block of a current picture to be inter-coded relative to a reference picture: Obtaining first motion information from a game engine, wherein the first motion information represents motion information of at least one sample between a current image and a reference image; determining second motion information based on the depth map and the camera parameters, wherein the second motion information represents motion information caused by camera motion between the current image and the reference image; as well as In response to determining that the difference between the first motion information and the second motion information is above a certain level, determining that the indication of camera motion information is not associated with the coding block.

20. The apparatus of claim 19, wherein the level is greater than zero.

21. The apparatus of any one of claims 19 or 20, wherein the one or more processors are further configured to: determining a number of non-camera motion samples in the coding block, wherein the non-camera motion samples have a difference between the first motion information and the second motion information above the level; and In response to determining that the number of non-camera motion samples is less than the number of samples, determining that an indication of camera motion information is associated with the coding block.

22. The apparatus of any one of claims 19 or 20, wherein the one or more processors are further configured to: Subsampling the coding block; determining a number of non-camera motion samples among samples of the subsampled coding block, wherein the non-camera motion samples have a difference between the first motion information and the second motion information above the level; and In response to determining that the number of non-camera motion samples is below the number of samples, determining that an indication of camera motion information is associated with the coding block.

23. The apparatus of any one of claims 19 to 22, wherein the camera parameters represent the position and characteristics of a game engine virtual camera that captures images of the game engine 2D rendered video.

24. An apparatus according to any one of claims 19 to 23, wherein the values ​​in the depth map represent the depth of samples of an image of a 2D rendered video by a gaming engine.

25. The apparatus of any one of claims 19 to 24, wherein the one or more processors are further configured to encode an indication of camera motion information associated with an encoding block.

26. The apparatus of any one of claims 19 to 25, wherein the one or more processors are further configured to, in response to determining that the indication of camera motion is associated with a coding block, stop testing to split the coding block into smaller coding blocks.

27. An apparatus according to any one of claims 19 to 25, wherein the one or more processors are further configured to: in response to determining that the indication of camera motion is not associated with the coding block, split the coding block into smaller coding blocks, and iteratively determine whether the indication of camera motion information is associated with the smaller coding blocks.

28. The apparatus of any one of claims 19 to 25, wherein the one or more processors are further configured to: in response to determining that the indication of camera motion information is associated with a coding block, select a subset of inter-frame prediction coding tools during encoding.

29. The apparatus of any one of claims 19 to 25, wherein to encode a coding block, the one or more processors are further configured to derive a CABAC context having an indication of camera motion associated with the coding block.

30. The apparatus of claim 19, wherein the one or more processors are further configured to: Get the depth value of the current sample from the depth map; A 3D point position corresponding to the current sample in the current image is determined based on the position of the current sample in the current image, the depth value of the current sample, and based on the 2D to 3D transformation specified by the camera parameters. Determine a 2D point location corresponding to the current sample in the reference image based on the 3D to 2D transformation specified by the camera parameters; and The second motion information of the current sample is determined as a displacement between a position of the current sample in the current image and a position of the current sample in the reference image.

31. An apparatus comprising a memory and one or more processors, wherein the one or more processors are configured to: Get the coding block in the current image; decoding an indication that camera motion information is associated with the coded block; and The coded block is decoded based on the indication.

32. The apparatus of claim 31 , wherein to decode a coding block, the one or more processors are further configured to derive a CABAC context having an indication of camera motion associated with the coding block, further comprising deriving a CABAC context having an indication of camera motion associated with the coding block.

33. A computer program product stored on a non-transitory computer-readable medium and comprising program code instructions for implementing the steps of the method according to at least one of claims 1 to 16 when executed by at least one processor.

34. A computer program comprising program code instructions for implementing the steps of the method according to at least one of claims 1 to 16 when executed by a processor.

35. A bitstream comprising information representing an encoded output generated according to one of the methods of any one of claims 1 to 14.

36. A non-transitory program storage device having encoded data representing an image block generated by the method of one of claims 1 to 14.

37. A non-transitory program storage device readable by a computer, tangibly embodying a program of instructions executable by the computer for performing the method of any one of claims 1 to 16.