Coding and decoding method or device for signaling indication of camera parameter
By introducing camera parameter syntax structure into the video bitstream, the problems of low encoding and decoding efficiency and high latency in cloud gaming systems are solved, and more efficient video encoding and decoding are achieved, which meets the needs of cloud gaming systems.
Patent Information
- Application Number
- CN202480007919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-20
- Filing Date
- 2024-01-16
- Publication Date
- 2025-08-22
AI Technical Summary
The existing video encoding methods have problems with low encoding and decoding efficiency and high latency in cloud gaming systems, especially when rendering game engine videos in 2D, it is difficult for existing tools to effectively utilize camera parameter information to improve encoding and decoding efficiency.
By introducing camera parameter syntax structures into the video bitstream, including position, orientation and characteristic information of the game engine virtual camera, screen-level encoding and decoding are performed, using this information to optimize the encoder's encoding decisions and improve the encoding and decoding efficiency.
It improves the efficiency of video encoding and decoding, reduces latency, enhances the consistency of the codec, and adapts to the needs of cloud gaming systems.
Smart Images

Figure CN120530631A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of European patent application 23305071.5 filed on January 20, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0003] At least one embodiment of the present invention generally relates to a method or apparatus for video encoding or decoding, and more specifically, to a method or apparatus that includes encoding / decoding camera parameters that provide information about the position, orientation, and characteristics of a game engine virtual camera that captures the footage. Background Art
[0004] To achieve high compression efficiency, image and video coding schemes typically employ prediction (including motion vector prediction) and transforms to exploit spatial and temporal redundancy in video content. Typically, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and the difference between the original image and the predicted image (usually represented as a prediction error or prediction residual) is then transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to entropy decoding, quantization, transform, and prediction.
[0005] To achieve codec gains, modern codec standards define increasingly complex tools and let the codec encoder decide the best tool to use. In the context of cloud gaming compression, minimizing latency is key, although recent encoders require intensive computational power that introduces latency between the rendering of game content and its encoding and decoding.
[0006] Existing methods for encoding and decoding show some limitations in the field of encoding and decoding 2D rendered video for game engines. Therefore, there is a need to improve the existing technology. Summary of the Invention
[0007] The shortcomings and deficiencies of the prior art are addressed and solved by the general aspects described herein.
[0008] According to a first aspect, a method is provided. The method includes performing video encoding on a syntax data element indicating whether a camera parameter syntax structure is present in a bitstream; and encoding at least one camera parameter syntax data structure representing camera parameters at a picture level in response to the presence of the camera parameter syntax structure, wherein the camera parameters provide information about a position, orientation, and characteristics of a game engine virtual camera that captures the picture.
[0009] According to another aspect, a second method is provided. The method includes: performing video decoding on a syntax data element indicating whether a camera parameter syntax structure is present in a bitstream; and in response to the presence of the camera parameter syntax structure, decoding at least one camera parameter syntax data structure representing camera parameters at a picture level, wherein the camera parameters provide information about a position, orientation, and characteristics of a game engine virtual camera that captured the picture.
[0010] According to another aspect, an apparatus is provided. The apparatus comprises one or more processors, wherein the one or more processors are configured to implement the method for video encoding according to any of its variations. According to another aspect, the apparatus for video encoding comprises means for implementing the method for video decoding according to any of its variations.
[0011] According to another aspect, another apparatus is provided. The apparatus comprises one or more processors, wherein the one or more processors are configured to implement the method for video decoding according to any of its variations. According to another aspect, the apparatus for video decoding comprises means for implementing the method for video decoding according to any of its variations.
[0012] According to another general aspect of at least one embodiment, there is provided an apparatus comprising the apparatus according to any of the decoding embodiments; and (i) an antenna configured to receive a signal comprising a video block, (ii) a band limiter configured to limit the received signal to a frequency band comprising the video block, or (iii) a display configured to display an output representing the video block.
[0013] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that contains data content generated according to any of the described encoding embodiments or variations.
[0014] According to another general aspect of at least one embodiment, there is provided a signal comprising video data generated according to any of the described encoding embodiments or variations.
[0015] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.
[0016] According to another general aspect of at least one embodiment, there is provided a computer program product comprising instructions that, when executed by a computer, cause the computer to implement any of the described encoding / decoding embodiments or variations.
[0017] These and other aspects, features and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments, which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In the drawings, several examples of embodiments are shown.
[0019] Figure 1 A block diagram is shown of an example apparatus in which various aspects of the embodiments may be implemented.
[0020] Figure 2 A block diagram of an embodiment of a video encoder is shown in which various aspects of the embodiments may be implemented.
[0021] Figure 3 A block diagram of an embodiment of a video decoder is shown in which various aspects of the embodiments may be implemented.
[0022] Figure 4 An example architecture of a cloud gaming system is shown in which various aspects of the embodiments may be implemented.
[0023] Figure 5 The principle of the pinhole camera model of the virtual camera in the cloud gaming system is shown.
[0024] Figure 6 Shown is the projection plane of a virtual camera in a cloud gaming system.
[0025] Figure 7 A general encoding method according to a general aspect of at least one embodiment is shown.
[0026] Figure 8 A general decoding method according to a general aspect of at least one embodiment is shown. DETAILED DESCRIPTION
[0027] Various embodiments relate to a video codec system, wherein in at least one embodiment, it is proposed to adapt video codec tools to a cloud gaming system. Different embodiments are proposed below, introducing some tool modifications to increase codec efficiency and improve codec consistency when processing 2D rendered game engine videos. Among them, encoding methods, decoding methods, encoding devices, decoding devices based on this principle are proposed. Although the present embodiments are presented in the context of a cloud gaming system, they can be applied to any system in which 2D video can be associated with camera parameters, such as video captured by a mobile device together with information of a sensor that allows determining the position and characteristics of the camera of the device capturing the video.
[0028] Furthermore, although principles are described in relation to a specific draft of the VVC (Universal Video Coding) or HEVC (High Efficiency Video Coding) specification or ECM (Enhanced Compression Model) reference software, aspects of the present invention are not limited to VVC or HEVC or ECM and may be applied, for example, to other standards and recommendations (whether pre-existing or developed in the future) and extensions of any such standards and recommendations (including VVC and HEVC and ECM). Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.
[0029] The abbreviations used here reflect the current state of video codec development and should therefore be considered examples of nomenclature that may be renamed at a later stage while still representing the same technology.
[0030] Figure 1 A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. System 100 can be implemented as a device including the various components described below, and is configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances and servers. The elements of system 100 can be implemented individually or in combination in a single integrated circuit, multiple ICs and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed over multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.
[0031] The system 100 includes at least one processor 110, which is configured to execute instructions loaded therein, for implementing various aspects described in the present application. The processor 110 may include embedded memory, input and output interfaces, and various other circuits well known in the art. The system 100 includes at least one memory 120 (e.g., volatile memory device and / or non-volatile memory device). The system 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drive, and / or optical disk drive. As a non-limiting example, the storage device 140 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0032] System 100 includes an encoder / decoder module 130, which is configured to process data to provide encoded video or decoded video, for example, and may include its own processor and memory. Encoder / decoder module 130 represents a module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. In addition, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated into processor 110 as a combination of hardware and software as known to those skilled in the art.
[0033] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described herein may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described herein. These stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0034] In several embodiments, memory internal to the processor 110 and / or encoder / decoder module 130 is used to store instructions and to provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or a storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations (such as HEVC or VVC).
[0035] Input to the elements of system 100 may be provided through various input devices as shown in block 105. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0036] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band, which in some embodiments may be referred to as a channel, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a stream of desired data packets. The RF section of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or down-converting to baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and by filtering, down-conversion and filtering to the frequency band of expectation again to perform frequency selection.Various embodiments rearrange the order of above-mentioned (and other) element, remove some in these elements, and / or add other elements of execution similar or different functions.Adding element can be included in and inserts element between existing element, for example, inserts amplifier and analog to digital converter.In various embodiments, the RF part comprises antenna.
[0037] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, in a separate input processing IC or processor 110, as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within processor 110, as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements (including, for example, the processor 110 and encoder / decoder 130 operating in conjunction with memory and storage elements) to process the data streams as needed for presentation on an output device.
[0038] The various components of the system 100 may be disposed within an integrated housing in which the various components may be interconnected and transmit data using suitable connection means 115 (e.g., an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board).
[0039] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented, for example, within a wired and / or wireless medium.
[0040] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received via a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments provide the streamed data to the system 100 using a set-top box that transmits the data via an HDMI connection of the input box 105. Still other embodiments provide the streamed data to the system 100 using an RF connection of the input box 105.
[0041] The system 100 can provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. In examples of various embodiments, the other peripheral devices 185 include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the system 100. In various embodiments, control signals are transmitted between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 100 via dedicated connections through the respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to the system 100 via the communication interface 150 using the communication channel 190. The display 165 and speakers 175 can be integrated into a single unit with the other components of the system 100 in an electronic device (e.g., a television). In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0042] For example, if the RF portion of input 105 is part of a separate set-top box, the display 165 and speaker 175 may alternatively be separate from one or more of the other components. In various embodiments where the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.
[0043] Figure 2 An example video encoder 200 is shown, such as a VVC (Versatile Video Coding) encoder. Figure 2 It is also possible to show an encoder in which the VVC standard is improved or an encoder that adopts a technique similar to VVC.
[0044] In this application, the terms "reconstruction" and "decoding" are used interchangeably, the terms "encoding" or "coding" are used interchangeably, and the terms "image", "picture" and "frame" are used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0045] Before being encoded, the video sequence may undergo pre-encoding processing (201), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components in order to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.
[0046] In encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is divided (202) and processed in units such as CUs. Each unit is encoded using, for example, intra or inter mode. When the unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) whether to use intra mode or inter mode to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. For example, a prediction residual is calculated by subtracting (210) the predicted block from the original image block.
[0047] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass the transform and quantization, i.e., encode the residual directly without applying a transform or quantization process.
[0048] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).
[0049] Figure 3 A block diagram of an example video decoder 300 is shown. In the decoder 300, a bitstream is decoded by decoder elements as described below. The video decoder 300 generally performs the same Figure 2 The encoding pass described in
[0044] is a decoding pass that is the inverse of the encoding pass described in
[0045] . The encoder 200 also typically performs video decoding as part of encoding the video data.
[0050] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoding information. The picture segmentation information indicates how the picture is segmented. The decoder can then divide (335) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined (355) with the prediction block to reconstruct the image block. The prediction block can be obtained (370) from intra-frame prediction (360) or motion compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).
[0051] The decoded picture may further undergo post-decoding processing (385), such as an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or performing an inverse remapping of the remapping process performed in the pre-encoding process (201). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0052] Figure 4 An example architecture of a cloud gaming system is shown, where a game engine can run on a cloud server. The gaming system can render a game scene based on player actions. The rendered game scene can be represented as a 2D video comprising a set of texture frames. For example, a video encoder can be used to encode the rendered game engine 2D video into a bitstream. The bitstream can be encapsulated by a transport protocol and can be sent to the player's device as a transport stream. The player's device can decapsulate and decode the transport stream and present the decoded 2D video representing the game scene to the player. Figure 4 As shown, additional information such as depth information, motion information, object ID, occlusion mask, or camera parameters may be obtained from the game engine (e.g., as output of the game engine) and provided to the cloud server (e.g., the cloud's encoder) as prior information.
[0053] The information described herein (such as depth information, motion information, camera parameters, or a combination thereof) can be used to improve the encoding of rendered game engine 2D video in a video processing device (e.g., the encoder side of a video codec). At least one embodiment involves signaling camera parameters (e.g., obtained from a game engine for a virtual camera in a 3D scene) as high-level syntax contained in a picture header. At least one embodiment proposes adding a camera parameter syntax data structure to transmit some information about the position, orientation, and characteristics of the game engine's virtual camera. Advantageously, these parameters are considered mandatory for the decoder. The camera parameters can be synchronized with the video frames, and if the camera parameters change, they can be updated for each frame. Advantageously, if these parameters remain the same for several consecutive frames, they can only be sent once.
[0054] In the following, at least one embodiment of a model of a virtual camera in a 3D scene capturing a picture is described in detail. According to at least one embodiment, the video to be encoded is composed of Figure 4 The 3D game engine generated in the cloud gaming system shown in FIG. Therefore, the frame is part of the video rendered in 2D by the game engine. However, the present principles are not limited to signaling virtual camera parameters in a cloud gaming system and can be applied to any parameters of a camera used for 2D video of a 3D scene.
[0055] Figure 5 The principle of the pinhole camera model of the virtual camera in the cloud gaming system is shown. The 3D engine uses a virtual camera 510 to project the 3D scene 520 onto a plane 530 to generate a 2D image. In the pinhole camera representation, the physical properties of the camera (focal length f, sensor size, field of view FOV...) can be used to calculate the projection matrix, which is the internal matrix of the camera. This matrix defines the point Pi (X, Y) in the 2D image, where the point P (X, Y, Z) in the 3D space is projected. In the following, this matrix is referred to as the camera projection matrix or internal matrix, and the 2D image is referred to as the game engine 2D rendered image. The position P (X, Y, Z) of the 3D object in the 3D world is known relative to the 3D world coordinate system. However, the camera performs its own translation relative to its own coordinate system ( Figure 5 Since the player can move around the 3D world, the virtual camera is not fixed to the origin of the 3D world. This means that before applying the camera projection of a 3D point P(X, Y, Z) (known relative to the 3D world coordinate system), the point is mapped into the camera coordinate system. Figure 5 The "world to camera" matrix EWtoC in performs this mapping. This matrix represents the rotation and translation of the camera relative to the 3D world coordinate system. The inverse transformation is Figure 5 The "camera to world" matrix ECtoW in . These two matrices are also called the camera's extrinsic matrices.
[0056] Figure 6 The projection planes of the virtual camera in the cloud gaming system are shown. In fact, unlike a physical camera that projects objects from a distance of 0 to infinity, the virtual camera of the game engine projects objects between two projection planes: a near plane 610 and a far plane 620. This means that these two planes represent the minimum and maximum depths for rendering: the near plane 610 is usually mapped to a depth of 0 and the far plane 620 is mapped to a depth of 1. However, according to one variant, the depth values associated with the far plane and the near plane can be represented inversely. The camera projection matrix depends on the position of the planes 610 and 620. The way to construct this matrix is described below. The calculation is usually performed using a 4×4 projection matrix in homogeneous coordinates. In one variant, the game engine provides 16 coefficients of the projection matrix for performing rendering.
[0057] Alternatively, you can extract some parameters from the game engine to calculate these matrix coefficients. For example, you can calculate the projection matrix as follows:
[0058]
[0059] Where FoV represents the vertical field of view.
[0060] H and V represent the width and height of the picture. Since H / V is the aspect ratio, in a variant, the aspect ratio is used as a camera parameter instead of the picture width and height.
[0061] Z N and Z F Indicates the positions of the near plane 610 and the far plane 620, such as Figure 6 The focal length f (modified when the camera zooms in or out) is shown in Figure 6 In fact, the focal length is directly related to the FoV parameter, such as Figure 5 Depending on the variant, one or another parameter can be used to calculate the matrix coefficients.
[0062] Zoom operations (modification of focal length or field of view) affect only two coefficients of the internal matrix:
[0063]
[0064] Modifications to the near and far plane positions only affect two other internal matrix coefficients:
[0065]
[0066] Therefore, in a variant embodiment, 2 sets of internal matrix coefficients are defined and signaled independently: one set related to the field of view and one set related to the projection plane.In yet another variant, the 2 sets can be updated separately.
[0067] According to yet another variant, instead of signaling the internal matrix coefficients, parameters representing rendering properties are transmitted to the decoder, such as FoV, near and far plane positions, etc. In this variant, the internal matrix coefficients are calculated by the decoder according to the equation of the matrix I.
[0068] As another alternative, the 16 matrix coefficients of the projection matrix PM (e.g. calculated by the encoder or provided by the game engine) are encoded separately in the picture header. Advantageously, this variant covers the general case of transmitting any projection matrix.
[0069] According to another variant, the game engine provides 16 coefficients of a world-to-camera matrix (called the extrinsic matrix). The virtual camera can be placed anywhere in the 3D world, i.e. not specifically at its origin, and according to any orientation. The translation and rotation of the virtual camera in the 3D world coordinate system are represented by the "world-to-camera" extrinsic matrix EWtoC or its inverse matrix the "camera-to-world" matrix ECtoW. In a variant, the game engine provides 16 extrinsic matrix coefficients for performing the rendering. The way these matrices are calculated is not described in detail herein, but the following 4×4 translation-scale-rotation (TSR) matrix indicates how it is composed.
[0070] TSR matrix composition:
[0071] The 9 Ri (R0, .., R8) coefficients represent the rotation of the camera around the 3 axes. In one variant, the Ri coefficients are calculated using the 3 rotation parameters of the camera (one for each axis). The 3Ti (T0, T1, T2) coefficients represent the translation of the camera relative to the 3D world origin. Camera translations are more frequent than rotations. Therefore, in one variant, 2 sets of extrinsic matrix coefficients are defined, one set related to rotations and one set related to translations. In yet another variant, the 2 sets can be updated separately, allowing the translation parameters to be modified independently of the rotation parameters.
[0072] As previously mentioned, the translation parameter provided by the game engine corresponds to the translation of the camera's system coordinates relative to the 3D world's system coordinates. This translation can be quite significant, as its maximum value corresponds to the size of the game's 3D world and is unknown (except to the game designer, who knows the size of their 3D world). Therefore, in a variant, the camera's translation relative to its previous position (indicating its displacement or position difference) is indicated rather than its absolute translation relative to the 3D world.
[0073] According to a general embodiment, at least one camera parameter syntax data structure representing camera parameters is signaled at the picture level, the camera parameters providing information about the position, orientation and characteristics of the game engine's virtual camera that captured the picture. Advantageously, the information about the position, orientation of the camera is provided as an internal set of matrix coefficients, while the information about the characteristics of the camera is provided by an external set of matrix coefficients. Thus, the game engine is able to project 3D points to 2D image points and vice versa. In order to project a 3D point to a 2D image, an (external) "world to camera" matrix and an (internal) camera projection matrix are required. Conversely, in order to find the 3D position of a 2D image point, an inverse camera projection matrix and a "camera to world" matrix are required. This "camera to world" matrix is the inverse of the "world to camera" matrix. As Figure 5 As shown, the four matrices (i.e., the inverse internal projection matrix [PM] -1 , extrinsic camera-to-world matrix [ECtoW], extrinsic world-to-camera matrix [EWtoC], and intrinsic projection matrix [PM]) represent the change of coordinate system or projection / deprojection is used to transform 3D points into 2D points and vice versa.
[0074] According to different variant embodiments described below with non-limiting examples of camera parameter semantics, all this information can be transmitted to the decoder as absolute values (real rotation and translation relative to the 3D world coordinate system) or relative values (e.g., the rotation and translation of the camera of picture n relative to its position at picture n-1).
[0075] Some parameters can be sent, while others can be inferred at the decoder side. For example, if the "world-to-camera" translation is transmitted, the "camera-to-world" translation is the opposite translation. If an intrinsic or extrinsic matrix is sent, a decoder with computational capabilities can, for example, compute the inverse matrix. In one variant, a flag is signaled to indicate that the inverse matrix is not signaled and needs to be computed at the decoder. In another variant, a single flag is used, for example, to indicate the presence or absence of the inverse intrinsic and / or extrinsic matrices in the bitstream. Alternatively, a separate flag is signaled to indicate the presence or absence of each element of the intrinsic or extrinsic matrix.
[0076] At least some embodiments relate to a method for encoding or decoding video using a syntax data element indicating the presence of a camera parameter syntax structure in a bitstream; wherein, in response to the presence of the camera parameter syntax structure, the bitstream also includes at least one camera parameter syntax data structure representing camera parameters at the picture level, wherein the camera parameters provide information about the position, orientation, and characteristics of a game engine virtual camera that captured the picture. Advantageously, this information about the camera parameters can be used by the encoder to make encoding decisions, improve encoding efficiency, and optionally transmitted to the decoder.
[0077] Figure 7A general encoding method 700 is shown in accordance with general aspects of at least one embodiment. Figure 7 The block diagram partially shows that for example Figure 4 or Figure 2 Modules of an encoder or encoding method implemented in an exemplary encoder.
[0078] according to Figure 7 In a preliminary step not shown in FIG, the game engine may generate at least one frame (texture image) of the 2D video, the rendered game engine 2D video, and auxiliary information. According to a non-limiting example, the side information may include camera parameters of a virtual camera capturing the game scene. According to a first step 710, a syntax data element (sps_camera_param_enabled_flag) is encoded that indicates whether a camera parameter syntax structure (gaming_camera_data) is present in the bitstream. This syntax data element may be signaled at the sequence level (e.g., in an SPS). In step 720, the syntax data element is tested. In response to the presence of a camera parameter syntax structure (yes), the method further includes a step 740 of encoding at the frame level at least one camera parameter syntax data structure (gaming_camera_data) representing camera parameters, wherein the camera parameters provide information about the position, orientation, and characteristics of the game engine virtual camera capturing the frame. In response to the absence of a camera parameter syntax structure (no), the method ends 730. According to another optional step, at least a portion of the frame of the 2D video is encoded using the camera parameters to improve encoding efficiency.
[0079] Figure 8 A general decoding method 800 is shown in accordance with general aspects of at least one embodiment. Figure 8 The block diagram partially represents a module of a decoder or decoding method, such as in Figure 4 or Figure 3 Modules implemented in an exemplary decoder of . Figure 8In a preliminary step not shown above, a bitstream encoding data representing at least one frame (texture image) of 2D video, the rendered game engine 2D video, and side information is received. According to a non-limiting example, the side information may include camera parameters for a virtual camera capturing the game scene. According to a first step 810, a syntax data element (sps_camera_param_enabled_flag) is decoded that indicates whether a camera parameter syntax structure (gaming_camera_data) is present in the bitstream. In step 820, the syntax data element is tested. In response to the presence of a camera parameter syntax structure (yes), the method further includes a step 840 of decoding at least one camera parameter syntax data structure (gaming_camera_data) representing camera parameters at the frame level, wherein the camera parameters provide information about the position, orientation, or characteristics of the game engine virtual camera capturing the frame. In response to the absence of a camera parameter syntax structure (no), the method ends 830. According to another optional step 850, the decoded camera parameters are used to decode at least a portion of the frame of the 2D video.
[0080] Various embodiments of the overall encoding or decoding method are described below .
[0081] According to at least one embodiment, a syntax data element indicates whether a camera parameter syntax structure is present in the bitstream. For example, an advanced flag (sps_camera_param_enabled_flag) is added to the sequence parameter set (SPS) to indicate that the decoder side uses camera parameters, which has the following syntax:
[0082]
[0083] According to at least one embodiment, at least one camera parameter syntax data structure is added at the picture level to represent camera parameters, wherein the camera parameters provide information about the position, orientation, and characteristics of the game engine virtual camera that captured the picture. For example, a syntax structure gaming_camera_data() is added to the picture header corresponding to the game engine's camera parameters, with the following syntax (the new syntax structure is shown in bold):
[0084]
[0085]
[0086]
[0087] According to a specific embodiment, a non-limiting example of a syntax structure is described in detail below. In this embodiment, the parameters are separated into internal and external parameters. In one variant, the external parameters provided to the decoder represent the absolute position of the game engine's camera relative to the 3D world system coordinates.
[0088] The syntax of game camera parameters can be defined as:
[0089]
[0090]
[0091]
[0092] In the syntax table defining "gaming_camera_data", the format of the camera parameters is indicated as FP for floating point. The way these parameters are encoded can depend on their use. For example, it can be 32-bit floating point (as commonly expressed in C), or double floating point encoded with 64 bits. It can also be a dedicated floating point encoding consisting of a sign, a mantissa, and an exponent, where the size of the mantissa depends on a fourth parameter that defines the precision. Finally, it can also be encoded as a 32-bit integer value instead of floating point.
[0093] The semantics of these parameters are:
[0094] intrinsic_param_focal_flag: indicates (when equal to 1) that the first set of intrinsic camera parameters related to the field of view is defined.
[0095] intrinsic_param plane_flag: Indicates (when equal to 1) that a second set of intrinsic camera parameters related to the game engine's far and near planes is defined.
[0096] extrinsic_param_rotation_flag: Indicates (when equal to 1) that the 9 extrinsic coefficients representing the camera rotation (world to camera extrinsic matrix) are defined.
[0097] extrinsic_param_translation_flag: Indicates (when equal to 1) that 3 extrinsic coefficients representing camera translation are defined.
[0098] inv_intrinsic_param_focal_flag: indicates (when equal to 1) that the first set of inverse intrinsic camera parameters related to the field of view is defined.
[0099] inv_intrinsic_param plane_flag: indicates (when equal to 1) that a second set of inverse intrinsic camera parameters related to the game engine's far and near planes is defined.
[0100] inv_extrinsic_param_rotation_flag: Indicates (when equal to 1) that 9 extrinsic coefficients representing camera rotation are defined.
[0101] inv_extrinsic_param_translation_flag
[0102] IF0: represents the first coefficient of the internal matrix of the game engine related to the horizontal field of view, named IF0 in the internal matrix I.
[0103] Alternatively, in another embodiment, this parameter may represent the aspect ratio camera parameter. Associated with the next parameter representing the field of view of the camera parameter, this parameter may be used to calculate the matrix coefficients.
[0104] IF1: represents the second coefficient of the internal matrix of the game engine related to the vertical field of view, named IF1 in the internal matrix I.
[0105] Alternatively, in another embodiment, this parameter may represent the field of view camera parameter. In association with the previous parameter representing the camera aspect ratio, this parameter may be used to calculate the two first matrix coefficients.
[0106] IP0: represents the first coefficient of the internal matrix of the game engine related to the plane position, named IP0 in the internal matrix I.
[0107] Alternatively, in another embodiment, this parameter may represent the near plane position Z N Camera parameters. Z represents the far plane position F This parameter is associated with the next parameter, which can be used to calculate the matrix coefficients.
[0108] IP1: represents the second coefficient of the internal matrix of the game engine related to the plane position, named IP1 in the internal matrix I.
[0109] Alternatively, in another embodiment, this parameter may represent the far plane position Z F Camera parameters. Z represents the near plane position N This parameter is associated with the previous parameter of , which can be used to calculate the matrix coefficients.
[0110] EWtoCR[j][k]: These 9 coefficients represent the rotation part of the world-to-camera matrix.
[0111] EWtoCT[k]: These three coefficients represent the translation part of the external world to camera matrix.
[0112] IIF0: represents the first coefficient of the inverse internal matrix of the game engine related to the horizontal field of view.
[0113] IIF1: Represents the second coefficient of the inverse internal matrix of the game engine related to the vertical field of view.
[0114] IIP0: represents the first coefficient of the inverse internal matrix of the game engine related to the position of the plane.
[0115] IIP1: represents the second coefficient of the inverse internal matrix of the game engine related to the plane position.
[0116] ECtoWR[j][k]: These 9 coefficients represent the rotation part of the extrinsic camera-to-world matrix.
[0117] ECtoWT[k]: These 3 coefficients represent the translation part of the extrinsic camera to world matrix.
[0118] According to at least one embodiment, an indication that the inverse inner and outer matrix coefficients are calculated at the decoder is added to the bitstream. Compared to the dedicated flag for the inverse matrix disclosed above, in this variation, a single flag is encoded to indicate the presence or absence of the inverse matrix. When the flag is set, all inverse matrices are calculated at the encoder and signaled to the decoder. When the flag is disabled, the inverse matrix is calculated at the decoder.
[0119] According to yet another variation, the camera parameter syntax data structure includes four extrinsic matrix coefficients (EWtoCq[k]) representing the quaternion rotation of the game engine's virtual camera in the world coordinate system. In this variation, the rotation portion of the external world-to-camera matrix is converted to four quaternion values, rather than the nine coefficients of the rotation matrix, and the quaternion values are signaled in the picture header. The quaternion values are then converted to a rotation matrix, such as described in JB Kuipers' "Quaternions and Rotation Sequences: A Primer with Applications to Orbits, Aerospace and Virtual Reality" (Chapter 5, Section 5.14, "Quaternions to Matrices," page 125).
[0120] The syntax of the game camera parameters defined above is modified as follows:
[0121]
[0122]
[0123] EWtoCq[k]: These 4 coefficients represent the quaternion representation of the rotation part of the external world to camera matrix.
[0124] ECtoWq[k]: These 4 coefficients represent the quaternion representation of the rotation part of the extrinsic camera to world matrix.
[0125] According to at least one embodiment, camera parameters provide information about the absolute position and orientation of the game engine's virtual camera that captured the frame. In one variation, the camera parameters provide information about the difference in position of the game engine's virtual camera between the frame and the previous frame, as well as the difference in orientation of the game engine's virtual camera between the frame and the previous frame. According to at least one embodiment, whether absolute or relative values are signaled is explicitly indicated by a flag, or derived from the frame type; for example, for an I frame, the value is absolute, while for a P frame, the value is relative. Therefore, in a variation, the camera parameter syntax structure also includes an indication that the camera parameters provide information about the difference in position of the game engine's virtual camera between the frame and the previous frame, as well as the difference in orientation of the game engine's virtual camera between the frame and the previous frame. In another variation, the camera parameters provide information about the difference in position of the game engine's virtual camera between the frame and the previous frame, as well as the difference in orientation of the game engine's virtual camera between the frame and the previous frame, as inferred from the frame type. In other words, the rotation and translation parameters may correspond to extrinsic matrix coefficients provided by the game engine (thus, relative to its 3D world coordinate system). In a variant embodiment, the provided extrinsic parameters do not represent the camera's position relative to the 3D world coordinate system, but rather represent the rotation and translation relative to the previous position of the virtual camera. Separate flags are encoded for rotation and translation to indicate whether the encoded parameters are relative to the previous frame or absolute. Based on this flag, the decoded rotation and translation parameters are processed on the decoder side. For example, if the flag is set, the current rotation parameter is determined by adding the decoded rotation parameter value to the previous frame's absolute rotation parameter value.
[0126] For example, you could add this subsection:
[0127]
[0128]
[0129] Here, RelativeWtoCR[j][k] and RelativeWtoCT[k] do not represent positions, but displacements relative to the previous position. Similarly, RelativeIF0 and RelativeIF1 do not represent focal lengths, but differences relative to the previous focal length value. This implies that the decoder requires the absolute values of the extrinsic and intrinsic parameters to perform decoding. This is controlled by sending an absolute value every few frames (usually every other I frame).
[0130] In a variant, a flag indicating whether the parameter is signaled relative to the previous frame or as an absolute value is inferred based on the frame type. For example, for the start of an I frame or a GDR frame, an absolute value is inferred.
[0131] In another variation, the above embodiments may be combined such that for a given frame n, the transmitted extrinsic parameters are relative to the 3D world position and for frame n+1, the transmitted extrinsic parameters are relative to the previous frame n.
[0132] In yet another variation of this embodiment, for example for an I frame of a new GOP, the camera position may be automatically initialized to 0 (no rotation, no translation). The transmission parameters for the next frame correspond to the displacement relative to the previous frame.
[0133] In the above embodiments, we considered only one camera. In some applications, two or more cameras may be required. According to at least one embodiment, a syntax data element indicates whether one or more instances of a camera parameter syntax data structure are present in the bitstream. For example, to render 3D stereoscopic content, two cameras are required. Instead of sending camera parameters for a single camera as described above, a list of camera parameters is sent. These parameters can be absolute, or one camera can be considered a master camera with absolute parameters, with the other camera parameters provided relative to that master camera.
[0134] Since intrinsic camera parameters do not change frequently, signaling flags and relative values in the picture header may not be effective. According to at least one embodiment, at least one camera parameter syntax data element is added at the sequence level, wherein the at least one camera parameter syntax data element at the sequence level includes a first set of intrinsic camera parameters related to the field of view and the far and near planes. Advantageously, the at least one camera parameter syntax data structure at the picture level includes a set of extrinsic matrix coefficients representing the rotation or translation of the game engine virtual camera. In yet another variation, a syntax data element is added at the sequence level, indicating an update to at least one camera parameter syntax data structure at the sequence or picture level representing the camera parameters. This embodiment proposes sending the intrinsic camera parameters relative to 3D world coordinates in the SPS syntax. Next, a flag indicating whether a change has occurred is encoded in the picture header. When the flag is set, only the difference is signaled. The intrinsic camera parameters are updated in the next SPS. As a variation of this embodiment, the intrinsic camera parameters can alternatively be signaled in a picture parameter set (PPS).
[0135] Additional Examples and Information
[0136] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined. In addition, terms such as "first", "second" and the like can be used to modify elements, components, steps, operations, etc. in various embodiments, for example, "first decoding" and "second decoding". Unless otherwise required, the use of these terms does not mean the sequencing of modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can, for example, be performed before, during, or in a time period overlapping with the second decoding.
[0137] The various methods and other aspects described in this application can be used to modify modules, e.g., Figure 2 and Figure 3 . In addition, aspects of the present invention are not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, as well as extensions of any such standards and recommendations. Unless otherwise specified or technically excluded, the aspects described in this application can be used alone or in combination.
[0138] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0139] Various implementations involve decoding. As used in this application, "decoding" can include, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description and is believed to be fully understood by those skilled in the art.
[0140] Various implementations involve encoding.In a manner similar to the discussion above regarding "decoding," "encoding," as used in this application, may include all or part of a process performed on an input video sequence, for example, to produce an encoded bitstream.
[0141] Note that the syntax elements used here are descriptive terms. Therefore, they do not exclude the use of other syntax element names.
[0142] The implementations and aspects described herein can be implemented as, for example, various pieces of information, such as syntax, that can be sent or stored. This information can be packaged or arranged in various ways, including, for example, those commonly found in video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or a slice header), or SEI message. Other approaches are also available, including, for example, those commonly found in system-level or application-level standards, such as placing the information in one or more of the following:
[0143] SDP (Session Description Protocol), a format for describing multimedia communication sessions for the purpose of session announcement and session invitation, such as that described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transport;
[0144] DASH MPD (Media Presentation Description) descriptors, such as those used in DASH and delivered over HTTP, are associated with a representation or a series of representations to provide additional features to the content presentation;
[0145] RTP header extensions, such as those used during RTP streaming;
[0146] ISO Base Media File Format, such as that used in OMAF, and using boxes, also referred to as "atoms" in some specifications, which are object-oriented building blocks defined by a unique type identifier and length;
[0147] HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can, for example, be associated with a version or a series of versions of the content to provide characteristics of the version or series of versions.
[0148] The implementations and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., a device or program). For example, a device can be implemented with appropriate hardware, software, and firmware. The method can be implemented in, for example, a device such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.
[0149] Reference to "one embodiment" or "an embodiment" or "an implementation" or "an implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations thereof in various places throughout this application are not necessarily all referring to the same embodiment.
[0150] Additionally, the present application may refer to “determining” various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, and retrieving information from a memory.
[0151] Furthermore, the present application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, and estimating information.
[0152] Additionally, this application may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information can include, for example, one or more of accessing information and retrieving information (e.g., from a memory device). Furthermore, during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is often involved in one way or another.
[0153] It should be understood that use of any of the following “ / ,” “and / or,” and “at least one of” (e.g., in the case of “A / B,” “A and / or B,” and “at least one of A and B”) is intended to encompass selection of only the first-listed option (A), or only the second-listed option (B), or both options (A and B). As a further example, in the case of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to include selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or all three options (A, B, and C). This can be extended to multiple items listed, as will be apparent to one of ordinary skill in this and related arts.
[0154] Furthermore, as used herein, the term "signaling" specifically refers to indicating certain information to a corresponding decoder. For example, in certain embodiments, an encoder signals a quantization matrix for dequantization. Thus, in one embodiment, the same parameters are used on both the encoder and decoder sides. Thus, for example, an encoder may send specific parameters to a decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters along with other parameters, signaling may be used instead of sending them (implicit signaling), simply allowing the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. While the foregoing refers to the verb form of the term "signaling," the term "signaling" may also be used herein as a noun.
[0155] As will be apparent to one of ordinary skill in the art, implementations can generate various signals that are formatted to carry information that can, for example, be stored or transmitted. The information can include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal can be formatted to carry a bit stream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.
Claims
1. A method comprising: encodes a syntax data element (sps_camera_param_enabled_flag) indicating whether a camera parameter syntax structure (gaming_camera_data) is present in the bitstream; as well as In response to the existence of the camera parameter syntax structure, the method further includes: At least one camera parameter syntax data structure (gaming_camera_data) representing camera parameters is encoded at the picture level, wherein the camera parameters provide information on the position, orientation and characteristics of the game engine virtual camera that captured the picture. 2 . The method of claim 1 , wherein the camera parameter syntax data structure comprises an indication (intrinsic_param_focal_flag) defining a first set of intrinsic camera parameters related to the field of view.
3. The method according to claim 1 or 2, wherein the camera parameter syntax data structure includes a first set of intrinsic camera parameters related to the field of view, the first set of intrinsic camera parameters related to the field of view including first intrinsic matrix coefficients related to the horizontal field of view (IF0) and second intrinsic matrix coefficients related to the vertical field of view (IF1).
4. The method according to claim 1 or 2, wherein the camera parameter syntax data structure includes a first set of intrinsic camera parameters related to the field of view, the first set of intrinsic camera parameters including camera parameters related to aspect ratio (H / V) and camera parameters related to field of view (FOV). The method of claim 1 , wherein the camera parameter syntax data structure includes an indication (intrinsic_param_plane_flag) defining a second set of intrinsic camera parameters related to a far plane and a near plane.
6. The method according to claim 1 or 5, wherein the camera parameter syntax data structure comprises a second set of intrinsic camera parameters associated with the far plane and the near plane, wherein the second set of intrinsic camera parameters associated with the far plane and the near plane comprises first coefficients of an intrinsic matrix associated with the plane position (IP0) and second coefficients of an intrinsic matrix associated with the plane position (IP1).
7. The method according to claim 1 or 5, wherein the camera parameter syntax data structure includes a second set of intrinsic camera parameters related to the far plane and the near plane, wherein the second intrinsic camera parameters related to the far plane and the near plane include parameters related to the far plane position (Z F ) and the camera parameters related to the near plane position (Z N ) related camera parameters.
8. The method of claim 1, wherein a camera parameter syntax data structure includes an indication (extrinsic_param_rotation_flag) defining a set of extrinsic matrix coefficients representing a rotation of the game engine virtual camera.
9. The method according to claim 1 or 8, wherein the camera parameter syntax data structure includes 9 external matrix coefficients (EWtoCR[j][k]) representing the rotation of the game engine virtual camera in the world coordinate system.
10. The method according to claim 1 or 8, wherein the camera parameter syntax data structure includes four external matrix coefficients (ECtoCq[k]) represented by the quaternion of the rotation of the game engine virtual camera in the world coordinate system.
11. The method of claim 1 , wherein a camera parameter syntax data structure includes an indication (extrinsic_param_translation_flag) defining a set of extrinsic matrix coefficients representing translation of the game engine virtual camera.
12. The method according to claim 1 or 11, wherein the camera parameter syntax data structure includes three external matrix coefficients (ECtoWT[k]) representing the translation of the game engine virtual camera in the world coordinate system.
13. The method of claim 1, wherein the camera parameters are encoded based on one of: 32-bit floating point; 64-bit double floating point; a combination of sign, mantissa, and exponent; and a 32-bit integer value.
14. The method of claim 1, wherein the camera parameter syntax structure (gaming_camera_data) further includes an indication to calculate inverse inner matrix coefficients and inverse outer matrix coefficients at a decoder.
15. The method according to claim 1, wherein the camera parameters provide information of the absolute position and absolute orientation of a game engine virtual camera that captures the picture.
16. The method of claim 1, wherein the camera parameters provide information of a difference in position of a game engine virtual camera between a screen and a previous screen and a difference in orientation of the game engine virtual camera between a screen and a previous screen.
17. A method according to claim 1, 15 or 16, wherein the camera parameter syntax structure also includes an indication of information that the camera parameters provide a position difference of the game engine virtual camera between the picture and the previous picture and an orientation difference of the game engine virtual camera between the picture and the previous picture.
18. A method according to claim 1, 15 or 16, wherein inferring the camera parameters from the screen type provides an indication of information of a difference in position of the game engine virtual camera between the screen and a previous screen and a difference in orientation of the game engine virtual camera between the screen and a previous screen.
19. The method of claim 1 , further comprising encoding at a sequence level at least one camera parameter syntax data element, the at least one camera parameter syntax data element at the sequence level comprising a first set of intrinsic camera parameters related to the field of view and to a far plane and a near plane.
20. The method of claim 1 or 19, further comprising encoding a syntax data element at a sequence level, the syntax data element indicating an updated instance of the at least one camera parameter syntax data structure at the sequence level representing camera parameters.
21. The method of claim 1, further comprising encoding a syntax data element indicating whether one or more instances of a camera parameter syntax data structure are present in the bitstream.
22. A method comprising: Decodes the syntax data element (sps_camera_param_enabled_flag) that indicates whether the camera parameter syntax structure (gaming_camera_data) is present in the bitstream; as well as In response to the existence of the camera parameter syntax structure, the method further includes: At least one camera parameter syntax data structure (gaming_camera_data) representing camera parameters is decoded at the picture level, wherein the camera parameters provide information on the position, orientation, and characteristics of the game engine virtual camera that captured the picture.
23. The method of claim 22, wherein the camera parameter syntax data structure comprises an indication (intrinsic_param_focal_flag) defining a first set of intrinsic camera parameters related to the field of view.
24. The method of claim 22 or 23, wherein the camera parameter syntax data structure comprises a first set of intrinsic camera parameters related to the field of view, the first set of intrinsic camera parameters related to the field of view comprising first intrinsic matrix coefficients related to the horizontal field of view (IF0) and second intrinsic matrix coefficients related to the vertical field of view (IF1).
25. The method of claim 22 or 23, wherein the camera parameter syntax data structure comprises a first set of internal camera parameters related to the field of view, the first set of internal camera parameters comprising camera parameters related to aspect ratio (H / V) and camera parameters related to field of view (FOV).
26. The method of claim 22, wherein the camera parameter syntax data structure includes an indication (intrinsic_param_plane_flag) defining a second set of intrinsic camera parameters related to a far plane and a near plane.
27. The method of claim 22 or 26, wherein the camera parameter syntax data structure comprises a second set of intrinsic camera parameters associated with the far plane and the near plane, wherein the second set of intrinsic camera parameters associated with the far plane and the near plane comprises first coefficients of an intrinsic matrix associated with the plane position (IP0) and second coefficients of an intrinsic matrix associated with the plane position (IP1).
28. The method of claim 22 or 26, wherein the camera parameter syntax data structure comprises a second set of internal camera parameters associated with a far plane and a near plane, the second set of internal camera parameters associated with the far plane and the near plane comprising camera parameters associated with a far plane position (ZF) and camera parameters associated with a near plane position (ZN).
29. The method of claim 22, wherein a camera parameter syntax data structure includes an indication (extrinsic_param_rotation_flag) defining a set of extrinsic matrix coefficients representing a rotation of the game engine virtual camera.
30. The method of claim 22 or 29, wherein the camera parameter syntax data structure comprises 9 extrinsic matrix coefficients (EWtoCR[j][k]) representing the rotation of the game engine virtual camera in the world coordinate system.
31. The method according to claim 22 or 29, wherein the camera parameter syntax data structure includes four external matrix coefficients (EWtoCq[k]) of the quaternion representation of the rotation of the game engine virtual camera in the world coordinate system.
32. The method of claim 22, wherein a camera parameter syntax data structure includes an indication (extrinsic_param_translation_flag) defining a set of extrinsic matrix coefficients representing translation of the game engine virtual camera.
33. The method of claim 22 or 32, wherein the camera parameter syntax data structure comprises three extrinsic matrix coefficients (ECtoWT[k]) representing the translation of the game engine virtual camera in the world coordinate system.
34. The method of claim 22, wherein the camera parameters are encoded based on one of: 32-bit floating point; 64-bit double floating point; a combination of sign, mantissa, and exponent; and a 32-bit integer value.
35. The method of claim 22, wherein the camera parameter syntax structure (gaming_camera_data) further includes an indication to calculate inverse inner matrix coefficients and inverse outer matrix coefficients at a decoder.
36. The method of claim 22, wherein the camera parameters provide information on the absolute position and absolute orientation of a game engine virtual camera that captures the image.
37. The method of claim 22, wherein the camera parameters provide information of a difference in position of the game engine virtual camera between a scene and a previous scene and a difference in orientation of the game engine virtual camera between a scene and a previous scene.
38. A method according to claim 22, 36 or 37, wherein the camera parameter syntax structure also includes an indication of information that the camera parameters provide a position difference of the game engine virtual camera between the picture and the previous picture and an orientation difference of the game engine virtual camera between the picture and the previous picture.
39. A method according to claim 22, 36 or 37, wherein inferring the camera parameters from the screen type provides an indication of information of a difference in position of the game engine virtual camera between the screen and the previous screen and a difference in orientation of the game engine virtual camera between the screen and the previous screen.
40. The method of claim 22, further comprising decoding at least one camera parameter syntax data element at a sequence level, the at least one camera parameter syntax data element at the sequence level comprising a first set of intrinsic camera parameters related to the field of view and to a far plane and a near plane.
41. The method of claim 22 or 40, further comprising decoding, at a sequence level, a syntax data element indicating an updated instance of the at least one camera parameter syntax data structure at a sequence level representing camera parameters.
42. The method of claim 22, further comprising decoding a syntax data element indicating whether one or more instances of a camera parameter syntax data structure are present in the bitstream.
43. The method of claim 22, wherein the footage is part of a game engine 2D rendered video.
44. A video encoding device comprising a processor, the processor being configured to: encodes a syntax data element (sps_camera_param_enabled_flag) that indicates whether a camera parameter syntax structure (gaming_camera_data) is present in the bitstream; and In response to the presence of the camera parameter syntax structure, at least one camera parameter syntax data structure (gaming_camera_data) representing camera parameters is encoded at the picture level, wherein the camera parameters provide information on the position, orientation and characteristics of the game engine virtual camera that captures the picture.
45. A video decoding device comprising a processor, the processor being configured to: decoding a syntax data element (sps_camera_param_enabled_flag) indicating whether a camera parameter syntax structure (gaming_camera_data) is present in the bitstream; and In response to the presence of the camera parameter syntax structure, at least one camera parameter syntax data structure (gaming_camera_data) representing camera parameters is decoded at the picture level, wherein the camera parameters provide information on the position, orientation and characteristics of the game engine virtual camera that captures the picture.
46. A computer program product stored on a non-transitory computer readable medium and comprising program code instructions for implementing the steps of the method according to at least one of claims 1 to 43 when executed by at least one processor.
47. A computer program comprising program code instructions for implementing the steps of the method according to at least one of claims 1 to 43 when executed by a processor.
48. A bitstream comprising information representing an encoded output generated by one of the methods according to any of claims 1 to 21.
49. A non-transitory program storage device having encoded data representing an image block generated by the method of one of claims 1 to 21.
50. A non-transitory program storage device readable by a computer, tangibly embodying a program of instructions executable by the computer for performing the method of any one of claims 1 to 43.