Method, apparatus, electronic device, and storage medium for video encoding and decoding

The conversion of 3D scene data to point cloud data for 2D projection and encoding with 6DoF metadata enhances virtual reality streaming by allowing flexible viewing angles and reducing data processing inefficiencies.

CN115514972BActive Publication Date: 2025-07-15TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211137156.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-23
Filing Date
2020-06-24
Publication Date
2025-07-15
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

In the prior art, virtual reality streaming cannot effectively support switching of observation positions and angles in other dimensions except the x-axis, y-axis, and z-axis, resulting in an increase in the amount of panoramic video and image data and low processing efficiency.

Method used

The volume data of the visual three-dimensional scene is converted into point cloud data and projected onto a two-dimensional image for encoding, generating a media file indicating six degrees of freedom, supporting observation position and angle-related processing.

Benefits of technology

By focusing on high-quality partial transmission of point cloud data, the transmission of unused partials is reduced, and the efficiency of the panoramic video system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115514972B_ABST
    Figure CN115514972B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, electronic device, and storage medium for video encoding and decoding. The method includes: obtaining volume data of at least one visual three-dimensional (3D) scene; converting the volume data into point cloud data; projecting the point cloud data onto a two-dimensional (2D) image based on a first viewing position and a first viewing angle; encoding the point cloud data projected onto the 2D image; and composing a media file that encapsulates metadata and the encoded point cloud data. The metadata indicates six degrees of freedom (6DoF) media, and further indicates at least one second viewing position on the 6DoF coordinate system other than the first viewing position, and at least one second viewing angle other than the first viewing angle at the second viewing position.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims priority to U.S. Provisional Application No. 62 / 868,797, filed on Jun. 28, 2019, and U.S. Patent Application No. 16 / 909,314, filed on Jun. 23, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to a set of advanced video coding techniques, including improved schemes for view-position and angle-dependent processing of point cloud data. Background Art

[0004] In the case where effective streaming for other dimensions is not allowed, the streaming of virtual reality, such as the streaming of images and audio, limits the user's viewing experience to an experience of a panoramic image, which allows the user to view different parts of the image from different angles in the x-axis, y-axis, and z-axis environments. The panoramic image is similar to a three-dimensional image, and the other dimensions can enable the user to experience this virtual reality from different viewing positions in the front / back, up / down, and left / right directions in addition to viewing from different angles at at least one of these positions.

[0005] Therefore, in the prior art, in addition to streaming at different angles in the x-axis, y-axis, and z-axis environments, effective streaming for other dimensions is not allowed, that is, the user cannot experience this virtual reality from different viewing positions in the front / back, up / down, and left / right directions. However, the user's experience of virtual reality from multiple dimensions will increase the amount of panoramic video and / or image data, thereby resulting in low processing efficiency of the panoramic video system. Corresponding technical solutions are urgently needed to solve such problems. Summary of the Invention

[0006] Embodiments of the present application include a method, an apparatus, an electronic device, and a storage medium for video coding.

[0007] The method for video coding provided by the embodiments of the present application includes: obtaining volume data of at least one visual three-dimensional (3D) scene; converting the volume data into point cloud data; projecting the point cloud data onto a two-dimensional (2D) image; encoding the point cloud data projected onto the 2D image; and composing a media file that encapsulates metadata and the encoded point cloud data, where the metadata indicates six degrees of freedom (6DoF) media.

[0008] The apparatus for video coding provided by the embodiments of the present application includes:

[0009] A selection module for obtaining volume data of at least one visual three-dimensional (3D) scene;

[0010] A conversion module for converting the volume data into point cloud data;

[0011] A projection module for projecting the point cloud data onto a two-dimensional (2D) image;

[0012] An encoding module for encoding the point cloud data projected onto the 2D image; and

[0013] A composition module for composing a media file that encapsulates metadata and the encoded point cloud data, the metadata indicating six degrees of freedom (6DoF) media.

[0014] An embodiment of the present application also provides a non-volatile computer-readable storage medium storing multiple instructions that can cause at least one processor to execute the method of the embodiment of the present application.

[0015] An embodiment of the present application also provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method of the embodiment of the present application when executing the program.

[0016] Through the technical solution of the embodiment of the present application, more effective processing can be performed on a specific part of the point cloud data, enabling the player to focus on an image of higher quality in the point cloud data than other parts without transmitting unused parts, thereby improving the efficiency of the panoramic video system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0018] Figure 1 is a simplified block diagram of a communication system according to an embodiment of the present disclosure;

[0019] Figure 2 is an example of the placement of a video encoder and a video decoder in a streaming environment according to an embodiment of the present disclosure;

[0020] Figure 3 is a functional block diagram of a video decoder according to an embodiment of the present application;

[0021] Figure 4 is a block diagram of a video encoder according to an embodiment of the present application;

[0022] Figure 5 is an intra prediction mode used in HEVC and JEM according to an embodiment of the present application;

[0023] Figure 6 are N reference levels of the intra direction mode according to an embodiment disclosed in the present application;

[0024] Figure 7 is an illustration of the DC mode PDPC weights at positions (0, 0) and (1, 0) within a 4×4 block according to an embodiment disclosed in the present application;

[0025] Figure 8 is an illustration of local luminance compensation according to an embodiment disclosed in the present application;

[0026] Figure 9A is an intra prediction mode used in HEVC according to an embodiment disclosed in the present application;

[0027] Figure 9B is an example of 87 intra prediction modes in VVC according to an embodiment disclosed in the present application;

[0028] Figure 10 is a simplified block diagram - style workflow diagram of exemplary window - related processing of the panoramic media application format according to an embodiment disclosed in the present application;

[0029] Figure 11A is a flowchart of a method for video coding according to an embodiment disclosed in the present application;

[0030] Figure 11B is a simplified block - style content flowchart of encoded point cloud data for observation position and angle - related processing according to an embodiment disclosed in the present application;

[0031] Figure 12 is a block diagram of a computer system according to an embodiment disclosed in the present application. Detailed Description

[0032] The features proposed in the embodiments of the present application discussed below can be used alone or in any combination. Further, the embodiments can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one embodiment, at least one processor executes a program stored in a non - volatile computer - readable medium.

[0033] Figure 1FIG. 0 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected by a network 105. For unidirectional transmission of data, a first terminal 103 encodes video data at a local location for transmission via the network 105 to another terminal 102. The second terminal 102 receives the encoded video data of the other terminal from the network 105, decodes the encoded data, and displays the restored video data. Unidirectional data transmission is relatively common in media service applications and the like.

[0034] Figure 1 Also shown is a second pair of terminal devices that perform bidirectional transmission of encoded video data, a third terminal (101) and a fourth terminal (104), and the bidirectional transmission may occur, for example, during a video conference. For bidirectional data transmission, each of the third terminal (101) and the fourth terminal (104) may encode video data collected at a local location (such as a video picture stream collected by the terminal device) for transmission via the network (105) to other terminals. Each of the third terminal (101) and the fourth terminal (104) may also receive the encoded video data transmitted by other terminals, may decode the encoded video data, and may display the restored video data on a local display device.

[0035] In Figure 1 the embodiment, the first terminal (103), the second terminal (102), the third terminal (101), and the fourth terminal (104) may be servers, personal computers, and smart phones, but the principles disclosed in this application are not limited thereto. The embodiments disclosed in this application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (105) represents any number of networks that transmit encoded video data between the first terminal (103), the second terminal (102), the third terminal (101), and the fourth terminal (104), including, for example, wired (wired) and / or wireless communication networks. The communication network (105) may exchange data in circuit-switched and / or packet-switched channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of the network (105) may be irrelevant to the operation disclosed in this application.

[0036] As an embodiment, Figure 2 FIG. 14 shows the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0037] A streaming system may include an acquisition subsystem (203), which may include a video source (201) such as a digital camera, and the video source creates an uncompressed video sample stream (213), for example. Compared with an encoded bitstream, this video sample stream (213) is emphasized as having a high data volume, and the video sample stream (213) can be processed by an encoder 202 coupled to the camera 201. The video encoder (202) may include hardware, software, or a combination thereof to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video bitstream 204 can be stored on a streaming server 205 for future use, and the encoded video bitstream 204 can be emphasized as having a lower data volume compared with the video sample stream. At least one streaming client 212 and 207 can access the streaming server (305) to retrieve copies (208) and (206) of the encoded video bitstream (204). The client (212) may include a video decoder (211), which decodes an incoming copy of the encoded video bitstream 208 and generates an output video sample stream (210) that can be presented on a display (209) or other presentation device (not depicted). The video bitstreams 204, 206, and 208 can be encoded according to certain video coding / compression standards. Examples of these standards are as described above and are further described herein.

[0038] Figure 3 is a functional block diagram of a video decoder (300) according to an embodiment disclosed in the present application.

[0039] A receiver (302) may receive one or at least two coded video sequences to be decoded by the video decoder (300); in the same or another embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence can be received from a channel (301), and the channel can be a hardware / software link to a storage device storing the encoded video data. The receiver (302) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that can be forwarded to their respective using entities (not labeled). The receiver (302) can separate the coded video sequence from other data. To prevent network jitter, a buffer memory (303) may be coupled between the receiver (302) and an entropy decoder / parser (304) (hereinafter referred to as "parser"). And when the receiver (302) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (303), or the buffer memory can be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (303) may also be required, and the buffer memory can be relatively large and may have an adaptive size.

[0040] The video decoder (300) may include a parser (304) to reconstruct symbols (313) from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of the video decoder (300), and potential information for controlling a display device such as a display device (e.g., display screen 312), which is not a component of the decoder (430) but can be coupled to the decoder. The control information for the display device can be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (304) can perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (304) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. Subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The entropy decoder / parser can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0041] The parser 304 can perform entropy decoding / parsing operations on the video sequence received from the buffer 303 to create symbols 313. The parser 304 can receive the encoded data and selectively decode specific symbols 313. Further, the parser 304 can determine whether to provide specific symbols 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0042] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (313) may involve at least two different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser (304) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (304) and at least two units below are not described.

[0043] In addition to the functional blocks already mentioned, the video decoder (300) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.

[0044] The first unit is the scaler / inverse transform unit (305). The scaler / inverse transform unit (305) receives the quantized transform coefficients as symbols (313) and control information from the parser (304), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (304) can output blocks including sample values, and the sample values can be input into the aggregator (310).

[0045] In some cases, the output samples of the scaler / inverse transform unit (305) can belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit (307). In some cases, the intra-picture prediction unit (307) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the (partially reconstructed) current picture 309. In some cases, the aggregator (310) adds the predictive information generated by the intra-picture prediction unit (307) to the output sample information provided by the scaler / inverse transform unit (305) based on each sample.

[0046] In other cases, the output samples of the scaler / inverse transform unit (305) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (306) can access the reference picture memory (308) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (313), these samples can be added by the aggregator (310) to the output of the scaler / inverse transform unit (305) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (306) obtaining the prediction samples from an address within the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol (313) for use by the motion compensation prediction unit (306), and the symbol (313) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory, a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.

[0047] The output samples of the aggregator (310) can be employed by various loop filtering techniques in the loop filter unit (311). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters can be used in the loop filter unit (311) as symbols (313) from the parser (304). However, in other embodiments, video compression techniques can also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0048] The output of the loop filter unit (311) can be a sample stream that can be output to the display device (312) and stored in the reference picture memory (557) for subsequent inter-picture prediction.

[0049] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (304)), the current picture buffer (309) can become part of the reference picture buffer (308), and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0050] The video decoder (300) can perform decoding operations according to, for example, the predefined video compression techniques recorded in the ITU-T H.265 standard. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used in the sense that the encoded video sequence follows the video compression technique or standard's syntax as documented in the video compression technique document or standard (especially its profiles). For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0051] In an embodiment, the receiver (302) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (300) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] Figure 4 is a block diagram of a video encoder (400) according to an embodiment disclosed in the present application.

[0053] The video encoder (400) may receive video samples from a video source (401) (not part of the encoder), and the video source may capture video images to be encoded by the video encoder (400).

[0054] The video source (401) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (401) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (401) may be a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or at least two samples depending on the sampling structure, color space, etc. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0055] According to an embodiment, a video encoder (400) may encode and compress pictures of a source video sequence into an encoded video sequence (410) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of a controller (402). In some embodiments, the controller (402) controls other functional units as described below and is functionally coupled to these units. For simplicity, couplings are not labeled in the figures. Parameters set by the controller (402) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can identify other suitable functions of the controller (402) that relate to optimizing the video encoder (400) for a particular system design.

[0056] Some video encoders operate in an encoding loop that is readily recognizable to those skilled in the art. As a simple description, the encoding loop consists of an encoding part of the encoder (402) (hereinafter referred to as the source encoder 530) (responsible for creating symbols based on input pictures to be encoded and reference pictures) and a (local) decoder (406) embedded in the video encoder. The decoder (406) reconstructs the symbols in the same way as a (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to the reference picture memory (405). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.

[0057] The operation of the “local” decoder (406) may be the same as that of the “remote” decoder, for example, as described in detail above in connection with Figure 3 However, briefly referring to Figure 4 , when symbols are available and the entropy encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (408) and the parser (304) can be losslessly performed, the entropy decoding part of the video decoder (300), including the channel (301), buffer (303), and parser (304), may not be fully implementable in the local decoder (406).

[0058] At this point, it can be observed that any decoder technology other than parsing / entropy decoding existing in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is inverse to the decoder technology described comprehensively. A more detailed description is only required in certain areas and is provided below.

[0059] As part of the operation, the source encoder (403) may perform motion compensated predictive coding. With reference to one or at least two previously encoded frames designated as "reference frames" in the video sequence, the motion compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine (407) encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame.

[0060] The local video decoder (406) may decode the encoded video data of the frame that can be designated as a reference frame based on the symbols created by the source encoder (403). The operation of the coding engine (407) may be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 4 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (406) replicates the decoding process that can be performed by the video decoder on the reference frame and may store the reconstructed reference frame in the reference picture cache (405). In this way, the video encoder (400) can locally store a copy of the reconstructed reference frame, which has the same content (in the absence of transmission errors) as the reconstructed reference frame to be obtained by the remote video decoder.

[0061] The predictor (404) may perform a prediction search for the coding engine (407). That is, for a new picture to be encoded, the predictor (404) may search the reference picture memory (405) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new frame. The predictor (404) may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (404), it can be determined that the input picture may have prediction references obtained from at least two reference pictures stored in the reference picture memory (405).

[0062] The controller (402) may manage the encoding operations of the encoder (403), including, for example, setting parameters and subgroup parameters for encoding video data.

[0063] The outputs of all the above functional units can be entropy-coded in an entropy encoder (408). The entropy encoder performs lossless compression on the symbols generated by the various functional units according to techniques well-known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., so as to convert the symbols into an encoded video sequence.

[0064] The transmitter (409) can buffer the encoded video sequence created by the entropy encoder (408) to prepare for transmission over a communication channel (411), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (409) can combine the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0065] The controller (402) can manage the operation of the video encoder (400). During encoding, the controller (405) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following frame types:

[0066] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.

[0067] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0068] A bi-predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0069] The source picture can typically be spatially subdivided into at least two sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0070] The video encoder (400) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (400) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0071] In an embodiment, the transmitter (409) can transmit additional data when transmitting the encoded video. The source encoder (403) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI (Supplementary Enhancement Information) messages, VUI (Visual Usability Information) parameter set fragments, etc.

[0072] Figure 5 The intra-frame prediction modes used in HEVC and JEM are shown. To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The additional directional modes in JEM on top of HEVC are Figure 5 depicted as dashed arrows, and the planar mode and the DC mode remain the same. These denser directional intra-frame prediction modes are applicable to all block sizes and to both luminance and chrominance intra-frame prediction. As Figure 5 shown, the directional intra-frame prediction modes identified by dashed arrows and associated with odd intra-frame prediction mode indices are called odd intra-frame prediction modes. The directional intra-frame prediction modes identified by solid arrows and associated with even intra-frame prediction mode indices are called even intra-frame prediction modes. In this document, the directional intra-frame prediction modes indicated by solid arrows or dashed arrows in Figure 5 are also called angular modes.

[0073] In JEM, a total of 67 intra prediction modes are used for luma intra prediction. To encode the intra mode, a Most Probable Mode (MPM) list of size 6 is established based on the intra modes of adjacent blocks. If the intra mode does not come from the MPM list, a flag is signaled to indicate whether the intra mode belongs to the selected modes. In JEM-3.0, there are 16 selected modes, which are uniformly selected in groups of four corner modes. In the standard proposals JVET-D0114 and JVET-G0060, 16 secondary MPMs are derived to replace the uniformly selected modes.

[0074] Figure 6 N reference levels of the intra direction mode are shown. There are block unit 611, Segment A 601, Segment B 602, Segment C 603, Segment D 604, Segment E 605, Segment F 606, the first reference layer 610, the second reference layer 609, the third reference layer 608, and the fourth reference layer 607.

[0075] In HEVC, JEM, and some other standards such as H.264 / AVC, the reference samples used to predict the current block are limited to the closest reference line (row or column). In the method of multi-reference line intra prediction, for the intra direction mode, the number of candidate reference lines (rows or columns) increases from 1 (i.e., the closest one) to N, where N is an integer greater than or equal to 1. Figure 2 Taking a 4×4 Prediction Unit (PU) as an example, the concept of the multi-reference line intra-directional prediction method is shown. The intra-directional mode selects any one of the N reference layers to generate a predictor. In other words, the predictor p(x, y) is generated from one of the reference samples S1, S2, …, and SN. A flag is signaled to indicate which reference layer is selected for the intra-directional mode. If N is set to 1, the intra-directional prediction method is the same as the traditional method in JEM2.0. In Figure 6 , the reference lines 610, 609, 608, and 607 are composed of six segments 601, 602, 603, 604, 605, and 606 and the reference sample in the upper left corner. In this document, the reference layer is also referred to as the reference line. The coordinates of the pixel in the upper left corner within the current block unit are (0, 0), and the coordinates of the pixel in the upper left corner in the first reference line are (-1, -1).

[0076] In JEM, for the luminance component, adjacent samples used for intra-prediction sample generation are filtered before the generation process. The filtering is controlled by a given intra-prediction mode and the size of the transform block. If the intra-prediction mode is DC or the size of the transform block is equal to 4×4, the adjacent samples are not filtered. If the distance between the given intra-prediction mode and the vertical mode (or horizontal mode) is greater than a predetermined threshold, the filtering process is enabled. The adjacent samples are filtered using a [1, 2, 1] filter and a bilinear filter.

[0077] The Position Dependent Intra Prediction Combination (PDPC) method is an intra-prediction method that calls a combination of unfiltered boundary reference samples and HEVC-style intra-prediction with filtered boundary reference samples. The calculation of each prediction sample pred[x][y] located at (x, y) is as follows:

[0078] pred[x][y] = (wL * R -1,y + wT * R x,-1 + wTL * R -1,-1 +(64 - wL - wT - wTL) * pred[x][y] + 32) >> 6 (Equation 2-1) where R x,-1 and R -1,y respectively represent the unfiltered reference samples located at the top and left of the current sample (x, y), and R -1,-1 represents the unfiltered reference sample located at the upper left corner of the current block. The weights are calculated according to the following formula,

[0079] wT = 32 >> ((y << 1) >> shift) (Equation 2-2)

[0080] wL = 32 >> ((x << 1) >> shift) (Equation 2-3)

[0081] wTL = -(wL >> 4) - (wT >> 4) (Equation 2-4)

[0082] shift = (log2(width) + log2(height) + 2) >> 2 (Equation 2-5)

[0083] Figure 7Illustration 700 shows the DC mode PDPC weights (wL, wT, wTL) at positions (0, 0) and (1, 0) within one of the 4×4 blocks. If PDPC is applied to DC mode, planar mode, horizontal mode, and vertical intra mode, no additional boundary filters such as the HEVC DC mode boundary filter or horizontal / vertical mode edge filters are required. Figure 7 Illustration shows the reference sample R of PDPC applied to the upper right diagonal mode x,-1 、R -1,y and R -1,-1 definitions. The predicted sample pred(x’, y’) is located at (x’, y’) within the prediction block. The coordinate x of the reference sample Rx,-1 is given by: x = x’ + y’ + 1, and similarly, the coordinate y of the reference sample R -1,y is given by: y = x’ + y’ + 1.

[0084] Figure 8 Illustration 800 shows local illumination compensation (LIC) and Figure 8 is a linear model based on the luminance change using the scaling factor a and the offset b. LIC is adaptively enabled or disabled for each coding unit (CU) encoded in an inter mode.

[0085] When LIC is applied to a CU, the least squares error method can be employed to derive the parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, as Figure 8 shown, the neighboring samples of the CU's quadratic samples (2:1 quadratic samples) and the corresponding samples in the reference picture (identified by the motion information of the current CU or sub-CU) are used. The IC parameters are derived and applied to each prediction direction separately.

[0086] When a CU is encoded in merge mode, the LIC flag is copied from the neighboring block in a manner similar to the copying of motion information in merge mode; otherwise, the LIC flag is signaled for the CU to indicate whether LIC is applied.

[0087] Figure 9A Illustration 900 shows the intra prediction modes used in HEVC. In HEVC, there are a total of 35 intra prediction modes, where mode 10 is the horizontal mode, mode 26 is the vertical mode, and modes 2, 18, and 34 are diagonal modes. The intra prediction mode is signaled by three most probable modes (MPM) and 32 remaining modes.

[0088] Figure 9BShows that in an embodiment of VVC, there are a total of 87 intra prediction modes, where mode 18 is the horizontal mode, mode 50 is the vertical mode, and modes 2, 34, and 66 are diagonal modes. Modes -1 to -10 and modes 67 to 76 are referred to as Wide-Angle Intra Prediction (WAIP) modes.

[0089] According to the PDPC expression, the predicted sample pred(x, y) at position (x, y) is predicted using a linear combination of intra prediction modes (DC, planar, angular) and reference samples:

[0090] pred(x,y) = (wL × R -1,y + wT × R x,-1 – wTL × R -1,-1 +(64 – wL – wT + wTL) × pred(x,y) + 32) >> 6

[0091] where R x,-1 、R -1,y represent the reference samples located above and to the left of the current sample (x, y) respectively, and R -1,-1 represents the reference sample located at the upper left corner of the current block.

[0092] For the DC mode, for a block with dimensions width and height, the weights are calculated as follows:

[0093] wT = 32 >> ((y << 1) >> nScale),

[0094] wL = 32 >> ((x << 1) >> nScale),

[0095] wTL = (wL >> 4) + (wT >> 4),

[0096] where nScale = (log2(width) – 2 + log2(height) – 2 + 2) >> 2, where wT represents the weighting factor of the reference sample in the above-mentioned reference line having the same horizontal coordinate as the current sample, wL represents the weighting factor of the reference sample in the left reference line having the same vertical coordinate as the current sample, and wTL represents the weighting factor of the reference sample at the upper left corner of the current block. nScale indicates the rate at which the weighting factor decreases along the axis (wL decreases from left to right or wT decreases from top to bottom), that is, the weighting factor decay rate. In the current design, this weighting factor decay rate is the same along the x-axis (from left to right) and the y-axis (from top to bottom). 32 represents the initial weighting factor of adjacent samples, and this initial weighting factor is also the weighting assigned to the top (left or upper left) of the sample at the upper left corner in the current CB. During the PDPC process, the weighting factor of adjacent samples should be equal to or less than this initial weighting factor.

[0097] For the planar mode, wTL = 0, while for the horizontal mode, wTL = wT, and for the vertical mode, wTL = wL. The PDPC weights can be calculated using only addition and shift. The value of Pred(x, y) is calculated in a single step using Equation 1.

[0098] Figure 10 FIG. 1000 shows a simplified block diagram working flow chart of an exemplary window-related process of an Omnidirectional Media Application Format (OMAF), which allows 360-degree virtual reality (VR360) streaming described in the OMAF.

[0099] At the acquisition block 1001, video data A is acquired. For example, in the case where the image data can represent a scene in VR360, the video data A can be data of at least two images and audio at the same moment. At the processing block 1003, the image B at the same moment is processed in one or more of the following ways i : stitching, mapping to a projection picture with respect to at least one virtual reality (VR) angle or other angle / viewpoint, and region-wise packing. In addition, metadata is created to facilitate the transmission and rendering processes, and this metadata indicates any one of such processing information and other information.

[0100] Regarding data D, at the image encoding block 1005, the projection picture is encoded into data E i , and the encoded projection picture is synthesized into a media file and a window-independent stream. At the video encoding block 1004, the video picture is encoded into data E v, the data E v is used as a single-layer bitstream. Regarding the data B a , at the audio encoding block 1002, the audio data can also be encoded as the data E a .

[0101] The data E a , E v E i , and the entire encoded bitstream F i and / or F can be stored in a (Content Delivery Network (CDN) / cloud) server. The data E a , E v , E i , and the entire encoded bitstream F i and / or F are typically (e.g., at the delivery block 1007 or otherwise) fully transmitted to the OMAF player 1020 and completely decoded by the decoder so that the display block 1016 presents at least one region of the decoded picture corresponding to the current window to the user for various metadata, file playback, and orientation / viewport metadata (e.g., the angle that the user can see the VR image device from the head / eye tracking module 1008 relative to the window specifications of the VR image device). The unique feature of VR360 is that only one window can be displayed at any specific time. The feature of VR360 for selective transmission according to the user's window (or any other criteria, such as the recommended window timing metadata) can be used to improve the performance of the panoramic video system. For example, window-related transmission can be achieved through tile-based video coding.

[0102] Similar to the above encoding block, according to an exemplary embodiment, the OMAF player 1020 can similarly unpack the file / segment of at least one of the data F' and / or F' i and the metadata to reverse at least one aspect of this encoding; decode the audio data E' i at the audio decoding block 1010, decode the video data E' v at the video decoding block 1013, and decode the image data E' i at the image decoding block 1014 to continue with the audio presentation of the data B' a at the audio presentation block 1011 and the image presentation of the data D' at the image presentation block 1015, and then output the display data A' i in VR360 format at the display block 1016 according to various metadata such as orientation / viewport metadata, and output the audio data A' sVarious metadata can affect the data decoding and presentation (rendering) process according to various tracks, languages, qualities, and views selected by the user of the OMAF player 1020. It should be understood that the processing order described herein is given by the exemplary embodiments and can be implemented according to other orders of other exemplary embodiments.

[0103] Figure 11A FIG. 1100A is a flowchart showing a video encoding method provided by an embodiment of the present disclosure. As Figure 11A shown, the method includes the following steps:

[0104] Step S111, obtaining volume data of at least one visual three-dimensional (3D) scene.

[0105] Step S112, converting the volume data into point cloud data.

[0106] Step S113, projecting the point cloud data onto a two-dimensional (2D) image.

[0107] Step S114, encoding the point cloud data projected onto the 2D image.

[0108] In some embodiments, encoding the point cloud data projected onto the 2D image includes: dividing the point cloud data into at least two partitions.

[0109] In some embodiments, encoding the point cloud data projected onto the 2D image includes: encoding the at least two partitions independently of each other.

[0110] Step S115, forming a media file that encapsulates metadata and the encoded point cloud data, the metadata indicating six degrees of freedom (6DoF) media.

[0111] In some embodiments, forming the media file includes: adding each encoded partition to the media file.

[0112] In some embodiments, the metadata further indicates layout information of the at least two partitions; or the at least two partitions include at least two 3D partitions in a six degrees of freedom (6DoF) coordinate system, and the metadata further indicates 3D positions of the 3D partitions in the six degrees of freedom (6DoF) coordinate system.

[0113] In some embodiments, the media file is transmitted to at least one of a cloud server and a media player, so that at least one of the cloud server and the media player extracts at least one specific partition from the media file according to the layout information of the at least two partitions; or the media file is transmitted to at least one of the cloud server and the media player, so that at least one of the cloud server and the media player extracts at least one specific 3D partition from the media file according to the 3D position.

[0114] In some embodiments, the metadata further indicates at least one viewing position on a six-degree-of-freedom 6DoF coordinate system and at least one angle at the at least one viewing position.

[0115] In some embodiments, the metadata includes 360-degree virtual reality data.

[0116] In some embodiments, the encoded point cloud data includes point cloud reconstruction metadata.

[0117] Through the embodiments of the present disclosure, more effective processing can be performed on specific parts of the point cloud data, so that the player can focus on images of higher quality in parts of the point cloud data than other parts, without transmitting unused parts, thereby improving the efficiency of the panoramic video system.

[0118] Figure 11B FIG. 1100B is a simplified block content flow chart showing the processing related to viewing position and angle of the encoded point cloud data, and the encoded point cloud data is related to acquisition / generation / encoding / decoding / rendering / display of six degrees of freedom media (referred to herein as: Video-based Point Cloud Coding, V-PCC). It should be understood that the described features can be used alone or in any combination, and the elements for encoding and decoding, etc. can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit), and according to an exemplary embodiment, at least one processor can execute a program stored in a non-volatile computer-readable medium.

[0119] FIG. 1100B shows an exemplary embodiment of the streaming of the encoded point cloud data according to V-PCC.

[0120] At the volume data acquisition block 1101, volume data of at least one visual three-dimensional 3D scene is acquired.

[0121] In some embodiments, a real-world visual scene or a computer-generated visual scene (or a combination thereof) can be acquired by a set of camera devices or synthesized by a computer into volume data.

[0122] At the conversion point cloud block 1102, the volume data is converted into point cloud data.

[0123] In some embodiments, the volume data in any format is converted into a (quantized) point cloud data format through image processing. For example, according to an exemplary embodiment, the data from the volume data can be the area data of some points to be converted into a point cloud, and the area data is data that extracts one or at least two of the following values described below from the volume data and related data into the desired point cloud format.

[0124] According to an exemplary embodiment, the volume data can be a 3D data set of a 2D image, for example, it can be a strip projected by a 2D projection of the 3D data set. According to an exemplary embodiment, the point cloud data format includes the representation of data points in at least one different space, which can be used to represent the volume data, and can provide improvements in terms of sample and data compression (such as time redundancy). For example, the point cloud data representation in the x, y, z format includes color values (such as RGB, etc.), brightness, intensity, etc. at each point of the cloud data, and can be used together with progressive decoding, polygon meshing, direct rendering, and the octree 3D representation of 2D quadtree data.

[0125] When projected onto the image block 1103, the acquired point cloud data is projected onto a 2D image, and the projected point cloud data is encoded into an image / video picture using V - PCC. The projected point cloud data can consist of attributes, geometric information, an occupancy map, and other metadata for point cloud data reconstruction. Other metadata includes, for example, Painter’s Algorithms, Ray Casting Algorithms, (3D) Binary Space Partitioning Algorithms, etc.

[0126] On the other hand, at the scene generator block 1109, the scene generator can generate some metadata for presenting and displaying 6 - Degree - of - Freedom (DoF) media according to the director's intention or the user's preference. Regarding the virtual experience within or at least based on the encoded point cloud data, other dimensions allow for forward / backward, up / down, and left / right movements. In addition, the 6DoF media includes 3D viewing scenarios such as 360VR, and the viewing scenario is viewed through rotational changes on the 3D axes X, Y, Z. The scene description metadata defines at least one scene, and the at least one scene consists of the encoded point cloud data and other media data (including VR360, light field, audio, etc.). The metadata is provided to at least one cloud server and / or file / segment encapsulation / de - encapsulation processing as indicated in Figure 11B and the related description.

[0127] At the video encoding block 1104, the point cloud data projected onto the two-dimensional (2D) image is encoded.

[0128] In some embodiments, the at least two partitions are encoded independently of each other.

[0129] At the image encoding block 1105, a media file is composed, which encapsulates the metadata and the encoded point cloud data, and the metadata indicates six degrees of freedom (6DoF) media.

[0130] In some embodiments, each encoded partition is added to the media file. Specifically, at the video encoding block 1104 and the image encoding block 1105, similar to the above video and image encoding (and it should be understood that audio encoding is also provided as above), the file / segment encapsulation block 1106 processes the encoded point cloud data to compose, according to a specific media container file format, the encoded point cloud data into a media file for file playback or a sequence of initialization segments and media segments for streaming. The specific media container file format can be, for example, at least one video container format, and a format that can be used relative to DASH described below, where such description represents an exemplary embodiment segment. The file container can also include scene description metadata into the file or segment, and the scene description metadata comes from the scene generator block 1109.

[0131] According to an exemplary embodiment, the file is encapsulated according to the scene description metadata so that it includes at least one viewing position and at least one viewing angle at each of the at least one viewing position in the 6DoF media one or more times, so as to transmit the file according to a request input by a user or a creator. Further, according to an exemplary embodiment, a segment of this file can include at least one part of this file, such as a part of the 6DoF media, which indicates a single viewing point and its angle one or more times; however, these are only exemplary embodiments and can be changed according to various conditions, such as the network, the capabilities and inputs of the user and the creator.

[0132] According to an exemplary embodiment, the point cloud data is divided into at least two 2D / 3D regions, and the 2D / 3D regions are encoded independently at least at one of the video encoding block 1104 and the image encoding block 1105. Then, each independently encoded partition of the point cloud data can be encapsulated as a track in the file and / or segment at the file / segment encapsulation block 1106. According to an exemplary embodiment, each point cloud track and / or metadata track can include some metadata useful for viewing position / angle related processing.

[0133] According to an exemplary embodiment, metadata useful for view / position / angle-related processing, such as metadata included in an encapsulated file and / or segment of a file / segment encapsulation block, includes at least one of the following: layout information of a 2D / 3D partition with an index, (dynamic) mapping information associating a 3D volume partition with at least one 2D partition (e.g., any one of a tile / slice group / strip / subpicture), 3D positions of each 3D partition on a 6DoF coordinate system, a list of representative viewing positions / angles, a list of selected viewing positions / angles corresponding to 3D volume partitions, indices of 2D / 3D partitions corresponding to the selected viewing positions / angles, quality (grade) information of each 2D / 3D partition, and rendering information of each 2D / 3D partition depending on each viewing position / angle. When requested, for example, by a user of a V-PCC player or indicated by a content creator for a user of a V-PCC player, invoking this metadata enables more efficient processing of a specific portion of the 6DoF media desired by this metadata, such that the V-PCC player can transmit higher-quality images focused on a portion of the 6DoF media than other portions, rather than transmitting unused portions of the media.

[0134] In some embodiments, the metadata further indicates layout information of the at least two partitions; the media file is transmitted to at least one of a cloud server and a media player, such that at least one of the cloud server and the media player extracts at least one specific partition from the media file according to the layout information of the at least two partitions.

[0135] In some embodiments, the at least two partitions include at least two three-dimensional (3D) partitions on a six-degree-of-freedom (6DoF) coordinate system, and the metadata further indicates 3D positions of the 3D partitions on the 6DoF coordinate system; the media file is transmitted to at least one of a cloud server and a media player, such that at least one of the cloud server and the media player extracts at least one specific 3D partition from the media file according to the 3D positions.

[0136] In some embodiments, the metadata includes 360-degree virtual reality data.

[0137] In some embodiments, the encoded point cloud data includes point cloud reconstruction metadata.

[0138] From the file / segment encapsulation block 1106, using a delivery mechanism, such as Dynamic Adaptive Streaming over HTTP (DASH), the file or at least one segment of the file is directly delivered to either the V-PCC player 1125 or the cloud server, e.g., at the cloud server block 1107. The cloud server can extract at least one track and / or at least one specific 2D / 3D partition from the file and merge at least two encoded point cloud data into one data.

[0139] According to the data of the position / viewpoint tracking block 1108, if the current viewing position and angle are defined on the client system in the 6DoF coordinate system, then the viewing position / angle metadata can be delivered from the file / segment encapsulation block 1106, or other processing of the viewing position / angle metadata can be performed at the cloud server block 1107 based on the file or segment already in the cloud server, so that the cloud server can extract the appropriate partition from the stored file, merge the extracted appropriate partitions (if necessary) according to the metadata from the client system. The client system has a V-PCC player 1125 and delivers the extracted data to the client as a file or segment.

[0140] For such data, at the file / segment decapsulation block 1109, the file decompressor processes the file or the received segment, extracts the encoded bitstream, and parses the metadata; at the video decoding and image decoding blocks, the encoded point cloud data is decoded, and the decoded point cloud data is reconstructed into point cloud data at the point cloud reconstruction block 1112. The reconstructed point cloud data can be displayed at the display block 1114, and / or the reconstructed point cloud data can be synthesized according to at least one type of scene description data from the scene generator block 1109 at the scene composition block 1113.

[0141] In view of the above, such an exemplary V-PCC process exhibits advantages over the V-PCC standard, including at least one of the following: the ability to partition at least two 2D / 3D regions as described, the ability to combine the compressed domains of the encoded 2D / 3D partitions into a single consistent encoded video bitstream, and the bitstream extraction ability to form a consistent encoded bitstream from the encoded 2D / 3D partitions of the encoded pictures, where the support for the V-PCC system is further improved by forming a container including the VVC bitstream to support a mechanism for including metadata carrying at least one of the above metadata.

[0142] Therefore, through the exemplary embodiments described herein, more effective processing can be performed on specific portions of the point cloud data through at least one of these technical solutions, enabling the player to focus on images of higher quality in the point cloud data than other portions, without transmitting unused portions, thereby improving the efficiency of the panoramic video system, that is, advantageously improving the above technical problems.

[0143] The embodiments of the present application also provide a video encoding device corresponding to the above video encoding method. The device includes:

[0144] A selection module, configured to obtain volume data of at least one visual three-dimensional (3D) scene;

[0145] A conversion module, configured to convert the volume data into point cloud data;

[0146] A projection module, configured to project the point cloud data onto a two-dimensional (2D) image;

[0147] An encoding module, configured to encode the point cloud data projected onto the 2D image; and

[0148] A composition module, configured to compose a media file, where the media file encapsulates metadata and the encoded point cloud data, and the metadata indicates six degrees of freedom (6DoF) media.

[0149] The encoding module further divides the point cloud data into at least two partitions.

[0150] In some embodiments, the encoding module further encodes the at least two partitions independently of each other.

[0151] In some embodiments, the composition module further composes the media file by adding each encoded partition to the media file.

[0152] In some embodiments, the metadata further indicates layout information of the at least two partitions; the device further includes a sending module, configured to transmit the media file to at least one of a cloud server and a media player, so that at least one of the cloud server and the media player extracts at least one specific partition from the media file according to the layout information of the at least two partitions.

[0153] In some embodiments, the at least two partitions include at least two three-dimensional (3D) partitions in a six degrees of freedom (6DoF) coordinate system, and the metadata further indicates the three-dimensional (3D) positions of the 3D partitions in the 6DoF coordinate system;

[0154] The device further includes a sending module configured to transfer the media file to at least one of a cloud server and a media player, so that at least one of the cloud server and the media player extracts at least one specific three-dimensional (3D) partition from the media file according to the 3D position.

[0155] The above technique can be implemented as computer software by computer-readable instructions and physically stored in at least one computer-readable medium, or implemented by one or at least two specially configured hardware processors. For example, Figure 12 FIG. 1200 shows a computer system that is suitable for implementing certain embodiments of the disclosed subject matter, such as an electronic device suitable for the embodiments of the present disclosure.

[0156] The computer software can be encoded in any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through methods such as decoding and microcode.

[0157] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0158] Figure 12 The components shown for the computer system 1200 are exemplary in nature and are not used to impose any limitations on the scope of use or functions of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system 1200.

[0159] The computer system 1200 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from one or at least two human users through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as voice, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human-machine interface device can also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still camera), and videos (e.g., two-dimensional videos, three-dimensional videos including stereoscopic videos).

[0160] The human-machine interface input device may include one or at least two of the following (only one is drawn): keyboard 1201, mouse 1202, touchpad 1203, touch screen 1210, joystick 1205, microphone 1206, scanner 1208, camera 1207.

[0161] The computer system 1200 may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or at least two human users through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (such as haptic feedback through the touch screen 1210 or the joystick 1205, but there may also be haptic feedback devices that are not used as input devices), audio output devices (such as, the speaker 1209, headphones (not shown)), visual output devices (such as, the screen 1210 including a cathode ray tube screen, a liquid crystal screen, a plasma screen, an organic light emitting diode screen, each of which has or does not have a touch screen input function, each of which has or does not have a haptic feedback function - some of which can output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0162] The computer system 1200 may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs with CD / DVD (CD / DVD ROM / RW) 1220 or similar media 1221, thumb drives 1222, removable hard disk drives or solid state drives 1223, traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.

[0163] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0164] The computer system 1200 may also include an interface 1299 to at least one communication network 1298. For example, the network 1298 may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. The network 1298 also includes local area networks such as Ethernet, wireless local area network, cellular network (GSM, 3G, 4G, 5G, LTE, etc.), and wired or wireless wide area digital networks for television (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), and so on. Certain networks 1298 typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1250 and 1251) (e.g., the USB port of the computer system 1200); other systems are typically integrated into the core of the computer system 1200 by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks 1298, the computer system 1200 can communicate with other entities. The communication may be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, such as via a local or wide area digital network to other computer systems. Each of the above-mentioned networks and network interfaces may use certain protocols and protocol stacks.

[0165] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core 1240 of the computer system 1200.

[0166] The core (440) may include one or at least two central processing units (CPUs) 1241, a graphics processing unit (GPU) 1242, a graphics adapter 1217, a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) 1243, a hardware accelerator 1244 for specific tasks, and so on. These devices, as well as a read-only memory (ROM) 1245, a random access memory 1246, an internal mass storage (e.g., an internal non-user-accessible hard disk drive, a solid-state drive, etc.) 1247, and so on, can be connected via a system bus 1248. In some computer systems, the system bus 1248 can be accessed in the form of one or at least two physical plugs for expansion by an additional central processing unit, a graphics processing unit, and so on. Peripheral devices can be directly attached to the system bus 1248 of the core or connected via a peripheral bus 1249. The architecture of the peripheral bus includes an external controller interface PCI, a universal serial bus USB, and so on.

[0167] The CPU 1241, GPU 1242, FPGA 1243, and accelerator 1244 can execute certain instructions, which, when combined, can constitute the aforementioned computer code. This computer code can be stored in the ROM 1245 or RAM 1246. Transitional data can also be stored in the RAM 1246, while permanent data can be stored in, for example, the internal mass storage 1247. Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or at least two of the CPU 1241, GPU 1242, mass storage 1247, ROM 1245, RAM 1246, etc.

[0168] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this application or may be of the kind well-known and available to those skilled in the computer software art.

[0169] By way of example and not limitation, a computer system having the architecture 1200, and in particular the core 1240, can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or at least two tangible computer-readable media. Such computer-readable media can be media associated with the aforementioned user-accessible mass storage and the specific memory of the non-volatile core 1240, such as the core internal mass storage 1247 or ROM 1245. The software implementing the various embodiments of this application can be stored in such devices and executed by the core 1240. Depending on specific needs, the computer-readable medium can include one or more storage devices or chips. The software can cause the core 1240, and in particular the processors therein (including the CPU, GPU, FPGA, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM 1246 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise embodied in circuitry (e.g., accelerator 1244) that can replace software or operate in conjunction with software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to the computer-readable medium can include circuitry (such as an integrated circuit (IC)) that stores and executes the software, circuitry that embodies the logic, or both. This application encompasses any suitable combination of hardware and software.

[0170] Although at least two exemplary embodiments have been described for the present application, various changes, permutations, and various equivalent replacements of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design a variety of systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and thus are within the spirit and scope of the present application.

Claims

1. A method for video encoding, characterized in that, The method includes: Obtaining volume data of at least one visual three-dimensional (3D) scene; Converting the volume data into point cloud data; Projecting the point cloud data onto a two-dimensional (2D) image; Dividing the point cloud data projected onto the 2D image into at least two 3D partitions on a six-degree-of-freedom (6DoF) coordinate system, and independently encoding the 3D partitions at least at one of a video coding block and an image coding block; At a file or segment encapsulation block, encapsulating each independently encoded 3D partition as a track in a file or segment, where the track includes metadata related to an observation position or angle, and the metadata includes: layout information of the 3D partitions with indexes, 3D positions of each 3D partition on the 6DoF coordinate system, a list of representative observation positions or angles, a list of selected observation positions or angles corresponding to the 3D partitions, indexes of the 3D partitions corresponding to the selected observation positions or angles, quality level information of each 3D partition, and rendering information of each 3D partition depending on each observation position or angle; and Transmitting, by the file or segment encapsulation block, the file or segment to a video-based point cloud coding (V-PCC) player.

2. The method for video coding according to claim 1, characterized in that, The metadata is 360-degree virtual reality data.

3. A method for video decoding, characterized in that, The method includes: Receiving a file or segment, where the point cloud data projected onto a two-dimensional (2D) image is divided into at least two 3D partitions on a six-degree-of-freedom (6DoF) coordinate system, and each independently encoded 3D partition is encapsulated as a track in the file or segment, and the track includes metadata related to an observation position or angle, and the metadata includes: layout information of the 3D partitions with indexes, 3D positions of each 3D partition on the 6DoF coordinate system, a list of representative observation positions or angles, a list of selected observation positions or angles corresponding to the 3D partitions, indexes of the 3D partitions corresponding to the selected observation positions or angles, quality level information of each 3D partition, and rendering information of each 3D partition depending on each observation position or angle; Processing the file or segment, extracting an encoded bitstream, and parsing the metadata; Decoding the encoded point cloud data, reconstructing the decoded point cloud data, and obtaining reconstructed point cloud data.

4. The method for video decoding according to claim 3, wherein The metadata is 360-degree virtual reality data.

5. A video encoding device, characterized in that, The apparatus includes: A selection module for obtaining volume data of at least one visual three-dimensional (3D) scene; A conversion module for converting the volume data into point cloud data; A projection module for projecting the point cloud data onto a two-dimensional (2D) image; An encoding module for dividing the point cloud data projected onto the 2D image into at least two 3D partitions on a six-degree-of-freedom (6DoF) coordinate system, and independently encoding the 3D partitions at least at one of a video coding block and an image coding block; and A composing module, configured to encapsulate each independently encoded 3D partition into a track in a file or segment at a file or segment encapsulation block, the track including metadata related to an observation position or angle, the metadata including: layout information of the 3D partitions with indexes, 3D positions of each 3D partition on a 6DoF coordinate system, a list of representative observation positions or angles, a list of selected observation positions or angles corresponding to the 3D partitions, indexes of the 3D partitions corresponding to the selected observation positions or angles, quality level information of each 3D partition, and rendering information of each 3D partition depending on each observation position or angle; and the file or segment encapsulation block transmits the file or segment to a video-based point cloud coding technology V-PCC player.

6. A video decoding device, characterized in that, The apparatus includes: A receiving module, configured to receive a file or segment, wherein point cloud data projected onto a two-dimensional 2D image is divided into at least two 3D partitions on a six-degree-of-freedom 6DoF coordinate system, and each independently encoded 3D partition is encapsulated into a track in the file or segment, the track including metadata related to an observation position or angle, the metadata including: layout information of the 3D partitions with indexes, 3D positions of each 3D partition on a 6DoF coordinate system, a list of representative observation positions or angles, a list of selected observation positions or angles corresponding to the 3D partitions, indexes of the 3D partitions corresponding to the selected observation positions or angles, quality level information of each 3D partition, and rendering information of each 3D partition depending on each observation position or angle; and A decoding module, configured to process the file or segment, extract the encoded bitstream, and parse the metadata; decode the encoded point cloud data, and reconstruct the decoded point cloud data to obtain the reconstructed point cloud data.

7. A method for storing or transmitting a video bitstream, characterized in that, The video bitstream is generated according to the video coding method as claimed in claim 1 or 2, or the video bitstream is decoded based on the video decoding method as claimed in claim 3 or 4.

8. A non-volatile computer-readable storage medium, characterized in that, Stores multiple instructions, which are executed by at least one processor to execute the video coding method as claimed in claim 1 or 2, generate a bitstream and store it.

9. An electronic device, characterized in that, Includes a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the computer program, the method as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • 6dof media consumption architecture using 2d video decoder

    US20190114830A1

  • Method and device for transmitting or receiving 6dof video using stitching and re-projection related metadata

    WO2019066191A1

  • Method, apparatus and stream for volumetric video format

    WO2019079032A1