Video bitstream processing method and apparatus, device, and storage medium
By generating video parameters and data packaging packages for multi-view videos, the problem that traditional packaging formats cannot handle multi-view videos is solved, and effective packaging and transmission storage of multi-view videos are realized.
Patent Information
- Application Number
- PCT/CN2025/079183
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-02-26
- Publication Date
- 2025-09-04
AI Technical Summary
Traditional encapsulation formats cannot effectively encapsulate multi-view videos, and cannot meet the transmission and storage needs of multi-view videos.
A video code stream processing method is provided, by obtaining the video encoding parameter set and encoding data of the multi-view video code stream, generating the video parameter packaging package and data packaging package of each viewpoint, and generating a video packaging file of multi-viewpoint video.
It realizes effective packaging of multi-view video, meets the transmission and storage needs of multi-view video, and improves the transmission efficiency and quality of immersive media content.
Smart Images

Figure CN2025079183_04092025_PF_FP_ABST
Abstract
Description
Video stream processing method, device, equipment and storage medium
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2024102393812, filed on March 1, 2024, entitled “Video Stream Processing Method, Device, Equipment and Storage Medium”, the entire text of which is incorporated herein by reference. Technical Field
[0003] The present application relates to the field of computer technology, specifically to the technical field of multimedia data processing, and more particularly to a method, apparatus, device, and storage medium for processing video streams. Background Art
[0004] With the continuous development of digital technology, immersive media is gradually becoming part of our daily lives. Immersive media uses technologies such as virtual reality (VR) to immerse users in virtual environments, providing an immersive experience. Multi-view video is a crucial component of immersive media, providing users with the most intuitive interactive experience.
[0005] Before a video player plays video content, a video production device can obtain and encode the video content to obtain a corresponding multi-view video stream. The video production device can then encapsulate the video stream and transmit the resulting encapsulated file to the video player, which can then play the corresponding video content based on the encapsulated file. However, traditional encapsulation formats can only encapsulate standard video, not multi-view video, and therefore cannot meet the transmission and storage requirements of multi-view video. Summary of the Invention
[0006] The embodiments of the present application provide a video stream processing method, apparatus, device, and storage medium that can meet the transmission and storage requirements of multi-viewpoint videos and are suitable for multi-viewpoint video transmission and playback scenarios.
[0007] In a first aspect, an embodiment of the present application provides a method for processing a video stream, comprising:
[0008] Acquire a multi-view video stream, wherein the multi-view video stream includes video encoding parameter sets and video encoding data corresponding to multiple viewpoints respectively;
[0009] Determining decoding parameter configuration information for each of the multiple viewpoints according to the video encoding parameter sets corresponding to the multiple viewpoints;
[0010] Generating video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information of each of the multiple viewpoints, and generating video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints;
[0011] A video encapsulation file corresponding to the multi-viewpoint video stream is generated according to the video parameter encapsulation packets and the video data encapsulation packets of the respective multiple viewpoints.
[0012] In a second aspect, an embodiment of the present application provides a method for processing a video stream, including:
[0013] Obtaining a video encapsulation file, wherein the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each viewpoint in a plurality of viewpoints;
[0014] The video encoding data in the video data encapsulation package is decoded according to the decoding parameter configuration information in the video parameter encapsulation packages of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file.
[0015] In a third aspect, an embodiment of the present application provides a video stream processing device, the device comprising:
[0016] An acquisition unit, configured to acquire a multi-view video stream, wherein the multi-view video stream includes a video encoding parameter set and video encoding data corresponding to a plurality of viewpoints respectively;
[0017] a determining unit, configured to determine decoding parameter configuration information of each of the multiple viewpoints according to the video encoding parameter sets corresponding to the multiple viewpoints;
[0018] A generation unit is configured to generate video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information of each of the multiple viewpoints, and to generate video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints; and to generate a video encapsulation file corresponding to the multi-viewpoint video stream based on the video parameter encapsulation packets and the video data encapsulation packets for each of the multiple viewpoints.
[0019] In a fourth aspect, an embodiment of the present application provides a video stream processing device, the device comprising:
[0020] an acquiring unit, configured to acquire a video encapsulation file, wherein the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each of the multiple viewpoints;
[0021] The decoding unit is used to decode the video encoding data in the video data encapsulation package according to the decoding parameter configuration information in the video parameter encapsulation package of each of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file.
[0022] In the fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor, a communication interface and a memory, wherein the processor, the communication interface and the memory are interconnected, wherein the memory stores an executable program code, and the processor is used to call the executable program code to execute the method of the first aspect, or to execute the method of the second aspect.
[0023] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, the computer executes the video stream processing method of the first aspect, or executes the video stream processing method of the second aspect.
[0024] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, the video stream processing method of the first aspect is implemented, or the video stream processing method of the second aspect is implemented.
[0025] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0027] FIG1 is a schematic diagram of the structure of a package file provided in an embodiment of the present application;
[0028] FIG2 is a schematic diagram of the architecture of a video stream processing system provided in an embodiment of the present application;
[0029] FIG3 is a flow chart of a method for processing a video stream according to an embodiment of the present application;
[0030] FIG4 is a timing diagram of encapsulating a multi-view video stream according to an embodiment of the present application;
[0031] FIG5 is another flow chart of a method for processing a video stream provided by an embodiment of the present application;
[0032] FIG6 is a schematic diagram of an application of a video code stream processing method provided by an embodiment of the present application in a video live broadcast scenario;
[0033] FIG7 is a schematic diagram of an application of a video stream processing method provided in an embodiment of the present application in a video on demand scenario;
[0034] FIG8 is a schematic structural diagram of a video stream processing device provided in an embodiment of the present application;
[0035] FIG9 is another structural diagram of a video stream processing device provided in an embodiment of the present application;
[0036] FIG10 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] It should be noted in advance that, in order to enable those skilled in the art to better understand the technical solutions proposed in the embodiments of the present application, the embodiments of the present application will be combined with one or more drawings to clearly and completely describe the implementation of the technical solutions proposed in the embodiments of the present application. In addition, the various drawings shown in the embodiments of the present application are only exemplary illustrations. For example, the execution order of the various steps in the drawings can be adaptively adjusted according to the actual application scenario. In addition, in the embodiments of the present application, the block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in the form of software, or these functional entities can be implemented in one or more hardware modules or integrated circuits, or these functional entities can be implemented in different networks and / or processor devices and / or microcontroller devices.
[0039] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0040] It should be noted that the term "plurality" used in this document refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0041] The embodiments of the present application relate to the process of encoding and packaging the media content of immersive media. Immersive media refers to a form of media that places users in a virtual or enhanced environment through technologies such as virtual reality (VR) or augmented reality (AR) to provide an immersive experience. It enables users to observe and interact with objects and scenes in the real or virtual world in a three-dimensional manner, enhancing the user's sense of participation and experience. The immersive media mentioned in the embodiments of the present application refers to a media file that can provide immersive media content, so that users immersed in the media content can observe objects and scenes in the real world in a three-dimensional (3D) manner. Among them, immersive media can include videos represented in 3D space (referred to as 3D videos for short), and 3D videos include not only the color and brightness used to record the picture, but also the distance and direction of each pixel in the picture, that is, 3D videos can include depth information and motion information of the scene, presenting video content with a sense of three-dimensionality and depth. Among them, 3D video uses the parallax of the images seen by the human left and right eyes to give users a three-dimensional feeling when watching video content. In this scenario, the human left and right eyes can each serve as a viewpoint (also called a perspective). 3D video can be understood as a video including two viewpoints, and a video including two or more viewpoints can be called a multi-view video.
[0042] Among them, the process of encoding the media content of immersive media includes the process of video encoding of multi-viewpoint video. Video encoding refers to the compression of video signals, so as to reduce the amount of data required to represent the video signal and reduce the amount of data transmitted and stored. Video encoding is performed by an encoder, and the object of video encoding is to form a sequence of images included in the video signal. In the field of video coding and decoding, the image (image) included in the video signal can be called a picture (picture) or an image frame / video frame (frame). The encoder in the encoding end can encode the image to obtain a multi-viewpoint video stream, which can include a video coding parameter set and video coding data. Then the encoding end can send the video coding stream to the decoding end, and the decoder at the decoding end decodes the video coding data based on the video coding parameters in the video coding stream to obtain a decoded video. Taking the international video coding standard HEVC (High Efficiency Video Coding) as an example, the following series of operations and processing are performed on the input video signal:
[0043] (1) Block Partition Structure
[0044] Image partitioning refers to dividing each image in a video signal's image sequence into multiple non-overlapping image regions. Images in a video signal can be divided into slices, which can be further divided into blocks. Video coding can be performed on a block-by-block basis. The concept of a block can be further expanded in different video coding standards. For example, in the High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) standards, the concept of a block can be expanded to a Coding Tree Unit (CTU). A CTU can also be partitioned using a quadtree to generate one or more Coding Units (CUs). A CU is the most basic element in the coding process and can be the basic unit for image partitioning and encoding. Optionally, a CU can be square or rectangular in shape. Video coding involves encoding CUs one by one, organizing them into a continuous video coding stream. A Coding Tree Unit (CTU) is a basic unit used in video coding standards to represent an image region. A CTU can be further divided into multiple Coding Units (CUs) for image encoding. A CTU is typically 64x64 pixels in size. A coding unit (CU) is a basic unit used to represent an image region in video coding standards and serves as the fundamental unit of encoding. A CU can be square or rectangular and is typically used to divide and encode an image.
[0045] (2) Predictive Coding
[0046] Predictive coding is a method of coding based on the correlation between discrete signals, using one or more previous signals to predict the next signal. Prediction methods include intra-frame prediction and inter-frame prediction.
[0047] Intra-frame prediction means that since an image frame contains many areas with the same or similar colors, the pixel values in the same frame can be used for prediction, that is, prediction is made within an image frame to reduce spatial redundancy.
[0048] b. Inter-frame prediction refers to the situation where a video contains many image frames with only slight changes. Therefore, the pixel values in adjacent frames (such as the previous image frame) can be used for prediction. For example, the motion trajectory of a video of a continuous motion action in a group of images can be predicted to reduce temporal redundancy.
[0049] Among them, both intra-frame prediction and inter-frame prediction include multiple coding modes. The encoder can determine the coding cost required for each intra-frame prediction coding mode and / or each inter-frame prediction coding mode for the currently encoded CU, and thus determine the coding mode corresponding to the minimum coding cost based on the coding cost, that is, the coding mode corresponding to the CU.
[0050] When encoding multi-view videos, in addition to the aforementioned intra-frame prediction and inter-frame prediction methods, coding techniques based on non-independent viewpoints can also be used. For example, Multiview Extension of High Efficiency Video Coding (MV-HEVC) also includes disparity compensation prediction, inter-view motion prediction, and inter-view redundancy prediction.
[0051] a. Disparity-compensated prediction is a type of inter-frame prediction. Since image frames from different viewpoints at the same moment are relatively similar, prediction can be performed based on the pixel values of image frames from the reference viewpoint at the same moment.
[0052] b. Inter-viewpoint motion prediction means that since multi-viewpoint video includes video signals shot from different angles at the same time and the same scene, the motion of objects presented from different viewpoints is similar, so prediction can be made based on the motion information encoded in the reference viewpoint.
[0053] c. Inter-view redundant prediction means that due to the similarity between image frames of multiple viewpoints, the image frame of the current viewpoint can be predicted based on the redundant information in the encoded viewpoints (including the reference viewpoint and the currently encoded viewpoint).
[0054] (3) Transform
[0055] Since the human eye is more sensitive to low-frequency information and relatively insensitive to high-frequency information, the image can be converted from the spatial domain to the frequency domain by means of transformation, so that the high-frequency and low-frequency parts of the image can be separated. Specifically, the transformation can be a discrete Fourier transform (DFT) or a discrete cosine transform (DCT). For the CU, the object of the transformation is the residual information between the original image of the CU and the constructed image unit. This is because the encoding end only transmits the residual information, so the residual information can continue to be compressed, that is, the residual information is transformed. That is, the encoding end can transform the residual information, such as DCT transformation, to obtain a transformation matrix, so that the decoding end can perform an inverse transformation based on the transformation matrix to obtain the residual information. Each element in the transformation matrix is called a transformation coefficient. The transformation coefficients in the upper left portion of the transformation matrix represent the low-frequency component, while the transformation coefficients in the lower right portion represent the high-frequency component. Thus, in the pixel string obtained by performing a "Z"-shaped scan of the transformation rectangle, the front portion represents the high-frequency component, while the back portion represents the low-frequency component, which is the low-frequency component. A transformation matrix is a matrix obtained by transforming the residual information of an image during the video encoding process. Each element in the transformation matrix is called a transformation coefficient and represents the frequency domain information of the image. Transformation matrices can be used to compress image data and reduce redundancy.
[0056] (4) Quantization
[0057] After obtaining the transform matrix, the encoder can remove high-frequency information, which is less sensitive to the human eye, through a method called quantization. Quantization involves dividing each element in the transform matrix by a value called the quantization step (QStep). The QStep is determined by the quantization parameter (QP). There is a one-to-one correspondence between the QP and QStep. After determining the QP value, the QStep value can be determined by a table lookup. The encoder can quantize the transform matrix and use adaptive quantization techniques to determine the QP value for each CU, thereby determining the QStep for each CU. The decoder can determine the QStep based on the QP and then perform inverse quantization based on the QStep. A smaller QP corresponds to a smaller QStep, resulting in finer quantization, which preserves more detail, reduces distortion, and achieves a higher bitrate. Similarly, a larger QP corresponds to a larger QStep, resulting in coarser quantization, resulting in greater distortion and a lower bitrate. It can be understood that quantization is a lossy compression process. The quantization step size is a value used during video encoding to remove high-frequency information, which is less sensitive to the human eye. The quantization step size is determined by the quantization parameter (QP), and there is a one-to-one relationship between QP and QStep. A larger quantization step size results in coarser quantization and higher compression rates, but also greater distortion.
[0058] (5) Loop Filtering
[0059] In-loop filtering refers to the process of performing inverse quantization and inverse transformation on a transformed and quantized signal when using inter-frame prediction. Compared to the original image, the reconstructed image is affected by the quantization process, that is, the effect of lossy compression. Some image information in the reconstructed image is different from that in the original image, which is the distortion caused by the reconstructed image. In order to reduce the impact of using the reconstructed image as a reference image on subsequent predictions, the reconstructed image can be filtered. Since this filtering operation is within the encoding loop, it can be called in-loop filtering. Specifically, the reconstructed image can be filtered by a filter. The filter can be, for example, a deblocking filter, a sample-adaptive offset (SAO) filter, a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, a collaborative filter, etc. It can be understood that by filtering the reconstructed image and then using it as a reference frame image for subsequent coded image prediction, the distortion caused by quantization can be reduced. In-loop filtering refers to a filter used to reduce the distortion of the reconstructed image during the video encoding process. The loop filter performs filtering on the reconstructed image within the encoding loop to reduce the distortion caused by quantization and improve video quality.
[0060] (6) Statistical Coding / Entropy Coding
[0061] Entropy coding refers to converting the symbols used to represent a video sequence into a compressed code stream for transmission or storage to remove information entropy redundancy, thereby achieving the purpose of compression. The input symbols are coding parameters, which can include quantized transform coefficients (residual information), motion vector data (coding mode information) and other additional information, such as marker information and header information for correct decoding. Among them, entropy coding methods mainly include variable length coding (VLC) and arithmetic coding. VLC can include Huffman coding and Shannon-Feno coding, for example. Arithmetic coding includes context-adaptive binary arithmetic coding (CABAC), exponential coding, etc. It can be understood that the entropy coding method adopted by the encoding end can be informed to the decoding end, and the decoding end can perform decoding based on the encoding method of the encoding end.
[0062] Among them, since the binary symbols in the bitstream obtained after entropy coding are relatively important, the possibility of a lost or erroneous symbol will cause the video to be unable to be correctly decoded. Therefore, HEVC adopts a two-layer architecture of the Video Coding Layer (VCL) and the Network Abstract Layer (NAL) to cope with different network environments and video applications. The network adaptation layer refers to the layer used to adapt to different network environments in the video coding standard. The network adaptation layer includes the functions of dividing and encapsulating video coding data, such as defining the encapsulation format of the data and adapting the encoded data to various network environments. VCL includes the core functions of video coding, such as image partitioning, prediction, transformation, quantization, loop filtering and entropy coding of video signals as mentioned above. NAL includes the functions of dividing and encapsulating the encoded data output by VCL, such as defining the encapsulation format of the data and adapting the bit string generated by VCL to various network environments.
[0063] Specifically, in the NAL layer, the encoder in the encoding end can divide the encoded data into multiple data segments according to the content characteristics of the encoded data, and put them into data packets of NAL units (Network Abstract Layer Unit, NALU, Chinese for Network Adaptation Layer Unit). NAL unit refers to the basic unit used to encapsulate video encoding data in the video coding standard. NALU is divided into a header and a payload. The NALU header can be used to identify the type of data it carries, and the payload is used to carry the encoded data (VCL NALU) and other related data (non-VCL NALU):
[0064] Coded data can include intra-coded pictures (I-frames) and predictive-coded pictures (P-frames). An intra-coded picture (I-frame) is a complete image frame that exists independently of other frames and can be decoded independently without relying on information from other frames. An I-frame is similar to a static image and can be considered a key frame in a video sequence. In video coding, an I-frame provides a reference point for the video sequence, facilitating decoder resynchronization and video content recovery. An I-frame can be understood as a key frame, meaning that the I-frame image is fully preserved in the coded data, and decoding requires only the data from that I-frame. A predictive-coded picture (P-frame) is a non-key frame that relies on a previous key frame (I-frame) or other P-frames for predictive coding. A P-frame stores the difference between the current frame and a reference frame. Decoding requires superimposing this difference information with a previously cached reference frame to obtain the complete image content. A P-frame can represent a non-keyframe and can be understood as representing the difference between a P-frame and a previous keyframe (or P-frame). Decoding requires superimposing the difference between the P-frame and a previously cached image (the previous keyframe or P-frame) to obtain the P-frame. A bidirectionally interpolated prediction frame (B-frame) contains not only the information of the current frame but also the difference between the previous and next frames. B-frames rely on the previous and next reference frames for bidirectional prediction coding. Decoding requires using the information from both previous and next reference frames to reconstruct the content of the current frame.
[0065] Other relevant data includes the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (Picture Parameter Set), and Supplemental Enhancement Information (SEI). The Video Parameter Set (VPS) is a set of parameters used in video coding standards to transmit video classification information. The VPS contains information such as the overall structure of the encoded video data, such as the video hierarchy and viewpoint information. The VPS is primarily used to transmit video classification information, including the overall structure of the encoded video data. A multi-view video stream contains only one VPS. The Sequence Parameter Set (SPS) is a set of parameters used in video coding standards to describe the global parameters of a set of encoded video sequences. The SPS contains basic information about the video sequence, such as resolution, frame rate, and coding standard. The SPS includes a set of global parameters for a Coded Video Sequence (CVS), which is a sequence composed of the structure of the encoded pixel data of the original video frame. The Picture Parameter Set (PPS) is a set of parameters used in video coding standards to describe the common parameters used for an image. The PPS contains specific settings for each image frame, such as the initial quantization parameter and blocking information. The PPS also includes common parameters used for an image, meaning that each image frame may have different settings, such as self-reference information, initial QP, and blocking information.
[0066] Therefore, NALU can be divided into 7 types: VPS, SPS, PPS, SEI, I frame, P frame and B frame. NALU can be divided by start code, and the structure of multi-view video stream can be: start code (Start Code) + VPS + start code + SPS + start code + PPS + start code + SEI + start code + I frame + start code + P frame + start code + P frame + start code + B frame.... The VPS, SPS, PPS, SEI, I frame, P frame and B frame in the multi-view video stream are respectively contained in one NALU.
[0067] A start code is a specific byte sequence in a video encoding stream that identifies the beginning of a NAL unit. This helps decoders correctly parse the video encoding stream. The process of encapsulating immersive media content includes encapsulating the multi-view video stream described above. This is because the multi-view video stream needs to be encapsulated in a file container according to a pre-defined file format to form a media file resource for easy storage and transmission.
[0068] This application uses the example of encapsulating a multi-view video stream using the streaming media format (Flash Video, FLV) to obtain an encapsulated file. The FLV encapsulated file obtained by encapsulating using the FLV encapsulation format has the advantages of small size, fast loading speed, and high quality of media content obtained by decapsulation at the decoding end. This application chooses the FLV format to encapsulate the multi-view video stream because the FLV format has the advantages of small size, fast loading speed, and high decoding quality. The FLV format supports the encapsulation of multi-view video streams and can meet the transmission and storage requirements of multi-view videos.
[0069] Please refer to Figure 1, which is a structural diagram of an encapsulated file provided by an embodiment of the present application. As shown in Figure 1, the FLV encapsulated file includes a file header (FLV header) and a file body (FLV body). The file header of the FLV encapsulated file is 9 bytes in size and includes the FLV file format identifier, version number, data type included in the FLV encapsulated file, etc. The file body of the FLV encapsulated file can be composed of a combination of multiple encapsulation packages (tags) and the size of the previous encapsulation package (PreviousTagSize) of 4 bytes. Among them, the "size of the previous encapsulation package" after the file header is 0, because the previous one is a file header rather than an encapsulation package, and the other "sizes of the previous encapsulation package" are not 0. The FLV encapsulated file format can be shown in Table 1:
[0070] Table 1
[0071] As shown in Table 1, in the FLV encapsulated file body, after the file header, there is the first previous encapsulation package size 0 (PreviousTagSize0), which is 0, followed by encapsulation package 1 (Tag1) and the previous encapsulation package size 1 (PreviousTagSize1), which represent the first encapsulation package (First tag) and the size of the first encapsulation package, followed by encapsulation package 2 (Tag2) and the previous encapsulation package size 2 (PreviousTagSize2), which represent the second encapsulation package and the size of the second encapsulation package, and so on, until the previous encapsulation package size N-1 (PreviousTagSizeN-1), which represents the size of the N-1th encapsulation package, and encapsulation package N (Tag2) and the previous encapsulation package size N (PreviousTagSizeN), which represent the last encapsulation package and the size of the last encapsulation package. Among them, the size of the encapsulation package includes the header of the encapsulation package and is in bytes. All encapsulation packages are arranged in sequence to obtain the encapsulation package sequence of the FLV encapsulated file as shown in Figure 1. UI32 can refer to an unsigned 32-bit integer type.
[0072] The encapsulation package of an FLV file may include a script data encapsulation package (script tag), a video encapsulation package (video tag), and an audio encapsulation package (audio tag). The script data encapsulation package is used to encapsulate the media stream's metadata, such as duration, width, height, and frame rate. An FLV file may not include the script data encapsulation package. The video encapsulation package is used to encapsulate multi-viewpoint video streams, and the audio encapsulation package is used to encapsulate audio streams.
[0073] As shown in Figure 1, the video encapsulation package of the FLV encapsulation file encapsulates a multi-view video code stream unit, and can also encapsulate multiple multi-view video code stream units. The multi-view video code stream unit can be the above-mentioned NALU, and the NALU can include the code stream data of a video frame (image frame). Specifically, taking the encoding end encoding the video based on Advanced Video Coding (AVC) as an example, the structure of the video encapsulation package can be shown in Table 2:
[0074] Table 2
[0075] As shown in Table 2, the first byte of the video encapsulation packet is divided into two parts. The upper 4 bits are used to indicate the frame type (Frame Type), which is specifically defined as: 1 is a key frame (a seekable frame in the AVC multi-view video stream); 2 is a non-key frame (inter frame) (a non-seekable frame in the AVC multi-view video stream); 3 is a disposable inter frame (applicable only to the H.263 video coding standard); 4 is a generated key frame (reserved for server use only); and 5 is a video information / command frame, which is used to refer to a frame containing information or control commands about the video stream. The lower 4 bits of the first byte are used to indicate the codec type. For example, codec IDs 1-7 correspond to the corresponding codec types. Specifically, codec ID 7 indicates AVC (H264).
[0076] When the codec type is AVC, that is, when the multi-view video stream is an AVC stream, the second byte of the video encapsulation packet, AVC packet type (AVCPacketType), indicates the data type stored in the video encapsulation packet. Specifically, 0 represents the AVC sequence header (AVC sequence header), which can be the AVC decoding parameter configuration information (AVCDecoderConfigurationRecord) obtained based on the AVC multi-view video stream. 1 represents AVC NALU, that is, video encoding data obtained based on AVC, and 2 indicates the end of the AVC sequence.
[0077] When the multi-view video code stream is an AVC stream, the composition time (CTS) is the difference between the presentation time (PTS) and the decoding time (DTS) of the AVC video frame. PTS can be used to indicate the time when the receiver displays the frame on the display, and the unit is 1 / 1000 second. DST is the timestamp during transmission and can be used to indicate the order of decoding. CTS is used to store the PTS information of the video frame, the time when the receiver displays the frame on the display. Since the AVC encoding method may contain B frames, when B frames exist, the values of PTS and DTS may not be equal, and PTS may not necessarily increase monotonically. Therefore, a field is required to store PTS information.
[0078] The type indicates the type of the field. For example, UB[4] indicates a 4-byte field that can be used to store integers or certain specific data types. For example, SI24 in Table 1 can be used to represent an identifier or code. UI8 can be used to represent an unsigned 8-bit integer type, that is, a 1-byte integer.
[0079] Since the video encapsulation package shown in Table 2 above cannot encapsulate the HEVC multi-view video stream, an extension is made based on Table 2. The structure of the extended video encapsulation package can be seen in Table 2. The structure of the video encapsulation package shown in Table 3 can support the encapsulation of the HEVC multi-view video stream, as shown in Table 3:
[0080] Table 3
[0081] As shown in Table 3, the differences from Table 2 include: in the first byte of the video encapsulation packet, (frame type line) the key frame represented by 1 can be used to represent the key frame of the HEVC multi-view video stream. Similarly, 2 can represent the non-key frame of the HEVC multi-view video stream. 12 is added to the encoding ID to indicate that the codec type is HEVC. In the second byte of the video encapsulation packet, when the encoding ID is 12, it is the HEVC packet type (HEVCPacketType), which includes: HEVC sequence header (HEVC sequence header), HEVC NALU and HEVC sequence end. In addition, when the encoding ID is 12, the synthesis time CTS can be used to represent the difference between the PTS and DTS of the HEVC video frame.
[0082] In addition, immersive media may also include audio content synchronized with the video content. The audio content may be audio-encoded and encapsulated to obtain an audio encapsulation package, as shown in the audio encapsulation package in FIG1 . Based on the video encapsulation package and the audio encapsulation package, an encapsulation file may be obtained.
[0083] Based on the above description, in the process of encapsulating a multi-view video stream, such as an HEVC multi-view video stream, it is found that the above encapsulation method is only applicable to the encapsulation of a single-view video stream, that is, a video stream including only one viewpoint, which can be understood as encapsulating a two-dimensional (2D) video and cannot be applied to the encapsulation scenario of a multi-view video. Based on this, an embodiment of the present application provides a video stream encapsulation solution, which can be applied to the encapsulation scenario of a multi-view video stream of an immersive media, such as a 3D video. Specifically, a video encapsulation device can obtain a multi-view video stream, which includes a video encoding parameter set and video encoding data corresponding to multiple viewpoints respectively. Then, the video encapsulation device can determine the decoding parameter configuration information of each viewpoint based on the video encoding parameter set corresponding to the multiple viewpoints, and generate a video parameter encapsulation package for each viewpoint based on the decoding parameter configuration information of each of the multiple viewpoints, and generate a video data encapsulation package based on the video encoding data corresponding to each of the multiple viewpoints. Finally, a video encapsulation file corresponding to the multi-view video stream can be generated based on the video parameter encapsulation package and the video data encapsulation package. A multi-view video stream refers to video encoding data that contains multiple viewpoints, while a standard video stream only contains video encoding data from a single viewpoint. This application focuses on the encapsulation and processing of multi-view video streams. This video encapsulation file is a file corresponding to multi-view videos, enabling the encapsulation of multi-view video streams, which helps meet the transmission and storage requirements of multi-view videos and facilitates the transmission of immersive media content.
[0084] The video stream processing solution proposed in the embodiment of this application involves cloud storage and other technologies, among which:
[0085] Cloud storage is a new concept that has been extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0086] Currently, storage systems utilize a method for creating logical volumes. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may consist of disks on a specific storage device or several storage devices. When a client stores data on a logical volume, it stores the data on a file system. The file system divides the data into multiple parts, each of which is an object. An object contains not only the data but also additional information such as the data identifier. The file system writes each object to the physical storage space of the logical volume and records the storage location of each object. Therefore, when a client requests access to data, the file system can provide access based on the storage location of each object.
[0087] The storage system allocates physical storage space to logical volumes by pre-dividing the physical storage space into stripes based on the estimated capacity of the objects to be stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the Redundant Array of Independent Disks (RAID) groupings. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.
[0088] Based on the above description, please refer to Figure 2, which is a schematic diagram of the architecture of a video code stream encapsulation system provided in an embodiment of the present application. As shown in Figure 2, the video code stream encapsulation system may include a video production device 201, a video encoding device 202, a video encapsulation device 203, and a video decapsulation device 204. Among them, the video production device 201 can be a device for producing media content of immersive media, and the video encoding device 202 can be a device for performing video encoding on the media content of immersive media produced by the video production device 201 to obtain a multi-view video code stream. The video encapsulation device 203 can encapsulate the multi-view video code stream output by the video encoding device 202 to obtain an encapsulated file. Furthermore, the video encapsulation device 203 can transmit the encapsulated file to the video decapsulation device 204, and the video decapsulation device 204 can decapsulate the encapsulated file to obtain a multi-view video code stream corresponding to the encapsulated file. Optionally, the video decapsulation device 204 can also play the video code stream obtained by the decapsulation process.
[0089] Video encapsulation equipment refers to equipment used to encapsulate multi-view video streams. Its main function is to encapsulate multi-view video streams and their related parameters into video files for easy storage and transmission. Video decapsulation equipment refers to equipment used to decapsulate video files. Its main function is to parse the video parameter encapsulation packets and video data encapsulation packets in the encapsulation files and decode and play the video content.
[0090] The video encapsulation device is responsible for encapsulating the multi-view video stream and its related parameters into a video encapsulation file. The specific steps include: obtaining the multi-view video stream, parsing the video stream to extract the video encoding parameter set and video encoding data, generating a video parameter encapsulation package and a video data encapsulation package, and combining them into a video encapsulation file. The video encapsulation device can also transmit the encapsulated file to the video decapsulation device. In some embodiments, the video encapsulation device and the video decapsulation device can be integrated into the same electronic device, such as a video production device or a video playback device. In this case, the functions of the video encapsulation device and the video decapsulation device can be implemented by the same processor or multiple processors working in collaboration.
[0091] Among them, the video production device 201 is directly or indirectly connected to the video encoding device 202 through a wired or wireless manner, the video encoding device 202 can be directly or indirectly connected to the video encapsulation device 203 through a wired or wireless manner, and the video encapsulation device 203 can be directly or indirectly connected to the video decapsulation device 204 through a wired or wireless manner. It should be noted that the number and form of the devices shown in Figure 2 are for example only and do not constitute a limitation on the embodiments of the present application. In actual applications, the video production device 201, the video encoding device 202 and the video encapsulation device 203 can be the same electronic device or different electronic devices. In actual applications, the video decapsulation device 204 can be multiple electronic devices, and this application does not limit the number of video decapsulation devices 204. Electronic devices can also be called computer devices.
[0092] The following describes the electronic devices included in the video stream processing system:
[0093] (1) Video production equipment 201
[0094] Video production device 201 is an electronic device capable of acquiring immersive media content, that is, an electronic device for acquiring multi-viewpoint video. Specific acquisition methods include capturing real-world sound and visual scenes through a capture device, and generating them through an electronic device. Specifically, the capture device can refer to a hardware component configured in video production device 201, for example, a capture device including a microphone, a camera, and various sensors, such as an activated radar sensor. The capture device can also be a device directly or indirectly connected to video production device 201 via a wired or wireless connection. For example, if video production device 201 is a server, the capture device can be a camera connected to the server, providing the video production device 201 with the ability to acquire immersive media content. The capture device can include a camera and a sensor device, and can also include an audio device for acquiring audio content synchronized with the video content. For example, the camera device can include a standard camera, a depth camera, a light field camera, etc. The sensor device can include a laser device, a radar device, etc. The audio device can include an audio sensor, a microphone, etc. The capture device is deployed at a specific location in a real space to capture video content and audio content synchronized with the video content in the space.
[0095] (2) Video encoding device 202
[0096] The video encoding device 202 is a device that can perform video encoding on the video signal (video content) obtained by the video production device 201. The video signal (video content) can be understood as video data of a color mode (Red, Green, Blue, RGB) / luminance-bandwidth-chrominance (YUV) video data. The video encoding device 202 can use a coding standard such as the HEVC coding standard to perform image partitioning, predictive coding, transformation, quantization, loop filtering, and entropy coding on the video signal obtained by the video production device 201. The specific processing procedures of image partitioning, predictive coding, transformation, quantization, loop filtering, and entropy coding can be referred to the previous description and will not be repeated here. It should be noted that when the video signal captured by the video production device 201 is a video signal corresponding to a multi-view video, when the HEVC coding standard is used to encode the video signal, on the one hand, a frame of the multi-view video can be spliced from video frames of multiple viewpoints, so the video encoding device 202 can first segment the video frames of the multiple viewpoints. On the other hand, the video encoding device 202 can adopt the MV-HEVC coding mode for predicting the video signal of the multi-view video, thereby obtaining a multi-view video stream. The multi-view video stream can also be referred to as a video stream, a video coding stream, a multi-view video stream, etc. The multi-view video stream mentioned in the embodiment of the present application can be an HEVC multi-view video stream. The video encoding device 202 can also encode the audio signal corresponding to the video signal to obtain an audio stream.
[0097] (3) Video packaging equipment 203
[0098] The video encapsulation device 203 encapsulates the multi-view video stream output by the video encoding device 202. After obtaining the multi-view video stream, the video encapsulation device 203 obtains the multi-view video stream corresponding to the multi-view video, including video coding parameter sets and video encoding data corresponding to each of the multiple viewpoints. The video encapsulation device 203 can determine decoding parameter configuration information for each viewpoint based on the video coding parameter sets corresponding to the multiple viewpoints in the multi-view video stream. Furthermore, the video encapsulation device 203 can generate a video parameter encapsulation package for each viewpoint based on the decoding parameter configuration information, and a video data encapsulation package based on the video encoding data corresponding to each of the multiple viewpoints. Finally, the video encapsulation device 203 can generate a video encapsulation file corresponding to the multi-view video stream based on the video parameter encapsulation package and the video data encapsulation package for each viewpoint. After the encapsulation process, the video encapsulation device 203 can obtain a video file resource, such as an FLV encapsulation file, for the multi-view video. The video encapsulation device 203 may also encapsulate the multi-view video stream and the audio stream according to a file format to obtain an encapsulated file, which is a media file resource of the immersive media content.
[0099] (4) Video decapsulation equipment 204
[0100] The video decapsulation device 204 can receive media file resources transmitted by the video encapsulation device 203, such as the above-mentioned video encapsulation file. The video encapsulation device 203 can transmit media file resources to the video decapsulation device 204 based on various transmission protocols. Among them, the transmission protocol can be, for example, Dynamic Adaptive Streaming over HTTP (DASH) protocol, Dynamic Bitrate Adaptive Transmission (HTTP Live Streaming, HLS) protocol, Smart Media Transport Protocol (SMTP), Transmission Control Protocol (TCP), Real-Time Messaging Protocol (RTMP), etc.
[0101] After the video decapsulation device 204 obtains a media resource file, such as a video encapsulation file, it can decapsulate the video encapsulation file using a process that is inverse to the encapsulation process. For example, based on the decoding parameter configuration information in the video parameter encapsulation package for each of the multiple viewpoints included in the video encapsulation file, it decapsulates the video encoding data in the video data encapsulation package included in the video encapsulation file to obtain the multi-view video stream corresponding to the video encapsulation file. The video decapsulation device 204 can decapsulate the media resource file to obtain a multi-view video stream and an audio stream. The multi-view video stream and audio stream can then be decoded using a decoding process that is inverse to the encoding process to obtain the video content and audio content.
[0102] In one implementation, the video decapsulation device 204 can also be a device for playing the media content of the immersive media. The media encapsulation file can also carry display parameter information. The video decapsulation device 204 can render the video and audio content based on the display parameter information obtained from the decoding process. For example, it can render the audio content and 3D images. After rendering, the 3D video composed of the 3D images can be played. The video decapsulation device 204 can include multiple displays, each of which can be used to display the video content corresponding to a viewpoint.
[0103] Among them, any electronic device among the above-mentioned multiple electronic devices (such as video production device 201, video encoding device 202, video packaging device 203 and video decapsulation device 204) can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart voice interaction device, smart home appliance, car terminal, VR device (such as VR glasses, VR helmet, etc.), etc., but not limited to this. The above-mentioned video production device 201, video encoding device 202, video packaging device 203 and video decapsulation device 204 can also be a server, for example, it can be an independent physical server, or it can be a server cluster or distributed system composed of multiple physical servers, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.
[0104] In one implementation, the aforementioned multiple multi-view video streams, decoding parameter configuration information for each viewpoint, video parameter encapsulation packages for each viewpoint, video data encapsulation packages, and video encapsulation files can all be stored on a blockchain, thereby preventing tampering with the multiple multi-view video streams, decoding parameter configuration information for each viewpoint, video parameter encapsulation packages for each viewpoint, video data encapsulation packages, and video encapsulation files. Blockchain is a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. It is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block.
[0105] It can be understood that the video code stream processing system described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0106] Based on the above-described video stream processing scheme and video stream processing system, embodiments of the present application provide a video stream processing method. The video stream processing method described in embodiments of the present application can be performed by an electronic device, which can be the video encapsulation device 203 in the video stream processing system shown in Figure 2. The video encapsulation device 203 can be the same electronic device as the video production device 201 and the video encoding device 202. Please refer to Figure 3, which is a schematic flow chart of a video stream processing method provided in embodiments of the present application. The video stream processing method includes the following steps S301-S304:
[0107] S301: Acquire a multi-view video stream, parse the multi-view video stream, and obtain video encoding parameter sets corresponding to multiple viewpoints and video encoding data corresponding to multiple viewpoints.
[0108] A multi-view video stream is a video stream generated by encoding video content captured from multiple angles of the same scene. Each viewpoint corresponds to a specific viewing angle. The multi-view video stream contains the video encoding data and video encoding parameter sets for multiple viewpoints, supporting the transmission and decoding of multi-view video.
[0109] In the embodiments of the present application, a multi-view video stream can be a video stream obtained by encoding video content, and can also be referred to as a video encoding stream. This application uses an HEVC multi-view video stream as an example. To transmit and store video content, a video encapsulation device can encapsulate the multi-view video stream to obtain an encapsulated file. In other words, the multi-view video stream is a multi-view video stream. The multi-view video stream can be received by the video encapsulation device from other devices, or it can be obtained by encoding video content itself. This application does not limit this.
[0110] A viewpoint can be understood as observing a scene from a specific angle, while multiple perspectives can be understood as observing the scene from multiple different angles within the scene. For example, a 3D video that includes two viewpoints can be called multi-viewpoint video. These two viewpoints can be a set of video sequences captured by a camera array simultaneously capturing the same scene from two angles, each viewed by the user's left and right eyes, respectively. This allows the user to experience a three-dimensional stereoscopic visual experience based on the parallax between the left and right eyes (the difference in visual lines between the two eyes).
[0111] In one possible implementation, after the video encapsulation device obtains the multi-viewpoint video stream, it can parse and process the multi-viewpoint video stream to obtain the viewpoint identifier, the video encoding parameter set and video encoding parameters corresponding to the viewpoint identifier; and determine the video encoding parameter sets and video encoding data corresponding to the multiple viewpoint identifiers according to the video encoding parameter sets and video encoding data respectively.
[0112] The multi-view video stream may include a video coding parameter set and video coding data obtained by encoding the video content. A video coding parameter set is a set of parameters generated during video content encoding and is used to decode the multi-view video stream to restore the video content. The video coding parameters are the specific parameters that make up the parameter set, and the video coding data is the data obtained by encoding the video frames. Based on the video coding parameters in the video coding parameter set, the video coding data can be decoded to restore the video content. The video coding data may be binary data obtained by encoding the video frames. This data contains the coding information of the video content and is used for transmitting and storing the video content.
[0113] Video encoding data is encoded data obtained by encoding the video frames (images) included in the video content. Based on the video encoding parameters, the video encoding parameters can be decoded to obtain the restored video content. The viewpoint identifier is an identifier used to identify the viewpoint. One viewpoint identifier corresponds to one viewpoint. The viewpoint identifier can be a layer ID (layerID). For example, taking two viewpoints as an example, a layerID of 0 can be used to represent one viewpoint, such as the left viewpoint, and a layerID of 1 can be used to represent the other viewpoint, such as the right viewpoint.
[0114] Since a viewpoint corresponds to a video captured at a specific angle within a scene, parsing the multi-view video stream can yield the video coding parameter set and video coding data corresponding to the viewpoint identifier. This means that the video coding parameter sets and video coding data for different viewpoints can be distinguished by the viewpoint identifier. After parsing the multi-view video stream, the video packaging device can determine whether the multi-view video stream is a standard multi-view video stream (i.e., a single-view multi-view stream) or a multi-view video stream for multiple viewpoints based on the number of viewpoint identifiers obtained. For example, if only one viewpoint identifier is included, it can be determined to be a single-view multi-view stream; otherwise, it can be determined to be a multi-view video stream for multiple viewpoints. If multiple viewpoint identifiers are obtained through parsing, the video packaging device can correspond the multiple viewpoint identifiers to video coding parameter sets, determining them as the video coding parameter sets corresponding to the multiple viewpoints. Similarly, the video coding data corresponding to the multiple viewpoint identifiers can be determined as the video coding data corresponding to the multiple viewpoint identifiers. Thus, the video coding parameter sets and video coding data corresponding to the multiple viewpoints included in the multi-view video stream are obtained.
[0115] Specifically, taking the HEVC multi-view video stream as an example, the structure of the multi-view video stream can be a start code (Start Code) + VPS + start code + SPS + start code + PPS + start code + SEI + start code + I frame + start code + P frame + start code + P frame + start code + B frame.... In the multi-view video stream, VPS, SPS, PPS, SEI, I frame, and P frame are respectively contained in a NALU, then the video encapsulation device can parse the multi-view video stream to obtain a NALU sequence, and the NALU header (header) may include an identifier for indicating the type of data loaded in the NALU. Exemplarily, the structure of the HEVC NALU header (header) can be as shown in Table 4:
[0116] Table 4
[0117] As shown in Table 4, the first bit of the first byte of the NALU header is a forbidden zero bit, and its value defaults to 0, indicating that the NALU is valid, otherwise it is invalid so that the receiver can correct errors or discard the NALU. The second to seventh bits of the first byte of the NALU header can be used to indicate the type of NALU (nal unit type), including 32 categories for indicating coded data (VCL NALU) and 32 categories for indicating other related data (non-VCL NALU). In other related data (non-VCL NALU), there are categories for indicating VPS, SPS, PPS and SEI respectively. The last bit of the first byte of the NALU header and the first five bits of the second byte are the NALU layer ID (nuh_layer_id), which defaults to 0 and can be used to indicate the identification of the viewpoint. The last three digits of the second byte of the NALU header are the NALU temporal layer number + 1 (nuh temporal id plus 1), which defaults to 1. The value minus 1 can be used to indicate the temporal layer number of the NALU.
[0118] The right side of Table 4 shows the character type. For example, f(1) represents a prohibited bit, occupying the first bit of the NALU header. u(6) represents an unsigned 6-bit integer, and u(3) represents an unsigned 3-bit integer.
[0119] Thus, the video encapsulation device can be based on the number of layer IDs (nuh_layer_id) in the NALU header. If only the value of nuh_layer_id is parsed to be 0, it is determined that the multi-view video stream includes only one viewpoint identifier, and the multi-view video stream is a multi-view video stream of a single viewpoint. When multiple values of nuh_layer_id are parsed, such as the value of nuh_layer_id is 0 and the value of nuh_layer_id is 1, it can be determined that the multi-view video stream includes multiple viewpoint identifiers, and the multi-view video stream is a multi-view video stream of multiple viewpoints. The value of nuh_layer_id can also be other values greater than 0, which is not limited in this application.
[0120] Furthermore, the video encapsulation device can determine the NALU type corresponding to each viewpoint identifier as a NALU carrying VPS, SPS, PPS and SEI as the video coding parameter set corresponding to each viewpoint identifier, that is, the video coding parameter set corresponding to each viewpoint. In addition, the NALU type corresponding to each viewpoint identifier is each type of NALU of VCL NALU, which is determined as the video coding data corresponding to each viewpoint identifier, that is, the video coding data corresponding to each viewpoint. Further, the video encapsulation device can determine the decoding parameter configuration information of each of the multiple viewpoints based on the video coding parameter sets corresponding to the multiple viewpoints.
[0121] S302: Determine decoding parameter configuration information for each of the multiple viewpoints according to the video encoding parameter sets corresponding to the multiple viewpoints.
[0122] In an embodiment of the present application, the video coding parameter sets corresponding to the multiple viewpoints can be obtained by parsing the multi-view video stream by the video encapsulation device, and the NALUs corresponding to the multiple viewpoint identifiers are used to load the video coding parameter sets. The decoding parameter configuration information (such as expressed as HEVCDecoderConfigurationRecord) stores the information necessary for decoding the HEVC multi-view video stream, and may include the parameter sets VPS, SPS and PPS obtained by parsing. The decoding parameter configuration information refers to a set of parameters stored in the video encapsulation file, which is used to guide the decoder to correctly decode the video encoding data. The decoding parameter configuration information includes preset decoding parameter identifiers and their corresponding parameter values, which are used to describe the characteristics and decoding requirements of the video encoding data. These preset decoding parameter identifiers include video coding parameter sets, decoder configuration information, viewpoint identifiers, etc. Among them, the multiple viewpoints include a main viewpoint and at least one secondary viewpoint, and the video coding parameter sets corresponding to the multiple viewpoints include the main video coding parameter set corresponding to the main viewpoint, and the secondary video coding parameter sets corresponding to each secondary viewpoint. It is understandable that the video coding parameters of different viewpoints are encoded separately, and the layer ID (nuh_layer_id) in the NALU header is used to distinguish and correspond. It should be noted that since a multi-view video stream only includes one VPS, the layer ID of the VPS is the viewpoint identifier of the main viewpoint.
[0123] The primary viewpoint refers to the video stream that serves as the base or primary perspective in a multi-view video. The primary viewpoint's video coding parameter set and video coding data are the foundational components of the multi-view video stream and typically contain essential video content. Correspondingly, the secondary viewpoint refers to the video streams from perspectives other than the primary viewpoint in a multi-view video, used to enhance the stereoscopic visual effect. Together with the primary viewpoint's video coding data, the secondary viewpoint's video coding parameter set and video coding data provide richer visual information.
[0124] The decoding parameter configuration information can be the main viewpoint decoding parameter information corresponding to the main viewpoint, which is a structure including a preset decoding parameter identifier (first preset decoding parameter identifier). The video encapsulation device can obtain the parameter value corresponding to the first preset decoding parameter identifier from the main video encoding parameter set based on the first preset decoding parameter identifier and fill it into the structure. Exemplarily, taking the 0th layer of the HEVC stream used by the main viewpoint as an example, this layer can be called the base layer (Base Layer), that is, the nuh_layer_id in the NALU header (nal_unit_header) is 0, and the NALU type is a NALU loaded with VPS, SPS and PPS, and the parameter value corresponding to the first preset decoding parameter identifier is obtained. Furthermore, according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, the main viewpoint decoding parameter information is determined. That is, the video encapsulation device can fill the parameter value corresponding to the first preset parsing parameter identifier into the structure, thereby determining the main viewpoint decoding parameter information.
[0125] Since the primary view decoding parameter information corresponding to the primary view can only store the parameter values in the parameter set corresponding to the primary view, it is not sufficient to use only the primary view decoding parameter information for multi-view video codecs. Therefore, the relevant video codec standards define the structure of the secondary view decoding parameter information. The secondary view decoding parameter information can be represented as LHEVCDecoderConfigurationRecord. The L in the secondary view decoding parameter information indicates that the layer (Layered) is used to store some or all parameters in the secondary view coding parameter set for each secondary view. Similar to the primary view decoding parameter information, the secondary view decoding parameter information is also a structure that includes a preset decoding parameter identifier (second preset decoding parameter identifier). Based on the second preset decoding parameter identifier, the video encapsulation device can obtain the parameter value corresponding to the second preset decoding parameter identifier from the secondary video coding parameter set corresponding to each secondary view.
[0126] Furthermore, the video encapsulation device can fill the obtained parameter value corresponding to the second preset decoding parameter identifier into the secondary viewpoint decoding parameter information. Exemplarily, taking the secondary viewpoint as 1, and the secondary viewpoint using the first layer of the HEVC stream as an example, this layer can be called a secondary stereo layer (Secondary Stereo Layer), for example, the nuh_layer_id in the NALU header (nal_unit_header) is 1, and the NALU type is a NALU loaded with SPS and PPS, and the parameter value corresponding to the second preset decoding parameter identifier is obtained. Specifically, the video encapsulation device determines the secondary viewpoint decoding parameter information based on the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier. That is, the video encapsulation device can fill the parameter value corresponding to the second preset parsing parameter identifier into the structure, thereby determining the secondary viewpoint decoding parameter information.
[0127] It should be noted that when there are multiple secondary viewpoints, the parameter value corresponding to the second preset parsing parameter identifier can be obtained from the secondary video encoding parameter set corresponding to the multiple secondary viewpoints. That is, it can be understood that the secondary viewpoint decoding parameter information can be used to store the parameter values of part or all of the parameters in the video encoding parameter set of other layers (each secondary viewpoint).
[0128] The primary video encoding parameter set is contained in multiple primary multi-view video bitstream units associated with the primary viewpoint. The multi-view video bitstream units are NALUs. The multiple primary multi-view video bitstream units associated with the primary viewpoint are NALUs whose viewpoint identifier is the viewpoint identifier of the primary viewpoint, and the video encoding parameters are loaded in the NALUs. For example, taking the viewpoint identifier of the primary viewpoint as 0, the multiple primary multi-view video bitstream units associated with the primary viewpoint are multiple NALUs whose nuh_layer_id in the NALU header (nal_unit_header) is 0. In other words, the primary video encoding parameter set is contained in multiple NALUs whose nuh_layer_id is 0 and which carry the VPS, SPS, and PPS. The primary viewpoint decoding parameter information structure also includes a first preset array for writing the multi-view video bitstream units (NALUs). The video encapsulation device can write the multiple primary multi-view video bitstream units into the first preset array to obtain a primary viewpoint decoding data array.
[0129] Specifically, the video encapsulation device can use the viewpoint identifier in the NALU header as the viewpoint identifier of the main viewpoint, and load the NALU of VPS, SPS and PPS (which can be known by the nal unit type in the NALU header) into the first preset array of the main viewpoint decoding parameter information. Thus, the video encapsulation device can write the parameter value corresponding to the first preset decoding parameter identifier and the structure of the main viewpoint decoding data array (the array of NALUs containing the main video encoding parameter set) as the main viewpoint decoding parameter information.
[0130] Specifically, the syntax structure of the main view decoding parameter information can refer to Table 5:
[0131] Table 5
[0132] Among them, the syntax and semantics shown in Table 5 are as follows: ConfigurationVersion = 1 indicates that the configuration version is 1, and the parameter general_profile_space indicates whether the current NALU belongs to a general profile space. For example, when the value is 1, it can be indicated that the NALU is in accordance with the standard HEVC profile requirements. When the value is 0, it can be indicated that the NALU belongs to a specific non-standard profile space, which can be used for a specific application or extension. The parameter general_tier_flag indicates the tier information of the NALU, such as the Baseline Tier, the Main Tier, and the High Tier. The parameter general_profile_idc indicates the identifier of the General Profile to which the NALU belongs. The parameter general_profile_compatibility_flags indicates the HEVC profile that the current NALU is compatible with, which enables the decoder to determine whether the NALU is compatible with a specific HEVC profile and tier.
[0133] The parameter general_constraint_indicator_flags indicates whether the current NAL unit meets a series of specific coding constraints. The so-called constraints are for certain characteristics or requirements of video coding to be met, thereby ensuring consistent video quality on different decoders and playback devices. Constraints may include intra-frame prediction constraints, quantization parameter constraints, color space constraints, chroma sampling format constraints, loop filter constraints, etc. The parameter general_level_idc indicates the level of HEVC coding used by the current NALU. The HEVC standard defines multiple different levels, each corresponding to specific coding parameters and performance requirements, such as bit rate, frame rate, resolution, etc. The parameter min_spatial_segmentation_idc indicates the identifier of the minimum spatial segmentation level supported by the encoder, which is associated with the SPS. HEVC can support dividing the frame into different regions and encoding each region independently, that is, applying different coding parameters to different regions to improve coding efficiency and video quality.
[0134] The parameter parallelismType indicates the parameters or identifiers of HEVC parallel processing. This parameter is associated with SPS. Parallel coding techniques may include slices, threads, etc. The parameter chromaFormat indicates the sampling format of the chroma components in the video sequence, such as 4:2:0, 4:2:2, and 4:4:4. The chroma components usually represent the color information of the image, and the chroma sampling format determines the sampling method and resolution of the chroma components relative to the luminance components. The parameter bitDepthLumaMinus8 is a parameter in SPS that indicates the bit depth of the luminance component (Luma) minus 8. Bit depth refers to the number of bits used for each pixel value, which determines the color range and accuracy of the image. In HEVC, the bit depth of the luminance component can be any value from 8 to 16 bits, and bits are saved by subtracting 8 and storing the difference. The parameter bitDepthchromaMinus8 is a parameter in SPS that indicates the bit depth of the chroma component (Chroma) minus 8. It also saves bits by subtracting 8 from its value and storing the difference. The bit depth of the luminance component determines its sampling accuracy, while the bit depth of the chroma component determines its color accuracy.
[0135] The reserved field is reserved for future use. The reserved field is optional, has a variable length, and its content should not be decoded or used during decoding. This field is reserved for future standard versions or applications to support possible extensions or other functions. All bits of this field are set to 1, such as reserved = '11111'b, reserved = '111111'b, etc. in Table 5, which are standard placeholders. The parameter avgFrameRate is the average frame rate (Average Frame Rate), which indicates the average display rate of frames in the video sequence and is configured in the SPS and VPS. The parameter constantFrameRate is a parameter in the SPS, indicating whether the video sequence has a constant frame rate. If yes, the parameter value is 1, otherwise it is 0. The parameter numTemporalLayers is used to indicate the number of temporal layers (Temporal Layers) in the video sequence, which is used to implement layered coding. Layered coding refers to encoding a video sequence into multiple different quality layers or temporal layers to support different network conditions and device capabilities. Temporal layers allow the encoder to encode video at different temporal resolutions, thus providing better temporal scalability.
[0136] The parameter temporalIdNested indicates whether nested temporal hierarchy is enabled. The temporal hierarchy is a hierarchical architecture used to describe the dependencies between video frames. The parameter lengthsizeMinusone indicates the size of the length field of the NALU minus one. This parameter is set in the SPS or VPS and applies to all NALU units. The parameter numOfArrays indicates the length of the array. The parameter array_completeness indicates whether the current array contains a complete set of parameters. If the parameter value is 0, it means that the current array is incomplete, there is an undefined field or unfilled bytes. If the parameter value is 1, it means that the current array is complete. The parameter NAL_unit_type indicates the type of NALU, such as SPS, PPS, etc. The parameter numNalus indicates the number of NALUs. The parameter nalUnitLength indicates the length of the current NALU inside the for loop. The parameter bit(8*nalUnitLength)nalUnit indicates the actual data of the current NALU, with a length of 8 times nalUnitLength bits.
[0137] It can be seen that the main viewpoint decoding parameter information (HEVCDecoderConfigurationRecord) includes multiple first preset decoding parameter identifiers (such as the parameter identifiers of each parameter in Table 5). The video encapsulation device can obtain the value corresponding to the first preset decoding parameter identifier from the main video encoding parameter set (VPS, SPS and PPS corresponding to the viewpoint identifier of the main viewpoint) based on the first preset decoding parameter identifier, and write the NALU containing the VPS, SPS and PPS corresponding to the viewpoint identifier of the main viewpoint into the first preset array (such as the array-related parameters in Table 5) to obtain the main viewpoint decoding parameter information.
[0138] In one possible implementation, the main view decoding parameter information may include, in addition to the video coding parameter set of the main view (base layer), namely, VPS, SPS, and PPS, the NALU of the SEI of the three-dimensional reference display information (three_dimensional_reference_displays_info), which may be referred to as 3D SEI. Specifically, the 3D SEI includes a preset supplementary parameter identifier. The parameter value corresponding to the preset supplementary parameter identifier (such as the view identifier) can be obtained from the main video coding parameter set, so that the SEI (i.e., 3D SEI) can be determined based on the preset supplementary parameter identifier and the parameter value corresponding to the preset supplementary parameter identifier. Further, the video packaging device can determine the main view decoding parameter information based on the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, the main view decoding data array, and the 3D SEI. Specifically, the NALU of the 3D SEI can be written into the decoded data array of the main view decoding parameter information.
[0139] Specifically, the syntax structure of 3D SEI can refer to Table 6:
[0140] Table 6
[0141] As shown in Table 6, the parameter prec_ref_display_width indicates the pre-decoded display width of the reference image, which is used to inform the decoder of the original display size of the reference image so that it can be correctly scaled or processed during decoding. The parameter ref_viewing_distance_flag indicates whether the viewing distance information of the reference view is present so that it can be properly processed and displayed during decoding. The parameter prec ref viewing dist indicates the pre-decoded viewing distance. The parameter num_ref_displays_minus1 represents the number of reference displays minus 1, that is, the display device or screen used to present the reference image minus 1. The above 3D SEI takes two viewpoints as an example, including the left viewpoint and the right viewpoint, which are represented by the parameters left view id[i] and right view id[i]. For example, in the NALU header, the left viewpoint is the main viewpoint, the layerID is 0, and the layerID of the right viewpoint is 1, then the value of left view id[i] is 0, and the value of right view id[i] is 1. In this loop, other viewpoints, such as multiple secondary viewpoints, can also be included.
[0142] The parameter exponent_ref_display_width[i] indicates the exponential portion of the reference display width, and the parameter mantissa_ref_display_width[i] indicates the mantissa portion of the reference display width. The parameter exponent_ref_viewing_distance[i] indicates the exponential portion of the reference viewing distance. The parameter mantissa_ref_viewing_distance[i] indicates the mantissa portion of the reference viewing distance. The parameter additional_shift_present_flag[i] indicates whether additional offset information is present. num_sample_shift_plus512[i]: Indicates an additional offset, typically used to adjust the size or position of the image. The parameter three_dimensional_reference_displays_extension_flag indicates whether a flag for three-dimensional reference display extension information is present. If this flag is true, additional parameters may be included to describe information about the three-dimensional reference display.
[0143] It should be noted that the above-mentioned 3D SEI stores the parameters required for playing multi-viewpoint videos, such as viewpoint ID, rendering accuracy, etc., among which the parameters left view id[i] and right view id[i] can be derived from VPS, and the values of the remaining parameters can be configured by the user. It is understandable that Table 6 is only an example of 3D SEI, and the main viewpoint decoding parameter information can also include other SEIs, which are not limited in this application. Thus, the video packaging device can determine the main viewpoint decoding parameter information (HEVCDecoderConfigurationRecord).
[0144] Similar to the main viewpoint, the secondary video coding parameter sets corresponding to each secondary viewpoint are respectively included in the multiple secondary multi-view video code stream units associated with each secondary viewpoint, that is, each secondary viewpoint corresponds to a viewpoint identifier, and the multiple secondary multi-view video code stream units associated with each secondary viewpoint are NALUs whose viewpoint identifiers are the viewpoint identifiers of each secondary viewpoint, and the video coding parameters, such as SPS and PPS, are loaded in the NALU. For example, taking one secondary viewpoint and the viewpoint identifier of the secondary viewpoint as 1 as an example, the multiple secondary multi-view video code stream units associated with the secondary viewpoint are multiple NALUs whose nuh_layer_id in the NALU header (nal_unit_header) is 1. In other words, the secondary video coding parameter set of the secondary viewpoint is included in multiple NALUs whose nuh_layer_id is 1 and which are loaded with SPS and PPS.
[0145] Among them, similar to the main viewpoint decoding parameter information (HEVCDecoderConfigurationRecord), the secondary viewpoint decoding parameter information (LHEVCDecoderConfigurationRecord) includes not only a second preset parameter identifier, but also a second preset array. The video encapsulation device can write multiple secondary multi-view video stream units associated with each secondary viewpoint into the second preset array to obtain a secondary viewpoint decoding data array, so that the secondary viewpoint decoding parameter information (LHEVCDecoderConfigurationRecord) can be determined based on the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the secondary viewpoint decoding data array. It should be noted that when the number of secondary viewpoints is multiple, the secondary viewpoint decoding parameter information (LHEVCDecoderConfigurationRecord) includes the video encoding parameters of each secondary viewpoint and the NALU associated with each secondary viewpoint.
[0146] Specifically, the syntax structure of the secondary view decoding parameter information (LHEVCDecoderConfigurationRecord) can refer to Table 7:
[0147] Table 7
[0148] As shown in Table 7, the secondary view decoding parameter information (LHEVCDecoderConfigurationRecord) is very similar in structure to the primary view decoding parameter information (HEVCDecoderConfigurationRecord), with some fields missing. ConfigurationVersion = 1 indicates that the configuration version is 1. Similarly, the parameter min_spatial_segmentation_idc indicates the identifier of the minimum spatial segmentation level supported by the encoder, the parameter parallelismType indicates the parameter or identifier for parallel processing with HEVC, and the parameter numTemporalLayers is used to indicate the number of temporal layers in the video sequence for hierarchical coding. The parameter temporalIdNested indicates whether nested temporal layers are enabled. The parameter lengthsizeMinusone indicates the size of the NALU length field minus one. The parameter numOfArrays indicates the length of the array, the parameter array_completeness indicates whether the current array contains a complete parameter set, and the parameter NAL_unit_type indicates the type of the NALU. The parameter numNalus indicates the number of NALUs. The parameter nalUnitLength indicates the length of the current NALU within the for loop. The parameter bit(8*nalUnitLength)nalUnit indicates the actual data of the current NALU, and its length is 8 times nalUnitLength bits.
[0149] It should be noted that although the first preset parameter identifier and the second preset parameter identifier have the same parameter identifier, the corresponding parameter values are different, the meanings expressed are also different, and the NALUs written into the array are also different. Taking the number of secondary viewpoints as 1 as an example, the numOfArrays array of the secondary viewpoint decoding parameter information (LHEVCDecoderConfigurationRecord) stores the video encoding parameter set and multiple NALUs corresponding to the secondary stereo layer (Secondary Stereo Layer), that is, the SPS and PPS of the secondary viewpoint. Thus, the video encapsulation device can determine the secondary viewpoint decoding parameter information (LHEVCDecoderConfigurationRecord). Among them, the decoding parameter configuration information of each viewpoint includes the main viewpoint decoding parameter information and the secondary viewpoint decoding parameter information.
[0150] S303 : Generate video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information for each of the multiple viewpoints, and generate video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints.
[0151] In an embodiment of the present application, a video parameter encapsulation package is a video encapsulation package used to encapsulate decoding parameter configuration information for each viewpoint. A video parameter encapsulation package can be an encapsulation package used to encapsulate a video encoding parameter set, typically including decoding parameter configuration information. It is used to store and transmit video encoding parameters within a video encapsulation file. A video encapsulation package includes a video parameter encapsulation package and a video data encapsulation package, and may also include an audio data encapsulation package. Taking the FLV encapsulation file format as an example, the video encapsulation package is the video tag within FLV.
[0152] A video data package encapsulates video encoding data, typically containing the encoded data of a video frame, and is used to store and transmit video content within a video encapsulation file. This package is also known as the video tag in FLV. Video parameter packages and video data packages encapsulate different content: a video parameter package encapsulates decoding parameter configuration information for each viewpoint, while a video data package encapsulates the video encoding data corresponding to each viewpoint.
[0153] In some embodiments, the video encapsulation device may generate video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information for each of the multiple viewpoints, and generate video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints. The generation of the video data encapsulation packets depends on the decoding parameter configuration information in the video parameter encapsulation packets.
[0154] It should be noted that, since the embodiment of the present application is for encapsulation of multi-viewpoint video, the secondary viewpoint decoding parameter information (LHEVCDecoderConfigurationRecord) is newly added. Taking the FLV encapsulation file format as an example, the design is based on Table 3. The structure of the video encapsulation package can be shown in Table 8:
[0155] Table 8
[0156] As shown in Table 8, taking the HEVC multi-view video stream as an example, the only difference from Table 3 is in the HEVC / AVC packet type. 0 represents the HEVC sequence header, which is the primary view decoding parameter information (HEVCDecoderConfigurationRecord) obtained above, which stores the VPS, SPS, PPS, and 3D SEI of the primary view (base layer). 3 represents the LHEVC sequence header, which is the secondary view decoding parameter information (LHEVCDecoderConfigurationRecord) obtained above, which can store the SPS and PPS of each secondary view (Secondary Stereo Layer).
[0157] In one possible implementation, the video packaging device may determine the main view packaging identifier corresponding to the main view decoding parameter information based on the correspondence between the preset packaging category and the packaging identifier, and the packaging category to which the main view decoding parameter information belongs. The correspondence between the preset packaging category and the packaging identifier can be referred to Table 8. As shown in Table 8, if the packaging category of the main view decoding parameter information is HEVC sequence header, the corresponding packaging identifier includes: frame type (Frame Type) is 1, encoding ID is 12, HEVC packet type is 0, and synthesis time is 0.
[0158] Furthermore, the video encapsulation device may write the main view decoding parameter information into the first video encapsulation package and update the encapsulation identifier of the first video encapsulation package to the main view encapsulation identifier, thereby obtaining a main view video parameter encapsulation package. That is, the video encapsulation device may write the HEVCDecoderConfigurationRecord into the first video encapsulation package (video tag) and, referring to the description in Table 8, update the frame type (Frame Type) to 1, the encoding ID to 12, the HEVC packet type to 0, and the composition time to 0, thereby obtaining a main view video parameter encapsulation package.
[0159] Similarly, the video encapsulation device can determine the secondary viewpoint encapsulation identifier corresponding to the secondary viewpoint decoding parameter information based on the corresponding relationship and the encapsulation category to which the secondary viewpoint decoding parameter information belongs, and write the secondary viewpoint decoding parameter information into the second video encapsulation package, and update the encapsulation identifier of the second video encapsulation package to the secondary viewpoint encapsulation identifier to obtain the secondary viewpoint video parameter encapsulation package. That is, the video encapsulation device can write LHEVCDecoderConfigurationRecord into the second video encapsulation package (video tag), and according to the corresponding relationship shown in Table 8, update the frame type (Frame Type) to 1, the encoding ID to 12, the HEVC packet type to 3, and the synthesis time to 0 to obtain the secondary viewpoint video parameter encapsulation package. Thus, the video parameter encapsulation packages of each viewpoint have been obtained, including the main viewpoint video parameter encapsulation package and the secondary viewpoint video parameter encapsulation package.
[0160] In a possible implementation, the video encapsulation device further needs to encapsulate the video encoding data corresponding to the multiple viewpoints to obtain a video data encapsulation package. The video encoding data corresponding to the multiple viewpoints include the main video encoding data corresponding to the main viewpoint and the secondary video encoding data corresponding to each secondary viewpoint. Take the multi-view video stream including M video frames (image frames) as an example for explanation, and the M video frames have a certain arrangement order in the multi-view video stream. Specifically, the video encapsulation device can write the main video encoding data and the secondary video encoding data corresponding to each secondary viewpoint into M video encapsulation packages in sequence according to the arrangement order of the M video frames in the multi-view video stream to obtain a video data encapsulation package, where M is an integer greater than 1.
[0161] Among them, the i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, and i is an integer greater than 1 and less than or equal to M. It can be understood that, compared with the ordinary HEVC multi-view video stream, the HEVC multi-view video stream of the multi-view video has the video encoding data of each viewpoint corresponding to the video frame for one video frame. Since the video encoding data of each viewpoint corresponding to the video frame is captured at different angles of the same scene at the same time, it is rendered and played at the same time during playback, and the timestamp is the same. Therefore, the video encoding data of the video frame corresponding to each viewpoint can be regarded as the complete video encoding data of the video frame, and encapsulated in a video encapsulation package (video tag), thereby obtaining a video data encapsulation package based on the video encapsulation package of each frame in the M frames.
[0162] Specifically, the primary video coded data is contained in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is respectively contained in M secondary video data stream units associated with each secondary viewpoint. Specifically, the i-th video package includes the i-th primary video data stream unit among the M primary video data stream units and at least one i-th secondary video data stream unit; the at least one i-th secondary video data stream unit is respectively the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
[0163] It can be understood that one video frame corresponds to a NALU, and each NALU is the video coding data of each viewpoint in the video frame. For example, taking the HEVC multi-view video stream with two viewpoints as an example, the two viewpoints can be the left viewpoint and the right viewpoint, and one video frame corresponds to two NALUs, one is the NALU of the Base Layer (left viewpoint), and the other is the NALU of the Secondary Stereo Layer (right viewpoint). Because the video coding data corresponding to the left and right viewpoints are obtained by encoding videos from different angles taken at the same time, the two NALUs can be put together during encapsulation and regarded as a complete frame. By analogy, in the case of including N viewpoints, each video encapsulation package includes N NALUs, and these N NALUs are the NALUs containing video coding data corresponding to each viewpoint of the same video frame, and N is an integer greater than 1.
[0164] Thus, after obtaining a video tag of each video frame according to the arrangement order of the M video frames, the video tag of each video frame is used as a video data packet.
[0165] S304: Generate a video encapsulation file corresponding to the multi-viewpoint video stream according to the video data encapsulation package and the video parameter encapsulation packages of the multiple viewpoints.
[0166] A video encapsulation file is a file that encapsulates video content and its associated parameter information in a specific format. It typically consists of a video parameter package and a video data package, used to store and transmit video data. For example, the FLV package file is a common video encapsulation format that encapsulates video and audio data for easy network transmission and playback. A video encapsulation file may consist of two parts: a video parameter package and a video data package. The video parameter package contains the parameter information required for decoding, while the video data package contains the encoded video data. During the encapsulation process, a video parameter package and a video data package are first generated for each viewpoint, and then these are sequentially combined to form the complete video encapsulation file.
[0167] In an embodiment of the present application, a video encapsulation device can generate a video encapsulation file corresponding to the multi-view video stream based on the video parameter encapsulation packages of each viewpoint, namely, the primary viewpoint video parameter encapsulation package and the secondary viewpoint video parameter encapsulation package, and the video data encapsulation package, namely, the video encapsulation package (video tag) of each video frame. Specifically, the video encapsulation device can splice the primary viewpoint video parameter encapsulation package, the secondary viewpoint video parameter encapsulation package, and the video encapsulation packages of each video frame in the video parameter encapsulation package according to the structure shown in FIG1 and Table 1, and generate the size of each encapsulation package to obtain the encapsulation file corresponding to the multi-view video stream.
[0168] Optionally, the video encapsulation device can also encapsulate the audio code stream to obtain an audio encapsulation package (audio tag), and splice the audio encapsulation package and the video encapsulation package (video parameter encapsulation package and video data encapsulation package) according to the structure shown in Figure 1 and Table 1 to obtain the encapsulation file corresponding to the media content.
[0169] Please also refer to Figure 4, which is a timing diagram of encapsulating a multi-view video stream provided by an embodiment of the present application. As shown in Figure 4, 401. A video encapsulation device can obtain a multi-view video stream. This multi-view video stream can be a multi-view video stream output by a video encoding device encoding the video signal of a video frame according to encoder parameters set by the user. 402. The video encapsulation device can parse the multi-view video stream and obtain the video coding parameter sets for each viewpoint, including the VPS, SPS, and PPS for each viewpoint. Taking two views as an example, the video coding parameter sets include the video coding parameter sets for the base layer (layerID is 0) and the secondary stereo layer (layerID is 1). 403. A determination is made as to whether the stream is a single-view HEVC multi-view stream or a multi-view HEVC multi-view stream. 404. If the result of the determination is a multi-view HEVC multi-view stream, SEI (3D SEI) is generated. The SEI may be generated according to the parameters and VPS (such as left_view_id and right_view_id) required for playing the multi-view video configured by the user.
[0170] 405. Generate secondary viewpoint decoding parameter information LHEVCDecoderConfigurationRecord. This secondary viewpoint decoding parameter information can be generated based on the secondary video coding parameter set corresponding to each secondary viewpoint. 406. If the judgment result is a single-viewpoint HEVC multi-viewpoint video stream, steps 404 and 405 can be skipped and the step of generating primary viewpoint decoding parameter information HEVCDecoderConfigurationRecord can be performed. Specifically, it can be generated based on the primary video coding parameter set corresponding to the primary viewpoint. If the multi-viewpoint video stream is a multi-viewpoint HEVC multi-viewpoint video stream, the primary viewpoint decoding parameter information can be generated based on the primary video coding parameter set corresponding to the primary viewpoint and the supplemental enhancement information. 407. Generate a primary viewpoint video parameter package. The primary viewpoint video parameter package can be written with the primary viewpoint decoding parameter information, and the package identifier of the video package is updated to the primary viewpoint package identifier, such as updating the frame type (Frame Type) to 1, the encoding ID to 12, the HEVC packet type to 0, and the synthesis time to 0.
[0171] 408. It can be determined again whether it is a single-view HEVC multi-view video stream or a multi-view HEVC multi-view video stream, or the judgment result of step 403 can be directly obtained. 409. In the case where the judgment result is a multi-view HEVC multi-view video stream, a secondary-view video parameter encapsulation packet is generated. The secondary-view video parameter encapsulation packet can be written with the secondary-view decoding parameter information, and the encapsulation identifier of the video encapsulation packet is updated to the view encapsulation identifier, such as the frame type (Frame Type) is updated to 1, the encoding ID is updated to 12, the HEVC packet type is updated to 3, and the synthesis time is updated to 0. 410. A video data encapsulation packet is generated. In the case where the multi-view video stream is a multi-view HEVC multi-view video stream, each video encapsulation packet in the video data encapsulation packet includes a NALU for each of the multiple viewpoints, and the NALU for each viewpoint corresponds to the same frame of video data. Thus, the video parameter encapsulation packet and the video data encapsulation packet are spliced to obtain the encapsulation file corresponding to the multi-view video stream.
[0172] In an embodiment of the present application, after obtaining a multi-view video stream including video encoding parameter sets and video encoding data corresponding to multiple viewpoints, decoding parameter configuration information for each viewpoint is determined based on the video encoding parameter sets corresponding to the multiple viewpoints, and then a video parameter encapsulation package for each viewpoint of the video is generated based on the decoding parameter configuration information of each viewpoint, and a video data encapsulation package is generated based on the video encoding data corresponding to each viewpoint. Finally, a video encapsulation file corresponding to the multi-view video stream can be generated based on the video parameter encapsulation package and the video data encapsulation package for each viewpoint. On the one hand, the encapsulation of the multi-view video stream can be achieved, which is conducive to meeting the transmission and storage requirements of the multi-view video, thereby facilitating the transmission of immersive media content. On the other hand, the multi-view video is encapsulated into an encapsulation file including a video data encapsulation package and a video parameter encapsulation package, which has a simple structure, fast loading speed, and high quality of the media content obtained by decapsulation.
[0173] The video stream processing method described above obtains a multi-view video stream containing video coding parameter sets and video coding data corresponding to multiple viewpoints. Next, the multi-view video stream is parsed to extract viewpoint identifiers, video coding parameter sets, and video coding data. Based on the parsing results, the video coding parameter set and video coding data for each viewpoint are determined. The multi-view video includes a primary viewpoint and at least one secondary viewpoint, and the corresponding video coding parameter sets are a primary video coding parameter set and a secondary video coding parameter set, respectively. A first preset decoding parameter identifier and its corresponding parameter value are obtained from the primary video coding parameter set, and a second preset decoding parameter identifier and its corresponding parameter value are obtained from the secondary video coding parameter set. Based on these parameter values, primary viewpoint decoding parameter information and secondary viewpoint decoding parameter information are generated. The primary video coding parameter set is contained in multiple primary multi-view video stream units associated with the primary viewpoint. These units are written into a first preset array to form a primary viewpoint decoding data array. The primary viewpoint decoding parameter information is determined by combining the first preset decoding parameter identifier and its parameter value. Furthermore, a preset supplementary parameter identifier and its parameter value are obtained from the primary video encoding parameter set to generate supplementary enhancement information, which is then incorporated into the primary viewpoint decoding parameter information. For the secondary viewpoint, multiple secondary multi-viewpoint video stream units associated with the secondary viewpoint are written into a second preset array to form a secondary viewpoint decoding data array. The secondary viewpoint decoding parameter information is then determined by combining the second preset decoding parameter identifier and its parameter value.
[0174] Subsequently, based on the preset correspondence between encapsulation categories and encapsulation identifiers, the primary-view encapsulation identifier corresponding to the primary-view decoding parameter information and the secondary-view encapsulation identifier corresponding to the secondary-view decoding parameter information are determined. The primary-view decoding parameter information is written into a first video encapsulation packet, and the encapsulation identifier is updated to the primary-view encapsulation identifier, thereby generating a primary-view video parameter encapsulation packet. The secondary-view decoding parameter information is written into a second video encapsulation packet, and the encapsulation identifier is updated to the secondary-view encapsulation identifier, thereby generating a secondary-view video parameter encapsulation packet. Next, based on the primary and secondary video encoded data, the video encoded data for each view is sequentially written into video encapsulation packets in the order of the video frames in the multi-view video stream, thereby generating video data encapsulation packets. Each video encapsulation packet contains data corresponding to a video frame in the primary and secondary video encoded data, as well as data corresponding to a video frame in the secondary video encoded data. Finally, the primary-view video parameter encapsulation packet, the secondary-view video parameter encapsulation packet, and the video data encapsulation packet are combined to generate a video encapsulation file corresponding to the multi-view video stream, which can be used for efficient transmission and storage of multi-view video.
[0175] The above-mentioned video stream processing method obtains a multi-view video stream and parses its encoding parameter set and encoding data to generate decoding parameter configuration information, which then generates a video parameter encapsulation package and a video data encapsulation package, ultimately generating a video encapsulation file. This method uses a systematic encapsulation process to structure the encoding parameters and data of multi-view video, enabling efficient storage and transmission of multi-view video. With clear decoding parameter configuration information, the decoder can accurately decode the content of each viewpoint, thereby supporting the transmission and playback of complex immersive media (such as 3D video) and improving the user experience.
[0176] By extracting viewpoint identifiers, video encoding parameter sets, and video encoding data, it is possible to distinguish video content from different viewpoints. This parsing method ensures that the data from each viewpoint is not mixed up during the encapsulation process, improving the accuracy and reliability of the encapsulation. Furthermore, by clarifying the video encoding parameters and video encoding data for each viewpoint, it provides accurate input for the subsequent generation of decoding parameter configuration information, further optimizing the encapsulation process.
[0177] The concepts of primary and secondary viewpoints are introduced, processing the primary and secondary video coding parameter sets separately and generating corresponding decoding parameter configuration information. The primary viewpoint typically contains the basic video content, while the secondary viewpoint provides additional viewing angle information. By processing the primary and secondary viewpoints separately, the storage and transmission of video content can be optimized, while providing a clear decoding path for the decoder, supporting efficient decoding and playback of multi-viewpoint videos.
[0178] Extract the main viewpoint decoding parameter information from the main video encoding parameter set and write it to an array. By writing the main viewpoint encoding parameter set into an array, the main viewpoint decoding parameters can be systematically managed, facilitating the subsequent generation of decoding parameter configuration information. This structured approach improves the efficiency and scalability of the encapsulation process, ensuring that the main viewpoint content can be quickly identified and processed during decoding.
[0179] Supplementary enhancement information (such as 3D SEI) is further introduced and incorporated into the primary viewpoint decoding parameter information. Supplementary enhancement information provides additional display parameters and viewpoint information, enhancing video display quality and user experience. By incorporating this information into the primary viewpoint decoding parameter information, decoding devices can optimize video rendering and display based on these parameters, enhancing immersion, especially in 3D video or immersive media.
[0180] By writing the secondary view encoding parameter set into an array and generating decoding parameter information, this approach ensures that the secondary view video content is correctly processed during decoding, supporting the integrity and consistency of multi-view video. This is particularly important for complex multi-view video content (such as 3D video), ensuring that the video data of each viewpoint is accurately restored during decoding.
[0181] Generates a video parameter package, including primary and secondary viewpoint identifiers. By generating independent identifiers for each viewpoint, the decoding parameter information for the primary and secondary viewpoints can be clearly distinguished. This packaging method improves the readability and compatibility of the packaged file, allowing decoding devices to quickly identify and process content from different viewpoints, thereby supporting efficient decoding and playback.
[0182] By encapsulating the primary and secondary viewpoint video data in frame order, video data packets are generated, ensuring temporal coherence during transmission and storage. This encapsulation method supports efficient decoding and playback, while reducing data redundancy and improving transmission efficiency. With clear frame order encapsulation, the decoder can quickly locate and process the data of each video frame, ensuring smooth video playback.
[0183] The structure of the video data package includes video data stream units for the primary and secondary viewpoints. By clearly defining the package structure, decoding devices can accurately extract and process the video data for each viewpoint. This structured packaging approach improves the compatibility and scalability of the packaged file, supports complex multi-viewpoint video content, and ensures that the video data for each viewpoint is correctly processed during decoding.
[0184] This embodiment of the present application provides a method for processing a video stream. The method described in this embodiment of the present application can be performed by an electronic device, which may be the video decapsulation device 204 in the video stream processing system shown in Figure 2. Please refer to Figure 5, which is another schematic flow chart of a video stream processing method provided in this embodiment of the present application. The video stream processing method includes the following steps S501-S502:
[0185] S501: Obtain a video encapsulation file, where the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each of a plurality of viewpoints.
[0186] In an embodiment of the present application, a video encapsulation file is a file encapsulating video content, and may also be a file encapsulating media content including video content and audio content. The video encapsulation file is decapsulated, and the media content corresponding to the video encapsulation file can be played. The video encapsulation file may be a video encapsulation file sent by a video encapsulation device to a video decapsulation device, such as a video decapsulation device obtaining a video encapsulation file pushed by a video encapsulation device by pulling a stream. A viewpoint is a perspective of observing a scene from a certain angle of a scene, and multiple perspectives can be understood as perspectives of observing the scene from multiple different angles in the scene. A video parameter encapsulation package is an encapsulation package for encapsulating a video encoding parameter set obtained by encoding the video content, and a video data encapsulation package is an encapsulation package for encapsulating encoded data obtained by encoding the video frames included in the video content.
[0187] In one possible implementation, a video decapsulation device may decapsulate the video encapsulation file and parse it to obtain a video parameter encapsulation package, which may be a video encapsulation package that encapsulates an HEVCDecoderConfigurationRecord. When the parsed video encapsulation package includes a 3D SEI, such as a NALU of the 3D SEI, the video decapsulation device may determine that the video encapsulation file is a video encapsulation file for a multi-view video. The 3D SEI includes viewpoint identifiers for each viewpoint, and the video decapsulation device may also determine whether the video encapsulation file is a video encapsulation file for a multi-view video based on whether the number of viewpoint identifiers is greater than a preset threshold, such as 1.
[0188] In another possible implementation, the video decapsulation device can decapsulate the video encapsulation file and parse it into a video parameter encapsulation package. If the parsed video parameter encapsulation package includes a video encapsulation package for encapsulating HEVCDecoderConfigurationRecord and a video encapsulation package for encapsulating LHEVCDecoderConfigurationRecord, such as based on the HEVC packet type identifier in the video encapsulation package (video tag) being 0 and 3, the video decapsulation device can determine that the video encapsulation file is a video encapsulation file for multi-viewpoint video.
[0189] When the video decapsulation device determines that the encapsulated file is a multi-view video encapsulation file, it can determine the video parameter encapsulation package and the video data encapsulation package for each viewpoint based on the video parameter encapsulation package and the video data encapsulation package corresponding to the viewpoint identifier of each viewpoint. This results in obtaining the video parameter encapsulation package and the video data encapsulation package for each of the multiple viewpoints included in the video encapsulation package. Furthermore, the video decapsulation device can decode the video encoded data in the video data encapsulation package based on the decoding parameter configuration information in the video parameter encapsulation package for each viewpoint, thereby obtaining the multi-view video stream corresponding to the video encapsulation file.
[0190] S502: Decode the video encoding data in the video data encapsulation package according to the decoding parameter configuration information in the video parameter encapsulation package of each viewpoint to obtain a multi-view video stream corresponding to the video encapsulation file.
[0191] In an embodiment of the present application, each viewpoint includes a primary viewpoint and at least one secondary viewpoint, and the decoding parameter configuration information in the video parameter encapsulation packet of each viewpoint includes primary viewpoint decoding parameter information and secondary viewpoint decoding parameter information, such as the above-mentioned HEVCDecoderConfigurationRecord (primary viewpoint decoding parameter information) and LHEVCDecoderConfigurationRecord (secondary viewpoint decoding parameter information). The primary viewpoint decoding parameter information and the secondary viewpoint decoding parameter information can be obtained by the video decapsulation device decapsulating the video parameter encapsulation packet of each viewpoint, specifically based on the decoding parameter configuration information corresponding to the multiple viewpoint identifiers obtained by parsing. The primary viewpoint decoding parameter information is associated with the primary viewpoint, and the secondary viewpoint decoding parameter information is associated with at least one secondary viewpoint.
[0192] The video data encapsulation packet is generated by the video encapsulation device based on the video encoding data corresponding to multiple viewpoints. After receiving the video encapsulation file, the video decapsulation device can decapsulate the video parameter encapsulation packet to obtain decoding parameter configuration information corresponding to the multiple viewpoint identifiers; decapsulate the video data encapsulation packet to obtain the video encoding data corresponding to the multiple viewpoints; and decode the video encoding data corresponding to each viewpoint based on the decoding parameter configuration information of each viewpoint, thereby obtaining the multi-view video stream corresponding to the video encapsulation file.
[0193] Furthermore, the video data encapsulation package includes video encoding data, specifically including primary video encoding data corresponding to a primary viewpoint and secondary video encoding data corresponding to each secondary viewpoint. The video data encapsulation package includes video encapsulation packages (video tags) corresponding to M video frames, respectively. The i-th video encapsulation package includes the video encoding data of the i-th video frame in the primary video encoding data and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, where i is an integer greater than 1 and less than or equal to M, and M is the number of video frames included in the video content. The primary video encoding data is contained in the M primary video data stream units associated with the primary viewpoint, and the secondary video encoding data corresponding to each secondary viewpoint is respectively contained in the M secondary video data stream units associated with each secondary viewpoint. That is to say, the above-mentioned i-th video package includes the i-th primary video data stream unit in the M primary video data stream units and at least one i-th secondary video data stream unit; the at least one i-th secondary video data stream unit is the i-th secondary video data stream unit in the M secondary video data stream units associated with each secondary viewpoint.
[0194] For example, taking three viewpoints as an example, a video encapsulation package includes 3 NALUs, which correspond to the three viewpoints respectively. These 3 NALUs are the video coding data of the 3 video frames corresponding to the same frame in the M-frame video frame. Thus, the video decoding device can decode the main video coding data based on the main viewpoint decoding parameter information, and decode the secondary video coding data corresponding to each secondary viewpoint according to the secondary viewpoint decoding parameter information to obtain the multi-viewpoint video stream corresponding to the video encapsulation file. That is, the video decoding device can decode the main video coding data (the NALU storing the video coding data of the main viewpoint) based on the main viewpoint decoding parameter information HEVCDecoderConfigurationRecord, and decode the secondary video coding data corresponding to each secondary viewpoint (the NALU storing the video coding data of each secondary viewpoint) based on the secondary viewpoint decoding parameter information LHEVCDecoderConfigurationRecord.
[0195] Since the primary viewpoint decoding parameter information and secondary viewpoint decoding parameter information are structures, they store preset decoding parameter identifiers and corresponding parameter values. Preset decoding parameter identifiers are specific identifiers in the video coding parameter set used to identify parameters required for decoding. Preset supplementary parameter identifiers are identifiers used to identify supplementary enhancement information. During the encoding process, these identifiers and their corresponding parameter values are written into the video coding parameter set to guide the decoder in correctly decoding the video data.
[0196] Specifically, the primary viewpoint decoding parameter information includes a first preset decoding parameter identifier and a parameter value corresponding to the first preset decoding parameter identifier, and the secondary viewpoint decoding parameter information includes a second preset decoding parameter identifier and a parameter value corresponding to the second preset decoding parameter identifier. Thus, the video decoding device can decode the primary video encoded data based on the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and decode the secondary video encoded data corresponding to each secondary viewpoint based on the parameter value corresponding to the second preset decoding parameter identifier and the second preset decoding parameter identifier, thereby obtaining a multi-viewpoint video stream corresponding to the video encapsulation file.
[0197] That is, the video decapsulation device can decode the NALU storing the video encoding data of the main viewpoint according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier in HEVCDecoderConfigurationRecord, and decode the NALU storing the video encoding data of each secondary viewpoint according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier in LHEVCDecoderConfigurationRecord to obtain the multi-viewpoint video stream corresponding to the video encapsulation file.
[0198] In addition, the main viewpoint decoding parameter information also includes multiple main multi-viewpoint video code stream units associated with the main viewpoint, and the secondary viewpoint decoding parameter information also includes multiple secondary multi-viewpoint video code stream units associated with each secondary viewpoint. The video decapsulation device can decode the main video encoded data according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the multiple main multi-viewpoint video code stream units, and decode the secondary video encoded data corresponding to each secondary viewpoint in the video data encapsulation package according to the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the multiple secondary multi-viewpoint video code stream units associated with each secondary viewpoint to obtain a multi-viewpoint video code stream corresponding to the video encapsulation file.
[0199] It can be understood that HEVCDecoderConfigurationRecord also includes multiple NALUs associated with the main viewpoint, including video coding parameter sets, such as VPS, SPS, and PPS. Similarly, LHEVCDecoderConfigurationRecord also includes multiple NALUs associated with each secondary viewpoint, including video coding parameter sets, such as SPS and PPS corresponding to each secondary viewpoint. The video decapsulation device can then decode the NALU storing the video coding data of the main viewpoint based on the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the multiple primary multi-viewpoint video stream units. And based on the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the multiple secondary multi-viewpoint video stream units associated with each secondary viewpoint, the NALU storing the video coding data of each secondary viewpoint in the video data encapsulation packet is decoded to obtain the multi-viewpoint video stream corresponding to the video encapsulation file.
[0200] Specifically, the decoding process of the video decapsulation device may include: after obtaining the encoded data, entropy decoding is first performed according to the decoding process opposite to the encoding process to obtain the quantized transform coefficients (residual information), motion vector data (coding mode information) and other additional information. Then, inverse quantization and inverse transformation are performed based on the transform coefficients, and the frequency domain is transformed into the time domain to obtain the residual signal. And according to the motion vector data (coding mode information), the prediction signal corresponding to each CU can be determined, and after adding the prediction signal to the residual signal, the reconstructed signal can be obtained. Finally, the reconstructed signal (the reconstructed value of the decoded image) is processed by loop filtering to obtain the final (restored) video content.
[0201] Since the multi-view video stream corresponding to the video encapsulation file is a multi-view video stream of multiple viewpoints, that is, a multi-view video stream including multiple viewpoints, the main view decoding parameter information also includes supplementary enhancement information (i.e., SEI), that is, the HEVCDecoderConfigurationRecord also includes 3D SEI, and the 3D SEI includes multiple viewpoint identifiers. Taking the multiple viewpoint identifiers as two viewpoint identifiers as an example, the 3D SEI may include left_view_id and right_view_id. The 3D SEI includes the display parameters required to play the multi-view video. Supplementary enhancement information refers to information added to the video encoding stream to provide additional display parameters or enhance the video playback effect. For example, the 3D SEI may include information such as viewpoint identifiers and display parameters. If the video decoding device includes multiple displays, the video decoding device can display the video content corresponding to the multi-view video stream of the j-th viewpoint identifier on the j-th display based on the display parameters and multiple viewpoint identifiers in the SEI, where j is a positive integer. That is, based on each viewpoint identifier and display parameter in the SEI, the decoded video content is rendered, and then the video content corresponding to each viewpoint identifier is displayed on each display.
[0202] Optionally, the video decapsulation device can parse the audio tag in the video encapsulation file (such as the FLV encapsulation file) to obtain an audio code stream, and then decode the audio code stream to obtain the audio content corresponding to the audio code stream. When the video decapsulation device plays the video content, the audio content corresponding to the video content can be played.
[0203] In an embodiment of the present application, after obtaining a multi-view video stream including video encoding parameter sets and video encoding data corresponding to multiple viewpoints, decoding parameter configuration information for each viewpoint is determined based on the video encoding parameter sets corresponding to the multiple viewpoints, and then a video parameter encapsulation package for each viewpoint of the video is generated based on the decoding parameter configuration information of each viewpoint, and a video data encapsulation package is generated based on the video encoding data corresponding to each viewpoint. Finally, a video encapsulation file corresponding to the multi-view video stream can be generated based on the video parameter encapsulation package and the video data encapsulation package for each viewpoint. On the one hand, the encapsulation of the multi-view video stream can be achieved, which is conducive to meeting the transmission and storage requirements of the multi-view video, thereby facilitating the transmission of immersive media content. On the other hand, the multi-view video is encapsulated into an encapsulation file including a video data encapsulation package and a video parameter encapsulation package, which has a simple structure, fast loading speed, and high quality of the media content obtained by decapsulation.
[0204] The above-mentioned video stream processing method recovers the multi-view video stream by obtaining the video encapsulation file and decoding the video encoding data according to the decoding parameter configuration information. This method, through a systematic decapsulation process, can efficiently recover the multi-view video stream from the encapsulation file, support the decoding and playback of multi-view videos, and provide users with a high-quality immersive media experience.
[0205] The decoding process is further refined, extracting decoding parameter configuration information and video encoding data by decapsulating the video parameter and video data packets. This step-by-step decapsulation approach reduces the possibility of errors and improves decoding reliability. By processing video parameters and video data separately, the decoder can more efficiently recover video content, ensuring the accuracy of the decoding process.
[0206] The concepts of primary and secondary viewpoints are introduced, and decoding parameter information for each viewpoint is processed separately. This ensures that video content from each viewpoint is correctly decoded. This processing approach supports complex multi-viewpoint video content and improves the user experience, especially in 3D video and immersive media.
[0207] The process of decoding the primary and secondary viewpoint video encoded data based on preset decoding parameter identifiers ensures efficient and accurate decoding. This identifier-based decoding method can quickly locate and process video data, reduce decoding delays, and improve decoding efficiency.
[0208] We further introduce multi-view video stream units for primary and secondary views and incorporate them into the decoding process. By considering these units, we enable more comprehensive processing of video content and support complex multi-view video structures. This approach ensures that video content from each viewpoint is correctly decoded and displayed, enhancing the user experience.
[0209] The structure of video data packets is clarified, including packets corresponding to multiple video frames. This clear packet structure enables decoding devices to accurately extract and process the data of each video frame. This structured encapsulation method improves the compatibility and scalability of the encapsulated files, supporting efficient decoding and playback.
[0210] The structure of the video data stream units for the primary and secondary viewpoints has been further clarified. By defining the video data stream units for each viewpoint, decoding devices can accurately process the video content of each viewpoint. This structured processing approach improves decoding efficiency and accuracy, ensuring that the video data for each viewpoint is correctly processed during decoding.
[0211] Supplemental enhancement information is introduced, and video content is displayed on different displays based on display parameters and viewpoint identifiers. By utilizing this supplemental enhancement information, the video display quality can be optimized, enhancing the user experience. This display method supports an immersive multi-viewpoint video experience, and is particularly suitable for 3D video or VR / AR applications, significantly enhancing the user's sense of immersion.
[0212] It is clarified that video encapsulation files are generated using the video bitstream processing method described above for encapsulating video encapsulation files. By ensuring consistency between the generation and decapsulation methods for video encapsulation files, the overall compatibility and reliability of the system can be improved. This consistency ensures a seamless encapsulation and decapsulation process, reducing the possibility of errors and data loss, thereby improving overall system performance.
[0213] The above content introduces the specific execution process of the video code stream processing method and the video code stream processing system suitable for implementing the video code stream processing method. The following will introduce the scenarios in which the video code stream processing method is applicable:
[0214] (1) Live video streaming scenario
[0215] Please refer to Figure 6, which is a schematic diagram of the application of a video code stream processing method provided by an embodiment of the present application in a video live broadcast scenario. As shown in Figure 6, in a video live broadcast scenario, a video encapsulation device can be an electronic device used by the anchor user. The video encapsulation device can provide the functions of video production, video encoding and video encapsulation for multi-viewpoint videos. A video live broadcast application can be run in the video encapsulation device. The video encapsulation device can refer to a device for encapsulating and processing video encoding data. It can encapsulate video encoding data and parameter sets into a video encapsulation file for easy storage and transmission. The video decapsulation device is an electronic device used by the audience user who watches the live broadcast of the anchor. The video decapsulation device can provide the function of decapsulating the video encapsulation file and then rendering and playing the multi-viewpoint video. The video decapsulation device can include multiple displays, each for displaying the video content corresponding to multiple viewpoints. The video decapsulation device can refer to a device for receiving and decapsulating video encapsulation files. It can parse the video parameter encapsulation package and the video data encapsulation package in the encapsulation file, and decode and play the video content.
[0216] First, the video encapsulation device can obtain a multi-view video stream and, based on the video encoding parameter sets corresponding to the multiple viewpoints in the multi-view video stream, determine the decoding parameter configuration information for each viewpoint. Furthermore, the video encapsulation device can generate a video parameter encapsulation package for each viewpoint based on the decoding parameter configuration information for each viewpoint, and generate a video data encapsulation package based on the video encoding data corresponding to each viewpoint. Finally, based on the video parameter encapsulation package and the video data encapsulation package for each viewpoint, a video encapsulation file corresponding to the multi-view video stream is generated. Optionally, the video encapsulation device can encode and encapsulate audio content corresponding to the video content corresponding to the multi-view video stream, and the resulting video encapsulation file can include the audio encapsulation package.
[0217] Furthermore, after obtaining the video encapsulation file, the video encapsulation device can push the video encapsulation file (FLV encapsulation file) to the cloud server through the RTMP transmission protocol through the push streaming tool. After the cloud server receives the video encapsulation file, it processes the video encapsulation file based on user needs, such as transcoding, denoising, enhancement, analysis, etc. Among them, the video encapsulation device can also push the video encapsulation file (FLV encapsulation file) to the content delivery network (CDN) through the RTMP transmission protocol through the push streaming tool. The cloud server can pull the video encapsulation file from the CDN and process the video encapsulation file based on user needs. Then, the video decapsulation device of the viewer user can pull the video encapsulation file through the video live broadcast application, decapsulate and decode the video encapsulation file, and render the decoded video content based on the display parameters, and then display (play) the video content of the video encapsulation file.
[0218] Optionally, the video decoding device can also decapsulate the audio package to obtain an audio stream, and then decode the audio stream to obtain the audio content corresponding to the video content, so that the audio content can be played synchronously with the video content. In this way, the use cases of FLV are further expanded to meet the transmission and storage requirements of multi-viewpoint video. In this way, the entire video cloud chain, from encoding, streaming, cloud processing, and distribution, can support multi-viewpoint video, which can meet the viewing, transmission, storage, and processing requirements of multi-viewpoint video.
[0219] (2) Video on demand scenario
[0220] Video on Demand (VOD) refers to the ability to play corresponding video content according to the requirements of the viewer user, and can also be understood as transmitting the content that the user wants to watch (click or select) to the requested user. Please refer to Figure 7, which is a schematic diagram of the application of a video stream processing method provided by an embodiment of the present application in a video on demand scenario. As shown in Figure 7, in a video on demand scenario, the video encapsulation device can be an electronic device of a platform or user that provides video content (such as immersive media content). The video encapsulation device can provide the functions of video production, video encoding, and video encapsulation for multi-viewpoint videos. The video unpacking device is an electronic device used by viewer users who use the on-demand function. The video unpacking device can provide the function of unpacking the video encapsulation file, and then rendering and playing the multi-viewpoint video. The video unpacking device can include multiple displays, each used to display the video content corresponding to multiple viewpoints.
[0221] Specifically, the video encapsulation device can obtain a multi-view video stream and, based on the video encoding parameter sets corresponding to the multiple viewpoints in the multi-view video stream, determine the decoding parameter configuration information for each viewpoint. Based on the decoding parameter configuration information for each of the multiple viewpoints, it can then generate a video parameter encapsulation package for each viewpoint, and a video data encapsulation package based on the video encoding data corresponding to each of the multiple viewpoints. Finally, based on the video parameter encapsulation package and the video data encapsulation package for each viewpoint, it can generate a video encapsulation file corresponding to the multi-view video stream. Optionally, the video encapsulation device can encode and encapsulate audio content corresponding to the video content corresponding to the multi-view video stream, and the resulting video encapsulation file can include the audio encapsulation package.
[0222] Furthermore, after obtaining the video encapsulation file, the video encapsulation device can store the video encapsulation file (FLV encapsulation file) in cloud storage, such as Tencent Cloud COS. A cloud server or other device that processes video encapsulation files can obtain the video encapsulation file in the cloud storage and process the video encapsulation file. Then, the video unpacking device can obtain the corresponding video encapsulation file from the cloud storage based on the video content selected by the user, and unpack and decode the video encapsulation file, and render the decoded video content based on the display parameters, and then display (play) the video content of the video encapsulation file. Figure 7 takes the example of a user wearing VR glasses to watch the video content corresponding to the video encapsulation file, which can include video content corresponding to two viewpoints respectively. The VR glasses can also play the audio content corresponding to the video content in the video encapsulation file. In this way, the transmission and storage requirements of multi-viewpoint videos can be met, and it is stored as an FLV encapsulation file with a simple structure and fast loading speed. The quality of the media content obtained by unpacking is high.
[0223] The above describes in detail the method of the embodiment of the present application. In order to facilitate better implementation of the above scheme of the embodiment of the present application, the device of the embodiment of the present application is provided below accordingly.
[0224] Please refer to Figure 8, which is a schematic diagram of the structure of a video stream processing device provided in an embodiment of the present application. The video stream processing device 80 can be used to perform the corresponding steps in the video stream processing method shown in Figure 3. The video stream processing device 80 includes the following units:
[0225] An acquiring unit 801 is configured to acquire a multi-view video stream, where the multi-view video stream includes video coding parameter sets and video coding data corresponding to multiple viewpoints.
[0226] A determining unit 802 is configured to determine decoding parameter configuration information for each of the multiple viewpoints based on the video encoding parameter sets corresponding to the multiple viewpoints;
[0227] The generation unit 803 is used to generate video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information of each of the multiple viewpoints, and to generate video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints; and to generate a video encapsulation file corresponding to the multi-viewpoint video stream based on the video parameter encapsulation packets and video data encapsulation packets for each of the multiple viewpoints.
[0228] In a possible implementation, the video code stream processing device 80 further includes:
[0229] The processing unit 804 is configured to parse the multi-view video stream to obtain multiple view identifiers, video encoding parameter sets corresponding to the multiple view identifiers, and video encoding data;
[0230] The determining unit 802 determines the video coding parameter sets and video coding data corresponding to the multiple viewpoints respectively according to the video coding parameter sets and video coding data corresponding to the multiple viewpoint identifiers respectively.
[0231] In a possible implementation, the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; the video coding parameter sets corresponding to the multiple viewpoints include a primary video coding parameter set corresponding to the primary viewpoint and secondary video coding parameter sets corresponding to each secondary viewpoint;
[0232] The determining unit 802 is configured to determine decoding parameter configuration information for each of the multiple viewpoints based on the video coding parameter sets corresponding to the multiple viewpoints, specifically for:
[0233] Obtaining a parameter value corresponding to a first preset decoding parameter identifier from a primary video encoding parameter set, and obtaining a parameter value corresponding to a second preset decoding parameter identifier from a secondary video encoding parameter set corresponding to each secondary viewpoint;
[0234] The primary viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and the secondary viewpoint decoding parameter information is determined according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier.
[0235] In one possible implementation, the primary video encoding parameter set is included in multiple primary multi-view video stream units associated with the primary view; the determining unit 802 is configured to determine the primary view decoding parameter information according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, specifically for:
[0236] Writing a plurality of main multi-view video stream units into a first preset array to obtain a main view decoded data array;
[0237] The main viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the main viewpoint decoding data array.
[0238] In a possible implementation, the acquiring unit 801 is further configured to acquire a parameter value corresponding to a preset supplementary parameter identifier from the main video encoding parameter set, and determine the supplementary enhancement information according to the preset supplementary parameter identifier and the parameter value corresponding to the preset supplementary parameter identifier;
[0239] The determining unit 802 is configured to determine the primary viewpoint decoding parameter information according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the primary viewpoint decoding data array, specifically for:
[0240] The main viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, the main viewpoint decoding data array and the supplemental enhancement information.
[0241] In one possible implementation, the secondary video coding parameter sets corresponding to the respective secondary views are respectively included in a plurality of secondary multi-view video stream units associated with the respective secondary views; the determining unit 802 is configured to determine the secondary view decoding parameter information according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier, specifically to:
[0242] Writing a plurality of secondary multi-view video stream units associated with each secondary view into a second preset array to obtain a secondary view decoding data array;
[0243] Secondary viewpoint decoding parameter information is determined according to the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the secondary viewpoint decoding data array.
[0244] In a possible implementation, the generating unit 803 is configured to generate video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information for each of the multiple viewpoints, specifically to:
[0245] Determine the main viewpoint package identifier corresponding to the main viewpoint decoding parameter information according to the corresponding relationship between the preset package category and the package identifier and the package category to which the main viewpoint decoding parameter information belongs;
[0246] Determining a secondary-view encapsulation identifier corresponding to the secondary-view decoding parameter information according to the corresponding relationship and the encapsulation category to which the secondary-view decoding parameter information belongs;
[0247] Writing the main viewpoint decoding parameter information into the first video encapsulation package, and updating the encapsulation identifier of the first video encapsulation package to the main viewpoint encapsulation identifier, to obtain a main viewpoint video parameter encapsulation package;
[0248] The secondary viewpoint decoding parameter information is written into the second video encapsulation packet, and the encapsulation identifier of the second video encapsulation packet is updated to the secondary viewpoint encapsulation identifier to obtain a secondary viewpoint video parameter encapsulation packet.
[0249] In a possible implementation, the multiple viewpoints include a main viewpoint and at least one secondary viewpoint; the video encoding data corresponding to the multiple viewpoints include main video encoding data corresponding to the main viewpoint and secondary video encoding data corresponding to each secondary viewpoint;
[0250] The generating unit 803 is configured to generate a video data encapsulation packet based on the video encoding data corresponding to each of the multiple viewpoints, specifically for:
[0251] According to the main video encoding data and the secondary video encoding data corresponding to each secondary viewpoint, the primary video encoding data and the secondary video encoding data corresponding to each secondary viewpoint are sequentially written into M video encapsulation packets in accordance with the arrangement order of M video frames in the multi-view video stream to obtain a video data encapsulation packet, where M is an integer greater than 1;
[0252] Among them, the i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, and i is an integer greater than 1 and less than or equal to M.
[0253] In one possible implementation, the primary video coded data is included in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is included in M secondary video data stream units associated with each secondary viewpoint respectively;
[0254] The i-th video encapsulation package includes the i-th main video data stream unit among the M main video data stream units, and at least one i-th secondary video data stream unit; the at least one i-th secondary video data stream unit is the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
[0255] According to one embodiment of the present application, the steps involved in the method shown in FIG3 can all be performed by the various units in the video stream processing device shown in FIG8. For example, step S301 shown in FIG3 is performed by the acquisition unit 801 shown in FIG8, step S302 is performed by the determination unit 802 shown in FIG8, and steps S303 and S304 are both performed by the generation unit 803 shown in FIG8.
[0256] According to one embodiment of the present application, the various units in the video code stream processing device 80 shown in Figure 8 can be individually or completely merged into one or several other units to form a unit, or one (or some) of the units can be further divided into multiple functionally smaller units to form a unit, which can achieve the same operation without affecting the realization of the technical effects of the embodiment of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the video code stream processing device 80 can also include other units. In actual applications, these functions can also be assisted by other units and can be implemented by multiple units working together. According to another embodiment of the present application, the video code stream processing device 80 shown in Figure 8 can be constructed and the video code stream processing method of the embodiment of the present application can be implemented by running a computer program (including program code) capable of executing the various steps involved in the corresponding method shown in Figure 3 on a general-purpose computing device such as a general-purpose computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory medium (RAM), and a read-only memory medium (ROM). The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the video encapsulation device of the video code stream processing system shown in FIG. 2 through the computer-readable storage medium, and run therein.
[0257] Please refer to Figure 9, which is a schematic diagram of the structure of a video stream processing device provided in an embodiment of the present application. The video stream processing device 90 can be used to perform the corresponding steps in the video stream processing method shown in Figure 5. The video stream processing device 90 includes the following units:
[0258] An acquiring unit 901 is configured to acquire a video encapsulation file, where the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each of a plurality of viewpoints;
[0259] The decoding unit 902 is configured to decode the video encoding data in the video data encapsulation packet according to the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints, and obtain a multi-view video stream corresponding to the video encapsulation file.
[0260] In one possible implementation, the decoding unit 902 is configured to decode the video coded data in the video data encapsulation packet according to the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file, specifically for:
[0261] Decapsulating the video parameter encapsulation packets of the multiple viewpoints to obtain decoding parameter configuration information corresponding to the multiple viewpoint identifiers;
[0262] Decapsulating the video data package to obtain video encoding data corresponding to the multiple viewpoint identifiers;
[0263] According to the decoding parameter configuration information corresponding to the multiple viewpoint identifiers, the video encoding data corresponding to the multiple viewpoints are decoded to obtain the multi-viewpoint video stream corresponding to the video encapsulation file.
[0264] In a possible implementation, the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints includes primary viewpoint decoding parameter information and secondary viewpoint decoding parameter information, the primary viewpoint decoding parameter information is associated with the primary viewpoint, and the secondary viewpoint decoding parameter information is associated with at least one secondary viewpoint; the video encoding data includes primary video encoding data corresponding to the primary viewpoint, and secondary video encoding data corresponding to each secondary viewpoint;
[0265] The decoding unit 902 is configured to decode the video coded data in the video data encapsulation packet according to the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file, specifically for:
[0266] The main video coded data is decoded according to the main viewpoint decoding parameter information, and the secondary video coded data corresponding to each secondary viewpoint is decoded according to the secondary viewpoint decoding parameter information to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
[0267] In a possible implementation, the primary viewpoint decoding parameter information includes a first preset decoding parameter identifier and a parameter value corresponding to the first preset decoding parameter identifier; the secondary viewpoint decoding parameter information includes a second preset decoding parameter identifier and a parameter value corresponding to the second preset decoding parameter identifier;
[0268] The decoding unit 902 is configured to decode the primary video coded data according to the primary viewpoint decoding parameter information, and to decode the secondary video coded data corresponding to each secondary viewpoint according to the secondary viewpoint decoding parameter information, to obtain a multi-viewpoint video stream corresponding to the video encapsulation file. Specifically, the decoding unit 902 is configured to:
[0269] The main video encoded data is decoded according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and the secondary video encoded data corresponding to each secondary viewpoint is decoded according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
[0270] In a possible implementation, the primary viewpoint decoding parameter information further includes a plurality of primary multi-viewpoint video stream units associated with the primary viewpoint, and the secondary viewpoint decoding parameter information further includes a plurality of secondary multi-viewpoint video stream units associated with each secondary viewpoint;
[0271] The decoding unit 902 is configured to decode the primary video coded data according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and decode the secondary video coded data corresponding to each secondary viewpoint according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier to obtain a multi-view video stream corresponding to the video encapsulation file, specifically for:
[0272] The primary video coded data is decoded according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the plurality of primary multi-viewpoint video stream units, and the secondary video coded data corresponding to each secondary viewpoint in the video data encapsulation packet is decoded according to the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the plurality of secondary multi-viewpoint video stream units associated with each secondary viewpoint to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
[0273] In a possible implementation, the video data encapsulation packet includes video encapsulation packets corresponding to M video frames respectively;
[0274] Among them, the i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, and i is an integer greater than 1 and less than or equal to M.
[0275] In one possible implementation, the primary video coded data is included in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is included in M secondary video data stream units associated with each secondary viewpoint respectively;
[0276] The i-th video encapsulation package includes the i-th main video data stream unit among the M main video data stream units, and at least one i-th secondary video data stream unit; the at least one i-th secondary video data stream unit is the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
[0277] In a possible implementation, the multi-view video stream corresponding to the video encapsulation file includes a multi-view video stream with multiple viewpoint identifiers, the main viewpoint decoding parameter information also includes supplementary enhancement information, and the supplementary enhancement information includes multiple viewpoint identifiers; the video stream processing device 90 further includes:
[0278] The display unit 903 is configured to display the video content corresponding to the multi-view video stream with the j-th viewpoint identifier on the j-th display according to the display parameters and the multiple viewpoint identifiers in the supplemental enhancement information, where j is a positive integer.
[0279] According to one embodiment of the present application, each step involved in the method shown in FIG5 can be performed by each unit in the video stream processing device shown in FIG9. For example, step S501 shown in FIG5 is performed by the acquisition unit 901 shown in FIG9, and step S502 is performed by the decoding unit 902 shown in FIG9.
[0280] According to one embodiment of the present application, the various units in the video code stream processing device 90 shown in Figure 9 can be individually or completely merged into one or several other units to form a composition, or one (or some) of the units can be further divided into multiple functionally smaller units to form a composition, which can achieve the same operation without affecting the realization of the technical effects of the embodiment of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the video code stream processing device 90 can also include other units. In actual applications, these functions can also be assisted by other units and can be implemented by multiple units working together. According to another embodiment of the present application, the video code stream processing device 90 shown in Figure 9 can be constructed and the video code stream processing method of the embodiment of the present application can be implemented by running a computer program (including program code) capable of executing the various steps involved in the corresponding method shown in Figure 5 on a general-purpose computing device such as a general-purpose computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory medium (RAM), and a read-only memory medium (ROM). The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the video decapsulation device of the video stream processing system shown in FIG. 2 through the computer-readable storage medium, and run therein.
[0281] Based on the description of the above-mentioned video stream processing method embodiment, the present application also discloses an electronic device. Referring to FIG10 , the electronic device 100 may include at least a processor 1001, an input device 1002, an output device 1003, and a memory 1004. The processor 1001, input device 1002, output device 1003, and memory 1004 within the electronic device 100 may be connected via a bus or other means. The electronic device may be used as a video encapsulation device or a video decapsulation device.
[0282] The above-mentioned memory 1004 is a memory device in the electronic device 100, which is used to store programs and data. It can be understood that the memory 1004 here can include both the built-in storage medium of the electronic device and, of course, the extended storage medium supported by the electronic device 100. The memory 1004 provides a storage space, which stores the operating system of the electronic device 100. In addition, computer programs (including program codes) are also stored in the storage space. It should be noted that the computer storage medium here can be a high-speed RAM memory; optionally, it can also be at least one computer storage medium away from the aforementioned processor, and the aforementioned processor can be called a central processing unit (CPU), which is the core and control center of the electronic device and is used to run the computer program stored in the above-mentioned memory 1004.
[0283] In one embodiment, the processor 1001 may load and execute a computer program stored in the memory 1004 to implement the corresponding steps of the method in the above-mentioned embodiment of the video stream processing method. Specifically, the processor 1001 loads and executes the computer program stored in the memory 1004 to:
[0284] Obtain a multi-view video stream, where the multi-view video stream includes video encoding parameter sets and video encoding data corresponding to multiple viewpoints respectively;
[0285] Determining decoding parameter configuration information for each of the multiple viewpoints according to the video encoding parameter sets corresponding to the multiple viewpoints;
[0286] Generating video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information for each of the multiple viewpoints, and generating video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints;
[0287] A video encapsulation file corresponding to the multi-viewpoint video stream is generated according to the video parameter encapsulation packets and video data encapsulation packets of the respective multiple viewpoints.
[0288] In one possible implementation, the processor 1001 loads and executes a computer program stored in the memory 1004, and is further configured to:
[0289] Parsing the multi-view video code stream to obtain multiple view identifiers, video encoding parameter sets corresponding to the multiple view identifiers, and video encoding data;
[0290] According to the video coding parameter sets and video coding data respectively corresponding to the multiple viewpoint identifiers, the video coding parameter sets and video coding data respectively corresponding to the multiple viewpoints are determined.
[0291] In one possible implementation, the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; the video coding parameter sets corresponding to the multiple viewpoints include a primary video coding parameter set corresponding to the primary viewpoint and secondary video coding parameter sets corresponding to each secondary viewpoint; the processor 1001 loads and executes a computer program stored in the memory 1004 for determining decoding parameter configuration information for each of the multiple viewpoints based on the video coding parameter sets corresponding to the multiple viewpoints, specifically for:
[0292] Obtaining a parameter value corresponding to a first preset decoding parameter identifier from a primary video encoding parameter set, and obtaining a parameter value corresponding to a second preset decoding parameter identifier from a secondary video encoding parameter set corresponding to each secondary viewpoint;
[0293] The primary viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and the secondary viewpoint decoding parameter information is determined according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier.
[0294] In one possible implementation, the primary video encoding parameter set is included in a plurality of primary multi-view video stream units associated with a primary viewpoint; the processor 1001 loads and executes a computer program stored in the memory 1004 for determining the primary viewpoint decoding parameter information based on a first preset decoding parameter identifier and a parameter value corresponding to the first preset decoding parameter identifier, specifically for:
[0295] Writing a plurality of main multi-view video stream units into a first preset array to obtain a main view decoded data array;
[0296] The main viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the main viewpoint decoding data array.
[0297] In one possible implementation, the processor 1001 loads and executes a computer program stored in the memory 1004, and is further configured to:
[0298] Obtaining a parameter value corresponding to a preset supplementary parameter identifier from a main video encoding parameter set, and determining supplementary enhancement information according to the preset supplementary parameter identifier and the parameter value corresponding to the preset supplementary parameter identifier;
[0299] Determining primary viewpoint decoding parameter information according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the primary viewpoint decoding data array includes:
[0300] The main viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, the main viewpoint decoding data array and the supplemental enhancement information.
[0301] In one possible implementation, the secondary video coding parameter sets corresponding to the respective secondary views are respectively included in a plurality of secondary multi-view video stream units associated with the respective secondary views; the processor 1001 loads and executes a computer program stored in the memory 1004, configured to determine the secondary view decoding parameter information according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier, specifically configured to:
[0302] Writing a plurality of secondary multi-view video stream units associated with each secondary view into a second preset array to obtain a secondary view decoding data array;
[0303] Secondary viewpoint decoding parameter information is determined according to the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the secondary viewpoint decoding data array.
[0304] In one possible implementation, the processor 1001 loads and executes a computer program stored in the memory 1004, configured to generate video parameter encapsulation packages for each of the multiple viewpoints based on the decoding parameter configuration information for each of the multiple viewpoints, specifically configured to:
[0305] Determine the main viewpoint package identifier corresponding to the main viewpoint decoding parameter information according to the corresponding relationship between the preset package category and the package identifier and the package category to which the main viewpoint decoding parameter information belongs;
[0306] Determining a secondary-view encapsulation identifier corresponding to the secondary-view decoding parameter information according to the corresponding relationship and the encapsulation category to which the secondary-view decoding parameter information belongs;
[0307] Writing the main viewpoint decoding parameter information into the first video encapsulation package, and updating the encapsulation identifier of the first video encapsulation package to the main viewpoint encapsulation identifier, to obtain a main viewpoint video parameter encapsulation package;
[0308] The secondary viewpoint decoding parameter information is written into the second video encapsulation packet, and the encapsulation identifier of the second video encapsulation packet is updated to the secondary viewpoint encapsulation identifier to obtain a secondary viewpoint video parameter encapsulation packet.
[0309] In a possible implementation, the multiple viewpoints include a main viewpoint and at least one secondary viewpoint; the video encoding data corresponding to the multiple viewpoints include main video encoding data corresponding to the main viewpoint and secondary video encoding data corresponding to each secondary viewpoint;
[0310] The processor 1001 loads and executes a computer program stored in the memory 1004, configured to generate a video data encapsulation packet based on the video encoding data corresponding to each of the multiple viewpoints, specifically configured to:
[0311] According to the main video encoding data and the secondary video encoding data corresponding to each secondary viewpoint, the primary video encoding data and the secondary video encoding data corresponding to each secondary viewpoint are sequentially written into M video encapsulation packets in accordance with the arrangement order of M video frames in the multi-view video stream to obtain a video data encapsulation packet, where M is an integer greater than 1;
[0312] Among them, the i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, and i is an integer greater than 1 and less than or equal to M.
[0313] In one possible implementation, the primary video coded data is included in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is included in M secondary video data stream units associated with each secondary viewpoint respectively;
[0314] The i-th video encapsulation package includes the i-th main video data stream unit among the M main video data stream units, and at least one i-th secondary video data stream unit; the at least one i-th secondary video data stream unit is the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
[0315] In one possible implementation, the processor 1001 may load and execute a computer program stored in the memory 1004 to implement the corresponding steps in the above-mentioned embodiment of another method for processing a video stream. Specifically, the processor 1001 loads and executes the computer program stored in the memory 1004 to:
[0316] Obtaining a video encapsulation file, where the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each of the multiple viewpoints;
[0317] The video encoding data in the video data encapsulation package is decoded according to the decoding parameter configuration information in the video parameter encapsulation packages of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file.
[0318] In one possible implementation, the processor 1001 loads and executes a computer program stored in the memory 1004, configured to decode the video coded data in the video data encapsulation packet according to the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints, and obtain a multi-view video stream corresponding to the video encapsulation file, specifically for:
[0319] Decapsulating the video parameter encapsulation packets of the multiple viewpoints to obtain decoding parameter configuration information corresponding to the multiple viewpoint identifiers;
[0320] Decapsulating the video data package to obtain video encoding data corresponding to the multiple viewpoint identifiers;
[0321] According to the decoding parameter configuration information corresponding to the multiple viewpoint identifiers, the video encoding data corresponding to the multiple viewpoints are decoded to obtain the multi-viewpoint video stream corresponding to the video encapsulation file.
[0322] In a possible implementation, the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints includes primary viewpoint decoding parameter information and secondary viewpoint decoding parameter information, the primary viewpoint decoding parameter information is associated with the primary viewpoint, and the secondary viewpoint decoding parameter information is associated with at least one secondary viewpoint; the video encoding data includes primary video encoding data corresponding to the primary viewpoint, and secondary video encoding data corresponding to each secondary viewpoint;
[0323] The processor 1001 loads and executes a computer program stored in the memory 1004, configured to decode the video coded data in the video data encapsulation package according to the decoding parameter configuration information in the video parameter encapsulation packages of the multiple viewpoints, and obtain a multi-view video stream corresponding to the video encapsulation file, specifically for:
[0324] The main video coded data is decoded according to the main viewpoint decoding parameter information, and the secondary video coded data corresponding to each secondary viewpoint is decoded according to the secondary viewpoint decoding parameter information to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
[0325] In a possible implementation, the primary viewpoint decoding parameter information includes a first preset decoding parameter identifier and a parameter value corresponding to the first preset decoding parameter identifier; the secondary viewpoint decoding parameter information includes a second preset decoding parameter identifier and a parameter value corresponding to the second preset decoding parameter identifier;
[0326] The processor 1001 loads and executes a computer program stored in the memory 1004, configured to decode the primary video encoded data according to the primary viewpoint decoding parameter information, and decode the secondary video encoded data corresponding to each secondary viewpoint according to the secondary viewpoint decoding parameter information, to obtain a multi-viewpoint video stream corresponding to the video encapsulation file, specifically configured to:
[0327] The main video encoded data is decoded according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and the secondary video encoded data corresponding to each secondary viewpoint is decoded according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
[0328] In a possible implementation, the primary viewpoint decoding parameter information further includes a plurality of primary multi-viewpoint video stream units associated with the primary viewpoint, and the secondary viewpoint decoding parameter information further includes a plurality of secondary multi-viewpoint video stream units associated with each secondary viewpoint;
[0329] The processor 1001 loads and executes a computer program stored in the memory 1004, configured to decode the primary video coded data according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and decode the secondary video coded data corresponding to each secondary viewpoint according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier, to obtain a multi-view video stream corresponding to the video encapsulation file, specifically configured to:
[0330] The primary video coded data is decoded according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the plurality of primary multi-viewpoint video stream units, and the secondary video coded data corresponding to each secondary viewpoint in the video data encapsulation packet is decoded according to the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the plurality of secondary multi-viewpoint video stream units associated with each secondary viewpoint to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
[0331] In a possible implementation, the video data encapsulation packet includes video encapsulation packets corresponding to M video frames respectively;
[0332] Among them, the i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, and i is an integer greater than 1 and less than or equal to M.
[0333] In one possible implementation, the primary video coded data is included in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is included in M secondary video data stream units associated with each secondary viewpoint respectively;
[0334] The i-th video encapsulation package includes the i-th main video data stream unit among the M main video data stream units, and at least one i-th secondary video data stream unit; the at least one i-th secondary video data stream unit is the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
[0335] In one possible implementation, the multi-view video code stream corresponding to the video encapsulation file includes a multi-view video code stream with multiple viewpoint identifiers, the main viewpoint decoding parameter information also includes supplemental enhancement information, and the supplemental enhancement information includes multiple viewpoint identifiers; the processor 1001 loads and executes the computer program stored in the memory 1004, and is further configured to:
[0336] According to the display parameters and multiple viewpoint identifiers in the supplemental enhancement information, the video content corresponding to the multi-viewpoint video stream with the j-th viewpoint identifier is displayed on the j-th display, where j is a positive integer.
[0337] In an embodiment of the present application, after obtaining a multi-view video stream including video encoding parameter sets and video encoding data corresponding to multiple viewpoints, decoding parameter configuration information for each viewpoint is determined based on the video encoding parameter sets corresponding to the multiple viewpoints, and then a video parameter encapsulation package for each viewpoint of the video is generated based on the decoding parameter configuration information of each viewpoint, and a video data encapsulation package is generated based on the video encoding data corresponding to each viewpoint. Finally, a video encapsulation file corresponding to the multi-view video stream can be generated based on the video parameter encapsulation package and the video data encapsulation package for each viewpoint. On the one hand, the encapsulation of the multi-view video stream can be achieved, which is conducive to meeting the transmission and storage requirements of the multi-view video, thereby facilitating the transmission of immersive media content. On the other hand, the multi-view video is encapsulated into an encapsulation file including a video data encapsulation package and a video parameter encapsulation package, which has a simple structure, fast loading speed, and high quality of the media content obtained by decapsulation.
[0338] It should be understood that in the embodiments of the present application, the processor 1001 may be a central processing unit (CPU), and the processor 1001 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0339] In an embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the steps performed in all the above embodiments can be executed.
[0340] An embodiment of the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. When the computer instructions are executed by a processor of an electronic device, the methods in all the above embodiments are executed.
[0341] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The above-described program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The above-described storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0342] The above disclosure is only a preferred embodiment of the present invention, and certainly cannot be used to limit the scope of the rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
[0343] It is also particularly important to note that when the above embodiments of this application are applied to specific products or technologies, if it is necessary to obtain user data, the user's permission or consent must be obtained, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0344] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0345] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A video stream processing method, performed by an electronic device, comprising: Acquire a multi-view video stream, parse the multi-view video stream, and obtain video encoding parameter sets corresponding to each of the multiple viewpoints and video encoding data corresponding to each of the multiple viewpoints; Determining decoding parameter configuration information for each of the multiple viewpoints according to the video encoding parameter sets corresponding to the multiple viewpoints; Generating video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information of each of the multiple viewpoints, and generating video data encapsulation packets based on the video encoding data corresponding to each of the multiple viewpoints; and A video encapsulation file corresponding to the multi-viewpoint video stream is generated according to the video data encapsulation package and the video parameter encapsulation packages of the multiple viewpoints.
2. The method according to claim 1, wherein parsing the multi-view video stream to obtain video coding parameter sets corresponding to each of the multiple viewpoints and video coding data corresponding to each of the multiple viewpoints comprises: Parsing the multi-view video stream to obtain a plurality of viewpoint identifiers, video encoding parameter sets corresponding to the plurality of viewpoint identifiers, and video encoding data; According to the video coding parameter sets and video coding data respectively corresponding to the multiple viewpoint identifiers, the video coding parameter sets and video coding data respectively corresponding to the multiple viewpoints are determined.
3. The method according to claim 1 or 2, wherein the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; and the video coding parameter sets corresponding to the multiple viewpoints include a primary video coding parameter set corresponding to the primary viewpoint and secondary video coding parameter sets corresponding to each secondary viewpoint. The determining, according to the video encoding parameter sets corresponding to the multiple viewpoints, the decoding parameter configuration information of each of the multiple viewpoints includes: Obtaining a parameter value corresponding to a first preset decoding parameter identifier from the primary video encoding parameter set, and obtaining a parameter value corresponding to a second preset decoding parameter identifier from the secondary video encoding parameter set corresponding to each secondary viewpoint; Determining primary viewpoint decoding parameter information according to the first preset decoding parameter identifier and a parameter value corresponding to the first preset decoding parameter identifier; Secondary viewpoint decoding parameter information is determined according to the second preset decoding parameter identifier and a parameter value corresponding to the second preset decoding parameter identifier.
4. The method according to claim 3, wherein the primary video coding parameter set is included in a plurality of primary multi-view video stream units associated with the primary view; The determining the main viewpoint decoding parameter information according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier includes: Writing the plurality of main multi-view video stream units into a first preset array to obtain a main view decoded data array; The main viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the main viewpoint decoding data array.
5. The method according to claim 4, further comprising: Obtaining a parameter value corresponding to a preset supplementary parameter identifier from the main video encoding parameter set, and determining supplementary enhancement information according to the preset supplementary parameter identifier and the parameter value corresponding to the preset supplementary parameter identifier; The determining the primary viewpoint decoding parameter information according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the primary viewpoint decoding data array includes: The main viewpoint decoding parameter information is determined according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, the main viewpoint decoding data array, and the supplemental enhancement information.
6. The method according to any one of claims 3 to 5, wherein the secondary video coding parameter sets corresponding to the respective secondary views are respectively included in a plurality of secondary multi-view video stream units associated with the respective secondary views; The determining the secondary viewpoint decoding parameter information according to the second preset decoding parameter identifier and a parameter value corresponding to the second preset decoding parameter identifier includes: Writing the plurality of secondary multi-view video stream units associated with the respective secondary viewpoints into a second preset array to obtain a secondary viewpoint decoded data array; The secondary viewpoint decoding parameter information is determined according to the second preset decoding parameter identifier, a parameter value corresponding to the second preset decoding parameter identifier, and the secondary viewpoint decoding data array.
7. The method according to any one of claims 3 to 6, wherein generating the video parameter encapsulation packets for each of the multiple viewpoints based on the decoding parameter configuration information for each of the multiple viewpoints comprises: Determining the main viewpoint package identifier corresponding to the main viewpoint decoding parameter information according to a correspondence between a preset package category and a package identifier and the package category to which the main viewpoint decoding parameter information belongs; determining a secondary-view encapsulation identifier corresponding to the secondary-view decoding parameter information according to the corresponding relationship and the encapsulation category to which the secondary-view decoding parameter information belongs; Writing the main viewpoint decoding parameter information into a first video encapsulation package, and updating the encapsulation identifier of the first video encapsulation package to the main viewpoint encapsulation identifier, to obtain a main viewpoint video parameter encapsulation package; The secondary viewpoint decoding parameter information is written into a second video encapsulation packet, and the encapsulation identifier of the second video encapsulation packet is updated to the secondary viewpoint encapsulation identifier, to obtain a secondary viewpoint video parameter encapsulation packet.
8. The method according to any one of claims 1 to 7, wherein the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; and the video encoding data corresponding to the multiple viewpoints include primary video encoding data corresponding to the primary viewpoint and secondary video encoding data corresponding to each secondary viewpoint; The generating of the video data encapsulation packet based on the video encoding data corresponding to each of the plurality of viewpoints includes: Writing the primary video coded data and the secondary video coded data corresponding to each secondary viewpoint into M video encapsulation packets in the order of arrangement of M video frames in the multi-view video stream to obtain the video data encapsulation packet, where M is an integer greater than 1; Among them, the i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, and i is an integer greater than 1 and less than or equal to M.
9. The method according to claim 8, wherein the primary video coded data is contained in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is respectively contained in M secondary video data stream units associated with each secondary viewpoint; in, The i-th video encapsulation package includes the i-th main video data stream unit among the M main video data stream units, and at least one i-th secondary video data stream unit; The at least one i-th secondary video data stream unit is respectively the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
10. A video stream processing method, performed by an electronic device, comprising: Obtaining a video encapsulation file, wherein the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each viewpoint in a plurality of viewpoints; The video encoding data in the video data encapsulation package is decoded according to the decoding parameter configuration information in the video parameter encapsulation packages of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file.
11. The method according to claim 10, wherein decoding the video coded data in the video data encapsulation packet according to the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints to obtain the multi-view video stream corresponding to the video encapsulation file comprises: Decapsulating the video parameter encapsulation packets of the multiple viewpoints to obtain decoding parameter configuration information corresponding to the multiple viewpoint identifiers; Decapsulating the video data package to obtain video encoding data corresponding to the multiple viewpoint identifiers; According to the decoding parameter configuration information corresponding to the multiple viewpoint identifiers, the video encoding data corresponding to the multiple viewpoints are decoded to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
12. The method according to claim 10 or 11, wherein the multiple viewpoints include a primary viewpoint and at least one secondary viewpoint; the decoding parameter configuration information in the video parameter encapsulation packets of the multiple viewpoints includes primary viewpoint decoding parameter information and secondary viewpoint decoding parameter information, the primary viewpoint decoding parameter information is associated with the primary viewpoint, and the secondary viewpoint decoding parameter information is associated with the at least one secondary viewpoint; the video encoding data includes primary video encoding data corresponding to the primary viewpoint and secondary video encoding data corresponding to each secondary viewpoint; The decoding process of the video encoding data in the video data encapsulation package according to the decoding parameter configuration information in the respective video parameter encapsulation packages of the multiple viewpoints to obtain the multi-view video stream corresponding to the video encapsulation file includes: The main video encoding data is decoded according to the main viewpoint decoding parameter information, and the secondary video encoding data corresponding to each secondary viewpoint is decoded according to the secondary viewpoint decoding parameter information to obtain a multi-viewpoint video code stream corresponding to the video encapsulation file.
13. The method according to claim 12, wherein the primary viewpoint decoding parameter information comprises a first preset decoding parameter identifier and a parameter value corresponding to the first preset decoding parameter identifier; and the secondary viewpoint decoding parameter information comprises a second preset decoding parameter identifier and a parameter value corresponding to the second preset decoding parameter identifier; The decoding process of the primary video coded data according to the primary viewpoint decoding parameter information, and the decoding process of the secondary video coded data corresponding to each secondary viewpoint according to the secondary viewpoint decoding parameter information, to obtain a multi-viewpoint video stream corresponding to the video encapsulation file, includes: The main video encoded data is decoded according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and the secondary video encoded data corresponding to each secondary viewpoint is decoded according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
14. The method according to claim 13, wherein the primary viewpoint decoding parameter information further includes a plurality of primary multi-viewpoint video stream units associated with the primary viewpoint, and the secondary viewpoint decoding parameter information further includes a plurality of secondary multi-viewpoint video stream units associated with each secondary viewpoint; The decoding process of the primary video coded data according to the first preset decoding parameter identifier and the parameter value corresponding to the first preset decoding parameter identifier, and the decoding process of the secondary video coded data corresponding to each secondary viewpoint according to the second preset decoding parameter identifier and the parameter value corresponding to the second preset decoding parameter identifier, to obtain a multi-viewpoint video stream corresponding to the video encapsulation file, includes: The primary video coded data is decoded according to the first preset decoding parameter identifier, the parameter value corresponding to the first preset decoding parameter identifier, and the multiple primary multi-viewpoint video stream units, and the secondary video coded data corresponding to each secondary viewpoint in the video data encapsulation packet is decoded according to the second preset decoding parameter identifier, the parameter value corresponding to the second preset decoding parameter identifier, and the multiple secondary multi-viewpoint video stream units associated with each secondary viewpoint, to obtain a multi-viewpoint video stream corresponding to the video encapsulation file.
15. The method according to claim 14, wherein the video data encapsulation packet comprises video encapsulation packets corresponding to M video frames respectively; in, The i-th video encapsulation package includes the video encoding data of the i-th video frame in the main video encoding data, and the video encoding data of the i-th video frame in the secondary video encoding data corresponding to each secondary viewpoint, where i is an integer greater than 1 and less than or equal to M.
16. The method according to claim 15, wherein the primary video coded data is contained in M primary video data stream units associated with the primary viewpoint, and the secondary video coded data corresponding to each secondary viewpoint is respectively contained in M secondary video data stream units associated with each secondary viewpoint; in, The i-th video encapsulation package includes the i-th main video data stream unit among the M main video data stream units, and at least one i-th secondary video data stream unit; The at least one i-th secondary video data stream unit is the i-th secondary video data stream unit among the M secondary video data stream units associated with each secondary viewpoint.
17. The method according to any one of claims 12 to 16, wherein the multi-view video stream corresponding to the video encapsulation file includes a multi-view video stream with multiple viewpoint identifiers, the main viewpoint decoding parameter information further includes supplemental enhancement information, and the supplemental enhancement information includes the multiple viewpoint identifiers; the method further comprising: According to the display parameters in the supplemental enhancement information and the multiple viewpoint identifiers, the video content corresponding to the multi-viewpoint video stream with the j-th viewpoint identifier is displayed on the j-th display, where j is a positive integer.
18. The method according to any one of claims 10-11, wherein the video encapsulation file is generated by using the video code stream processing method according to any one of claims 1-9.
19. A video stream processing device, comprising: An acquisition unit, configured to acquire a multi-view video stream, wherein the multi-view video stream includes a video encoding parameter set and video encoding data corresponding to a plurality of viewpoints respectively; a determining unit, configured to determine decoding parameter configuration information of each of the multiple viewpoints according to the video encoding parameter sets corresponding to the multiple viewpoints; a generating unit, configured to generate video parameter encapsulation packets for each of the plurality of viewpoints based on the respective decoding parameter configuration information of the plurality of viewpoints, and generate video data encapsulation packets based on the respective video encoding data corresponding to the plurality of viewpoints; And generating a video encapsulation file corresponding to the multi-viewpoint video stream according to the video parameter encapsulation packets and the video data encapsulation packets of the respective multiple viewpoints.
20. A video stream processing device, comprising: an acquiring unit, configured to acquire a video encapsulation file, wherein the video encapsulation file includes a video parameter encapsulation package and a video data encapsulation package for each of the multiple viewpoints; The decoding unit is used to decode the video encoding data in the video data encapsulation package according to the decoding parameter configuration information in the video parameter encapsulation package of each of the multiple viewpoints to obtain a multi-view video stream corresponding to the video encapsulation file.
21. An electronic device comprising a processor, a communication interface, and a memory, wherein the processor, the communication interface, and the memory are interconnected, wherein: The memory stores an executable program code, and the processor is used to call the executable program code to execute the video stream processing method according to any one of claims 1 to 9, or to implement the video stream processing method according to any one of claims 10 to 18.
Citation Information
Patent Citations
3D video coding transmission method and apparatus
CN102984548A
Method for generating data stream for providing 3-dimensional multimedia service and method for receiving the data stream
CN104822071A
Code stream, encapsulation method thereof, decoding method and device
CN108616748A
AVS2 based code stream encapsulation method, code stream, and decoding method and device
CN108696471A
Method for transmitting video, apparatus for transmitting video, method for receiving video, and apparatus for receiving video
US20210329216A1