Encoding method, decoding method, encoder, decoder, and storage medium

WO2026199127A1PCT designated stage Publication Date: 2026-10-01ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084495
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2025084495_01102026_PF_FP_ABST
    Figure CN2025084495_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present application are an encoding method, a decoding method, an encoder, a decoder, and a storage medium. The decoding method comprises: determining first indication information of a first unit in a first bitstream, wherein the first indication information is used to indicate a level of the first unit, and the first unit comprises the first indication information and reconstruction information; parsing the first unit of which the level meets a first condition, to determine reconstruction information corresponding to decoded video data; and on the basis of the reconstruction information, determining a reconstructed video corresponding to the decoded video data.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding / decoding methods, encoders, decoders, and storage media Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and in particular to an encoding and decoding method, encoder, decoder, and storage medium. Background Technology

[0002] Video Coding for Machines (VCM) currently organizes the video bitstream using multiple sub-streams. One sub-stream is the kernel video sub-stream obtained by the kernel encoder in the VCM encoder, and another sub-stream is the reconstruction information sub-stream generated by the VCM encoder based on its preprocessing operations on the video. The kernel video sub-stream is decoded by the kernel decoder in the VCM decoder to obtain the decoded video, and the reconstruction information sub-stream is processed by the VCM decoder to obtain the reconstructed information. The decoded video, based on the reconstructed information, undergoes post-reconstruction processing to obtain the reconstructed video. This reconstructed video retains key semantic information and can achieve sufficient task accuracy for machine tasks. However, there is still room for improvement in the decoding efficiency of the decoder for the NAL unit. Summary of the Invention

[0003] In a first aspect, embodiments of this application provide a decoding method applied to a decoder. The method includes: determining first indication information of a first unit in a first bitstream, the first indication information indicating the level of the first unit; the first unit containing the first indication information and reconstruction information; parsing the first unit whose level satisfies a first condition to determine reconstruction information of the corresponding decoded video data; and determining the reconstructed video of the corresponding decoded video data based on the reconstruction information.

[0004] Secondly, embodiments of this application provide an encoding method applied to an encoder. The method includes: determining reconstruction information and first indication information of encoded video data; generating a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the level of the first unit; and generating a first bitstream based on the first unit.

[0005] Thirdly, embodiments of this application provide an encoder, which includes: a first determining module configured to determine reconstruction information and first indication information of encoded video data; a first generating module configured to generate a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the level of the first unit; and a first encoding module configured to generate a first bitstream based on the first unit.

[0006] Fourthly, embodiments of this application provide an encoder, which includes a first memory and a first processor, wherein: the first memory is used to store a computer program that can run on the first processor; and the first processor is used to execute the method described in the second aspect when running the computer program.

[0007] Fifthly, embodiments of this application provide a decoder, which includes: a second determining module configured to determine first indication information of a first unit in a first bitstream, the first indication information being used to indicate the level of the first unit; the first unit containing the first indication information and reconstruction information; a decoding module configured to parse the first unit whose level satisfies a first condition and determine reconstruction information of the corresponding decoded video data; and a reconstruction module configured to determine the reconstructed video of the corresponding decoded video data based on the reconstruction information.

[0008] In a sixth aspect, embodiments of this application provide a decoder, which includes a second memory and a second processor, wherein: the second memory is used to store a computer program that can run on the second processor; and the second processor is used to execute the method described in the first aspect when running the computer program.

[0009] In a seventh aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect or the method described in the second aspect.

[0010] Eighthly, embodiments of this application provide a computer program product, including a computer program or instructions that, when executed by a processor, implement the method described in the first aspect or the method described in the second aspect.

[0011] In a ninth aspect, embodiments of this application provide a computer-readable storage medium having a bitstream stored thereon, the bitstream being generated by performing the steps of the encoding method as described in the second aspect.

[0012] It is understood that in the embodiments of this application, at the encoding end, a first indication information indicating the level is introduced in the reconstruction information unit (i.e., the first unit), so that the first unit in the bit stream not only contains reconstruction information but also level information. This allows the decoding end to select a portion of the first units of the level for decoding, and only the first units that meet the first condition can be parsed, thereby avoiding the parsing of useless first units, thus improving the decoding efficiency of the decoding end and saving the computational and power consumption of the decoding end. Attached Figure Description

[0013] Figure 1 is a schematic diagram of the V3C stream structure provided in an embodiment of this application;

[0014] Figure 2 is a schematic diagram of a video encoding and decoding system according to an embodiment of this application;

[0015] Figure 3 is a schematic block diagram of the system composition of an encoder provided in an embodiment of this application;

[0016] Figure 4 is a schematic block diagram of a decoder system provided in an embodiment of this application;

[0017] Figure 5 is a schematic diagram of the implementation flow of the decoding method provided in the embodiments of this application;

[0018] Figure 6 is a schematic diagram of the implementation process for determining decoded video data provided in an embodiment of this application;

[0019] Figure 7 is a schematic diagram of the implementation flow of the encoding method provided in the embodiments of this application;

[0020] Figure 8 is a schematic diagram of the implementation process for generating the second bitstream provided in an embodiment of this application;

[0021] Figure 9 is a schematic diagram of a VCM stream structure provided in an embodiment of this application;

[0022] Figure 10 is a schematic diagram of a VCM stream structure provided in an embodiment of this application;

[0023] Figure 11 is a schematic diagram of the composition structure of the encoder provided in the embodiment of this application;

[0024] Figure 12 is a schematic diagram of the hardware structure of the encoder provided in an embodiment of this application;

[0025] Figure 13 is a schematic diagram of the composition structure of the decoder provided in the embodiment of this application;

[0026] Figure 14 is a schematic diagram of the hardware structure of the decoder provided in an embodiment of this application;

[0027] Figure 15 is a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application. Detailed Implementation

[0028] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0030] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0031] It should also be noted that the descriptions such as "first," "second," and "third" appearing in the embodiments of this application do not have a specific meaning (such as no order, nor do they indicate a special limitation on the number of devices in the embodiments of this application), but are merely for the purpose of clearly describing the embodiments of this application and do not constitute any limitation on the embodiments of this application.

[0032] This application can be applied to the fields of image encoding and decoding, video encoding and decoding, hardware video encoding and decoding, dedicated circuit video encoding and decoding, and real-time video encoding and decoding. For example, the solution of this application can be combined with audio video coding standards (AVS), such as H.264 / Audio Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), Visual Volumetric Video-based Coding (V3C), Video Coding for Machine (VCM), and Feature Coding for Machine (FCM). Alternatively, the solutions in this application can be incorporated into other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), which include Scalable Video Codec (SVC) and Multi-View Video Codec (MVC) extensions. It should be understood that the technology in this application is not limited to any particular codec standard or technology.

[0033] A high-degree-of-freedom immersive coding system can be broadly divided into the following stages based on the task flow: data acquisition, data organization and expression, data encoding and compression, data decoding and reconstruction, and data synthesis and rendering, ultimately presenting the target data to the user. To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies or terms of the embodiments of this application are described below. The following relevant technologies or terms are optional solutions and can be arbitrarily combined with the technical solutions of the embodiments of this application, all of which fall within the protection scope of the embodiments of this application.

[0034] In related technology one, H.264, H.265, and H.266 encapsulate encoded data in the form of Network Abstraction Layer (NAL) units, and use a 24-bit start code prefix to distinguish NAL units in the bitstream, or rely on system layer identifiers to distinguish NAL units in the bitstream. This enables fast scanning and random access to each NAL unit. In related technology one, a NAL unit type is defined for each type of data packet. The advantage of this is that it allows the decoder or system layer to quickly scan and identify the type of each NAL unit. However, the inventors of this application found in their research and analysis that the drawback of this technology is that each new NAL unit requires a NAL unit type number, and the number of NAL unit types it can support is limited. It cannot effectively support situations containing multiple sub-bitstreams, each containing its own unique NAL unit type. Moreover, since all NAL units need to be identified, it increases the scanning burden on the decoder or system layer.

[0035] In related technology 2, the video bitstream encoded by V3C (Visual Volumetric Video-based Coding) contains multiple parallel sub-bitstreams, such as atlas sub-bitstreams, attribute video sub-bitstreams, and geometry video sub-bitstreams.

[0036] The data in these sub-streams are also encapsulated in the form of NAL units. However, unlike related technologies, the NAL units in the sub-streams are not distinguished by the start code prefix, but by the size of the data packet of each NAL unit identified in the sub-stream. In this way, the decoder can still achieve fast scanning and random access to the data in the sub-streams by scanning from the head of the sub-stream.

[0037] Figure 1 is a schematic diagram of the V3C bitstream structure provided in an embodiment of this application. The V3C bitstream includes: the V3C parameter set (V3C_parameter_set()) of V3C_VPS may include ptl_profile_toolset_idc. If ptl_profile_toolset_idc is 128 / 129 / 130, it indicates that the current bitstream simultaneously contains both VPCC extend and MIV main bitstreams.

[0038] The Atlas sequence parameter set (Atlas_sequence_parameter_set_rbsp()) in NAL_ASPS of the V3C_AD stitched sub-bitstream (Atlas_sub_bitstream()) can include asps_vpcc_extension_present_flag and asps_miv_extension_present_flag. When ptl_profile_toolset_idc is 128 / 129 / 130, asps_vpcc_extension_present_flag is true (i.e., 1), and asps_miv_extension_present_flag is also true (i.e., 1).

[0039] The ACL NAL unit type (ACL_NAL_unit_type) in V3C_AD's Atlas_sub_bitstream() includes hybrid stitching information. For example, the atlas tile data unit (atlas_tile_data_unit()) can include atdu_type_flag. If atdu_type_flag is yes (i.e., 1), it indicates that the current tile belongs to a point cloud tile; if atdu_type_flag is no (i.e., 0), it indicates that the current tile belongs to a multi-view video tile.

[0040] Furthermore, the sub-tack information data (patch_information_data) includes sub-tack data units (patch_data_unit). If atdu_type_flag is negative and asps_miv_extension_present_flag is positive, it indicates that the current sub-tack is implemented using a multi-view video coding standard. If atdu_type_flag is positive, it indicates that the current sub-tack is implemented using a point cloud video decoding standard.

[0041] The data in the V3C_AD splicing diagram sub-bit stream (also called sub-code stream) is also encapsulated in the form of NAL units. However, unlike existing technologies, the NAL units in the sub-code stream are not distinguished by the start prefix code, but by the packet size of the data in each NAL unit identified in the code stream. In this way, the decoder can still achieve fast scanning and random access to the data in the sub-code stream by scanning from the head of the sub-code stream.

[0042] The video sub-bitstreams include: V3C_GVD video sub-bitstream(), V3C_AVD video sub-bitstream(), V3C_CAD video sub-bitstream(), and V3C_PVD video sub-bitstream().

[0043] In addition, to distinguish different sub-stream data within the video stream, the sub-stream data is split and encapsulated in V3C units. Each V3C unit records the type of sub-stream data it carries. The V3C units in the video stream are distinguished based on the data packet size of each V3C unit identified within the stream.

[0044] Related technology two addresses the shortcomings of related technology one by implementing a two-layer data unit encapsulation. However, the inventors of this application discovered during their research and analysis that related technology two suffers from the following drawback: the data units of the video stream and sub-stream are distinguished by the size of the data packets identified within the stream. This requires the decoder or system layer to start scanning from the beginning of the video stream or sub-stream when accessing a data packet, since the starting position of each data packet is calculated by summing the sizes of the preceding data packets. This is feasible in file storage scenarios, where the decoder or system layer can obtain the complete video stream and start scanning from its beginning. However, in real-time transmission scenarios, when packet loss or errors occur during video stream transmission, the decoder or system layer cannot discard erroneous data and scan to the next complete and error-free data packet to begin correct decoding because the size of each data packet cannot be obtained. This prevents the decoding from achieving random access functionality.

[0045] Video Coding for Machines (VCM) currently organizes video bitstreams using multiple sub-streams. One sub-stream is the kernel video sub-stream obtained from the kernel encoder in the VCM encoder, and another sub-stream is the reconstruction information sub-stream generated by the VCM encoder based on its preprocessing operations on the video. The kernel video sub-stream is decoded by the kernel decoder in the VCM decoder to obtain the decoded video, and the reconstruction information sub-stream is processed by the VCM decoder to obtain the reconstructed information. The decoded video is then reconstructed based on the reconstructed information to obtain the reconstructed video. This reconstructed video retains key semantic information and can achieve sufficient task accuracy for machine tasks.

[0046] Currently, to avoid disrupting the existing characteristics of the kernel codec and ensure VCM's compatibility with existing kernel codecs, the kernel video sub-stream retains its original characteristics and is encapsulated in the form of NAL units. Since the reconstruction information sub-stream introduces a new data unit type, VCM currently uses a bitstream format similar to V3C to organize the VCM video bitstream. However, V3C's limitations prevent the VCM video bitstream from effectively supporting mainstream scenarios requiring real-time encoding and decoding, such as surveillance, autonomous driving, and smart manufacturing. Therefore, this technical solution designs a new bitstream structure and encoding / decoding method to address the problems of existing VCMs.

[0047] This application provides an encoding method, namely, determining reconstruction information and first indication information of encoded video data; generating a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the level of the first unit; and generating a first bitstream based on the first unit. This application also provides a decoding method, namely, determining first indication information of a first unit in the first bitstream, wherein the first indication information is used to indicate the level of the first unit; the first unit includes the first indication information and reconstruction information; parsing a first unit whose level satisfies a first condition to determine reconstruction information of the corresponding decoded video data; and determining the reconstructed video of the corresponding decoded video data based on the reconstruction information.

[0048] It is understandable that at the encoding end, the first indication information of the level is introduced into the reconstruction information unit (i.e., the first unit), so that the first unit in the bit stream not only contains reconstruction information, but also level information. This allows the decoding end to select a portion of the first units of the level for decoding, and only the first units that meet the first condition of the level can be parsed. This avoids the parsing of useless first units, thereby improving the decoding efficiency of the decoding end and saving the computational and power consumption of the decoding end.

[0049] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0050] The encoding involved in this application embodiment is mainly video encoding and decoding. For ease of understanding, the video encoding and decoding system involved in this application embodiment will first be introduced with reference to Figure 2. Figure 2 is a schematic diagram of a video encoding and decoding system involved in this application embodiment. It should be noted that Figure 2 is only an example, and the video encoding and decoding system of this application embodiment includes, but is not limited to, the one shown in Figure 2. As shown in Figure 2, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (can be understood as compression) video data to generate a bitstream, and transmits the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain the decoded video data.

[0051] The encoding device 110 in this application embodiment can be understood as a device with video encoding function, and the decoding device 120 can be understood as a device with video decoding function. That is, the encoding device 110 and the decoding device 120 in this application embodiment include a wider range of devices, such as smartphones, desktop computers, mobile computing devices, laptops (e.g., laptop computers), tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, etc.

[0052] In some embodiments, encoding device 110 may transmit encoded video data (such as a bitstream) to decoding device 120 via channel 130. Channel 130 may include one or more media and / or means capable of transmitting encoded video data from encoding device 110 to decoding device 120.

[0053] In one example, channel 130 includes one or more communication media that enable encoding device 110 to transmit encoded video data directly to decoding device 120 in real time. In this example, encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum; optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.

[0054] In another example, channel 130 includes a storage medium that can store video data encoded by encoding device 110. The storage medium includes various local access data storage media, such as optical discs, DVDs, flash memory, etc. In this example, decoding device 120 can retrieve the encoded video data from this storage medium.

[0055] In another example, channel 130 may include a storage server that can store the video data encoded by encoding device 110. In this example, decoding device 120 can download the stored encoded video data from the storage server. Optionally, the storage server can store the encoded video data and transmit the encoded video data to decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0056] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter. In some embodiments, in addition to the video encoder 112 and the input interface 113, the encoding device 110 may also include a video source 111.

[0057] Video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.

[0058] Video encoder 112 encodes video data from video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the encoding information of the pictures or picture sequences in the form of a bitstream. The encoding information may include encoded image data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax element structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. A syntax element structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0059] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data can also be stored on a storage medium or a storage server for subsequent retrieval by the decoding device 120.

[0060] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122. In some embodiments, in addition to the input interface 121 and the video decoder 122, the decoding device 120 may also include a display device 123.

[0061] The input interface 121 includes a receiver and / or a modem. The input interface 121 can receive encoded video data through channel 130.

[0062] The video decoder 122 is used to decode the encoded video data to obtain the decoded video data, and transmit the decoded video data to the display device 123.

[0063] Display device 123 displays the decoded video data. Display device 123 may be integrated with decoding device 120 or external to decoding device 120. Display device 123 may include various display devices, such as liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or other types of display devices.

[0064] Furthermore, Figure 2 is merely an example, and the technical solutions of this application embodiment are not limited to Figure 2. For example, the technology of this application can also be applied to one-sided video encoding or one-sided video decoding.

[0065] This application provides a network architecture for a video encoding / decoding system that includes decoding and encoding methods. The decoder or encoder in this application can be an electronic device, or the electronic device may include a decoder or encoder. In other words, the electronic device in this application has video encoding / decoding capabilities and generally includes a video encoder and a video decoder.

[0066] Figure 3 is a schematic block diagram of an encoder system according to an embodiment of this application. As shown in Figure 3, the encoder 100 may include a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control and analysis unit 107, a filtering unit 108, an encoding unit 109, and a Decoded Picture Buffer (DPB) unit 110, etc. Here, the input of the encoder 100 can be a video composed of a series of images or a single static image, and the output of the encoder 100 can be a bitstream (also called a "bitstream") representing a compressed version of the input video. The images in the input video can be segmented into one or more Coding Tree Units (CTUs). For example, an image can be divided into multiple tiles, and a tile can be further divided into one or more bricks. Here, a tile or a brick may include one or more complete and / or partial CTUs.

[0067] The filtering unit 108 can implement deblocking filtering and Sample Adaptive Offset (SAO) filtering, while the encoding unit 109 can implement header information encoding and Context-based Adaptive Binary Arithmetic Coding (CABAC). For the input raw video signal, the coding tree unit... The partitioning of a video coding unit (CTU) yields a video coding block. The residual sample information obtained after intra-frame or inter-frame prediction is then transformed by the transform and quantization unit 101. This transformation includes converting the residual information from the sample domain to the transform domain and quantizing the resulting transform coefficients to further reduce the bit rate. Intra-frame estimation unit 102 and intra-frame prediction unit 103 perform intra-frame prediction on the video coding block. Specifically, intra-frame estimation unit 102 and intra-frame prediction unit 103 determine the intra-frame prediction mode to be used to encode the video coding block. Motion compensation unit 104 and motion estimation unit 105 perform inter-frame prediction coding of the received video coding block relative to one or more blocks in one or more reference frames to provide temporal prediction information. The motion estimation performed by motion estimation unit 105 is a process of generating motion vectors, which can estimate the motion of the video coding block. Then, motion compensation unit 104 uses the motion vectors determined by motion estimation unit 105 to generate motion vectors. The motion compensation is performed. After determining the intra-prediction mode, the intra-prediction unit 103 is also used to provide the selected intra-prediction data to the coding unit 109, and the motion estimation unit 105 also sends the calculated motion vector data to the coding unit 109. In addition, the inverse transform and inverse quantization unit 106 is used to reconstruct the video coding block, reconstruct the residual block in the sample domain, and remove the block artifacts by the filter control analysis unit 107 and the filtering unit 108. Then, the reconstructed residual block is added to a predictive block in the frame of the decoding image buffer unit 110 to generate the reconstructed video coding block. The coding unit 109 is used to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-prediction mode and output the bitstream of the video signal. The decoding image buffer unit 110 is used to store the reconstructed video coding block for prediction reference. As video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoding image buffer unit 110.

[0068] Furthermore, encoder 100 may be a first memory having a first processor and a computer program for recording. When the first processor reads and runs the computer program, encoder 100 reads the input video and generates a corresponding bitstream. Alternatively, encoder 100 may also be a computing device having one or more chips. These units, implemented as integrated circuits on the chips, have connection and data exchange functions similar to the corresponding units in Figure 3.

[0069] Figure 4 is a schematic block diagram of a decoder system provided in an embodiment of this application. As shown in Figure 4, the decoder 120 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205, and a decoded image buffer unit 206, etc. Here, the input of the decoder 120 is a bitstream representing a compressed version of a video or a still image, and the output of the decoder 120 can be a decoded video composed of a series of images or a decoded still image.

[0070] The decoding unit 201 can perform header information decoding and CABAC decoding, while the filtering unit 205 can perform deblocking filtering and SAO filtering. After the input video signal is encoded, a bitstream of the video signal is output. This bitstream is input into the decoder 120, first passing through the decoding unit 201 to obtain the decoded transform coefficients. These transform coefficients are then processed by the inverse transform and inverse quantization unit 202 to generate residual blocks in the sample domain. The intra-frame prediction unit 203 can generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and data from previously decoded blocks in the current frame or image. The motion compensation unit 204 determines the prediction information for the video decoding block by analyzing motion vectors and other associated syntax elements, and uses this prediction information... The measurement information is used to generate a predictive block of the video block being decoded; the decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 with the corresponding predictive block generated by the intra-frame prediction unit 203 or the motion compensation unit 204; the decoded video signal is passed through the filtering unit 205 to remove block artifacts, which can improve video quality; then the decoded video block is stored in the decoding image buffer unit 206, which stores reference images for subsequent intra-frame prediction or motion compensation, and is also used for the output of the video signal, thus obtaining the recovered original video signal.

[0071] Furthermore, the decoder 120 may be a second memory having a second processor and a computer program for recording. When the first processor reads and runs the computer program, the decoder 120 reads the input bitstream and generates the corresponding decoded video. Alternatively, the decoder 120 may also be a computing device having one or more chips. These units, implemented as integrated circuits on the chips, have similar connection and data exchange functions to the corresponding units in Figure 4.

[0072] This application provides a decoding method that is applied to a decoder.

[0073] Figure 5 is a schematic diagram of the implementation flow of the decoding method provided in the embodiment of this application. As shown in Figure 5, the method may include the following steps 501 to 503:

[0074] Step 501: Determine the first indication information of the first unit in the first bitstream. The first indication information is used to indicate the level of the first unit. The first unit contains the first indication information and reconstruction information.

[0075] Step 502: Analyze the first unit that satisfies the first condition at the parsing level, and determine the reconstruction information of the corresponding decoded video data;

[0076] Step 503: Based on the reconstruction information, determine the reconstructed video corresponding to the decoded video data.

[0077] It is understood that in the embodiments of this application, the first indication information of the level is introduced into the reconstruction information unit (i.e., the first unit), so that the first unit in the bit stream not only contains reconstruction information, but also level information. Thus, when the decoding end needs to select a portion of the first units of the level for decoding, only the first units that meet the first condition of the level can be parsed, thereby avoiding the parsing of useless first units, thereby improving the decoding efficiency of the decoding end and saving the computational and power consumption of the decoding end.

[0078] The following sections will describe further optional implementation methods for each of the above steps, as well as related terms.

[0079] Step 501: Determine the first indication information of the first unit in the first bitstream. The first indication information is used to indicate the level of the first unit. The first unit contains the first indication information and reconstruction information.

[0080] In some embodiments, the first indication information is used to indicate at least one of the following levels of the first unit:

[0081] (1) Resolution level, wherein the first unit corresponding to the resolution level is used to determine the reconstructed video at the corresponding resolution;

[0082] (2) Viewpoint hierarchy, wherein the first unit corresponding to the viewpoint hierarchy is used to determine the reconstructed video of the corresponding viewpoint;

[0083] (3) Spatial region level, wherein the first unit corresponding to the spatial region level is used to determine the reconstructed video of the corresponding spatial region.

[0084] It should be understood that different resolution levels correspond to different resolutions. In other words, the reconstruction information in the first unit of different resolution levels is used to determine the reconstructed video at their respective resolutions.

[0085] Different viewpoint levels correspond to different viewpoints. In other words, the reconstruction information in the first unit of different viewpoint levels is used to determine the reconstructed video of their respective viewpoints.

[0086] Different spatial region levels correspond to different spatial regions. In other words, the reconstruction information in the first unit of different spatial region levels is used to reconstruct the video of their respective spatial regions.

[0087] In this embodiment, the first indication information is not limited to indicating at least one of the aforementioned levels. In some embodiments, the level indicated by the first indication information in the first unit is consistent with the level indicated by the second indication information in the corresponding second unit. The second unit includes second indication information and coded video data, whereby the second indication information indicates the level of the second unit. For example, the second unit is a kernel video NAL unit.

[0088] It should be noted that the so-called second unit corresponding to the first unit refers to the second unit corresponding to the decoded video data to which the reconstructed information in the first unit can be applied. In one possible implementation, the second unit corresponding to the first unit means that the value of the third syntax element in the first unit is equal to or equivalent to the value of the fourth syntax element in the second unit. "Equivalent" means that the transformed value of the third syntax element is equal to the value of the fourth syntax element, or vice versa. For example, in VCM, the third and fourth syntax elements are prd_picture_order_cnt_lsb.

[0089] The above embodiments describe first indication information used to indicate at least one of resolution level, viewpoint level, and spatial region level, and the embodiments include a variety of combination schemes;

[0090] In one of the combination schemes, the first indication information is used to indicate the resolution level of the first unit. Further, in one possible implementation, the resolution level of the first unit can be indicated by the value of a first syntax element, where the value of the first syntax element belongs to the first indication information, and different values ​​of the syntax element correspond to different resolution levels.

[0091] In the second combination scheme, the first indication information is used to indicate the viewpoint level of the first unit. Further, in one possible implementation, the viewpoint level of the first unit can be indicated by the value of a first syntax element, where the value of the first syntax element belongs to the first indication information, and different values ​​of the syntax element correspond to different viewpoint levels.

[0092] In combination scheme three, the first indication information is used to indicate the spatial region hierarchy of the first unit. Further, in one possible implementation, the spatial region hierarchy of the first unit can be indicated by the value of a first syntax element, where the value of the first syntax element belongs to the first indication information, and different values ​​of the syntax element correspond to different spatial region hierarchies.

[0093] It should be noted that in combination schemes one, two, and three, the first syntax element corresponding to the three types of levels can be different syntax elements or the same syntax element. For an implementation where the first syntax element is the same, the level type indicated by the first syntax element can be specified at a higher level. The level type indicated at the higher level can be a resolution level, a viewpoint level, or a spatial region level.

[0094] In combination scheme four, the first indication information is used to indicate the resolution level and viewpoint level of the first unit. Further, in one possible implementation, the resolution level and viewpoint level of the first unit can be indicated by the values ​​of two syntax elements respectively; in another possible implementation, the resolution level and viewpoint level can also be indicated by the value of a single syntax element.

[0095] In combination scheme five, the first indication information is used to indicate the resolution level and spatial region level of the first unit. Further, in one possible implementation, the resolution level and spatial region level of the first unit can be indicated by the values ​​of two syntax elements respectively; in another possible implementation, the resolution level and spatial region level can also be indicated by the value of a single syntax element.

[0096] In combination scheme six, the first indication information is used to indicate the viewpoint level and spatial region level of the first unit. Further, in one possible implementation, the viewpoint level and spatial region level of the first unit can be indicated by the values ​​of two syntax elements respectively; in another possible implementation, the viewpoint level and spatial region level can also be indicated by the value of a single syntax element.

[0097] In combination scheme seven, the first indication information is used to indicate the resolution level, viewpoint level, and spatial region level of the first unit. Further, in one possible implementation, the resolution level, viewpoint level, and spatial region level of the first unit can be indicated by the values ​​of three syntax elements respectively; in another possible implementation, the resolution level, viewpoint level, and spatial region level can also be indicated by the value of a single syntax element.

[0098] In one or more of the above embodiments, further, in some embodiments, the first indication information is also used to indicate the temporal level of the first unit, and the first unit corresponding to the temporal level is used to determine the reconstructed video of the corresponding temporal level. It should be understood that this embodiment is combined with the above combination schemes one to seven, where the first indication information is used to indicate at least one level among resolution level, viewpoint level, and spatial region level, and the first indication information is also used to indicate the temporal level.

[0099] It should be understood that different temporal levels correspond to different video frames. In other words, the reconstruction information in the first unit corresponding to different temporal levels is used to determine the reconstructed video of the video frames at their respective temporal levels. For example, if the system layer or decoder receives 60 frames of images per second (i.e., 60 video frames), and the decoder's decoding capability is to decode 20 frames of images per second, then the temporal levels are divided into 3 levels. The first temporal level corresponds to frames 1 to 20 (i.e., the decoded video data of frames 1 to 20), the second temporal level corresponds to frames 21 to 40 (i.e., the decoded video data of frames 21 to 40), and the third temporal level corresponds to frames 41 to 60 (i.e., the decoded video data of frames 41 to 60).

[0100] In one possible implementation, the first indication information includes the value of a second syntax element, which indicates the temporal level. Different values ​​of the second syntax element correspond to different temporal levels.

[0101] As mentioned earlier, the first unit contains reconstruction information and first indication information. A first unit can be understood as a data packet. In some embodiments, the first indication information is in the header information (i.e., packet header) of the first unit, and the reconstruction information is in the payload information of the first unit; that is, the reconstruction information is the payload data of the first unit. Thus, since the hierarchical information (i.e., the first indication information) is in the header information of the first unit, it is convenient for the system layer or decoder to quickly filter out the first unit that needs to be parsed based on the hierarchical information.

[0102] Exemplarily, in some embodiments, the first unit is a reconstruction information NAL unit. It should be noted that in this application, the reconstruction information NAL unit can also be described as a reconstruction data NAL unit or a reconstruction data NAL packet, etc. The second unit can be described as a video data unit. Exemplarily, in some embodiments, the second unit is a kernel video NAL unit, a video data NAL unit, a kernel video NAL packet, or a video data NAL packet.

[0103] Taking the first unit as the reconstruction information NAL unit as an example, Table 1 shows the syntax structure of the header information of the first unit.

[0104] Table 1

[0105] As shown in Table 1, `vcm_nal_unit_layer_id` is an example of the first syntax element, representing the layer number of the reconstructed NAL unit. This layer number indicates the resolution layer, viewpoint layer, or spatial region layer, etc. `vcm_nal_temporal_id_plus1` is an example of the second syntax element, where the value of `vcm_nal_temporal_id_plus1` minus 1 represents the temporal layer number (also known as the temporal layer) of the reconstructed NAL unit. At least two layer numbers, including the layer number and the temporal layer number, are parsed from the header (i.e., packet header) of the reconstructed NAL unit through decoding. Depending on playback requirements, it is selected whether to skip the parsing of certain reconstructed NAL units with certain layer numbers and / or temporal layer numbers.

[0106] In one possible implementation, `vcm_nal_unit_layer_id` and `vcm_nal_temporal_id_plus1` can be used in combination. The system layer or decoder first obtains the `vcm_nal_unit_layer_id` of the kernel video NAL unit and the reconstruction information NAL unit. Based on playback requirements (e.g., the sufficiency of decoding computing resources, the user-selected viewpoint, or the video space region viewed by the user), it selects reconstruction information NAL units with certain specific layer numbers, decodes them, and outputs them. When decoding computing resources decrease, the temporal layer number can be further obtained from these undecoded reconstruction information NAL units with specific layer numbers, skipping the decoding of reconstruction information NAL units with high temporal layer numbers, thus adapting to the reduction in decoding computing resources by reducing the decoding frame rate.

[0107] The advantages of prioritizing the selection of NAL units for reconstruction information based on layer numbering followed by temporal layer numbering are understandable: 1) For videos with multiple viewpoint or spatial region levels, users typically select certain layers based on their habits and preferences, which needs to be prioritized. When decoding computing resources are insufficient, the selection of videos from different temporal layers is mainly used to adapt to the decoding computing resources; 2) For videos with multiple resolution levels, selecting different temporal or resolution layers can adapt to the decoding computing resources, but prioritizing the selection of different resolution layers while keeping the temporal layer unchanged ensures the time accuracy requirements of the user's machine task execution. This is because machine task networks can usually support the effective recognition of objects at different resolutions. Furthermore, if NAL units for reconstruction information are selected first based on temporal layer numbering, when decoding computing resources are insufficient, the decoder needs to switch between videos at different resolution levels, different viewpoint levels, or different spatial region levels. This can lead to discontinuities in the output video content, such as fluctuating video resolution, constant viewpoint switching, and continuous changes in spatial regions.

[0108] Step 502: Analyze the first unit that satisfies the first condition at the parsing level, and determine the reconstruction information of the corresponding decoded video data.

[0109] It can be understood that the first unit that satisfies the first condition at the parsing level means that the first unit in the first bitstream is selectively parsed, and the first unit in the first bitstream that satisfies the first condition at the parsing level is parsed.

[0110] As mentioned earlier, the first indication information of the first unit indicates at least one of the resolution level, viewpoint level, and spatial region level, or the first indication information of the first unit indicates at least one of the resolution level, viewpoint level, and spatial region level, as well as the temporal level.

[0111] In some embodiments, the first condition includes a first sub-condition, which includes at least one of the following: (1) the resolution level is a specific resolution level; (2) the viewpoint level is a specific viewpoint level; (3) the spatial region level is a specific spatial region level.

[0112] It is understood that, based on the above embodiments, step 502 includes multiple combination schemes; wherein, in combination scheme one, the first unit with a resolution level of a specific resolution level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme two, the first unit with a viewpoint level of a specific viewpoint level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme three, the first unit with a spatial region level of a specific spatial region level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme four, the first unit with both a resolution level and a viewpoint level of a specific resolution level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme five, the first unit with both a resolution level and a spatial region level of a specific resolution level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme six, the first unit with both a viewpoint level and a spatial region level of a specific spatial region level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme seven, the first unit with a resolution level of a specific resolution level, a viewpoint level of a specific viewpoint level, and a spatial region level of a specific spatial region level is analyzed to determine the reconstruction information of the corresponding decoded video data.

[0113] In one possible implementation, a specific level can be represented by a specific value of a syntax element. Exemplarily, in some embodiments, the first sub-condition includes the first syntax element having a value equal to a first value, which indicates a specific resolution level.

[0114] In other embodiments, the first sub-condition includes the first syntax element having a value equal to a first value, which indicates a specific viewpoint level.

[0115] In some other embodiments, the first sub-condition includes the first syntax element having a value equal to a first value, which is used to indicate a specific spatial region hierarchy.

[0116] Furthermore, in some embodiments, the first condition includes a first sub-condition and a second sub-condition; wherein the second sub-condition includes: the time domain level is a specific time domain level.

[0117] As mentioned earlier, step 502 includes multiple combination schemes. Here, in conjunction with the first condition, it also includes an embodiment where the temporal level is a specific temporal level. Further, in combination scheme one, the first unit with both a resolution level and a temporal level of a specific resolution level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme two, the first unit with both a viewpoint level and a temporal level of a specific viewpoint level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme three, the first unit with both a spatial region level and a temporal level of a specific spatial region level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme four, the first unit with both a resolution level, a viewpoint level, and a temporal level of a specific resolution level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme five, the first unit with a resolution level of a specific resolution level, a spatial region level of a specific spatial region level, and a temporal domain level of a specific temporal domain level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme six, the first unit with a viewpoint level of a specific viewpoint level, a spatial region level of a specific spatial region level, and a temporal domain level of a specific temporal domain level is analyzed to determine the reconstruction information of the corresponding decoded video data. In combination scheme seven, the first unit with a resolution level of a specific resolution level, a viewpoint level of a specific viewpoint level, a spatial region level of a specific spatial region level, and a temporal domain level of a specific temporal domain level is analyzed to determine the reconstruction information of the corresponding decoded video data.

[0118] In one possible implementation, a specific level can be represented by a specific value of a syntax element. Exemplarily, in some embodiments, the second sub-condition includes the second syntax element having a value equal to a second value, which indicates a specific temporal level.

[0119] It is understood that, in the above embodiments, a possible implementation of step 502 is described, in which the first unit whose layer satisfies the first sub-condition and the second sub-condition is parsed to determine the reconstruction information of the corresponding decoded video data. That is, the first unit in the first bitstream is parsed selectively, the first unit in the first bitstream whose layer satisfies the first sub-condition and the second sub-condition is parsed, and the first unit whose layer does not satisfy the first sub-condition and the second sub-condition is not parsed.

[0120] In other embodiments, the first unit whose layer satisfies the first sub-condition can be parsed first to determine the reconstruction information of the corresponding decoded video data. During the parsing of the first unit whose layer satisfies the first sub-condition, if it is necessary to parse the first unit whose temporal layer is a specific temporal layer, then the first unit whose layer satisfies the second sub-condition is parsed from the unparsed first units that satisfy the first sub-condition. That is, in this embodiment, the method further includes: parsing the first unit whose temporal layer is a specific temporal layer from the unparsed first units that satisfy the first condition (here, the first condition includes the first sub-condition but not the second sub-condition) to determine the reconstruction information of the corresponding decoded video data.

[0121] Therefore, the advantages of prioritizing the selection of the first unit based on the first sub-condition followed by the second sub-condition are: 1) For videos with multiple viewpoint or spatial region levels, users typically select certain levels based on their habits and preferences, which should be prioritized. When decoding computing resources are insufficient, the selection of videos at different temporal levels is mainly used to adapt to the decoding computing resources; 2) For videos with multiple resolution levels, selecting different temporal or resolution levels can adapt to the decoding computing resources, but prioritizing the selection of different resolution levels while keeping the temporal level unchanged can ensure the time accuracy requirements of users performing machine tasks. This is because machine task networks can usually support the effective recognition of objects at different resolutions. Furthermore, if the first unit is selected based on the second sub-condition first, when decoding computing resources are insufficient, the decoder needs to switch between videos at different resolution levels, different viewpoint levels, or different spatial region levels. This can lead to discontinuities in the output video content, such as fluctuating video resolution, constant viewpoint switching, and continuous changes in spatial regions.

[0122] Step 503: Based on the reconstruction information, determine the reconstructed video corresponding to the decoded video data.

[0123] In some embodiments, step 503 can be performed on the corresponding decoded video data based on the reconstruction information to determine the reconstructed video.

[0124] It can be understood that at the encoding end, the encoder first performs preprocessing operations on the original video (such as removing the background, reducing the resolution, etc.) to obtain a processed video, and then encodes and compresses it to obtain a second bitstream. This second bitstream consists of multiple second units, each containing encoded video data. Additionally, the encoder obtains reconstruction information from the preprocessing operations for the decoder to perform reconstruction processing on the decoded video data (i.e., post-reconstruction processing operations, such as restoring the background, increasing the resolution, etc.), and encapsulates this reconstruction information and first indication information into a first bitstream. This first bitstream consists of multiple first units, each containing reconstruction information. Therefore, the reconstruction information mentioned in step 503 is the information used for post-reconstruction processing of the decoded video data. For example, the reconstruction information includes the resolution of the original video frames, region of interest (ROI) information, and other information related to sample values ​​and image content adjustment / restoration. After performing post-reconstruction processing on the corresponding decoded video data based on the reconstruction information, the result is a video that closely resembles the original video at the encoding end.

[0125] In some embodiments, the first unit can be sequence-level or image-level, meaning the reconstruction information can be sequence-level or image-level. Sequence-level reconstruction information can also be described as sequence-level reconstruction data or video sequence-level reconstruction data. Image-level reconstruction information can also be described as image-level reconstruction data. Therefore, the decoded video data corresponding to the reconstruction information mentioned in step 503 refers to the decoded video data in the sequence corresponding to the reconstruction information or the decoded video data of the corresponding image.

[0126] In some embodiments, as shown in FIG6, the decoding method further includes the following steps 601 and 602:

[0127] Step 601: Determine the second indication information of the second unit in the second bitstream. The second indication information is used to indicate the level of the second unit. The second unit contains the second indication information and encoded video data.

[0128] Step 602: Analyze the second unit whose layer satisfies the first condition, and determine the decoded video data.

[0129] It is understood that, in this embodiment of the application, in order to avoid the system layer or decoder being unable to quickly filter the first unit when filtering the second unit, and in order to support providing different reconstruction information for different levels of video, not only does the second unit need to record the level information (i.e., the second indication information), but the first unit also needs to record the level information (i.e., the first indication information). In this way, the system layer or decoder can manage the second unit and the first unit synchronously. When it is necessary to select a portion of the second units for parsing / decoding, the first units that meet the same level conditions are also selected synchronously for parsing / decoding, thereby avoiding the parsing of useless first units. At the same time, the corresponding level of reconstruction information (first unit) can be provided for the second units of different levels respectively.

[0130] In some embodiments, the second indication information is used to indicate at least one of the following levels of the second unit: (1) resolution level; (2) viewpoint level; (3) spatial region level.

[0131] Furthermore, in some embodiments, the second indication information is used to indicate the time-domain level of the second unit.

[0132] In some embodiments, the second indication information corresponding to the second unit whose hierarchy satisfies the first condition is the same as or equivalent to the first indication information corresponding to the first unit whose hierarchy satisfies the first condition. It should be noted that the second unit refers to the encoded video data unit corresponding to the first unit.

[0133] It can be understood that the second indication information corresponding to the second unit that satisfies the first condition at the same level is the same as the first indication information corresponding to the first unit that satisfies the first condition at the same level, which means that the values ​​of the syntax elements used to indicate the same level are equal.

[0134] It can be understood that the second indication information corresponding to the second unit that satisfies the first condition at the level is equivalent to the first indication information corresponding to the first unit that satisfies the first condition at the level. This means that the values ​​of the syntax elements used to indicate the same level are equal after transformation. For example, the value of the syntax element used to indicate a certain level in the first indication information is equal to the value of the syntax element used to indicate the same level in the second indication information after transformation.

[0135] In some embodiments, the second indication information is in the header information of the second unit, and the encoded video data is in the payload information of the second unit. Exemplarily, in some embodiments, the second unit is a kernel video NAL unit.

[0136] This application provides an encoding method applied to an encoder.

[0137] Figure 7 is a schematic diagram of the implementation flow of the encoding method provided in the embodiment of this application; as shown in Figure 7, the method includes the following steps 701 to 703:

[0138] Step 701: Determine the reconstruction information and first indication information of the encoded video data;

[0139] Step 702: Generate a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the hierarchy of the first unit;

[0140] Step 703: Generate a first bitstream based on the first unit.

[0141] It is understood that in the embodiments of this application, the first indication information of the level is introduced into the reconstruction information unit (i.e., the first unit), so that the first unit in the bit stream not only contains reconstruction information, but also level information. Thus, when the decoding end needs to select a portion of the first units of the level for decoding, only the first units that meet the first condition of the level can be parsed, thereby avoiding the parsing of useless first units, thereby improving the decoding efficiency of the decoding end and saving the computational and power consumption of the decoding end.

[0142] The following sections will describe further optional implementation methods for each of the above steps, as well as related terms.

[0143] Step 701: Determine the reconstruction information and first indication information of the encoded video data.

[0144] It is understandable that at the encoding end, the encoder first performs preprocessing operations on the original video (such as removing the background, reducing the resolution, etc.) to obtain a processed video, and then encodes and compresses it to obtain a second bitstream. This second bitstream consists of multiple second units, each containing encoded video data. Additionally, the encoder obtains reconstruction information from the preprocessing operations for the decoder to perform reconstruction processing on the decoded video data (i.e., post-reconstruction processing operations, such as restoring the background, increasing the resolution, etc.), and encapsulates this reconstruction information into a first bitstream. This first bitstream consists of multiple first units, each containing reconstruction information. Therefore, the reconstruction information mentioned in step 701 is the information used for post-reconstruction processing of the decoded video data. For example, the reconstruction information includes the resolution of the original video frames, region of interest (ROI) information, and other information related to sample values ​​and image content adjustment / restoration. After performing post-reconstruction processing on the corresponding decoded video data based on the reconstruction information, the result is a video that closely resembles the original video at the encoding end.

[0145] In some embodiments, the first unit can be sequence-level or image-level, meaning the reconstruction information can be sequence-level or image-level. Sequence-level reconstruction information can also be described as sequence-level reconstruction data or video sequence-level reconstruction data. Image-level reconstruction information can also be described as image-level reconstruction data.

[0146] In some embodiments, determining the first indication information in step 701 may further include: obtaining second indication information from the second unit corresponding to the reconstructed information, the second indication information being used to indicate the hierarchy of the second unit; and determining the first indication information based on the second indication information.

[0147] For example, in some embodiments, the first indication information is the same as or equivalent to the second indication information.

[0148] In this context, the first and second indications are identical, which can be understood as indicating that the values ​​of grammatical elements at the same level are equal. The first and second indications being identical can also be understood as indicating that the values ​​of grammatical elements at the same level are equal after transformation. For example, the value of a grammatical element at a certain level in the first indication, after transformation, is equal to the value of a grammatical element at the same level in the second indication.

[0149] Step 702: Generate a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the hierarchy of the first unit.

[0150] It should be understood that the first unit contains the reconstruction information and the first indication information, and the reconstruction information and the first indication information can be encapsulated into the first unit. In one possible implementation, the reconstruction information and the first indication information can be encapsulated into a NAL unit. In this implementation, the NAL unit can be described as a reconstruction information NAL unit, a reconstruction data NAL unit, or a reconstruction data NAL package, etc.

[0151] In some embodiments, the first indication information is used to indicate at least one of the following levels of the first unit:

[0152] (1) Resolution level, wherein the first unit corresponding to the resolution level is used to determine the reconstructed video at the corresponding resolution;

[0153] (2) Viewpoint hierarchy, wherein the first unit corresponding to the viewpoint hierarchy is used to determine the reconstructed video of the corresponding viewpoint;

[0154] (3) Spatial region level, wherein the first unit corresponding to the spatial region level is used to determine the reconstructed video of the corresponding spatial region.

[0155] For example, in some embodiments, the first indication information includes the value of a first syntax element, the value of which is used to indicate one of the following levels: (1) the resolution level; (2) the viewpoint level; (3) the spatial region level.

[0156] It should be understood that different resolution levels correspond to different resolutions. At the decoding end, the reconstruction information in the first unit of each resolution level is used to determine the reconstructed video at its respective resolution.

[0157] Different viewpoint levels correspond to different viewpoints. At the decoding end, the reconstruction information in the first unit of each viewpoint level is used to determine the reconstructed video for their respective viewpoints.

[0158] Different spatial region levels correspond to different spatial regions. At the decoding end, the reconstruction information in the first unit of each spatial region level is used to reconstruct the video for its respective spatial region.

[0159] In this embodiment, the first indication information is not limited to indicating at least one of the aforementioned levels. In some embodiments, the level indicated by the first indication information in the first unit is consistent with the level indicated by the second indication information in the corresponding second unit. The second unit includes second indication information and coded video data, whereby the second indication information indicates the level of the second unit. For example, the second unit is a kernel video NAL unit.

[0160] It should be noted that the so-called second unit corresponding to the first unit refers to the second unit corresponding to the decoded video data to which the reconstructed information in the first unit can be applied. In one possible implementation, the second unit corresponding to the first unit means that the value of the third syntax element in the first unit is equal to or equivalent to the value of the fourth syntax element in the second unit. "Equivalent" means that the transformed value of the third syntax element is equal to the value of the fourth syntax element, or vice versa. For example, in VCM, the third and fourth syntax elements are prd_picture_order_cnt_lsb.

[0161] The above embodiments describe first indication information used to indicate at least one of resolution level, viewpoint level, and spatial region level, and the embodiments include a variety of combination schemes;

[0162] In one of the combination schemes, the first indication information is used to indicate the resolution level of the first unit. Further, in one possible implementation, the resolution level of the first unit can be indicated by the value of a first syntax element, where the value of the first syntax element belongs to the first indication information, and different values ​​of the syntax element correspond to different resolution levels.

[0163] In the second combination scheme, the first indication information is used to indicate the viewpoint level of the first unit. Further, in one possible implementation, the viewpoint level of the first unit can be indicated by the value of a first syntax element, where the value of the first syntax element belongs to the first indication information, and different values ​​of the syntax element correspond to different viewpoint levels.

[0164] In combination scheme three, the first indication information is used to indicate the spatial region hierarchy of the first unit. Further, in one possible implementation, the spatial region hierarchy of the first unit can be indicated by the value of a first syntax element, where the value of the first syntax element belongs to the first indication information, and different values ​​of the syntax element correspond to different spatial region hierarchies.

[0165] It should be noted that in combination schemes one, two, and three, the first syntax element corresponding to the three types of levels can be different syntax elements or the same syntax element. For an implementation where the first syntax element is the same, the level type indicated by the first syntax element can be specified at a higher level. The level type indicated at the higher level can be a resolution level, a viewpoint level, or a spatial region level.

[0166] In combination scheme four, the first indication information is used to indicate the resolution level and viewpoint level of the first unit. Further, in one possible implementation, the resolution level and viewpoint level of the first unit can be indicated by the values ​​of two syntax elements respectively; in another possible implementation, the resolution level and viewpoint level can also be indicated by the value of a single syntax element.

[0167] In combination scheme five, the first indication information is used to indicate the resolution level and spatial region level of the first unit. Further, in one possible implementation, the resolution level and spatial region level of the first unit can be indicated by the values ​​of two syntax elements respectively; in another possible implementation, the resolution level and spatial region level can also be indicated by the value of a single syntax element.

[0168] In combination scheme six, the first indication information is used to indicate the viewpoint level and spatial region level of the first unit. Further, in one possible implementation, the viewpoint level and spatial region level of the first unit can be indicated by the values ​​of two syntax elements respectively; in another possible implementation, the viewpoint level and spatial region level can also be indicated by the value of a single syntax element.

[0169] In combination scheme seven, the first indication information is used to indicate the resolution level, viewpoint level, and spatial region level of the first unit. Further, in one possible implementation, the resolution level, viewpoint level, and spatial region level of the first unit can be indicated by the values ​​of three syntax elements respectively; in another possible implementation, the resolution level, viewpoint level, and spatial region level can also be indicated by the value of a single syntax element.

[0170] In some embodiments, in addition to one or more of the above embodiments, the first indication information is further used to indicate the temporal level of the first unit, and the first unit corresponding to the temporal level is used to determine the reconstructed video of the corresponding temporal level.

[0171] Furthermore, in some embodiments, the first indication information includes the value of a second syntax element, the value of which is used to indicate the time domain level.

[0172] It should be understood that this embodiment is incorporated into the above-described combination schemes one to seven, wherein the first indication information is used to indicate at least one of the resolution level, viewpoint level, and spatial region level, and the first indication information is also used to indicate the temporal level.

[0173] It should be understood that different temporal levels correspond to different video frames. In other words, the reconstruction information in the first unit corresponding to different temporal levels is used to determine the reconstructed video of the video frames at their respective temporal levels. For example, if the system layer or decoder receives 60 frames of images per second (i.e., 60 video frames), and the decoder's decoding capability is to decode 20 frames of images per second, then the temporal levels are divided into 3 levels. The first temporal level corresponds to frames 1 to 20 (i.e., the decoded video data of frames 1 to 20), the second temporal level corresponds to frames 21 to 40 (i.e., the decoded video data of frames 21 to 40), and the third temporal level corresponds to frames 41 to 60 (i.e., the decoded video data of frames 41 to 60).

[0174] As mentioned earlier, the first unit contains reconstruction information and first indication information. A first unit can be understood as a data packet. Exemplarily, in some embodiments, the first indication information is in the header information (i.e., packet header) of the first unit, and the reconstruction information is in the payload information of the first unit. Thus, since the hierarchical information (i.e., the first indication information) is in the header information of the first unit, it is convenient for the system layer or decoder to quickly filter out the first unit that needs to be parsed based on the hierarchical information.

[0175] For example, the first unit is a reconstruction information NAL unit. It should be noted that in this application, the reconstruction information NAL unit can also be described as a reconstruction data NAL unit or a reconstruction data NAL package, etc. Reconstruction information can be described as reconstruction data.

[0176] Taking the first unit as the reconstruction information NAL unit as an example, Table 2 shows the syntax structure of the header information of the first unit.

[0177] Table 2

[0178] As shown in Table 2, vcm_nal_unit_layer_id is an example of the first syntax element, which represents the layer number of the reconstruction information NAL unit. This layer number is used to indicate the resolution layer, viewpoint layer, or spatial region layer, etc. vcm_nal_temporal_id_plus1 is an example of the second syntax element, where the value of vcm_nal_temporal_id_plus1 minus 1 represents the temporal layer number (also known as the time domain layer) of the reconstruction information NAL unit.

[0179] Step 703: Generate a first bitstream based on the first unit.

[0180] In this embodiment of the application, the first bitstream can also be described as a reconstruction information sub-bitstream. In one example, the first unit can be a reconstruction information NAL unit, and the first bitstream can be described as a reconstruction information sub-bitstream; the second unit can be a kernel video NAL unit, and the second bitstream can be described as a kernel video sub-bitstream.

[0181] In some embodiments, as shown in FIG8, the encoding method further includes the following steps 801 to 803:

[0182] Step 801: Determine the encoded video data and second indication information of the original video.

[0183] In some embodiments, the second indication information is in the header information of the second unit, and the encoded video data is in the payload information of the second unit. Exemplarily, in some embodiments, the second unit is a kernel video NAL unit.

[0184] In one possible implementation, the encoder performs preprocessing operations (such as removing the background, reducing the resolution, etc.) on the input video (i.e., multiple consecutive frames of raw images) to obtain the processed video, and then uses a kernel encoder to encode and compress it to obtain a kernel video bitstream. The kernel video bitstream consists of multiple kernel video NAL units; a kernel video NAL unit contains one or more frames of encoded video data, or a kernel video NAL unit contains one or more slices of encoded video data.

[0185] Step 802: Generate a second unit based on the encoded video data and the second indication information; the second indication information is used to indicate the level of the second unit, or it can be described as the second indication information being used to indicate the level of the encoded video data.

[0186] It should be understood that the second unit contains the encoded video data and the second indication information, and the encoded video data and the second indication information can be encapsulated into the second unit. In one possible implementation, the encoded video data and the second indication information can be encapsulated into a NAL unit. In this implementation, the NAL unit can be described as a kernel video NAL unit, an encoded video data NAL unit, or an encoded video data NAL packet, etc.

[0187] In some embodiments, the second indication information is used to indicate at least one of the following levels of the second unit: (1) resolution level; (2) viewpoint level; (3) spatial region level.

[0188] Furthermore, in some embodiments, the second indication information is used to indicate the time-domain level of the second unit.

[0189] Step 803: Generate a second bitstream based on the second unit.

[0190] In this application embodiment, the second bitstream can also be described as a video data sub-bitstream. In one example, the first unit can be a kernel video NAL unit, and the second bitstream can be described as a kernel video sub-bitstream.

[0191] It is understood that, in this embodiment of the application, in order to avoid the system layer or decoder being unable to quickly filter the first unit when filtering the second unit, and in order to support providing different reconstruction information for different levels of video, not only does the second unit need to record the level information (i.e., the second indication information), but the first unit also needs to record the level information (i.e., the first indication information). In this way, the system layer or decoder can manage the second unit and the first unit synchronously. When it is necessary to select a portion of the second units for parsing / decoding, the first units that meet the same level conditions are also selected synchronously for parsing / decoding, thereby avoiding the parsing of useless first units. At the same time, the corresponding level of reconstruction information (first unit) can be provided for the second units of different levels respectively.

[0192] In one possible implementation, the decoding method includes: obtaining VCM units from the bitstream; obtaining the type of the VCM unit from the VCM unit; if the VCM unit is a reconstruction information sub-bitstream unit, obtaining reconstruction information NAL units from the VCM unit (when the reconstruction information sub-bitstream unit contains multiple NAL units, the data packet size of the reconstruction information NAL units is obtained from the VCM unit to obtain each reconstruction information NAL unit separately), and then obtaining reconstruction information (such as original resolution, region of interest (ROI) information, and other information related to pixel values ​​and image content adjustment / restoration, to achieve a similarity to the original video at the encoding end) from the reconstruction information NAL units; extracting the hierarchical information of the NAL unit from the reconstruction information NAL unit, which identifies which temporal layer or scalable the NAL unit belongs to. Layers, etc.; when only a portion of the layer data needs to be decoded, the decoder selects the corresponding layer's NAL units from the bitstream and decodes and plays them. If the VCM unit is a kernel video sub-bitstream unit, the kernel video NAL unit is obtained from the VCM unit. When the kernel video sub-bitstream unit contains multiple NAL units, the data packet size of the kernel video NAL unit is obtained from the VCM unit to obtain each kernel video NAL unit separately. Then, the decoded kernel video (such as residual values, transform coefficients, motion vectors, etc.) is obtained from the kernel video NAL unit. Based on the reconstruction information, the decoded kernel video is reconstructed to obtain the reconstructed video, which can be used to complete machine tasks to achieve high machine task accuracy. It should be noted that the layer information is placed in the header information of the NAL unit, and the header information of each NAL unit includes the layer information. The header information of the NAL unit is to help the decoder quickly select the NAL units to be decoded when decoding begins. It should also be noted that a frame of image may be encapsulated into one or more kernel video NAL units; the reconstruction data / reconstruction information of a frame of image can be encapsulated into a reconstruction information NAL unit.

[0193] In one possible implementation, the encoding method includes: the encoder performs preprocessing operations (such as removing the background, reducing the resolution, etc.) on the input video to obtain the processed video, and uses a kernel encoder to encode and compress it to obtain a kernel video bitstream, which is composed of multiple kernel video NAL units.

[0194] The encoder encapsulates the kernel video bitstream into at least one VCM unit, identifies the type of the VCM unit as the kernel video sub-bitstream type, and records the unit data packet size of each kernel video NAL unit in the VCM unit to distinguish the boundaries of each kernel video NAL unit.

[0195] The encoder obtains reconstruction information from the preprocessing operation for the decoder's post-reconstruction operation. This reconstruction information can guide the reconstruction of the kernel decoded video. The encoder encapsulates the reconstruction information into a reconstruction information bitstream, which contains at least one reconstruction information NAL unit. The reconstruction information bitstream is then encapsulated into at least one VCM unit. The type of the VCM unit is identified as the reconstruction information sub-bitstream type. In the VCM unit, the unit data packet size of each reconstruction information NAL unit is recorded to distinguish the boundaries of each reconstruction information NAL unit. The encoder obtains the hierarchical information from the kernel video NAL unit corresponding to the reconstruction information and puts this information into the header information of the reconstruction information NAL unit so that the hierarchical information of the reconstruction information NAL unit and the kernel video NAL unit of the same image is consistent or equivalent.

[0196] The encoder encapsulates information such as the type of kernel decoder to be used by the decoder into VCM units, which are of type VPS. The encoder organizes the VPS type VCM units, the reconstructed information sub-stream type VCM units, the kernel video sub-stream type VCM units, and other possible types of VCM units to obtain the VCM video stream.

[0197] It should be noted that the above operations do not necessarily have to be performed in order. The encoding and encapsulation of VCM units of the reconstruction information sub-stream type and VCM units of the kernel video sub-stream type can be performed alternately to achieve low-latency encoding.

[0198] The following examples illustrate possible implementation schemes of the encoding and decoding methods described in one or more of the above embodiments.

[0199] Video Coding for Machines (VCM) currently organizes video bitstreams using multiple sub-streams. One sub-stream is the kernel video sub-stream obtained by the kernel encoder in the VCM encoder, and another sub-stream is the reconstruction information sub-stream generated by the VCM encoder based on its preprocessing operations on the video. The kernel video sub-stream is decoded by the kernel decoder in the VCM decoder to obtain the decoded video, and the reconstruction information sub-stream is processed by the VCM decoder to obtain the reconstructed information. The decoded video, based on the reconstructed information, undergoes reconstruction processing (i.e., post-reconstruction processing corresponding to the preprocessing operations at the encoder, which is the opposite and inverse of the preprocessing operations at the encoder) to obtain the reconstructed video. This reconstructed video retains key semantic information and can achieve sufficient task accuracy for machine tasks.

[0200] Currently, to avoid disrupting the existing characteristics of the kernel codec and ensure VCM's compatibility with existing kernel codecs, the kernel video sub-stream retains its original characteristics and is encapsulated in the form of NAL units. Since the reconstruction information sub-stream introduces a new data unit type, VCM currently uses a V3C-like bitstream structure to organize the VCM video stream. However, V3C's limitations prevent the VCM video stream from effectively supporting mainstream scenarios requiring real-time encoding and decoding, such as surveillance, autonomous driving, and smart manufacturing. Therefore, this technical solution designs a new bitstream structure and encoding / decoding method to address the problems of existing VCMs. The following is a general description of this technical solution.

[0201] The VCM video stream format remains similar to V3C, consisting of multiple sub-streams, each encapsulated using NAL units. Different sub-streams are further divided and encapsulated into different types of VCM units. Unlike V3C, to support rapid recovery decoding capabilities after random access or transmission errors in real-time encoding, VCM units use start codes to identify the unit's segmentation boundaries within the VCM video stream. Each VCM unit has a start code prefix, which can be the same 24-bit data used in H.266, or a novel start code prefix. This allows the decoder to identify the start position of a VCM unit in real-time decoding using the start code and the end position using the next start code. This enables the correct decoding in the event of random access or packet loss / errors. For sub-streams within a VCM unit, if a sub-stream contains multiple NAL units, this scheme can also identify the boundaries of each NAL unit by indicating the packet size. The reason for not using a start code is twofold. First, the start code prefix of the NAL unit overlaps with the start code prefix of the VCM unit, making it impossible for the decoder to distinguish whether the unit after the start code prefix is ​​a NAL unit or a VCM unit. Second, the VCM unit can already be read from beginning to end by the decoder through the identifier of the start code. Using the data packet size to identify the boundary of the NAL unit can take up less data than the start code and does not affect the fast scanning of the NAL unit.

[0202] Figure 9 is a schematic diagram of a VCM stream structure provided in an embodiment of this application; as shown in Figure 9, the VCM video stream includes:

[0203] The Video Parameter Set (VCM) Unit includes the VCM header information (VCM_VPS) and the VCM payload (vcm_parameter_set()).

[0204] Restoration Data Units (VCMs) include VCM header information (VCM_RAP_RSD) and VCM payload (restoration_data_unit()) that support random access, and VCM header information (VCM_RSD) and VCM payload (restoration_data_unit()) that do not support random access.

[0205] The kernel video VCM (Coded Video Data Units) include kernel video VCM header information (VCM_RAP_CVD) and VCM payload (coded_video_data()) that support random access, and kernel video VCM header information (VCM_CVD) and VCM payload (coded_video_data()) that do not support random access.

[0206] The payload of the reconstruction information VCM unit includes video sequence-level reconstruction information (Sequence Restoration Data) and image-level reconstruction information (Picture Restoration Data). The video sequence-level reconstruction information (Sequence Restoration Data) specifically includes the video sequence-level reconstruction information NAL header (VCM_NAL_SRD) and payload (sequence_restoration_data_rbsp()), while the image-level reconstruction information (Picture Restoration Data) specifically includes the image-level reconstruction information NAL header (VCM_NAL_PRD) and payload (picture_restoration_data_rbsp()).

[0207] The workload of kernel video VCM units that support random access includes NAL units that support random access (VVC IRAP NAL Unit) and NAL units that do not support random access (VVC non-IRAP NAL Unit). The workload of kernel video VCM units that do not support random access includes NAL units that do not support random access (VVC non-IRAP NAL Unit).

[0208] Figure 10 is a schematic diagram of a VCM bitstream structure provided in an embodiment of this application. As shown in Figure 10, the corresponding decoding method is described as follows: Obtain VCM units from the bitstream; obtain the type of VCM unit from the VCM unit; if the VCM unit is a reconstruction information sub-bitstream unit, obtain the reconstruction information NAL unit from the VCM unit (when the reconstruction information sub-bitstream unit contains multiple NAL units, obtain the data packet size of the reconstruction information NAL unit from the VCM unit to obtain each NAL unit separately), and then obtain reconstruction information (such as original resolution, region of interest (ROI) information, and other information related to pixel values ​​and image content adjustment / restoration, to achieve a similarity to the original video at the encoding end) from the reconstruction information NAL unit; extract the hierarchical information of the NAL unit from the reconstruction information NAL unit, which identifies which temporal layer or scalable the NAL unit belongs to. The decoder is configured as follows: Layer (either an extended layer or a capping layer); When only a portion of the layer's data needs to be decoded, the decoder selects the corresponding layer's reconstruction information NAL units from the bitstream and decodes and plays them; If the VCM unit is a kernel video sub-bitstream unit, the kernel video NAL units are obtained from the VCM unit. When the kernel video sub-bitstream unit contains multiple NAL units, the data packet size of the kernel video NAL units is obtained from the VCM unit to obtain each kernel video NAL unit separately, and then the decoded kernel video (such as residual values, transform coefficients, motion vectors, etc.) is obtained from the kernel video NAL units; Based on the reconstruction information, the decoded kernel video is reconstructed to obtain the reconstructed video, which can be used to complete machine tasks to achieve high machine task accuracy.

[0209] In some embodiments, the hierarchy information is placed in the header information of the NAL unit, and the header information of each NAL unit includes the hierarchy information; the header information of the NAL unit is to help the decoder quickly select the NAL units that need to be decoded when decoding begins.

[0210] In some embodiments, a frame of image may be encapsulated into one or more kernel video NAL units; the reconstruction data / reconstruction information of a frame of image may be encapsulated into a NAL unit.

[0211] The corresponding encoding method is described as follows:

[0212] The encoder performs preprocessing operations on the input video (such as removing the background and reducing the resolution) to obtain the processed video, and then uses the kernel encoder to encode and compress it to obtain the kernel video bitstream, which is composed of multiple kernel video NAL units.

[0213] The encoder encapsulates the kernel video bitstream into at least one VCM unit, identifies the type of the VCM unit as the kernel video sub-bitstream type, and records the unit data packet size of each kernel video NAL unit in the VCM unit to distinguish the boundaries of each kernel video NAL unit.

[0214] The encoder obtains reconstruction information from the preprocessing operation for the decoder's post-reconstruction operation. This reconstruction information can guide the reconstruction of the kernel decoded video. The encoder encapsulates the reconstruction information into a reconstruction information bitstream, which contains at least one reconstruction information NAL unit. The reconstruction information bitstream is then encapsulated into at least one VCM unit. The type of the VCM unit is identified as the reconstruction information sub-bitstream type. In the VCM unit, the unit data packet size of each reconstruction information NAL unit is recorded to distinguish the boundaries of each reconstruction information NAL unit. The encoder obtains the hierarchical information from the kernel video NAL unit corresponding to the reconstruction information and puts this information into the header information of the reconstruction information NAL unit so that the hierarchical information of the reconstruction information NAL unit and the kernel video NAL unit of the same image is consistent or equivalent.

[0215] The encoder encapsulates information such as the type of kernel decoder to be used by the decoder into VCM units, which are of type VPS. The encoder organizes the VPS type VCM units, the reconstructed information sub-stream type VCM units, the kernel video sub-stream type VCM units, and other possible types of VCM units to obtain the VCM video stream.

[0216] The above operations do not necessarily have to be performed in sequence. The encoding and encapsulation of VCM units of the reconstruction information sub-stream type and VCM units of the kernel video sub-stream type can be performed alternately to achieve low-latency encoding.

[0217] The syntax element structure of a VCM unit is as follows:

[0218] The syntax structure for the header and payload is as follows:

[0219] The syntax element structure of the VCM unit load is as follows:

[0220] The payload of a VCM unit is used to carry VCM parameter sets (VPS), reconstructed data (RSD), or encoded video data (CVD). The payload coded_video_data() in the encoded video data packet corresponds to data units (e.g., NAL units as defined in ISO / IEC 23008-2 or ISO / IEC 23090-3) that can be decoded by the appropriate video decoder indicated by the configuration file defined in the VCM parameter set.

[0221] The method for identifying the data packet size in the reconstruction information NAL unit within the reconstruction information sub-stream is shown in the table below:

[0222] The rsd_nal_unit_size_minus1 record the data packet size of the reconstruction information NAL unit vcm_nal_unit().

[0223] The syntax structure of the reconstruction information NAL unit in the reconstruction information sub-stream is as follows:

[0224] The syntax structure of the header information is as follows:

[0225] The data of rbsp_byte is determined by vcm_nal_unit_type, which is not the key to this technical solution and will not be discussed further here.

[0226] The kernel video NAL unit `nal_unit()` of the kernel video sub-stream maintains the NAL unit syntax structure conforming to the standard specifications of its kernel codec. For example, when the kernel codec uses the H.266 codec, the kernel video NAL unit `nal_unit()` should be an H.266 compliant NAL unit. The method for identifying the packet size of the kernel video NAL unit in the kernel video sub-stream is as follows:

[0227] Here, `cvd_nal_unit_size_precision_bytes_minus1` represents the bit width used by the `cvd_nal_unit_size` syntax element, `cvd_reserved_zero_5bits` is for byte-aligned data padding, and `cvd_nal_unit_size` is the packet size for each `nal_unit()`. The reason the packet size of the kernel video NAL unit is not represented using a fixed bit width similar to that of the reconstruction information NAL unit is that the data volume of the reconstruction information NAL unit is usually small, and a fixed bit width is sufficient to identify the packet size. However, the data volume of the kernel video NAL unit can vary greatly depending on whether it is parameter set or image encoding data, and a fixed bit width is insufficient to identify the packet size.

[0228] To provide a more detailed description of the overall content of this technical solution, further details are provided below.

[0229] The syntax element structure of the VCM parameter set is as follows:

[0230] The grade level information is as follows:

[0231] The payload data of the VCM NAL unit can be sequence-level reconstruction information, image-level reconstruction information, supplementary enhancement information, termination information, etc., depending on its type information.

[0232] The syntax element structure of sequence-level reconstruction information is as follows:

[0233] The syntax element structure of image-level reconstruction information is as follows:

[0234] Among them, prd_picture_order_cnt_lsb is used to map kernel video NAL units to reconstruction information NAL units. That is, the value of this syntax element can determine which kernel video NAL unit is used for the reconstruction post-processing operation of the decoded video data of which kernel video NAL units.

[0235] The syntax element structure for auxiliary enhancement information is as follows:

[0236] The syntax element structure of the bitstream terminator for the reconstructed information NAL is as follows:

[0237] The syntax element structure for byte alignment is as follows:

[0238] The above syntax elements are illustrated below:

[0239] The `vuh_unit_type` indicates the VCM unit type specified in the table below. Values ​​marked as reserved are reserved for future use by ISO / IEC and should not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document should ignore such reserved unit types.

[0240] ●vuh_vps_id specifies the value of vps_vcm_parameter_set_id for the effective VCM parameter set. The value of vuh_vps_id should be in the range of 0 to 15.

[0241] ● When vuh_reserved_zero_23bits appears, its value should be equal to 0 in a bitstream conforming to this version of this document. Other values ​​for vuh_reserved_zero_23bits are reserved by ISO / IEC for future use. The decoder should ignore the value of vuh_reserved_zero_23bits.

[0242] ● `vuh_reserved_zero_27bits`: When present, its value should be 0 in bitstreams conforming to this version of this document. Other values ​​for `vuh_reserved_zero_27bits` are reserved by ISO / IEC for future use. The decoder should ignore the value of `vuh_reserved_zero_27bits`.

[0243] ●coded_video_data(numBytes) contains a portion of video unit streams of size numBytes, presented as an ordered stream of bytes or bits, where the positions of unit boundaries can be determined based on the organization pattern of the video unit stream. The format of this video unit stream is identified by ptl_profile_codec_group_idc.

[0244] The ptl_profile_codec_group_idc identifier is shown in the table below:

[0245] ●vps_log2_max_restoration_data_frame_order_cnt_lsb_minus4 indicates the recording rules for the time information of image-level reconstruction data, which are used to deduce the effective time of the reconstruction data according to the rules based on the syntax elements in the image-level reconstruction data.

[0246] ● `srd_spatial_resampling_enabled_flag` indicates whether the decoder can use spatial sampling tools to reconstruct the decoded video, and also indicates whether spatial sampling parameters can be recorded in the reconstructed data. ● `srd_retargeting_enabled_flag` indicates whether the decoder can use region retargeting tools to reconstruct the decoded video, and also indicates whether region retargeting parameters can be recorded in the reconstructed data.

[0247] ●srd_temporal_restoration_enabled_flag indicates whether the decoder can use time-sampling tools to reconstruct the decoded video, and also indicates whether time-sampling parameters can be recorded in the reconstructed data.

[0248] ●srd_bit_depth_shift_enabled_flag indicates whether the decoder can use data bit width offset tools to reconstruct the decoded video, and also indicates whether the relevant parameters of data bit width offset can be recorded in the reconstructed data.

[0249] ●rbsp_byte[i] is the i-th byte of the RBSP. The RBSP is specified as an ordered sequence of bytes as follows: The RBSP contains a string of data bits (SODB) as follows:

[0250] ■If SODB is empty (i.e., its length is zero bits), RBSP is also empty.

[0251] ■ Otherwise, the RBSP contains the following SODB:

[0252] 1) The first byte of the RBSP contains the first (most important, leftmost) eight bits of the SODB; the next byte of the RBSP contains the next eight bits of the SODB, and so on, until there are fewer than eight bits remaining in the SODB.

[0253] 2) The rbsp_trailing_bits() syntax structure exists after SODB as follows:

[0254] a) The first (most important, leftmost) bit of the last RBSP byte contains the remaining bits of SODB (if any).

[0255] b) The next bit consists of a single bit equal to 1 (i.e., rbsp_stop_one_bit).

[0256] c) When rbsp_stop_one_bit is not the last bit of the byte alignment byte, there are one or more bits equal to 0 (i.e., instances of rbsp_alignment_zero_bit) that cause byte alignment.

[0257] Syntax structures with these RBSP attributes are indicated in the syntax table with the suffix "_rbsp". These structures are carried within the VCM NAL package as the contents of rbsp_byte[i] data bytes. The association between RBSP syntax structures and the VCM NAL package is as specified in [the relevant documentation].

[0258] ●vcm_nal_forbidden_zero_bit should be equal to 0.

[0259] ●vcm_nal_unit_type specifies the type of RBSP data structure contained in the VCM NAL unit, as specified in the table below.

[0260] The VCM NAL package type identifiers and classifications are shown in the table below:

[0261] VCM NAL units with the identifier VCM_NAL_UNSPEC and nal_unit_type have unspecified semantics and should not affect the decoding procedures specified in this document. VCM NAL unit types with the identifier VCM_NAL_UNSPEC may be used depending on the application. No decoding procedures are specified for these vcm_nal_unit_type values ​​in this document. Because different applications may use these VCM NAL unit types for different purposes, special care is required in encoder design for generating VCM NAL units with these vcm_nal_unit_type values ​​and in decoder design for interpreting the contents of VCM NAL units with these vcm_nal_unit_type values. No management is defined for these values ​​in this document. These vcm_nal_unit_type values ​​may only apply to the context of use cases where “conflicts” (i.e., different definitions of the meaning of the contents of VCM NAL units with the same vcm_nal_unit_type value) are not important, impossible, or managed—for example, in applications or transport specifications that control the definition or management of the bitstream distribution. For purposes other than determining the amount of data in the bitstream decoding unit, the decoder should ignore (i.e., remove and discard from the bitstream) the contents of all VCM NAL units that use the reserved value of nal_unit_type.

[0262] ●vcm_nal_temporal_id indicates the temporal level of the packet.

[0263] ●vcm_nal_reserved_6bits should be equal to 0 in bitstreams conforming to this version of this document. Other values ​​for vcm_nal_reserved_6bits are reserved by ISO / IEC for future use. The decoder should ignore the value of vcm_nal_reserved_6bits.

[0264] ●prd_vps_id indicates the number of the video parameter set or sequence-level reconstruction data referenced by the image-level reconstruction data.

[0265] ●prd_frame_order_cnt_lsb indicates the time number of the image-level reconstruction data. The effective time of the image-level reconstruction data can be calculated based on this time number and the derivation rules of time information.

[0266] ●prd_spatial_resampling_enabled_flag indicates whether the decoder can use spatial sampling tools to reconstruct video images. The display time of the video image should correspond to the effective time of the reconstructed data. It also indicates whether spatial sampling parameters can be recorded in the reconstructed data.

[0267] ●prd_retargeting_enabled_flag indicates whether the decoder can use region retargeting tools to reconstruct video images. The display time of the video image and the effective time of the reconstructed data should correspond. It also indicates whether region retargeting-related parameters can be recorded in the reconstructed data.

[0268] ●prd_temporal_restoration_enabled_flag indicates whether the decoder can use time-sampling tools to reconstruct video images. The display time of the video image should correspond to the effective time of the reconstructed data. It also indicates whether time-sampling parameters can be recorded in the reconstructed data.

[0269] ●prd_bit_depth_shift_enabled_flag indicates whether the decoder can use a data bit width offset tool to reconstruct the video image. The display time of the video image should correspond to the effective time of the reconstructed data. It also indicates whether the relevant parameters of the data bit width offset can be recorded in the reconstructed data.

[0270] ●The value of rbsp_stop_one_bit should be 1.

[0271] ●The value of rbsp_alignment_zero_bit should be 0.

[0272] In one embodiment, the bitstream also includes the decoding capabilities required by the bitstream, the decoding methods required by the video unit streams, and information on the reconstruction tools that can be used for the reconstruction operation, as shown in the table below. This information can be recorded in vcm_parameter_set().

[0273] in:

[0274] ●ptl_tier_flag and ptl_level_idc specify the level of decoding capabilities required by the bitstream;

[0275] ●ptl_profile_codec_group_idc specifies the type of video unit stream, that is, the video decoding method and level that can handle this type of bitstream need to be used to decode the video unit stream to obtain the decoded image, such as the several types specified in the table below.

[0276] The ptl_profile_restoration_idc specifies the combination of tools that need to be used for the bitstream to be decoded. For example, the bitstream may require a combination of one or more tools such as spatial sampling, region relocation, temporal sampling, and data bit width offset for reconstruction.

[0277] The above describes the basic methods for decoding using reconstructed data and video data. The innovations of this application are as follows.

[0278] In one implementation, the kernel video data NAL packet header also records the layer number of the NAL packet, such as nuh_layer_id and nuh_temporal_id_plus1 in a VVC NAL packet or nuh_layer_id and nuh_temporal_id_plus1 in a HEVC NAL packet. Using layer numbers provides the system layer or decoding method with the advantage of quickly filtering NAL packets of different layers. When only a portion of the video data, rather than all layers, needs to be played, the system layer or decoding method can quickly filter NAL packets based on the layer number, achieving the effect of saving computational power. To avoid the system layer or decoding method being unable to quickly filter reconstructed data NAL packets when filtering video data NAL packets, and to support providing different reconstructed data for different video layers, the reconstructed data NAL packet also needs to record its layer number, which should be consistent with or equivalent to the layer number of its corresponding video data NAL packet.

[0279] Taking the aforementioned method of using the reconstructed data NAL package as an example, the syntax structure of the reconstructed data NAL package header is as follows:

[0280] Here, `vcm_nal_unit_layer_id` represents the layer number of the reconstructed NAL unit (multi-resolution layer, multi-viewpoint layer, or multi-spatial region layer, etc.), and `vcm_nal_temporal_id_plus1` minus 1 represents the temporal layer number of the reconstructed NAL packet. The decoding operation parses the layer number and temporal layer number from the packet header, and selects whether to skip the parsing of certain reconstructed NAL packets with specific layer or temporal layer numbers based on playback requirements.

[0281] Layer numbering and temporal layer numbering can be used in combination. The system layer or decoder first obtains the layer numbers of the kernel video NAL units and the reconstructed data NAL units. Based on playback requirements (e.g., the sufficiency of decoding computing resources, the user-selected viewpoint, or the video space region viewed by the user), it selects NAL units with certain layer numbers for decoding and output. When decoding computing resources decrease, the temporal layer numbers in the NAL units can be further obtained, skipping the decoding of NAL units with high temporal layer numbers, thus reducing the decoding frame rate to accommodate the decrease in decoding computing resources.

[0282] The advantages of prioritizing NAL unit selection based on layer numbering followed by temporal layer numbering are understandable: 1) For videos with multiple viewpoint or spatial region levels, users typically select certain layers based on their habits and preferences, which needs to be prioritized. When decoding computing resources are insufficient, the selection of videos from different temporal layers is primarily used to adapt to the decoding computing resources. 2) For videos with multiple resolution levels, selecting different temporal or resolution layers can adapt to the decoding computing resources, but prioritizing different resolution layers while keeping the temporal layer unchanged ensures the time accuracy requirements of machine tasks, as machine task networks can typically support the effective recognition of objects at different resolutions. Furthermore, if NAL units are selected first based on temporal layer numbering, when decoding computing resources are insufficient, the decoder needs to switch between videos at different resolution levels, different viewpoint levels, or different spatial region levels. This leads to discontinuities in the output video content, such as fluctuating video resolution, constant viewpoint switching, and continuous changes in spatial regions.

[0283] It's understandable that introducing layer numbering to the reconstructed data NAL packets allows the system layer or decoder to synchronously manage both video data NAL packets and reconstructed data NAL packets. When it's necessary to select a subset of layers for decoding, the parsing of unnecessary reconstructed data NAL packets can be avoided. Simultaneously, different reconstructed data can be provided for different layers of video data.

[0284] In this embodiment, a layer number is introduced into the reconstructed data NAL packet to indicate the layer of the reconstructed data NAL packet. This layer is consistent with or equivalent to the layer of the corresponding video data NAL packet. When a video data NAL packet at a certain layer is skipped or selected, the reconstructed data NAL packet at the same layer can also be skipped or selected.

[0285] Other embodiments of this technical solution include, for example, each VCM unit contains only one reconstruction information NAL unit or kernel video NAL unit. The advantage is that the encoder can encode and encapsulate data in real time, that is, each image can be encapsulated into a set of VCM units to carry reconstruction information NAL units and kernel video NAL units without waiting for the encoding of more images. The disadvantage is that more VCM unit header information consumes additional data.

[0286] Another embodiment currently defines only a few types for the VCM unit, such as VPS (VCM Parameter Set), RSD (ReStoration Data), RSD_RAP (Reconstruction Information with Random Access Support), CVD (Coded Video Data), and CVD_RAP (Kernel Video Data with Random Access Support). In other implementations, the class identifier of the VCM unit and the category identifier of the NAL unit can be combined to form a richer VCM unit type definition. This reuses the category identifier of the NAL unit, which can quickly parse more type information in the header of the VCM unit without incurring additional header information data consumption.

[0287] In another embodiment, the NAL unit in the sub-bitstream currently uses the packet size to identify the NAL unit segmentation boundary. In other implementations, the start code can also be used to identify the NAL unit segmentation boundary. In this case, in order to avoid the conflict between the start code of the NAL unit and the start code of the upper-level VCM unit, two different start code prefixes can be used. For example, the VCM unit uses a 32-bit start code prefix, while the NAL unit uses a 24-bit start code prefix.

[0288] In addition, this technical solution can also be used in FCM (Feature coding for machine) because FCM adopts a bitstream structure similar to VCM. For example, VCM units are extended to FCM units, VCM NAL units are extended to FCM NAL units, and kernel video NAL units are still used in FCM.

[0289] FCM units include the following types: Global Vision Model Parameter Set (VMPS), Reconstruction Data (RSD), and Coded Video Data (CVD). When FCM and VCM share a stream structure, the stream unit includes units from both FCM and VCM. A unit type identification method is shown in the table below.

[0290] The type indicated by fuh_unit_type can be shown in the table below.

[0291] Among them, the FCM_VMPS type cell contains information for the reconstruction of feature data, and this information is recorded in the bitstream in the form shown in the table below.

[0292] Here, `vmps_vision_model_parameter_set_id` represents the VMPS number. Multiple different VMPSs are allowed in the bitstream, each using a different number, and can be referenced by the reconstruction of video features at different times. `img_wid` and `img_hei` represent the original resolution of the source video for the feature data in the video bitstream. `scaled_img_wid` and `scaled_img_hei` represent the resolution of the image after the feature extraction part of the source video has been scaled by the intelligent task network. This is because the intelligent task network performs preprocessing before analyzing the source video. `total_numer_of_input` represents the total number of frames of video features in the bitstream.

[0293] In one implementation, the decoding method obtains various types of FCM NAL units from the units of the reconstructed data in the bitstream. Examples of FCM NAL unit types are shown in the table below.

[0294] The `fcm_nal_unit_type` indicates the type of FCM NAL unit: `FCM_NAL_FSPS` represents a sequence-level NAL unit, `FCM_NAL_FPPS` represents an image-level NAL unit, `FCM_NAL_EOSS` represents the end unit of the reconstructed data, `FCM_NAL_SEI` represents auxiliary enhancement information of the reconstructed data, `FCM_NAL_RSV` represents a reserved unit type, and `FCM_NAL_UNSPEC` represents a reserved unit type that can be defined by the user.

[0295] In one implementation, the decoding method obtains sequence-level and image-level parameters for reconstructing features from units of reconstructed data in the bitstream. Examples of the syntax structures of these parameters are shown in the table below.

[0296] The sequence-level feature parameter set feat_seq_parameter_set_rbsp contains:

[0297] 1. Parameter set number: fsps_feat_seq_parameter_set_id;

[0298] 2. The number of layers for reconstructing the feature data is num_ori_feat_layers, and the width, height, and number of channels of each feature layer are num_ori_feat_wid[i], num_ori_feat_hei[i], and num_ori_feat_chan[i].

[0299] 3. The width (fused_feat_wid) and height (fused_feat_hei) of the decoded fusion feature obtained by the kernel decoder;

[0300] 4. The kernel decoder's switch inner_decoding_bypass_flag and type information inner_coding_idx. inner_coding_idx can indicate the type of the kernel decoder, such as VVC, HEVC, or AVC.

[0301] 5. Switches for tools that can be used during the decoding process, such as the dequantization switch dequant_bypass_flag, the feature repacking switch unpacking_bypass_flag, the feature reconstruction switch feat_restoration_bypass_flag, the temporal upsampling switch temporal_upsampling_enable_flag, the reconstructed feature adjustment switch restored_feat_refine_flag, and the fused feature adjustment switch fused_feat_refine_flag, etc.

[0302] 6. Feature reconstruction information: feat_restoration_info contains information about the network model used for feature reconstruction. One network model is the default feature reconstruction network, indexed by feat_restoration_weight_idx. Features reconstructed using this network also need to be cropped in width and height according to pad_size_min. The other network model is obtained by the decoding method based on fcm_decoder_info_sei_id. It can be obtained from the encoding end or from certain links.

[0303] 7. `restored_feat_refine_refresh_period` and `fused_feat_refine_refresh_period` record the periods of reconstruction feature adjustment and fusion feature adjustment, respectively.

[0304] The image-level parameter set feature_pic_parameter_set_rbsp contains:

[0305] 1. The parameter set number fpps_feat_pic_parameter_set_id and the sequence-level parameter set number fpps_feat_seq_parameter_set_id referenced by this parameter set;

[0306] 2. The parameters restored_feat_std and restored_feat_mean used for reconstructed feature adjustment, and the parameters fused_feat_std and fused_feat_mean used for fused feature adjustment.

[0307] In one embodiment, after obtaining the decoded fusion features based on the above information, the decoding method uses the feature reconstruction network specified by feat_restoration_info to decompose the fusion features and obtain reconstructed features. These reconstructed features contain multiple layers of sub-features at different scales, which can be used in networks after feature extraction in machine intelligence tasks to obtain task analysis results.

[0308] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps; or steps from different embodiments may be combined into a new technical solution.

[0309] Based on the same inventive concept as the foregoing embodiments, this application provides an encoder; Figure 11 is a schematic diagram of the composition structure of the encoder provided in this application. As shown in Figure 11, the encoder 110 may include a first determining module 1101, a first generating module 1102, and a first encoding module 1103, wherein:

[0310] The first determining module 1101 is configured to determine the reconstruction information and the first indication information of the encoded video data;

[0311] The first generation module 1102 is configured to generate a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the hierarchy of the first unit;

[0312] The first encoding module 1103 is configured to generate a first bitstream based on the first unit.

[0313] In some embodiments, the first determining module 1101 is further configured to determine the encoded video data and the second indication information of the original video; the first generating module 1102 is further configured to generate a second unit based on the encoded video data and the second indication information; the second indication information is used to indicate the hierarchy of the second unit; and the first encoding module 1103 is further configured to generate a second bitstream based on the second unit.

[0314] The description of the encoder embodiments above is similar to the description of the encoding / decoding method embodiments above, and has similar beneficial effects. For technical details not disclosed in the encoder embodiments of this application, please refer to the description of the encoding / decoding method embodiments of this application for understanding.

[0315] Figure 12 is a schematic diagram of the hardware structure of the encoder provided in an embodiment of this application. As shown in Figure 12, the encoder 120 may include: a first communication interface 1201, a first memory 1202, and a first processor 1203; the various components are coupled together through a first bus system 1204. It can be understood that the first bus system 1204 is used to realize the connection and communication between these components. In addition to a data bus, the first bus system 1204 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the first bus system 1204 in Figure 12.

[0316] The first communication interface 1201 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0317] The first memory 1202 is used to store computer programs that can run on the first processor 1203;

[0318] The first processor 1203 is configured to execute the steps of the encoding method described in the embodiments of this application when running the computer program.

[0319] It is understood that the first memory 1202 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The first memory 1202 of the system and method described in this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0320] The first processor 1203 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the first processor 1203 or by instructions in software form. The first processor 1203 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the encoding method disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the first memory 1202. The first processor 1203 reads the information in the first memory 1202 and completes the steps of the above encoding method in conjunction with its hardware.

[0321] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof. For software implementation, the technology described in this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions described in this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0322] Alternatively, as another embodiment, the first processor 1203 is further configured to execute the method described in any of the foregoing encoding method embodiments when running the computer program.

[0323] Based on the same inventive concept as the foregoing embodiments, Figure 13 is a schematic diagram of the composition structure of the decoder provided in the embodiment of this application. As shown in Figure 13, the decoder 130 may include a second determining module 1301, a decoding module 1302, and a reconstruction module 1303; wherein:

[0324] The second determining module 1301 is configured to determine first indication information of a first unit in the first bitstream, wherein the first indication information is used to indicate the level of the first unit; the first unit includes the first indication information and reconstruction information.

[0325] Decoding module 1302 is configured as the first unit whose parsing level meets the first condition, and determines the reconstruction information of the corresponding decoded video data;

[0326] The reconstruction module 1303 is configured to determine the reconstructed video corresponding to the decoded video data based on the reconstruction information.

[0327] In some embodiments, the decoding module 1302 is further configured to parse the first unit whose temporal level is a specific temporal level from the first unit that has never been parsed and whose level satisfies the first condition, and determine the reconstruction information of the corresponding decoded video data.

[0328] In some embodiments, the second determining module 1301 is further configured to determine second indication information of a second unit in the second bitstream, the second indication information being used to indicate the level of the second unit; the second unit includes the second indication information and encoded video data; the decoding module 1302 is further configured to parse the second unit whose level satisfies the first condition and determine the decoded video data.

[0329] The description of the decoder embodiments above is similar to the description of the encoding / decoding method embodiments above, and has similar beneficial effects. For technical details not disclosed in the decoder embodiments of this application, please refer to the description of the encoding / decoding method embodiments of this application for understanding.

[0330] Understandably, in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular component. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.

[0331] Figure 14 is a schematic diagram of the hardware structure of the decoder provided in an embodiment of this application. As shown in Figure 14, the decoder 140 may include: a second communication interface 1401, a second memory 1402, and a second processor 1403; the various components are coupled together through a second bus system 1404. It is understood that the second bus system 1404 is used to realize the connection and communication between these components. In addition to a data bus, the second bus system 1404 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the second bus system 1404 in Figure 14.

[0332] The second communication interface 1401 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0333] The second memory 1402 is used to store computer programs that can run on the second processor 1403;

[0334] The second processor 1403 is configured to, when running the computer program, perform:

[0335] Obtain the block vector information of the reference block of the current block, wherein the block vector information includes the first block vector and / or block vector related information;

[0336] The block vector of the current block is determined based on the block vector information of the reference block;

[0337] The predicted value of the current block is determined based on the block vector of the current block.

[0338] Alternatively, as another embodiment, the second processor 1403 is also configured to perform the method described in any of the foregoing embodiments when running the computer program.

[0339] It is understood that the second memory 1402 has similar hardware functions to the first memory 1202, and the second processor 1403 has similar hardware functions to the first processor 1403; details will not be elaborated here.

[0340] Figure 15 is a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application. As shown in Figure 15, the encoding / decoding system 150 may include an encoder 1501 and a decoder 1502.

[0341] In this embodiment, encoder 1501 can be any of the encoders described in the foregoing embodiments, and decoder 1502 can be any of the decoders described in the foregoing embodiments.

[0342] In some embodiments, this application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program implements the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, it implements the decoding method as described in any of the foregoing embodiments.

[0343] In some embodiments, this application also provides a computer program product, including a computer program or instructions. When executed by a processor, the computer program or instructions implement the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program or instructions implement the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, they implement the decoding method as described in any of the foregoing embodiments.

[0344] In some embodiments, this application also provides a computer program that, when executed by a processor, implements the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program or instructions implement the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, implement the decoding method as described in any of the foregoing embodiments.

[0345] In some embodiments, this application also provides a computer-readable storage medium storing a bitstream thereon. The bitstream is generated by performing the steps of the encoding method as described in any of the foregoing embodiments.

[0346] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0347] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0348] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0349] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0350] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0351] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0352] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0353] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0354] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined to obtain new product embodiments without conflict. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined to obtain new method embodiments or device embodiments without conflict.

[0355] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A decoding method applied to a decoder, the method comprising: Determine the first indication information of the first unit in the first bitstream, wherein the first indication information is used to indicate the level of the first unit; The first unit contains the first indication information and the reconstruction information; The first unit whose parsing level satisfies the first condition is used to determine the reconstruction information of the corresponding decoded video data. Based on the reconstruction information, the reconstructed video corresponding to the decoded video data is determined.

2. The method according to claim 1, wherein, The first indication information is used to indicate at least one of the following levels of the first unit: Resolution hierarchy, wherein the first unit corresponding to the resolution hierarchy is used to determine the reconstructed video at the corresponding resolution; Viewpoint hierarchy, wherein the first unit corresponding to the viewpoint hierarchy is used to determine the reconstructed video of the corresponding viewpoint; The spatial region hierarchy, wherein the first unit corresponding to the spatial region hierarchy is used to determine the reconstructed video of the corresponding spatial region.

3. The method according to claim 2, wherein, The first condition includes a first sub-condition, which includes at least one of the following: The resolution level is a specific resolution level; The viewpoint hierarchy is a specific viewpoint hierarchy; The spatial region hierarchy is a specific spatial region hierarchy.

4. The method according to claim 2 or 3, wherein, The first indication information includes the value of a first syntax element, which is used to indicate one of the following levels: The resolution levels; The viewpoint hierarchy; The spatial region hierarchy.

5. The method according to claim 4, wherein, The first subcondition includes the first syntax element having a value equal to a first value, which indicates a specific resolution level, a specific viewpoint level, or a specific spatial region level.

6. The method according to any one of claims 3-5, wherein, The first indication information is also used to indicate the temporal level of the first unit, and the first unit corresponding to the temporal level is used to determine the reconstructed video of the corresponding temporal level.

7. The method according to claim 6, wherein, The first condition includes a first sub-condition and a second sub-condition; wherein the second sub-condition includes: the time domain level is a specific time domain level.

8. The method according to claim 6, wherein, The method further includes: The first unit that has never been parsed and whose level satisfies the first condition is the first unit that parses the time domain level as a specific time domain level, and the reconstruction information of the corresponding decoded video data is determined.

9. The method according to any one of claims 6-8, wherein, The first indication information includes the value of the second syntax element, which is used to indicate the time domain level.

10. The method according to claim 9, wherein, The second subcondition includes the second syntax element having a value equal to a second value, which indicates a specific time-domain level.

11. The method according to any one of claims 1-10, wherein, The first indication information is in the header information of the first unit, and the reconstruction information is in the load information of the first unit.

12. The method according to claim 11, wherein, The first unit is the reconstruction information NAL unit.

13. The method according to any one of claims 1-12, wherein, The method further includes: A second indication information is determined for a second unit in the second bitstream, the second indication information being used to indicate the level of the second unit; the second unit contains the second indication information and encoded video data; The second unit whose parsing level satisfies the first condition is used to determine the decoded video data.

14. The method according to claim 13, wherein, The second indication information corresponding to the second unit that satisfies the first condition at the hierarchy level is the same as or equivalent to the first indication information corresponding to the first unit that satisfies the first condition at the hierarchy level.

15. The method according to claim 13 or 14, wherein, The second indication information is in the header information of the second unit, and the encoded video data is in the payload information of the second unit.

16. The method according to claim 15, wherein, The second unit is the kernel video NAL unit.

17. An encoding method applied to an encoder, the method comprising: Determine the reconstruction information and primary indication information of the encoded video data; A first unit is generated based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the hierarchy of the first unit; A first bitstream is generated based on the first unit.

18. The method according to claim 17, wherein, Determining the first indication information includes: Second indication information is obtained from the second unit corresponding to the reconstructed information, and the second indication information is used to indicate the level of the second unit; The first instruction information is determined based on the second instruction information.

19. The method according to claim 18, wherein, The first instruction information is the same as or equivalent to the second instruction information.

20. The method according to any one of claims 17-19, wherein, The first indication information is used to indicate at least one of the following levels of the first unit: Resolution hierarchy, wherein the first unit corresponding to the resolution hierarchy is used to determine the reconstructed video at the corresponding resolution; Viewpoint hierarchy, wherein the first unit corresponding to the viewpoint hierarchy is used to determine the reconstructed video of the corresponding viewpoint; The spatial region hierarchy, wherein the first unit corresponding to the spatial region hierarchy is used to determine the reconstructed video of the corresponding spatial region.

21. The method according to claim 20, wherein, The first indication information includes the value of a first syntax element, and the value of the first syntax element is used to indicate one of the following levels: The resolution levels; The viewpoint hierarchy; The spatial region hierarchy.

22. The method according to claim 20 or 21, wherein, The first indication information is also used to indicate the temporal level of the first unit, and the first unit corresponding to the temporal level is used to determine the reconstructed video of the corresponding temporal level.

23. The method according to claim 22, wherein, The first indication information includes the value of the second syntax element, which is used to indicate the time domain level.

24. The method according to any one of claims 17-23, wherein, The first indication information is in the header information of the first unit, and the reconstruction information is in the load information of the first unit.

25. The method according to claim 24, wherein, The first unit is the reconstruction information NAL unit.

26. The method according to any one of claims 18-25, wherein, The method further includes: Determine the encoded video data and second indication information of the original video; A second unit is generated based on the encoded video data and the second indication information; the second indication information is used to indicate the hierarchy of the second unit. A second bitstream is generated based on the second unit.

27. The method according to claim 26, wherein, The second indication information is in the header information of the second unit, and the encoded video data is in the payload information of the second unit.

28. The method according to claim 27, wherein, The second unit is the kernel video NAL unit.

29. An encoder, comprising: The first determining module is configured to determine the reconstruction information and the first indication information of the encoded video data; The first generation module is configured to generate a first unit based on the reconstruction information and the first indication information; wherein the first indication information is used to indicate the hierarchy of the first unit. The first encoding module is configured to generate a first bitstream based on the first unit.

30. An encoder, comprising a first memory and a first processor; wherein: The first memory is used to store computer programs that can run on the first processor; The first processor is configured to perform the method as described in any one of claims 17 to 28 when running the computer program.

31. A decoder, comprising: The second determining module is configured to determine first indication information of the first unit in the first bitstream, wherein the first indication information is used to indicate the level of the first unit; The first unit contains the first indication information and the reconstruction information; The decoding module is configured as the first unit whose parsing level meets the first condition, and determines the reconstruction information of the corresponding decoded video data. The reconstruction module is configured to determine the reconstructed video corresponding to the decoded video data based on the reconstruction information.

32. A decoder, comprising a second memory and a second processor; wherein: The second memory is used to store computer programs that can run on the second processor; The second processor is configured to perform the method as described in any one of claims 1 to 16 when running the computer program.

33. A computer-readable storage medium, wherein, The computer-readable storage medium stores the bitstream generated by the encoding method as described in any one of claims 17 to 29.

34. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 16, or the method as described in any one of claims 17 to 29.