Encoding method, decoding method, encoder, decoder, and storage medium

WO2026199139A1PCT designated stage Publication Date: 2026-10-01ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084540
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2025084540_01102026_PF_FP_ABST
    Figure CN2025084540_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an encoding method, a decoding method, an encoder, a decoder, and a storage medium. The decoding method comprises: determining first indication information of a first unit in a bitstream, the first indication information being used for indicating the type of the first unit; when the first indication information indicates that the first unit is of a first type, parsing the first unit to determine reconstruction information of a first image / image feature; and performing reconstruction processing on the first image / image feature on the basis of the reconstruction information to determine a reconstructed image / reconstructed image feature.
Need to check novelty before this filing date? Find Prior Art

Description

An encoding / decoding method, encoder, decoder, and storage medium Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and in particular to an encoding and decoding method, encoder, decoder and storage medium. Background Technology

[0002] Video encoding and decoding technologies organize video streams using multiple sub-streams. For example, Video Coding for Machines (VCM) uses multiple sub-streams: one is the kernel video sub-stream obtained from the kernel encoder in the VCM encoder, and another is the reconstruction information sub-stream generated by the VCM encoder based on its preprocessing operations on the video. The kernel video sub-stream is decoded by the kernel decoder in the VCM decoder to obtain the decoded video, and the reconstruction information sub-stream is processed by the VCM decoder to obtain the reconstructed information. The decoded video, based on the reconstructed information, undergoes reconstruction processing to obtain the reconstructed video. This reconstructed video retains key semantic information and can achieve sufficient task accuracy for machine tasks.

[0003] The encoded data in each sub-stream is encapsulated in the form of Network Abstraction Layer (NAL) units, and the data packets in the sub-stream can be distinguished by the size of the data packets identified in the stream. Currently, the VCM video stream decoding method needs optimization. Summary of the Invention

[0004] This application provides an encoding / decoding method, an encoder, a decoder, and a storage medium.

[0005] In a first aspect, this application provides a decoding method applied to a decoder, the method comprising:

[0006] Determine the first indication information of the first unit in the bitstream, wherein the first indication information is used to indicate the type of the first unit;

[0007] If the first indication information indicates that the first unit is of the first type, the first unit is parsed to determine the reconstruction information of the first image / image features;

[0008] Based on the reconstruction information, the first image / image features are reconstructed to determine the reconstructed image / reconstructed image features.

[0009] Secondly, embodiments of this application provide an encoding method applied to an encoder, the method comprising:

[0010] Preprocess the original image to determine the reconstruction information of the first image / image features;

[0011] The reconstructed information is encoded to obtain the first unit;

[0012] Determine the first indication information of the first unit, wherein the first indication information is used to indicate the type of the first unit;

[0013] Add the first indication information to the first unit to generate a bitstream.

[0014] Thirdly, embodiments of this application provide an encoder, which includes a first processing unit and an encoding unit, wherein...

[0015] The first processing unit is configured to preprocess the original image to determine the reconstruction information of the first image / image features;

[0016] The encoding unit is configured to encode the reconstructed information to obtain a first unit;

[0017] The encoding unit is further configured to determine first indication information of the first unit, the first indication information being used to indicate the type of the first unit;

[0018] The encoding unit is further configured to add the first indication information to the first unit to generate a bitstream.

[0019] Fourthly, embodiments of this application provide an encoder, which includes a first memory and a first processor; wherein,

[0020] A first memory for storing computer programs that can run on a first processor;

[0021] The first processor is used to execute the method described in the second aspect when running a computer program.

[0022] Fifthly, embodiments of this application provide a decoder, which includes a decoding unit and a second processing unit; wherein:

[0023] The decoding unit is configured to determine first indication information of a first unit in the bitstream, wherein the first indication information is used to indicate the type of the first unit;

[0024] The decoding unit is further configured to parse the first unit and determine the reconstruction information of the first image / image features when the first indication information indicates that the first unit is of the first type;

[0025] The second processing unit is further configured to perform reconstruction processing on the first image / image features based on the reconstruction information to determine the reconstructed image / reconstructed image features.

[0026] Sixthly, embodiments of this application provide a decoder, which includes a second memory and a second processor; wherein,

[0027] The second memory is used to store computer programs that can run on the second processor;

[0028] The second processor is used to execute the methods described in the first aspect when running a computer program.

[0029] In a seventh aspect, embodiments of this application provide a computer-readable storage medium that stores a bitstream generated by such encoding method.

[0030] Eighthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed, implements the method as described in the first aspect or the method as described in the second aspect.

[0031] This application provides an encoding / decoding method, encoder, decoder, and storage medium. For a first unit carrying reconstructed information, different types of first units can be indicated by first indication information, which can provide the system layer or decoder with the ability to quickly filter the first unit, avoid wasting decoding resources on first units that cannot be decoded or do not need to be decoded, and improve decoding efficiency. Attached Figure Description

[0032] Figure 1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of this application;

[0033] Figure 2 is a schematic block diagram of a video encoder involved in an embodiment of this application;

[0034] Figure 3 is a schematic block diagram of a video decoder involved in an embodiment of this application;

[0035] Figure 4 shows a schematic diagram of the network architecture of an encoding / decoding system provided in an embodiment of this application;

[0036] Figure 5 is a schematic diagram of the V3C stream structure provided in an embodiment of this application;

[0037] Figure 6 is a schematic flowchart of a decoding method provided in an embodiment of this application;

[0038] Figure 7 is a schematic diagram of a VCC stream structure in an embodiment of this application;

[0039] Figure 8 is a schematic diagram of another VCC stream structure in an embodiment of this application;

[0040] Figure 9 is a schematic flowchart of an encoding method provided in an embodiment of this application;

[0041] Figure 10 is a schematic diagram of the composition structure of an encoder provided in an embodiment of this application;

[0042] Figure 11 is a schematic diagram of the specific hardware structure of an encoder provided in an embodiment of this application;

[0043] Figure 12 is a schematic diagram of the composition structure of a decoder provided in an embodiment of this application;

[0044] Figure 13 is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of this application;

[0045] Figure 14 is a schematic diagram of the composition structure of an encoding and decoding system provided in an embodiment of this application. Detailed Implementation

[0046] This application can be applied to the fields of image encoding and decoding, video encoding and decoding, hardware video encoding and decoding, dedicated circuit video encoding and decoding, and real-time video encoding and decoding. For example, the solution of this application can be combined with audio video coding standards (AVS), such as H.264 / Audio Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), Visual Volumetric Video-based Coding (V3C), Video Coding for Machine (VCM), and Feature Coding for Machine (FCM). Alternatively, the solutions in this application can be incorporated into other proprietary or industry standards, including ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, and ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), which include Scalable Video Codec (SVC) and Multi-View Video Codec (MVC) extensions. It should be understood that the technology in this application is not limited to any particular codec standard or technology.

[0047] The high-degree-of-freedom immersive coding system can be roughly divided into the following stages according to the task line: data acquisition, data organization and expression, data encoding and compression, data decoding and reconstruction, data synthesis and rendering, and finally presenting the target data to the user.

[0048] The encoding involved in the embodiments of this application is mainly video encoding and decoding. For ease of understanding, the video encoding and decoding system involved in the embodiments of this application will be introduced first with reference to Figure 1.

[0049] Figure 1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of this application. It should be noted that Figure 1 is only an example, and the video encoding and decoding system of this application includes, but is not limited to, the one shown in Figure 1. As shown in Figure 1, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (can be understood as compressing) video data to generate a bitstream, and transmits the bitstream to the decoding device. The decoding device decodes the bitstream generated by the encoding device to obtain decoded video data.

[0050] The encoding device 110 in this application embodiment can be understood as a device with video encoding function, and the decoding device 120 can be understood as a device with video decoding function. That is, the encoding device 110 and the decoding device 120 in this application embodiment include a wider range of devices, such as smartphones, desktop computers, mobile computing devices, laptops (e.g., laptop computers), tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, etc.

[0051] In some embodiments, encoding device 110 may transmit encoded video data (such as a bitstream) to decoding device 120 via channel 130. Channel 130 may include one or more media and / or means capable of transmitting encoded video data from encoding device 110 to decoding device 120.

[0052] In one example, channel 130 includes one or more communication media that enable encoding device 110 to transmit encoded video data directly to decoding device 120 in real time. In this example, encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum; optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.

[0053] In another example, channel 130 includes a storage medium that can store video data encoded by encoding device 110. The storage medium includes various local access data storage media, such as optical discs, DVDs, flash memory, etc. In this example, decoding device 120 can retrieve the encoded video data from the storage medium.

[0054] In another example, channel 130 may include a storage server that can store the video data encoded by encoding device 110. In this example, decoding device 120 can download the stored encoded video data from the storage server. Optionally, the storage server can store and transmit the encoded video data to decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0055] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0056] In some embodiments, the encoding device 110 may include a video source 111 in addition to the video encoder 112 and the input interface 113.

[0057] The video source 111 may include at least one of a video capture device (e.g., a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.

[0058] Video encoder 112 encodes video data from video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the encoding information of the pictures or picture sequences in the form of a bitstream. The encoding information may include encoded image data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax element structures. The SPS may contain parameters applied to one or more sequences. The PPS may contain parameters applied to one or more pictures. A syntax element structure refers to a set of zero or more syntax elements arranged in a specified order within the bitstream.

[0059] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data can also be stored on a storage medium or a storage server for subsequent retrieval by the decoding device 120.

[0060] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122. In some embodiments, in addition to the input interface 121 and the video decoder 122, the decoding device 120 may also include a display device 123.

[0061] The input interface 121 includes a receiver and / or a modem. The input interface 121 can receive encoded video data through channel 130.

[0062] The video decoder 122 is used to decode the encoded video data to obtain the decoded video data, and transmit the decoded video data to the display device 123.

[0063] Display device 123 displays the decoded video data. Display device 123 may be integrated with decoding device 120 or external to decoding device 120. Display device 123 may include various display devices, such as liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or other types of display devices.

[0064] Furthermore, Figure 1 is merely an example, and the technical solutions of this application are not limited to Figure 1. For example, the technology of this application can also be applied to one-sided video encoding or one-sided video decoding.

[0065] The video coding framework involved in the embodiments of this application is described below.

[0066] Referring to Figure 2, it shows a schematic block diagram of an encoder provided in an embodiment of this application. As shown in Figure 2, the encoder (specifically a "video encoder") 200 may include a transform and quantization unit 101, an intra-frame estimation unit 102, an intra-frame prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse transform and inverse quantization unit 106, a filter control and analysis unit 107, a filtering unit 108, an encoding unit 109, and a decoded image buffer unit 140, etc. Among them, the filtering unit 108 can implement deblocking filtering and sample adaptive offset (SAO) filtering, and the encoding unit 109 can implement header information encoding and context-based adaptive binary arithmetic coding (CABAC).For the input raw video signal, a video coding block can be obtained by partitioning it through a Coding Tree Unit (CTU). Then, the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transform and quantization unit 101, including transforming the residual information from the pixel domain to the transform domain and quantizing the resulting transform coefficients to further reduce the bit rate. The intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to perform intra-frame prediction on the video coding block. Specifically, the intra-frame estimation unit 102 and the intra-frame prediction unit 103 are used to determine the intra-frame prediction mode to be used to encode the video coding block. The motion compensation unit 104 and the motion estimation unit 105 are used to perform inter-frame prediction coding of the received video coding block relative to one or more blocks in one or more reference frames to provide time prediction information. The motion estimation performed by the motion estimation unit 105 is a process of generating motion vectors, which can estimate the motion of the video coding block. Then, the motion compensation unit 104 is used to perform the motion estimation based on the motion vectors determined by the motion estimation unit 105. The motion compensation is performed. After determining the intra-prediction mode, the intra-prediction unit 103 is also used to provide the selected intra-prediction data to the coding unit 109, and the motion estimation unit 105 also sends the calculated motion vector data to the coding unit 109. In addition, the inverse transform and inverse quantization unit 106 is used to reconstruct the video coding block, reconstruct the residual block in the pixel domain, and remove the block artifacts by the filter control analysis unit 107 and the filtering unit 108. Then, the reconstructed residual block is added to a predictive block in the frame of the decoding image buffer unit 140 to generate the reconstructed video coding block. The coding unit 109 is used to encode various coding parameters and quantized transform coefficients. In the CABAC-based coding algorithm, the context content can be based on adjacent coding blocks and can be used to encode information indicating the determined intra-prediction mode and output the bitstream of the video signal. The decoding image buffer unit 140 is used to store the reconstructed video coding block for prediction reference. As video image encoding proceeds, new reconstructed video encoding blocks are continuously generated, and these reconstructed video encoding blocks are stored in the decoding image buffer unit 140.

[0067] Referring to Figure 3, it shows a schematic block diagram of a decoder provided in an embodiment of this application. As shown in Figure 3, the decoder (specifically a "video decoder") 200 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra-frame prediction unit 203, a motion compensation unit 204, a filtering unit 205, and a decoded image buffer unit 206, etc., wherein the decoding unit 201 can perform header information decoding and CABAC decoding, and the filtering unit 205 can perform deblocking filtering and SAO filtering. After the input video signal undergoes the encoding process shown in Figure 1, the output video signal bitstream is generated. This bitstream is input into the decoder 200, first passing through the decoding unit 201 to obtain the decoded transform coefficients. These transform coefficients are then processed by the inverse transform and inverse quantization unit 202 to generate residual blocks in the pixel domain. The intra-frame prediction unit 203 can generate prediction data for the current video decoding block based on the determined intra-frame prediction mode and data from previously decoded blocks in the current frame or image. The motion compensation unit 204 determines the prediction information for the video decoding block by analyzing motion vectors and other associated syntax elements, and uses... The prediction information is used to generate a predictive block of the video block being decoded; the decoded video block is formed by summing the residual block from the inverse transform and inverse quantization unit 202 with the corresponding predictive block generated by the intra-prediction unit 203 or the motion compensation unit 204; the decoded video signal is passed through the filtering unit 205 to remove block artifacts, which can improve video quality; then the decoded video block is stored in the decoding image buffer unit 206, which stores reference images for subsequent intra-prediction or motion compensation, and is also used for the output of the video signal, thus obtaining the recovered original video signal.

[0068] Furthermore, this application embodiment also provides a network architecture for an encoding / decoding system including an encoder and a decoder. Figure 4 shows a schematic diagram of such a network architecture. As shown in Figure 4, the network architecture includes one or more electronic devices 13 to 1N and a communication network 01. The electronic devices 13 to 1N can perform video interaction through the communication network 01. During implementation, the electronic devices can be various types of devices with video encoding / decoding capabilities. For example, the electronic devices may include smartphones, tablets, personal computers, personal digital assistants, navigators, digital phones, video phones, televisions, sensing devices, servers, etc., without specific limitations. Additionally, the decoder or encoder described in this application embodiment can be the aforementioned electronic device.

[0069] It should be noted that the embodiments of this application can be applied to encoders, decoders, or even both encoders and decoders, but the embodiments of this application are not specifically limited.

[0070] The above describes the basic flow of a video codec under a block-based hybrid coding framework. With the development of technology, some modules or steps of this framework or flow may be optimized. This application is applicable to the basic flow of a video codec under this block-based hybrid coding framework, but is not limited to this framework and flow.

[0071] Based on some related technologies of the aforementioned video encoding and decoding systems, the encoded video bitstream contains multiple parallel sub-bitstreams. For example, the V3C (Visual Volumetric Video-based Coding) encoded video bitstream contains multiple parallel sub-bitstreams, such as atlas sub-bitstreams, attribute video sub-bitstreams, and geometry video sub-bitstreams. The data in these sub-bitstreams is encapsulated in the form of NAL units.

[0072] Figure 5 is a schematic diagram of the V3C bitstream structure provided in an embodiment of this application. The V3C bitstream includes: the V3C parameter set (V3C_parameter_set()) of V3C_VPS may include ptl_profile_toolset_idc. If ptl_profile_toolset_idc is 128 / 129 / 130, it indicates that the current bitstream simultaneously contains both VPCC extend and MIV main bitstreams.

[0073] The Atlas sequence parameter set (Atlas_sequence_parameter_set_rbsp()) in NAL_ASPS of the V3C_AD stitched sub-bitstream (Atlas_sub_bitstream()) can include asps_vpcc_extension_present_flag and asps_miv_extension_present_flag. When ptl_profile_toolset_idc is 128 / 129 / 130, asps_vpcc_extension_present_flag is true (i.e., 1), and asps_miv_extension_present_flag is also true (i.e., 1).

[0074] The ACL NAL unit type (ACL_NAL_unit_type) in V3C_AD's Atlas_sub_bitstream() includes hybrid stitching information. For example, the atlas tile data unit (atlas_tile_data_unit()) can include atdu_type_flag. If atdu_type_flag is yes (i.e., 1), it indicates that the current tile belongs to a point cloud tile; if atdu_type_flag is no (i.e., 0), it indicates that the current tile belongs to a multi-view video tile.

[0075] Furthermore, the sub-tack information data (patch_information_data) includes sub-tack data units (patch_data_unit). If atdu_type_flag is negative and asps_miv_extension_present_flag is positive, it indicates that the current sub-tack is implemented using a multi-view video coding standard. If atdu_type_flag is positive, it indicates that the current sub-tack is implemented using a point cloud video decoding standard.

[0076] The data in the V3C_AD splicing diagram sub-bit stream (also called sub-code stream) is also encapsulated in the form of NAL units. It can be distinguished by the data packet size of each NAL unit identified in the code stream, or by the start prefix code of the NAL unit.

[0077] The video sub-bitstreams include: V3C_GVD video sub-bitstream(), V3C_AVD video sub-bitstream(), V3C_CAD video sub-bitstream(), and V3C_PVD video sub-bitstream().

[0078] In addition, to distinguish different sub-stream data within the video stream, the sub-stream data is split and encapsulated in V3C units. Each V3C unit records the type of sub-stream data it carries. The V3C units in the video stream are distinguished based on the data packet size of each V3C unit identified within the stream.

[0079] For example, Video Coding for Machines (VCM) currently organizes video bitstreams into multiple sub-streams. One sub-stream is the kernel video sub-stream obtained by the kernel encoder in the VCM encoder, and another sub-stream is the reconstruction information sub-stream generated by the VCM encoder based on its preprocessing operations on the video. The kernel video sub-stream is decoded by the kernel decoder in the VCM decoder to obtain the decoded video, and the reconstruction information sub-stream is processed by the VCM decoder to obtain the reconstructed information. The decoded video is then reconstructed based on the reconstructed information to obtain the reconstructed video. This reconstructed video retains key semantic information and can achieve sufficient task accuracy for machine tasks.

[0080] Currently, to avoid disrupting the existing characteristics of the kernel codec and ensure VCM's compatibility with existing kernel codecs, the kernel video sub-stream retains its original characteristics and is encapsulated in the form of NAL units. Since the reconstruction information sub-stream introduces a new data unit type, VCM currently uses a bitstream format similar to V3C to organize the VCM video bitstream. However, V3C's limitations prevent the VCM video bitstream from effectively supporting mainstream scenarios requiring real-time encoding and decoding, such as surveillance, autonomous driving, and smart manufacturing. Therefore, this technical solution designs a new bitstream structure and encoding / decoding method to address the problems of existing VCMs.

[0081] For example, feature coding for machine (FCM) video streams also include video data sub-streams and reconstructed data sub-streams.

[0082] The data in these sub-streams is encapsulated in the form of NAL units. Each type of data packet defines a NAL unit type. The advantage of this is that it allows the decoder or system layer to quickly scan and identify the type of each NAL unit. However, its drawback is that each new NAL unit requires a NAL unit type number, limiting the number of NAL unit types it can support. This makes it ineffective for scenarios with multiple sub-streams, each containing its own unique NAL unit type. Furthermore, since all NAL units need to be identified, it increases the scanning burden on the decoder or system layer.

[0083] Based on this, embodiments of this application provide an encoding / decoding method, an encoder, a decoder, and a storage medium. For a first unit carrying reconstructed information, different types of first units can be quickly identified through first indication information. For example, a first unit with a first type can be quickly identified, thereby parsing the first unit and avoiding the waste of decoding resources caused by decoding a first unit that cannot be decoded or does not need to be decoded, thus improving decoding efficiency.

[0084] The above-mentioned related technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application. The embodiments of this application include at least some of the following contents.

[0085] In one embodiment of this application, referring to Figure 6, a flowchart of a decoding method provided by an embodiment of this application is shown. The decoding method is applied to a decoder, which is used to analyze and parse the received video bitstream to obtain a reconstructed image. The decoder includes, but is not limited to, system layer and kernel decoders.

[0086] As shown in Figure 6, the method may include:

[0087] S601: Determine the first indication information of the first unit in the bitstream, the first indication information being used to indicate the type of the first unit;

[0088] S602: If the first indication information indicates that the first unit is of the first type, the first unit is parsed to determine the reconstruction information of the first image / image features;

[0089] S603: Based on the reconstruction information, perform reconstruction processing on the first image / image features to determine the reconstructed image / reconstructed image features;

[0090] The bitstream can be the video bitstream obtained after the encoder performs preprocessing and encapsulation operations on the input video, or it can be a sub-bitstream within the video bitstream. For example, in the V3C coding standard, the video bitstream includes atlas sub-bitstreams, attribute video sub-bitstreams, and geometry video sub-bitstreams. As another example, the VCM video bitstream includes a kernel video sub-bitstream obtained from the kernel encoder in the VCM encoder, and a reconstructed data sub-bitstream generated by the VCM encoder based on the reconstructed information obtained from its preprocessing operations on the video. Similarly, the FCM video bitstream also includes video data sub-bitstreams and reconstructed data sub-bitstreams.

[0091] The bitstream includes one or more first units. In some embodiments, the first unit is a data packet or data unit carrying reconstruction information. Exemplarily, the first unit may be a Network Abstraction Layer (NAL) unit carrying reconstruction information.

[0092] In some embodiments, determining first indication information of a first unit in a bitstream includes: determining first indication information from the first unit of the bitstream.

[0093] The first unit comprises two parts: header information and payload data. In some embodiments, first indication information is located in the header information of the first unit. The header information of the first unit is scanned to determine the first indication information, and the type of the first unit is quickly identified based on the first indication information, thereby determining whether to further parse the payload portion of the first unit. This avoids wasting decoding resources caused by decoding the first unit, which is either undecoding or unnecessary, and improves decoding efficiency.

[0094] Accordingly, determining the first indication information of the first unit in the bitstream may include: identifying one or more first units in the bitstream and determining the first indication information of the first unit from the header information of the first unit.

[0095] In some embodiments, the value of the first indication information is used to determine whether the first unit is of the first type, so as to determine whether to parse the first unit.

[0096] For example, the type of the first unit is determined based on the value of the first indication information and the mapping relationship, wherein the mapping relationship includes at least two mapping relationships between the values ​​of the indication information and the type of the first unit.

[0097] In some embodiments, the method further includes: skipping the step of parsing the first unit if the first indication information indicates that the first unit is of the second type.

[0098] In some embodiments, the first indication information is used to indicate at least two types of first units, each type of first unit carrying different types of reconstruction information, and each type of first unit corresponding to specific parsing conditions in different usage scenarios. By classifying the first units carrying reconstruction information in a more detailed manner, appropriate decoding processing can be performed for different types of first units, thereby improving decoding efficiency.

[0099] For example, at least two types of first units include one or more sequence-level first units and one or more image-level first units, wherein the sequence-level first unit carries sequence-level reconstruction information, and the image-level first unit may carry reconstruction information of the entire image or carry sub-image-level reconstruction information.

[0100] In some embodiments, the first indication information is used to indicate at least two types of image-level first units. The image-level first unit can carry reconstruction information of the entire image, or it can carry reconstruction information of a sub-image level. For example, a sub-image can be a partial image region such as a coding tree unit (CTU), a coding unit (CU), or a region of interest (ROI).

[0101] For example, at least two types include: a first type for indicating that the first unit is parsed if a first condition is met; and a second type for indicating that the first unit is not parsed if a second condition is met.

[0102] The first condition can be understood as the parsing condition of the first unit of the first type, and the second condition can be understood as the non-parsing condition (skipped condition) of the first unit of the second type.

[0103] In some embodiments, if a first condition is met and the first indication information indicates that the first unit is of a first type, the first unit is parsed to determine the reconstruction information of the first image / image features; if a second condition is met and the first indication information indicates that the first unit is of a second type, the first unit is not parsed.

[0104] In this application embodiment, the first type includes at least one of the following: random access point type, resolvable type, and hierarchical access point type.

[0105] In some embodiments, the method may further include: acquiring a random access event; and in response to the random access event, determining a first unit of a random access point type based on the random access time and first indication information, so as to parse the first unit of the random access point type. For example, based on the random access time and the first indication information, a first unit of a random access point type closest to the random access time is determined, where the random access point type can be a specific random access point type or any random access point type.

[0106] The random access point type is used to indicate the type of the first unit that is scanned and resolved in response to a random access event. That is, for the first unit of the random access point type, the first condition includes obtaining a random access event, which can be triggered by the user or by the system itself.

[0107] For example, the random access point type includes at least one of the following: a first subtype, further used to indicate other first units after the first unit in the parsed bitstream; a second subtype, further used to indicate a portion of the first units after the first unit in the unparsed bitstream; and a third subtype, further used to indicate one or more first units in the parsed bitstream associated with the first unit.

[0108] For example, the first subtype can be Instantaneous Decoding Refresh (IDR) type, marked as VCM_NALU_RSD_IDR, indicating that the type of the reconstructed data NAL packet is IDR, and the type of the corresponding video data NAL packet is also IDR or equivalent. When decoding and playback need to start from a certain point in time in the video stream, the decoding operation scans the video stream and finds the video data NAL packet of the closest IDR type to that point in time. Decoding the video data begins from this point. In addition, the nearest VCM_NALU_RSD_IDR type reconstructed data NAL packet is found, and the reconstructed data in it is parsed to reconstruct the previously decoded video data, thus obtaining the reconstructed video data.

[0109] For example, the second subtype can be Clean Random Access (CRA), denoted as VCM_NALU_RSD_CRA, which is similar in function to the VCM_NALU_RSD_IDR type. The video data NAL packet corresponding to this type of reconstructed data NAL packet is of type CRA or equivalent. The decoding operation of this type of reconstructed data NAL packet is similar to the aforementioned operation.

[0110] In another implementation, the first subtype and the second subtype can be merged into one type, for example, both called Intra Random Access Point (IRAP) type, labeled VCM_NALU_RSD_IRAP, indicating that the type of video data NAL packet corresponding to the reconstructed data NAL packet of this type is IRAP type.

[0111] For example, the third subtype can be a progressive decoding refresh (GDR) type, denoted as VCM_NALU_RSD_GDR. The VCM_NALU_RSD_GDR type is similar in function to the VCM_NALU_RSD_IDR type; the video data NAL packet corresponding to this type of reconstructed data NAL packet is of type GDR. When decoding and playback need to begin from a specific point in the video stream, the decoding operation scans the video stream and finds the nearest GDR-type video data NAL packet to that point in time. Using this as a starting point, it finds one or more other related GDR-type video data NAL packets and begins decoding the video data. Additionally, it finds the corresponding VCM_NALU_RSD_GDR-type reconstructed data NAL packet, parses it to obtain the reconstructed data, and uses this data to reconstruct the previously decoded video data, thus obtaining the reconstructed video data.

[0112] In another implementation, the first subtype and the third subtype can be merged into one type, for example, both called Intra Random Access Point (IRAP) type, marked as VCM_NALU_RSD_IRAP, indicating that the type of video data NAL packet corresponding to the reconstructed data NAL packet of this type is IRAP type.

[0113] In another implementation, the first subtype, the second subtype, and the third subtype can be merged into one type.

[0114] In some embodiments, a resolvable type is used to indicate the type of non-random access point that can be correctly resolved in response to a random access event.

[0115] For example, the resolvable type includes at least one of the following: a fourth subtype, used to indicate parsing the current first unit, wherein the reconstructed image / reconstructed image features corresponding to the current first unit are output or displayed before the reconstructed image / reconstructed image features corresponding to the first unit of the random access point type; a fifth subtype, used to indicate parsing the current first unit, wherein the reconstructed image / reconstructed image features corresponding to the current first unit are output or displayed after the reconstructed image / reconstructed image features corresponding to the first unit of the random access point type. The current first unit can be understood as the first unit indicated by the first indication information.

[0116] For example, the fourth subtype can be a Random Access Decodable Leading (Picture) (RADL) type, marked as VCM_NALU_RSD_RADL, indicating that the video data NAL packet corresponding to this reconstructed data NAL packet is of type RADL. When the decoding operation receives a RADL type video data NAL packet after the first IRAP type (i.e., any random access point type) video data NAL packet, since the RADL type video data NAL packet only depends on the IRAP type video data NAL packet and subsequent video data NAL packets, it can be correctly decoded. Therefore, the decoding operation decodes the RADL type video data NAL packet. Additionally, the decoding operation finds and decodes the VCM_NALU_RSD_RADL type reconstructed data NAL packet after the first VCM_NALU_RSD_IRAP type reconstructed data NAL packet. The reconstructed image of the RADL type video data packet should be output or displayed before the reconstructed image of the IRAP type video data packet.

[0117] For example, the fifth subtype can be of type TRAIL (tail (image)), marked as VCM_NALU_RSD_TRAIL, indicating that the video data NAL packet corresponding to this reconstructed data NAL packet is of type TRAIL. The decoding operation can correctly decode the TRAIL type video data NAL packet, and also decode the corresponding reconstructed data NAL packet, thus completing the reconstruction of the decoded video data.

[0118] In some embodiments, the tiered access point type is used to indicate that, in response to a tier switching event, a first unit is resolved to enter the target tier from the first unit. That is, for a first unit of the tiered access point type, the first condition includes obtaining a tier switching event. The tier switching event can be triggered by the user or automatically by the system.

[0119] For example, the hierarchical access point type includes a sixth subtype for indicating that the first unit is an access point of the target hierarchical level.

[0120] For example, the sixth subtype can be Step-wise Temporal Sub-layer Access (STSA), denoted as VCM_NALU_RSD_STSA. This type of video data NAL packet allows the decoder to start from a specific time point or image and, according to a certain order and rules, progressively decode images or data from higher temporal layers to achieve more refined processing of video content or presentation of different quality levels.

[0121] For example, the hierarchical access point type includes: a seventh subtype, used to indicate that the first unit is the final access point of the target hierarchy; and an eighth subtype, used to indicate that the first unit is an intermediate access point of the target hierarchy.

[0122] The seventh subtype can be Temporal Sub-layer Access (TSA), denoted as VCM_NALU_RSD_TSA, and the eighth subtype can be Step-wise Temporal Sub-layer Access (STSA), denoted as VCM_NALU_RSD_STSA. This type of video data NAL packet allows the decoder to start from a specific time point or image and, according to a certain order and rules, correctly decode TSA and / or STSA type video data NAL packets, progressively decoding images or data at higher temporal layers to achieve more refined processing of video content or presentation of different quality levels.

[0123] In some embodiments, the method further includes: acquiring a tier switching event; and in response to the tier switching event, determining a first unit of the tier access point type based on first indication information, so as to parse the first unit of the tier access point type.

[0124] In some embodiments, the second type can be the type of the image-level first unit. Exemplarily, the second type includes at least a random access skip-before type. That is, in response to a random access event, the parsing of a portion of the image's reconstruction information NAL unit and kernel video NAL unit is skipped to save decoding resources.

[0125] The Random Access Skipped Leading (Picture) (RASL) type, marked as VCM_NALU_RSD_RASL, indicates that the video data NAL packet corresponding to this reconstructed data NAL packet is of type RASL. When the decoding operation receives a RASL type video data NAL packet after the first IRAP type (i.e., any random access point type) video data NAL packet, because the RASL type video data NAL packet depends on the video data NAL packets preceding the IRAP type video data NAL packet, and these video data NAL packets have not been received, the decoding operation skips the decoding of the RASL type video data NAL packet. In addition, the decoding operation also skips the decoding of VCM_NALU_RSD_RASL type reconstructed data NAL packets after the first VCM_NALU_RSD_IRAP type reconstructed data NAL packet.

[0126] In some embodiments, the first indication information is located in the header information of the first unit. For example, the first indication information is located in the header information of the Reconstruction Information (NAL) unit.

[0127] For example, the syntax element structure of the reconstructed information NAL unit header information is as follows:

[0128] rbsp_byte[i] is the i-th byte of RBSP. rbsp_byte represents the reconstruction information in the NAL unit. The data type of rbsp_byte is determined by vcm_nal_unit_type.

[0129] The reconstruction information NAL package also includes more finely categorized image-level reconstruction data NAL packages. For example, the header of a reconstruction data NAL package records a NAL category identifier, which identifies the type of video image to be decoded in the video NAL package corresponding to the reconstruction data recorded in that NAL package. A specific syntax example and decoding operation description are as follows:

[0130] In some embodiments, the first indication information is located in the header information of the associated unit of the first unit.

[0131] For example, the associated unit includes the upper-level unit to which the first unit belongs. For instance, the upper-level unit to which the first unit belongs can be a reconstruction information VCM unit, a reconstruction information FCM unit, or a reconstruction information V3C unit. As another example, the upper-level unit to which the first unit belongs can also be a VCM unit, which can include the first unit and a second unit. The first unit is a data packet or data unit carrying encoded reconstruction information, and the second unit is a data packet or data unit carrying encoded image information. For example, the first unit can be a NAL unit carrying reconstruction information, and the second unit can be a NAL unit carrying image information.

[0132] For example, the association unit includes a second unit, which includes image information of the first image / image features. It is understood that by establishing a connection between the reconstructed information first unit and the corresponding kernel video second unit, and reusing the second indication information of the kernel video second unit to simultaneously indicate the type of both the first and second units, codeword consumption is reduced.

[0133] It should be noted that the second unit specifically includes the encoded image data of the first image / image features, and the decoded image data of the first image / image features is obtained by parsing the second unit.

[0134] In some embodiments, the method further includes: determining second indication information for a second unit in the bitstream, the second indication information indicating the type of the second unit; if the second indication information indicates that the second unit is a third type, parsing the second unit to determine a first image / image feature, the third type being consistent with or equivalent to the first type; if the second indication information indicates that the second unit is a fourth type, not parsing the second unit, the fourth type being consistent with or equivalent to the second type of the first unit.

[0135] The third type is used to indicate that the second unit should be parsed if the third condition is met; the fourth type is used to indicate that the second unit should not be parsed if the fourth condition is met.

[0136] The third type and the first type can be the same type or equivalent types. For example, the first type is a random access point type, and the third type can be a random access point type or a more granular random access point subtype.

[0137] Type 4 and Type 2 can be the same type or equivalent types. For example, both Type 2 and Type 4 are random access skip-prerequisite types.

[0138] The third condition can be understood as the analytic condition of the second unit of the third type, and the fourth condition can be understood as the non-analytic condition (skipped condition) of the second unit of the fourth type. The third condition and the first condition are the same or equivalent conditions. The fourth condition and the second condition are the same or equivalent conditions.

[0139] It should be noted that the first unit carries the reconstruction information of the first image / image features. The first unit specifically includes encoded reconstruction information / data, and parsing the first unit yields decoded reconstruction information / data. The second unit carries the image information of the first image / image features. The second unit specifically includes encoded image information / data, and parsing the second unit yields decoded image information / data.

[0140] In some embodiments, the method may further include: determining third indication information of a third unit in the bitstream, the third indication information being used to indicate the type of the third unit; and obtaining one or more first units from the third unit if the third indication information indicates that the third unit is a reconstruction information type.

[0141] In some embodiments, the method may further include: obtaining one or more second units from the third unit when the third indication information indicates that the third unit is of the image information type.

[0142] In some embodiments, the method further includes: parsing the third unit to obtain one or more video parameter sets when the third indication information indicates that the third unit is a parameter set type. That is, the video parameter sets are directly carried in the third unit, enabling rapid extraction and parsing of the video parameter sets.

[0143] If the third unit is a reconstruction information type, then one or more first units are obtained from the third unit; if the third unit is an image information type, then one or more second units are obtained from the third unit; if the third unit is a parameter set type, then one or more video parameter sets are obtained from the third unit.

[0144] For example, the third unit can be a VCM unit, an FCM unit, or a V3C unit.

[0145] By adopting the above technical solution, for the first unit carrying the reconstruction information, different types of first units can be quickly identified through the first indication information. For example, the first unit with the first type can be quickly identified, thereby parsing the first unit and avoiding the waste of decoding resources caused by decoding the first unit that cannot be decoded or does not need to be decoded, thus improving decoding efficiency.

[0146] Furthermore, taking VCM as an example, the decoding method provided in this application embodiment is further illustrated. VCM currently organizes the video bitstream in the form of multiple sub-bitstreams. One sub-bitstream is the kernel video sub-bitstream obtained by the kernel encoder in the VCM encoder, and the other sub-bitstream is the reconstruction information sub-bitstream generated by the VCM encoder based on the reconstruction information obtained from its preprocessing operations on the video. The kernel video sub-bitstream is decoded by the kernel decoder in the VCM decoder to obtain the decoded video, and the reconstruction information sub-bitstream is processed by the VCM decoder to obtain the reconstruction information. The decoded video is then reconstructed based on the reconstruction information to obtain the reconstructed video. This reconstructed video retains key semantic information and can achieve sufficient task accuracy on machine tasks.

[0147] The decoding method provided in this application embodiment, when applied to a VCM, includes: obtaining VCM units from the bitstream; obtaining the type of the VCM unit from the VCM unit; if the VCM unit is a reconstruction information sub-bitstream unit, obtaining reconstruction information NAL units from the VCM unit; when the reconstruction information sub-bitstream unit contains multiple NAL units, obtaining the data packet size of the NAL unit from the VCM unit to obtain each NAL unit separately, and then obtaining reconstruction information from the NAL unit; extracting the type information of the NAL unit from the NAL unit, which identifies whether the NAL unit is reconstruction data of a random access point, reconstruction data of a randomly accessed skipped image, or reconstruction data of a randomly accessed decodeable image; if a random access point occurs... Machine access is performed by searching for the nearest random access point image NAL unit and random access point reconstruction data NAL unit, skipping the random access skip image NAL unit and random access skip image reconstruction data NAL unit. If the VCM unit is a kernel video sub-stream unit, the kernel video NAL unit is obtained from the VCM unit. When the kernel video sub-stream unit contains multiple NAL units, the data packet size of the NAL unit is obtained from the VCM unit to obtain each NAL unit separately, and then the decoded kernel video is obtained from the NAL unit. Based on the reconstruction information, the decoded kernel video is reconstructed to obtain the reconstructed video, which can be used to complete machine tasks to obtain high machine task accuracy.

[0148] Figure 7 is a detailed structural diagram of a VCM stream structure in an embodiment of this application. The VCM video stream includes:

[0149] The Video Parameter Set (VCM) Unit includes the VCM header information (VCM_VPS) and the VCM payload (vcm_parameter_set()).

[0150] Restoration Data Units (VCMs) include VCM header information (VCM_RAP_RSD) and VCM payload (restoration_data_unit()) that support random access, and VCM header information (VCM_RSD) and VCM payload (restoration_data_unit()) that do not support random access.

[0151] The kernel video VCM (Coded Video Data Units) include kernel video VCM header information (VCM_RAP_CVD) and VCM payload (coded_video_data()) that support random access, and kernel video VCM header information (VCM_CVD) and VCM payload (coded_video_data()) that do not support random access.

[0152] The payload of the reconstruction information VCM unit includes video sequence-level reconstruction information (Sequence Restoration Data) and image-level reconstruction information (Picture Restoration Data). The video sequence-level reconstruction information (Sequence Restoration Data) specifically includes the video sequence-level reconstruction information NAL header (VCM_NAL_SRD) and payload (sequence_restoration_data_rbsp()), while the image-level reconstruction information (Picture Restoration Data) specifically includes the image-level reconstruction information NAL header (VCM_NAL_PRD) and payload (picture_restoration_data_rbsp()).

[0153] The workload of kernel video VCM units that support random access includes NAL units that support random access (VVC IRAP NAL Unit) and NAL units that do not support random access (VVC non-IRAP NAL Unit). The workload of kernel video VCM units that do not support random access includes NAL units that do not support random access (VVC non-IRAP NAL Unit).

[0154] Figure 8 is a schematic diagram of another VCM stream structure in an embodiment of this application. The corresponding VCM stream decoding method is described as follows: Obtain VCM units from the stream; obtain the type of the VCM unit from the VCM unit; if the VCM unit is a reconstruction information sub-stream unit, obtain the reconstruction information NAL unit from the VCM unit. When the reconstruction information sub-stream unit contains multiple NAL units, obtain the data packet size of the NAL unit from the VCM unit to obtain each NAL unit separately, and then obtain the reconstruction information from the NAL unit; extract the type information of the NAL unit from the NAL unit. This type information identifies whether the NAL unit is reconstruction data of a random access point, reconstruction data of a randomly accessed skipped image, or reconstruction data of a randomly accessed decodeable image. According to the data, if random access occurs, the nearest random access point image NAL unit and random access point reconstruction data NAL unit are searched, and the random access skip image NAL unit and random access skip image reconstruction data NAL unit are skipped. If the VCM unit is a kernel video sub-stream unit, the kernel video NAL unit is obtained from the VCM unit. When the kernel video sub-stream unit contains multiple NAL units, the data packet size of the NAL unit is obtained from the VCM unit to obtain each NAL unit separately, and then the decoded kernel video is obtained from the NAL unit. According to the reconstruction information, the decoded kernel video is reconstructed to obtain the reconstructed video, which can be used to complete machine tasks to obtain high machine task accuracy.

[0155] For example, the syntax element structure of a VCM unit is as follows:

[0156] Wherein, numBytesInVCMUnit represents the data packet size of the current VCM unit.

[0157] The syntax element structure for VCM unit header information and payload is as follows:

[0158] Among them, vuh_unit_type (third indicator information) indicates the data type of the VCM unit. For example, the data type unit of the VCM unit defines several types such as VPS parameter set (VPS), reconstruction information or reconstruction data (RSD), reconstruction information that supports random access (RSD_RAP), kernel video data (CVD), and kernel video data that supports random access (CVD_RAP).

[0159] `vuh_vps_id` represents the index of a VCM unit of type VPS that is referenced by certain types of VCM units. `vuh_vps_id` specifies the value of `vps_vcm_parameter_set_id` for the effective VCM parameter set. The value of `vuh_vps_id` should be in the range of 0 to 15, but the actual value is between 0 and 4. `vuh_reserved_zero_23bits` and `vuh_reserved_zero_27bits` are reserved bits.

[0160] Here, numBytesInVCMUnitPayload represents the size of the payload data of the current VCM unit. The payload data of the VCM unit is parsed based on the value of numBytesInVCMUnitPayload and the data type of the VCM unit. When vuh_unit_type represents VCM_VPS, the VPS is obtained by parsing the payload data of the VCM unit; when vuh_unit_type represents VCM_RSD or VCM_RAP_RSD, the reconstruction information is obtained by parsing the payload data of the VCM unit; when vuh_unit_type represents VCM_CVD or VCM_RAP_CVD, the kernel video is obtained by parsing the payload data of the VCM unit.

[0161] The payload of a VCM unit is used to carry VCM parameter sets (VPS), reconstructed data (RSD), or encoded video data (CVD). The payload coded_video_data() in the encoded video data packet corresponds to data units (e.g., NAL units as defined in ISO / IEC 23008-2 or ISO / IEC 23090-3) that can be decoded by the appropriate video decoder indicated by the configuration file defined in the VCM parameter set.

[0162] For example, the syntax element structure of the VCM parameter set is as follows:

[0163] In one embodiment, the bitstream also includes the decoding capabilities required by the bitstream, the decoding methods required by the video unit streams, and information on the reconstruction tools that can be used for the reconstruction operation, as shown in the table below. This information can be recorded in `vcm_parameter_set()` using `profile_tier_level()` as follows:

[0164] in:

[0165] ptl_tier_flag and ptl_level_idc specify the level of decoding capability required by the bitstream;

[0166] ptl_profile_codec_group_idc specifies the type of video unit stream, that is, the video decoding method and level that can handle this type of bitstream should be used to decode the video unit stream to obtain the decoded image, such as the several types specified in the table below.

[0167] ptl_profile_restoration_idc specifies the combination of tools that need to be used for the bitstream to be decoded. For example, the bitstream may need to be reconstructed using one or more tools such as spatial sampling, region relocation, temporal sampling, and data bit width offset.

[0168] vps_vcm_parameter_set_id indicates the identifier of the video parameter set.

[0169] vps_log2_max_restoration_data_picture_order_cnt_lsb_minus4 indicates the recording rules for the time information of image-level reconstruction data, which are used to deduce the effective time of the reconstruction data according to the rules based on the syntax elements in the image-level reconstruction information.

[0170] The method for identifying the data packet size of the NAL unit in the reconstruction information sub-stream is as follows:

[0171] Here, rsd_nal_unit_size_minus1 plus 1 represents the data packet size of the reconstructed information NAL unit vcm_nal_unit(), rsd_nal_unit_size_minus1 uses a fixed bit width (e.g., 1 byte), numBytes represents the number of bits to be decoded in the reconstructed information VCM unit, minus rsd_nal_unit_size_minus1+1 represents subtracting the data packet size of vcm_nal_unit(), and subtracting 1 represents subtracting the number of bits occupied by rsd_nal_unit_size_minus1.

[0172] The syntax element structure of the reconstruction information NAL unit in the reconstruction information sub-codestream is as follows:

[0173] The syntax element structure of the NAL unit header information is as follows:

[0174] rbsp_byte[i] is the i-th byte of RBSP. rbsp_byte represents the reconstruction information in the NAL unit. The data type of rbsp_byte is determined by vcm_nal_unit_type.

[0175] The kernel video NAL unit nal_unit() of the kernel video sub-stream maintains the NAL unit syntax element structure of the standard specification that its kernel codec conforms to. For example, when the kernel codec uses the H.266 codec, the kernel video NAL unit nal_unit() should be an H.266 compliant NAL unit.

[0176] vcm_nal_temporal_id indicates the temporal level of the packet.

[0177] vcm_nal_reserved_6bits should be equal to 0 in bitstreams conforming to this version of this document. Other values ​​for vcm_nal_reserved_6bits are reserved by ISO / IEC for future use. The decoder should ignore the value of vcm_nal_reserved_6bits.

[0178] The method for identifying the packet size of the kernel video NAL unit in the kernel video sub-stream is as follows:

[0179] Here, `cvd_nal_unit_size_precision_bytes_minus1` plus 1 indicates the number of bits occupied by `cvd_nal_unit_size` (e.g., in bytes). `cvd_nal_unit_size` uses a variable bit width, and `cvd_reserved_zero_5bits` is for byte-aligned data padding. `cvd_nal_unit_size` records the packet size of the kernel video NAL unit `nal_unit()`. The reason why the packet size of the kernel video NAL unit is not represented by a fixed bit width similar to that of the reconstruction information NAL unit is that the data volume of the reconstruction information NAL unit is usually small, and a fixed bit width is sufficient to identify the packet size. However, the data volume of the kernel video NAL unit may vary greatly due to parameter sets or image encoding data, and a fixed bit width is insufficient to identify the packet size.

[0180] numBytes represents the number of bits to be decoded in the kernel video VCM unit, i.e., the second value. Subtracting cvd_nal_unit_size means subtracting the data packet size of nal_unit(). Subtracting cvd_nal_unit_size_precision_bytes_minus1+1 means subtracting the number of bits occupied by cvd_nal_unit_size.

[0181] For example, the NAL category identifier recorded in the header of the reconstruction information NAL packet includes: reconstruction information includes video sequence-level reconstruction data (SRSD), image-level reconstruction data (PRSD), auxiliary enhancement information (SEI), and reconstruction information end (EOSS).

[0182] The syntax element structure of sequence-level reconstruction information is as follows:

[0183] The syntax element structure of image-level reconstruction information is as follows:

[0184] The syntax element structure for auxiliary enhancement information is as follows:

[0185] The syntax element structure of the bitstream terminator for the reconstructed information NAL is as follows:

[0186] The syntax element structure for byte alignment is as follows:

[0187] The above syntax elements are illustrated below:

[0188] The `vuh_unit_type` indicates the VCM unit type specified in the table below. Values ​​marked as reserved are reserved for future use by ISO / IEC and should not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document should ignore such reserved unit types.

[0189] The `vuh_vps_id` specifies the value of `vps_vcm_parameter_set_id` for the effective VCM parameter set. The value of `vuh_vps_id` should be in the range of 0 to 15.

[0190] When vuh_reserved_zero_23bits appears, its value should be equal to 0 in a bitstream conforming to this version of this document. Other values ​​for vuh_reserved_zero_23bits are reserved by ISO / IEC for future use. The decoder should ignore the value of vuh_reserved_zero_23bits.

[0191] When `vuh_reserved_zero_27bits` appears, its value should be equal to 0 in bitstreams conforming to this version of this document. Other values ​​for `vuh_reserved_zero_27bits` are reserved by ISO / IEC for future use. The decoder should ignore the value of `vuh_reserved_zero_27bits`.

[0192] `coded_video_data(numBytes)` contains a portion of video unit streams of size `numBytes`, presented as an ordered stream of bytes or bits. The positions of the unit boundaries can be determined based on the organization pattern of the video unit stream. The format of this video unit stream is identified by `ptl_profile_codec_group_idc`.

[0193] vps_log2_max_restoration_data_frame_order_cnt_lsb_minus4 indicates the recording rules for the time information of image-level reconstruction data, which are used to deduce the effective time of the reconstruction data according to the rules based on the syntax elements in the image-level reconstruction information.

[0194] The srd_spatial_resampling_enabled_flag indicates whether the decoder can use spatial sampling tools to reconstruct the decoded video, and also indicates whether spatial sampling parameters can be recorded in the reconstructed data.

[0195] The `srd_retargeting_enabled_flag` indicates whether the decoder can use region retargeting tools to reconstruct the decoded video, and also indicates whether region retargeting parameters can be recorded in the reconstructed data.

[0196] The srd_temporal_restoration_enabled_flag indicates whether the decoder can use time-sampling tools to reconstruct the decoded video, and also indicates whether time-sampling parameters can be recorded in the reconstructed data.

[0197] The `srd_bit_depth_shift_enabled_flag` indicates whether the decoder can use data bit width shift tools to reconstruct the decoded video, and also indicates whether the relevant parameters of the data bit width shift can be recorded in the reconstructed data.

[0198] rbsp_byte[i] is the i-th byte of the RBSP. The RBSP is specified as an ordered sequence of bytes as follows:

[0199] An RBSP contains a sequence of data bits (SODB) as follows:

[0200] If SODB is empty (i.e., its length is zero bits), RBSP is also empty.

[0201] Otherwise, the RBSP contains the SODB as follows:

[0202] 1) The first byte of the RBSP contains the first (most important, leftmost) eight bits of the SODB; the next byte of the RBSP contains the next eight bits of the SODB, and so on, until there are fewer than eight bits remaining in the SODB.

[0203] 2) The rbsp_trailing_bits() syntax structure exists after SODB as follows:

[0204] a) The first (most important, leftmost) bit of the last RBSP byte contains the remaining bits of SODB (if any).

[0205] b) The next bit consists of a single bit equal to 1 (i.e., rbsp_stop_one_bit).

[0206] c) When rbsp_stop_one_bit is not the last bit of the byte alignment byte, there are one or more bits equal to 0 (i.e., instances of rbsp_alignment_zero_bit) that cause byte alignment.

[0207] Syntax structures with these RBSP attributes are indicated in the syntax table with the suffix "_rbsp". These structures are carried within the VCM NAL package as the contents of rbsp_byte[i] data bytes. The association between RBSP syntax structures and the VCM NAL package is as specified in [Error! Reference source not found].

[0208] vcm_nal_forbidden_zero_bit should be equal to 0.

[0209] vcm_nal_unit_type (i.e., the first indication information) is used to indicate the type of the reconstruction information NAL unit. In one embodiment, the type of the RBSP data structure contained in the VCM NAL unit can be specified as specified in the table below.

[0210] VCM NAL units with the identifier VCM_NAL_UNSPEC and nal_unit_type have unspecified semantics and should not affect the decoding procedures specified in this document. VCM NAL unit types with the identifier VCM_NAL_UNSPEC may be used depending on the application. No decoding procedures are specified for these vcm_nal_unit_type values ​​in this document. Because different applications may use these VCM NAL unit types for different purposes, special care is required in encoder design for generating VCM NAL units with these vcm_nal_unit_type values ​​and in decoder design for interpreting the contents of VCM NAL units with these vcm_nal_unit_type values. No management is defined for these values ​​in this document. These vcm_nal_unit_type values ​​may only apply to the context of use cases where “conflicts” (i.e., different definitions of the meaning of the contents of VCM NAL units with the same vcm_nal_unit_type value) are not important, impossible, or managed—for example, in applications or transport specifications that control the definition or management of the bitstream distribution. For purposes other than determining the amount of data in the bitstream decoding unit, the decoder should ignore (i.e., remove and discard from the bitstream) the contents of all VCM NAL units that use the reserved value of nal_unit_type.

[0211] In another embodiment, the reconstruction information NAL packet also includes more finely categorized image-level reconstruction data NAL packets. For example, the header of the reconstruction data NAL packet records a NAL category identifier, which identifies the type of video image to be decoded in the video NAL packet corresponding to the reconstruction data recorded in that NAL packet. A specific syntax example and decoding operation description are as follows:

[0212] The type VCM_NALU_RSD_IDR indicates that the video data NAL packet corresponding to the reconstructed data NAL packet is of type IDR (Instantaneous Decoding Refresh). When decoding and playback need to start from a certain point in time in the video stream, the decoding operation scans the video stream and finds the video data NAL packet of type IDR closest to that point in time. Decoding the video data begins from this point. In addition, the nearest reconstructed data NAL packet of type VCM_NALU_RSD_IDR is found, and the reconstructed data in it is parsed to reconstruct the previously decoded video data.

[0213] 1) The type VCM_NALU_RSD_IDR indicates that the video data NAL packet corresponding to the reconstructed data NAL packet is of type IDR (Instantaneous Decoding Refresh). When decoding and playback need to start from a certain point in time in the video bitstream, the decoding operation scans the video bitstream and finds the video data NAL packet of type IDR closest to that point in time. Decoding the video data starts from this point. In addition, the nearest reconstructed data NAL packet of type VCM_NALU_RSD_IDR is found, and the reconstructed data in it is parsed to reconstruct the previously decoded video data.

[0214] 2) Type VCM_NALU_RSD_CRA is similar in function to type VCM_NALU_RSD_IDR. The video data NAL packet corresponding to this type of reconstructed data NAL packet is of type CRA (Clean Random Access). The decoding operation of this type of reconstructed data NAL packet is similar to the aforementioned operation. In another implementation, type VCM_NALU_RSD_IDR and type VCM_NALU_RSD_CRA can be combined into a single type VCM_NALU_RSD_IRAP, indicating that the video data NAL packet corresponding to this type of reconstructed data NAL packet is of type IRAP (Intra Random Access Point).

[0215] 3) Type VCM_NALU_RSD_GDR is similar in function to type VCM_NALU_RSD_IDR. The video data NAL packet corresponding to this type of reconstructed data NAL packet is of type GDR (gradual decoding refresh). The decoding operation of this type of reconstructed data NAL packet is similar to the aforementioned operation. In another implementation, type VCM_NALU_RSD_GDR and type VCM_NALU_RSD_IDR can be combined into a single type VCM_NALU_RSD_IRAP, indicating that the video data NAL packet corresponding to this type of reconstructed data NAL packet is of type IRAP (Intra Random Access Point).

[0216] 4) The type VCM_NALU_RSD_RASL indicates that the video data NAL packet corresponding to this reconstructed data NAL packet is of type RASL (Random Access Skipped Leading (Picture)). When the decoding operation receives a RASL type video data NAL packet after the first IRAP type video data NAL packet, since the RASL type video data NAL packet depends on the video data NAL packets preceding the IRAP type video data NAL packet, and these video data NAL packets have not been received, the decoding operation skips the decoding of the RASL type video data NAL packet. In addition, the decoding operation also skips the decoding of VCM_NALU_RSD_RASL type reconstructed data NAL packets after the first VCM_NALU_RSD_IRAP type reconstructed data NAL packet.

[0217] 5) The type VCM_NALU_RSD_RADL indicates that the video data NAL packet corresponding to this reconstructed data NAL packet is of type RADL (Random Access Decodable Leading (Picture)). When the decoding operation receives a RADL type video data NAL packet after the first IRAP type video data NAL packet, since the RADL type video data NAL packet only depends on the IRAP type video data NAL packet and the video data NAL packets that follow, it can be correctly decoded. Therefore, the decoding operation decodes the RADL type video data NAL packet. In addition, the decoding operation finds and decodes the VCM_NALU_RSD_RADL type reconstructed data NAL packet after the first VCM_NALU_RSD_IRAP type reconstructed data NAL packet. The reconstructed image of the RADL type video data packet should be output or displayed before the reconstructed image of the IRAP type video data packet.

[0218] 6) The type VCM_NALU_RSD_TRAIL indicates that the video data NAL packet corresponding to the reconstructed data NAL packet is of type TRAIL (tail (image)). The decoding operation can correctly decode the TRAIL type video data NAL packet, and also decode the corresponding reconstructed data NAL, thus completing the reconstruction of the decoded video data.

[0219] 7) The type VCM_NALU_RSD_TSA indicates that the video data NAL packet corresponding to this reconstructed data NAL packet is of type TSA (Temporal Sub-layer Access) or STSA (Step-wise Temporal Sub-layer Access). This type of video data NAL packet allows the decoder to start from a specific time point or image and decode higher temporal layers of images or data step by step according to a certain order and rules, so as to achieve more refined processing of video content or presentation of different quality levels. The decoding operation can correctly decode TSA or STSA type video data NAL packets, and also decode the corresponding reconstructed data NAL, completing the reconstruction of the decoded video data.

[0220] It should be noted that the decoding method provided in this application embodiment can also be used for FCM. The FCM unit includes the following types: global vision model parameter set (VMPS), reconstruction data (RSD), and coded video data (CVD), etc.

[0221] Furthermore, taking FCM as an example, the decoding method provided in this application embodiment will be further illustrated. The method for identifying the data packet size of the FCM reconstruction information NAL unit in the reconstruction information sub-bitstream, the syntax element structure of the reconstruction information NAL unit in the reconstruction information sub-bitstream, the syntax element structure of the reconstruction information NAL unit header information, and the method for identifying the data packet size of the kernel video NAL unit in the kernel video sub-bitstream are described in the above embodiments.

[0222] The decoding method applied to FCM includes: obtaining FCM units from the bitstream; obtaining the type of the FCM unit from the FCM unit; if the FCM unit is a reconstruction information sub-bitstream unit, then obtaining the reconstruction information NAL unit from the FCM unit; when the reconstruction information sub-bitstream unit contains multiple NAL units, obtaining the data packet size of the NAL unit from the FCM unit to obtain each NAL unit separately, and then obtaining the reconstruction information from the NAL unit; extracting the type information of the NAL unit from the NAL unit, which identifies whether the NAL unit is reconstruction data of a random access point, reconstruction data of a randomly accessed skipped image, or reconstruction data of a randomly accessed decodeable image; if a random access occurs, searching for the nearest random access time. The system includes NAL units for access point images and NAL units for random access point reconstruction data, as well as NAL units for skipped random access images and NAL units for reconstructed random access images. If the FCM unit is a kernel video sub-stream unit, the kernel video NAL unit is obtained from the FCM unit. When the kernel video sub-stream unit contains multiple NAL units, the data packet size of the NAL unit is obtained from the FCM unit to obtain each NAL unit separately. Then, the decoded image features are obtained from the NAL units. Based on the reconstruction information, the decoded image frequency features are reconstructed to obtain reconstructed image features. These reconstructed features contain multiple layers of sub-features at different scales, which can be used in the network after feature extraction in machine intelligence tasks to obtain task analysis results.

[0223] The syntax element structure of the FCM unit is as follows:

[0224] Here, numBytesInFCMUnit represents the data packet size of the current VCM unit.

[0225] The syntax element structure for FCM header information and payload is as follows:

[0226] Among them, fuh_unit_type (third indicator information) indicates the data type of the FCM unit. For example, the data type unit of the FCM unit defines several types such as VMPS parameter set, RSD (feature reconstruction information), RSD_RAP (feature reconstruction information that supports random access), CVD (kernel video feature data), and CVD_RAP (kernel video feature data that supports random access).

[0227] The data type indicated by fuh_unit_type can be shown in the table below.

[0228] The FCM unit of type FCM_VMPS contains a set of visual model parameters, which can be recorded in the bitstream in the form shown in the table below.

[0229] Here, `vmps_vision_model_parameter_set_id` represents the VMPS number. Multiple different VMPSs are allowed in the bitstream, each using a different number, and can be referenced by video feature reconstruction at different times. `img_wid` and `img_hei` represent the original resolution of the source video for the feature data in the video bitstream. `scaled_img_wid` and `scaled_img_hei` represent the resolution of the image after the feature extraction part of the source video has been scaled by the intelligent task network. This is because the intelligent task network performs preprocessing before analyzing the source video. `total_numer_of_input` represents the total number of frames of video features in the bitstream.

[0230] In one implementation, the decoding method obtains various types of FCM NAL units from the units of the reconstructed data in the bitstream, including: examples of FCM NAL unit types are shown in the table below.

[0231] The `fcm_nal_unit_type` indicates the type of FCM NAL unit: `FCM_NAL_FSPS` represents a sequence-level NAL unit, `FCM_NAL_FPPS` represents an image-level NAL unit, `FCM_NAL_EOSS` represents the end unit of the reconstructed data, `FCM_NAL_SEI` represents auxiliary enhancement information of the reconstructed data, `FCM_NAL_RSV` represents a reserved unit type, and `FCM_NAL_UNSPEC` represents a reserved unit type that can be defined by the user. This application's embodiments classify image-level NAL units into more detailed types, including IDR, CRA, GDR, TSA, RADL, RASL, and TRAIL types. The system layer or decoder can quickly locate and parse the reconstructed data NAL packet corresponding to the video data NAL packet based on these types.

[0232] In one implementation, the decoding method obtains sequence-level parameters and image-level parameters for reconstructing features from the NAL units of the reconstructed data in the bitstream. An example of the syntax structure of the sequence-level parameters is shown in the table below.

[0233] The syntax structure of image-level parameters is shown in the table below.

[0234] The sequence-level feature parameter set feat_seq_parameter_set_rbsp contains:

[0235] 1. Parameter set number: fsps_feat_seq_parameter_set_id;

[0236] 2. The number of layers for reconstructing the feature data is num_ori_feat_layers, and the width, height, and number of channels of each feature layer are num_ori_feat_wid[i], num_ori_feat_hei[i], and num_ori_feat_chan[i].

[0237] 3. The width (fused_feat_wid) and height (fused_feat_hei) of the decoded fusion feature obtained by the kernel decoder;

[0238] 4. The kernel decoder's switch inner_decoding_bypass_flag and type information inner_coding_idx. inner_coding_idx can indicate the type of the kernel decoder, such as VVC, HEVC, or AVC.

[0239] 5. Switches for tools that can be used during the decoding process, such as the dequantization switch dequant_bypass_flag, the feature repacking switch unpacking_bypass_flag, the feature reconstruction switch feat_restoration_bypass_flag, the temporal upsampling switch temporal_upsampling_enable_flag, the reconstructed feature adjustment switch restored_feat_refine_flag, and the fused feature adjustment switch fused_feat_refine_flag, etc.

[0240] 6. Feature reconstruction information: feat_restoration_info contains information about the network model used for feature reconstruction. One network model is the default feature reconstruction network, indexed by feat_restoration_weight_idx. Features reconstructed using this network also need to be cropped in width and height according to pad_size_min. The other network model is obtained by the decoding method according to fcm_decoder_info_sei_id. It can be obtained from the encoding end or from certain links.

[0241] 7. `restored_feat_refine_refresh_period` and `fused_feat_refine_refresh_period` record the periods of reconstruction feature adjustment and fusion feature adjustment, respectively.

[0242] The image-level parameter set feature_pic_parameter_set_rbsp contains:

[0243] 1. The parameter set number fpps_feat_pic_parameter_set_id and the sequence-level parameter set number fpps_feat_seq_parameter_set_id referenced by this parameter set;

[0244] 2. The parameters restored_feat_std and restored_feat_mean used for reconstructed feature adjustment, and the parameters fused_feat_std and fused_feat_mean used for fused feature adjustment.

[0245] In one embodiment, after obtaining the decoded fusion features based on the above information, the decoding method uses the feature reconstruction network specified by feat_restoration_info to decompose the fusion features and obtain reconstructed features. These reconstructed features contain multiple layers of sub-features at different scales, which can be used in networks after feature extraction in machine intelligence tasks to obtain task analysis results.

[0246] The decoding method provided in this application embodiment indicates different types of first units by using first indication information for the first unit carrying reconstruction information. This can provide the system layer or decoder with the ability to quickly filter the first unit, avoid wasting decoding resources on first units that cannot be decoded or do not need to be decoded, and improve decoding efficiency.

[0247] In another embodiment of this application, referring to FIG9, a flowchart of an encoding method provided by an embodiment of this application is shown. As shown in FIG9, the method may include:

[0248] S901: Perform preprocessing on the original image to determine the reconstruction information of the first image / image features;

[0249] S902: Encode the reconstructed information to obtain the first unit;

[0250] S903: Determine the first indication information of the first unit, the first indication information being used to indicate the type of the first unit;

[0251] S904: Add first instruction information to the first unit to generate a bitstream.

[0252] In some embodiments, the first indication information is used to indicate at least two types of first units.

[0253] In some embodiments, the first indication information is used to indicate at least two types of image-level first units.

[0254] In some embodiments, at least two types include: a first type for indicating that the first unit is parsed if a first condition is met; and a second type for indicating that the first unit is not parsed if a second condition is met.

[0255] In some embodiments, the first type includes at least one of the following: random access point type, resolvable type, and hierarchical access point type.

[0256] In some embodiments, the random access point type is used to indicate the parsing of the first unit in response to a random access event.

[0257] In some embodiments, the random access point type includes at least two random access point subtypes.

[0258] For example, the random access point type includes at least one of the following: a first subtype, further used to indicate other first units after the first unit in the parsed bitstream; a second subtype, further used to indicate a portion of the first units after the first unit in the unparsed bitstream; and a third subtype, further used to indicate one or more first units in the parsed bitstream associated with the first unit.

[0259] For example, the first subtype can be an Instant Decode Refresh (IDR) type, the second subtype can be a Clean Random Access (CRA) type, and the third subtype can be a Progressive Decode Refresh (GDR) type.

[0260] In some embodiments, a resolvable type is used to indicate that the current first unit is correctly resolvable in response to a random access event, and that the current first unit is located after the first unit of the random access point type.

[0261] In some embodiments, the resolvable type includes at least one of the following: a fourth subtype, used to indicate parsing the first unit, wherein the reconstructed image / reconstructed image feature corresponding to the first unit is output or displayed before the reconstructed image / reconstructed image feature corresponding to the first unit of the random access point type; a fifth subtype, used to indicate parsing the first unit, wherein the reconstructed image / reconstructed image feature corresponding to the first unit is output or displayed after the reconstructed image / reconstructed image feature corresponding to the first unit of the random access point type.

[0262] For example, the fourth subtype can be a Random Access Decodeable Front (RADL) type, and the fifth subtype can be a Trail type.

[0263] In some embodiments, the tiered access point type is used to indicate that, in response to a tier switching event, the first unit is resolved to enter the target tier from the first unit.

[0264] For example, the hierarchical access point type includes a sixth subtype for indicating that the first unit is an access point at the target hierarchical level. For instance, the sixth subtype could be a Stepwise Temporal Sublayer Access Point (STSA).

[0265] For example, the hierarchical access point type includes: a seventh subtype, used to indicate that the first unit is the final access point of the target hierarchy; and an eighth subtype, used to indicate that the first unit is an intermediate access point of the target hierarchy. For example, the seventh subtype can be a Time Domain Sublayer Access Point (TSA), and the eighth subtype can be a Stepwise Time Domain Sublayer Access Point (STSA).

[0266] In some embodiments, the second type includes at least the Random Access Skip Before (RASL) type.

[0267] In some embodiments, the first indication information is further used to indicate a sequence-level first unit. That is, at least two types of first units include one or more sequence-level first units and one or more image-level first units, wherein the sequence-level first unit carries sequence-level reconstruction information, and the image-level first unit may carry reconstruction information of the entire image or carry sub-image-level reconstruction information.

[0268] In some embodiments, the first indication information is located in the header information of the first unit. Accordingly, the first indication information is added to the header information of the first unit to generate the bitstream.

[0269] In some embodiments, the first indication information is located in the header information of the associated unit of the first unit. The associated unit may include the upper-level unit to which the first unit belongs. The associated unit may include a second unit, which carries image information of the first image / image features. Accordingly, the first indication information is added to the header information of the associated unit of the first unit to generate a bitstream. In some embodiments, the first unit is a Network Abstraction Layer (NAL) unit carrying reconstruction information.

[0270] In some embodiments, determining the first indication information of the first unit includes: determining the first indication information of the first unit based on the second indication information of the second unit, wherein the second indication information is used to indicate the type of the second unit and the second unit carries image information of the first image / image features.

[0271] In some embodiments, the method may further include: preprocessing the original image to determine a first image / image feature; encoding the first image / image feature to obtain a second unit; and adding second indication information to the second unit based on the type of the first image / image feature to generate a bitstream, wherein the second indication information is used to indicate the type of the second unit.

[0272] In some embodiments, the method may further include: generating a third unit based on one or more second units; adding third indication information of image information type to the third unit to generate a bitstream.

[0273] In some embodiments, the method may further include: if the second indication information indicates that the second unit is of a third type, instructing the decoder to parse the second unit and determine the first image / image features, wherein the third type is consistent with or equivalent to the first type of the first unit; if the second indication information indicates that the second unit is of a fourth type, instructing the decoder not to parse the second unit, wherein the fourth type is consistent with or equivalent to the second type of the first unit.

[0274] In some embodiments, the second unit is a Network Abstraction Layer (NAL) unit that carries image information.

[0275] In some embodiments, the method may further include: generating a third unit based on one or more first units; adding third indication information to the third unit to indicate the type of reconstruction information, thereby generating a bitstream.

[0276] For example, the encoding method provided in this application embodiment, when applied to VCM, includes:

[0277] The encoder performs preprocessing operations on the input video to obtain the processed video, and then uses the kernel encoder to encode and compress it to obtain the kernel video sub-stream, which is composed of multiple kernel NAL units.

[0278] The encoder encapsulates the kernel video sub-stream into at least one VCM unit, identifies the type of the VCM unit as the kernel video sub-stream type, and records the unit data packet size of each kernel NAL unit in the VCM unit to distinguish the boundaries of each kernel NAL unit.

[0279] The encoder obtains reconstruction information from the preprocessing operation for the reconstruction processing operation of the decoder. This reconstruction information can guide the reconstruction of the kernel decoded video. The encoder encapsulates the reconstruction information into a reconstruction information bitstream, which contains at least one reconstruction information NAL unit. The reconstruction information bitstream is then encapsulated into at least one VCM unit. The type of the VCM unit is identified as the reconstruction information sub-bitstream type. In the VCM unit, the size of the unit data packet for each reconstruction information NAL unit is recorded to distinguish the boundaries of each reconstruction information NAL unit. The encoder obtains the type information from the kernel video NAL unit corresponding to the reconstruction information and puts this information into the header information of the reconstruction information NAL unit so that the type information and hierarchical information of the reconstruction NAL unit and the kernel video NAL unit of the same image are consistent or equivalent.

[0280] The encoder encapsulates information such as the type of kernel decoder that the decoder should use into a VCM unit, which is of type VPS.

[0281] The encoder organizes VCM units of type VPS, VCM units of type reconstructed information sub-stream, VCM units of type kernel video sub-stream, and other possible types of VCM units to obtain a VCM video stream.

[0282] The above operations do not necessarily have to be performed in sequence. They can also be performed alternately to encode and encapsulate VCM units of the reconstruction information sub-stream type and VCM units of the kernel video sub-stream type, in order to achieve low-latency encoding.

[0283] For example, the encoding method provided in this application embodiment, when applied to VCM, includes:

[0284] The encoder performs preprocessing operations on the input video features to obtain processed video features, and then uses a kernel encoder to encode and compress them to obtain a video feature sub-bitstream, which is composed of multiple kernel NAL units.

[0285] The encoder encapsulates the video feature sub-stream into at least one FCM unit, identifies the type of the FCM unit as the kernel video sub-stream type, and records the unit data packet size of each kernel NAL unit in the FCM unit to distinguish the boundaries of each kernel NAL unit.

[0286] The encoder obtains reconstruction information from the preprocessing operation for the reconstruction processing operation of the decoder. This reconstruction information can guide the reconstruction of video features in the kernel decoding. The encoder encapsulates the reconstruction information into a reconstruction information bitstream, which contains at least one reconstruction information NAL unit. The reconstruction information bitstream is then encapsulated into at least one FCM unit. The type of the FCM unit is identified as the reconstruction information sub-bitstream type. The unit data packet size of each reconstruction information NAL unit is recorded in the FCM unit to distinguish the boundaries of each reconstruction information NAL unit. The encoder obtains the type information from the kernel video NAL unit corresponding to the reconstruction information and puts this information into the header information of the reconstruction information NAL unit so that the type information and hierarchical information of the reconstruction NAL unit and the kernel video NAL unit of the same image are consistent or equivalent.

[0287] The encoder encapsulates information such as the type of kernel decoder that the decoder should use into an FCM unit, which is of type VPS.

[0288] The encoder organizes FCM units of VMPS type, FCM units of reconstructed information sub-stream type, FCM units of kernel video sub-stream type, and other possible types of FCM units to obtain FCM video feature stream;

[0289] The above operations do not necessarily have to be performed in sequence. The encoding and encapsulation of FCM units of the reconstruction information sub-stream type and FCM units of the kernel video sub-stream type can be performed alternately to achieve low-latency encoding.

[0290] By adopting the above technical solution, for the first unit carrying reconstruction information, different types of first units can be indicated by the first indication information, which can provide the system layer or decoder with the ability to quickly filter the first unit, avoid the waste of decoding resources by the first unit that cannot be decoded or does not need to be decoded, and improve decoding efficiency.

[0291] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, referring to FIG10, a schematic diagram of the composition structure of an encoder provided in an embodiment of this application is shown. As shown in FIG10, the encoder 1000 may include a first processing unit 1001 and an encoding unit 1002; wherein,

[0292] The first processing unit is configured to preprocess the original image to determine the reconstruction information of the first image / image features;

[0293] The encoding unit is configured to encode the reconstructed information to obtain the first unit;

[0294] The encoding unit is further configured to determine first indication information of the first unit, the first indication information being used to indicate the type of the first unit;

[0295] The encoding unit is also configured to add first indication information to the first unit to generate a bitstream.

[0296] Understandably, each functional unit of the encoder also performs the encoding method of any of the foregoing embodiments.

[0297] Understandably, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular one. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional module.

[0298] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0299] Therefore, embodiments of this application provide a computer-readable storage medium applied to an encoder 1000, the computer-readable storage medium storing a computer program, which, when executed by a first processor, implements the method of any of the foregoing embodiments.

[0300] This application provides a computer-readable storage medium that stores a bitstream generated by an encoding method such as described above.

[0301] Based on the composition of the encoder 1000 and the computer-readable storage medium, referring to Figure 11, a schematic diagram of the specific hardware structure of the encoder 1000 provided in this application embodiment is shown. As shown in Figure 11, the encoder 1000 may include: a first communication interface 1101, a first memory 1102, and a first processor 1103; the various components are coupled together through a first bus system 1104. It is understood that the first bus system 1104 is used to realize the connection and communication between these components. In addition to a data bus, the first bus system 1104 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the first bus system 1104 in Figure 11.

[0302] The first communication interface 1101 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0303] The first memory 1102 is used to store computer programs that can run on the first processor 1103;

[0304] The first processor 1103 is used to execute the following when running computer programs:

[0305] Preprocess the original image to determine the reconstruction information of the first image / image features;

[0306] The reconstructed information is encoded to obtain the first unit;

[0307] Determine the first indication information of the first unit, the first indication information being used to indicate the type of the first unit;

[0308] Add first instruction information to the first unit to generate the bitstream.

[0309] It is understood that the first memory 1102 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The first memory 1102 of the system and method described in this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0310] The first processor 1103 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the first processor 1103 or by instructions in software form. The first processor 1103 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the first memory 1102. The first processor 1103 reads the information in the first memory 1102 and completes the steps of the above method in conjunction with its hardware.

[0311] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions of this application, or combinations thereof. For software implementation, the technology of this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions of this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0312] Alternatively, as another embodiment, the first processor 1103 is also configured to execute any of the methods in the foregoing embodiments when running a computer program.

[0313] This embodiment provides an encoder in which, for a first unit carrying reconstructed information, different types of first units are quickly identified through first indication information. For example, a first unit with a first type is quickly identified, thereby parsing the first unit and avoiding the waste of decoding resources caused by decoding a first unit that cannot be decoded or does not need to be decoded, thus improving decoding efficiency.

[0314] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, referring to FIG12, a schematic diagram of the composition structure of a decoder 1200 provided in an embodiment of this application is shown. As shown in FIG12, the decoder 1200 may include: a decoding unit 1201 and a second processing unit 1202; wherein,

[0315] The decoding unit is configured to determine first indication information of a first unit in the bitstream, wherein the first indication information is used to indicate the type of the first unit;

[0316] The decoding unit is also configured to parse the first unit and determine the reconstruction information of the first image / image features when the first indication information indicates that the first unit is of the first type;

[0317] The second processing unit is further configured to perform reconstruction processing on the first image / image features based on the reconstruction information to determine the reconstructed image / reconstructed image features.

[0318] Understandably, each functional unit of the decoder also performs the decoding method of any of the aforementioned embodiments.

[0319] Based on the composition of the decoder 1200 and the computer-readable storage medium, Figure 13 illustrates a schematic diagram of the specific hardware structure of the decoder 1200 provided in this embodiment. As shown in Figure 13, the decoder 1200 may include: a second communication interface 1301, a second memory 1302, and a second processor 1303; the various components are coupled together through a second bus system 1304. It is understood that the second bus system 1304 is used to realize the connection and communication between these components. In addition to a data bus, the second bus system 1304 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the second bus system 1304 in Figure 13.

[0320] The second communication interface 1301 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0321] The second memory 1302 is used to store computer programs that can run on the second processor 1303;

[0322] The second processor 1303 is used to execute the following when running computer programs:

[0323] Determine the first indication information of the first unit in the bitstream, the first indication information being used to indicate the type of the first unit;

[0324] If the first indication information indicates that the first unit is of the first type, the first unit is parsed to determine the reconstruction information of the first image / image features;

[0325] Based on the reconstruction information, the first image / image features are reconstructed to determine the reconstructed image / reconstructed image features.

[0326] Alternatively, as another embodiment, the second processor 1303 is also configured to execute any of the methods in the foregoing embodiments when running a computer program.

[0327] It is understood that the second memory 1302 has similar hardware functions to the first memory 1102, and the second processor 1303 has similar hardware functions to the first processor 1103; these will not be described in detail here.

[0328] This embodiment provides a decoder in which, for a first unit carrying reconstructed information, different types of first units are quickly identified through a first indication information. For example, a first unit with a first type is quickly identified, thereby parsing the first unit and avoiding the waste of decoding resources caused by decoding a first unit that cannot be decoded or does not need to be decoded, thus improving decoding efficiency.

[0329] In another embodiment of this application, referring to FIG14, a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application is shown. As shown in FIG14, the encoding / decoding system 1400 may include an encoder 1401 and a decoder 1402.

[0330] In this embodiment, encoder 1401 can be any of the encoders in the foregoing embodiments, and decoder 1402 can be any of the decoders in the foregoing embodiments.

[0331] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0332] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0333] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined to obtain new product embodiments without conflict. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined to obtain new method embodiments or device embodiments without conflict.

[0334] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims. Industrial applicability

[0335] This application provides an encoding / decoding method, encoder, decoder, and storage medium. The method includes: determining first indication information of a first unit in a bitstream, the first indication information indicating the type of the first unit; if the first indication information indicates that the first unit is of a first type, parsing the first unit to determine reconstruction information of a first image / image feature; and performing reconstruction processing on the first image / image feature based on the reconstruction information to determine a reconstructed image / reconstructed image feature. Thus, for a first unit carrying reconstruction information, the first indication information enables rapid identification of different types of first units, for example, rapidly identifying a first unit possessing a first type, thereby parsing the first unit and avoiding the waste of decoding resources caused by decoding first units that are undecodingable or unnecessary to decode, thus improving decoding efficiency.

Claims

1. A decoding method applied to a decoder, wherein, include: Determine the first indication information of the first unit in the bitstream, wherein the first indication information is used to indicate the type of the first unit; If the first indication information indicates that the first unit is of the first type, the first unit is parsed to determine the reconstruction information of the first image / image features; Based on the reconstruction information, the first image / image features are reconstructed to determine the reconstructed image / reconstructed image features.

2. The method of claim 1, wherein, The first indication information is used to indicate at least two types of first units.

3. The method of claim 2, wherein, The first indication information is used to indicate at least two types of image-level first units.

4. The method of claim 2 or 3, wherein, The at least two types include: A first type, used to indicate that the first unit is parsed if a first condition is met; The second type is used to indicate that the first unit is not parsed if the second condition is met.

5. The method according to any one of claims 1 to 4, wherein, The first type includes at least one of the following: random access point type, resolvable type, and hierarchical access point type.

6. The method of claim 5, wherein, The random access point type is used to indicate the parsing of the first unit in response to a random access event.

7. The method of claim 5, wherein, The random access point type includes at least two random access point subtypes.

8. The method of claim 5, wherein, The random access point type includes at least one of the following: instant decode refresh type, clean random access type, and progressive decode refresh type.

9. The method of claim 5, wherein, The resolvable type is used to indicate that in response to a random access event, and the current first unit is located after the first unit of the random access point type, it can be correctly resolved.

10. The method of claim 5, wherein, The resolvable type includes at least one of the following: random access decodable front type, tail type.

11. The method of any one of claims 5-10, wherein, The method further includes: Get random access events; In response to a random access event, based on the random access time and the first indication information, a first unit of the random access point type is determined to parse the first unit of the random access point type.

12. The method of claim 5, wherein, The tiered access point type is used to indicate that, in response to a tier switching event, the first unit is parsed to enter the target tier from the first unit.

13. The method of claim 5, wherein, The hierarchical access point type includes at least one of the following: time-domain sub-layer access point type, and progressive time-domain sub-layer access point type.

14. The method of any one of claims 5, 12, or 13, wherein, The method further includes: Get the level switching event; In response to a tier switching event, based on the first indication information, a first unit of the tier access point type is determined to parse the first unit of the tier access point type.

15. The method of claim 5, wherein, The second type includes at least the random access skip-precedence type.

16. The method of claim 3, wherein, The first indication information is also used to indicate the first unit of the sequence level.

17. The method of any one of claims 1-16, wherein, The method further includes: If the first indication information indicates that the first unit is of the second type, skip the step of parsing the first unit.

18. The method of any one of claims 1-17, wherein, The first indication information is located in the header information of the first unit.

19. The method of any one of claims 1-17, wherein, The first indication information is located in the header information of the associated unit of the first unit.

20. The method of claim 19, wherein, The associated unit includes the upper-level unit to which the first unit belongs.

21. The method of claim 20, wherein, The associated unit includes a second unit, which carries image information of the first image / image features.

22. The method of any one of claims 1-21, wherein, The first unit is the Network Abstraction Layer (NAL) unit that carries reconstruction information.

23. The method of claim 1, wherein, The method includes: Determine the second indication information of the second unit in the bitstream, the second indication information being used to indicate the type of the second unit; When the second indication information indicates that the second unit is of the third type, the second unit is parsed to determine the first image / image features, and the third type is consistent with or equivalent to the first type. If the second indication information indicates that the second unit is of the fourth type, the second unit is not parsed, and the fourth type is consistent with or equivalent to the second type of the first unit.

24. The method of claim 23, wherein, The second unit is the Network Abstraction Layer (NAL) unit that carries image information.

25. The method of claim 1, wherein, The method further includes: The third indication information of the third unit in the bitstream is determined, and the third indication information is used to indicate the type of the third unit; If the third indication information indicates that the third unit is of the reconstruction information type, one or more of the first units are obtained from the third unit.

26. The method of claim 25, wherein, The method further includes: If the third indication information indicates that the third unit is of image information type, one or more second units are obtained from the third unit.

27. An encoding method applied to an encoder, wherein, include: Preprocess the original image to determine the reconstruction information of the first image / image features; The reconstructed information is encoded to obtain the first unit; Determine the first indication information of the first unit, wherein the first indication information is used to indicate the type of the first unit; Add the first indication information to the first unit to generate a bitstream.

28. The method of claim 27, wherein, The first indication information is used to indicate at least two types of first units.

29. The method of claim 28, wherein, The first indication information is used to indicate at least two types of image-level first units.

30. The method of claim 28 or 29, wherein, The at least two types include: A first type, used to indicate that the first unit is parsed if a first condition is met; The second type is used to indicate that the first unit is not parsed if the second condition is met.

31. The method of any one of claims 27-30, wherein, The first type includes at least one of the following: random access point type, resolvable type, and hierarchical access point type.

32. The method of claim 31, wherein, The random access point type is used to indicate the parsing of the first unit in response to a random access event.

33. The method of claim 31, wherein, The random access point type includes at least two random access point subtypes.

34. The method of claim 31, wherein, The random access point type includes at least one of the following: instant decode refresh type, clean random access type, and progressive decode refresh type.

35. The method of claim 31, wherein, The resolvable type is used to indicate that in response to a random access event, and the current first unit is located after the first unit of the random access point type, it can be correctly resolved.

36. The method of claim 31, wherein, The resolvable type includes at least one of the following: random access decodable front type, tail type.

37. The method of claim 31, wherein, The tiered access point type is used to indicate that, in response to a tier switching event, the first unit is parsed to enter the target tier from the first unit.

38. The method of claim 31, wherein, The hierarchical access point type includes at least one of the following: time-domain sub-layer access point type, and progressive time-domain sub-layer access point type.

39. The method of claim 31, wherein, The second type includes at least the random access skip-precedence type.

40. The method of claim 29, wherein, The first indication information is also used to indicate the first unit of the sequence level.

41. The method of any one of claims 27-40, wherein, The first indication information is located in the header information of the first unit.

42. The method of any one of claims 27-40, wherein, The first indication information is located in the header information of the associated unit of the first unit.

43. The method of claim 42, wherein, The associated unit includes the upper-level unit to which the first unit belongs.

44. The method of claim 43, wherein, The associated unit includes a second unit, which carries image information of the first image / image features.

45. The method of any one of claims 27-44, wherein, The first unit is the Network Abstraction Layer (NAL) unit that carries reconstruction information.

46. The method of claim 27, wherein, The determination of the first indication information of the first unit includes: Based on the second indication information of the second unit, the first indication information of the first unit is determined. The second indication information is used to indicate the type of the second unit, and the second unit carries image information of the first image / image features.

47. The method according to claim 46, wherein, Preprocess the original image to determine the first image / image features; The first image / image features are encoded to obtain the second unit; Based on the type of the first image / image feature, the second indication information is added to the second unit to generate a bitstream, wherein the second indication information is used to indicate the type of the second unit.

48. The method of claim 47, wherein, The method further includes: Generate a third unit based on one or more second units; Add third indication information of image information type to the third unit to generate a bitstream.

49. The method of claim 46, wherein, The method includes: when the second indication information indicates that the second unit is of a third type, instructing the decoder to parse the second unit and determine the first image / image features, wherein the third type is consistent with or equivalent to the first type of the first unit; If the second indication information indicates that the second unit is of the fourth type, the decoder is instructed not to parse the second unit, wherein the fourth type is consistent with or equivalent to the second type of the first unit.

50. The method of claim 49, wherein, The second unit is the Network Abstraction Layer (NAL) unit that carries image information.

51. The method of claim 27, wherein, The method further includes: Generate a third unit based on one or more first units; Add third indication information to the third unit to indicate the type of reconstruction information in order to generate a bitstream.

52. An encoder, comprising a first processing unit and an encoding unit; wherein: The first processing unit is configured to preprocess the original image to determine the reconstruction information of the first image / image features; The encoding unit is configured to encode the reconstructed information to obtain a first unit; The encoding unit is further configured to determine first indication information of the first unit, the first indication information being used to indicate the type of the first unit; The encoding unit is further configured to add the first indication information to the first unit to generate a bitstream.

53. An encoder, comprising a first memory and a first processor; wherein: The first memory is used to store computer programs that can run on the first processor; The first processor is configured to perform the method as described in any one of claims 27 to 51 when running the computer program.

54. A decoder, comprising a decoding unit and a second processing unit; wherein: The decoding unit is configured to determine first indication information of a first unit in the bitstream, wherein the first indication information is used to indicate the type of the first unit; The decoding unit is further configured to parse the first unit and determine the reconstruction information of the first image / image features when the first indication information indicates that the first unit is of the first type; The second processing unit is further configured to perform reconstruction processing on the first image / image features based on the reconstruction information to determine the reconstructed image / reconstructed image features.

55. A decoder, comprising a second memory and a second processor; wherein: The second memory is used to store computer programs that can run on the second processor; The second processor is configured to perform the method as described in any one of claims 1 to 26 when running the computer program.

56. A computer readable storage medium, wherein, The computer-readable storage medium stores the bitstream generated by the encoding method as described in any one of claims 27 to 51.

57. A computer readable storage medium, wherein, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 26, or the method as described in any one of claims 27 to 51.