Methods, apparatuses, electronic devices, and storage media for video encoding and video decoding

JP7698746B2Active Publication Date: 2025-06-25ZTE CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023578781
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-30
Filing Date
2022-06-29
Publication Date
2025-06-25
Estimated Expiration
2042-06-29

AI Technical Summary

Benefits of technology

を有する。この装置は、ソフトウェア及び/又はハードウェアにより実現されてもよく、具体的には、画像取得モジュール610とビデオ符号化モジュール620とを含む。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698746000015
    Figure 0007698746000015
  • Figure 0007698746000016
    Figure 0007698746000016
  • Figure 0007698746000017
    Figure 0007698746000017
Patent Text Reader

Abstract

The present application provides a method, apparatus, electronic device and storage medium for video encoding and decoding, the method for video encoding including the steps of obtaining (110) a video image, which is an image of at least one frame of a video, and performing weighted predictive coding on the video image to generate an image code stream, the weighted predictive coding utilizing at least one set of weighted prediction identification information and parameters (120).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application is filed based on a Chinese patent application with an application number of 202110875430.8 and an application date of July 30, 2021, claims the priority of the Chinese patent application, and incorporates all the contents of the Chinese patent application by reference into this application.

[0002] This application relates to the technical field of image processing, and particularly relates to methods, devices, electronic devices, and storage media for video encoding and video decoding.

Background Art

[0003] With the rapid development of digital media technology, video has become an important transmission medium. Video-based communication, entertainment, and learning have gradually integrated into the daily lives of the public, becoming increasingly familiar and close to ordinary users. In currently common video formats, in many cases, a gradation effect is added to the start or end of the beginning part of one theme content, and the in / out effect of the screen is used to give viewers a more natural and comfortable viewing experience.

[0004] To reduce the pressure on network bandwidth caused by data transmission, video encoding / decoding technology has become an important research content in the multimedia field. When performing video encoding, the inter-frame prediction technology can effectively remove the redundancy of temporal domain data and significantly reduce the coding rate of video transmission. However, in a video sequence, when luminance change scenes such as fade-in, fade-out, lens aperture modulation, and overall or local light source changes appear, it is difficult to achieve an ideal data compression effect with conventional inter-frame motion estimation and motion compensation. Consequently, when there are local content blocks with luminance changes in actual encoding, the determination result of the optimization model usually is to adopt intra-frame prediction encoding for all, resulting in a significant reduction in video encoding efficiency. Therefore, in order to improve the encoding effect in the luminance change scenes of content, weighted prediction technology may be used in video encoding. When a luminance change is detected, the encoding side needs to determine the luminance change weight and offset of the current image with respect to the reference image and generate the corresponding weighted prediction frame through a luminance compensation operation.

[0005] Currently, in the H.264 / AVC standard, weighted prediction technology has already been proposed. When applying weighted prediction, there are two modes: explicit weighted prediction and implicit weighted prediction. In the case of implicit weighted prediction, all model parameters are fixed. That is, it is agreed that the same weighted prediction parameters are adopted on the encoding side and the decoding side, and there is no need to transmit the parameters from the encoding side, thus reducing the pressure of coded stream transmission and improving the transmission efficiency. However, since the weighted prediction parameters in the implicit mode are fixed, when applied to inter-frame unidirectional prediction, due to the change distance between the current frame and the reference frame, the prediction effect of the fixed weight is not desirable. That is, the explicit mode is applied to both inter-frame unidirectional prediction and bidirectional prediction, while the implicit mode is only applied to inter-frame bidirectional prediction. In the case of the explicit mode, the encoding side needs to determine the weighted prediction parameters and label them in the data header information. The decoding side needs to read the corresponding weighted prediction parameters from the coded stream to decode the image normally. Three prediction parameters, namely weight, offset, and log weight denominator, are related to weighted prediction. Here, in order to avoid floating-point operations, on the encoding side, it is necessary to expand the weight, that is, introduce the log weight denominator, and on the decoding side, it is necessary to shrink it by the corresponding multiple.

[0006] Taking the H.266 / VVC standard as an example, a series of parameter information of weighted prediction may be included in the picture header or the slice header. Note that each luminance component or chrominance component of the reference image has independent weighted prediction parameters. When a complex luminance change scene appears in the video content, different weighted prediction parameters can be arranged for different slice regions in the image. However, there are many restrictions on the slice division method, such as the slice division during encoding affecting the data transmission of the network abstraction layer. It is difficult for the weighted prediction technology adapted to the slice layer to flexibly cope with various luminance change situations. On the other hand, in the current standardization technology, only one set of weighted prediction parameters (that is, one set of weights and offsets that cooperate with each other) is set for each reference image. When the current entire image has a completely uniform luminance change form, only the previous closest reference image and its weighted prediction parameters are selected, and a good prediction effect can be achieved. However, in order to appropriately consider other forms of content luminance change scenes, when performing inter-frame encoding of the video, as shown in FIG. 1, a plurality of different reference images may be selected according to the conventional standard, and weights and offsets suitable for the reference images may be artificially arranged for the content of different regions in the current image. Especially in the case of media content with a large amount of data such as super high-definition video and panoramic video, considering that usually only a small number of decoded reference images can be stored in the buffer area on the decoding side, in actual applications, the scheme shown in FIG. 1 can only be applied to partial luminance change scenes.

[0007] In short, the video encoding scheme in the conventional video encoding standardization technology has great limitations in its feasibility and flexibility when encoding a complex graphics luminance change scene in the video. Summary of the Invention Problems to be Solved by the Invention

[0008] The main objective of the embodiments of this application is to propose a video encoding method, apparatus, electronic device, and storage medium that can achieve flexible video encoding in a complex graphics brightness change scene, improve the efficiency of video encoding, and reduce the impact of graphics brightness changes on encoding efficiency.

Means for Solving the Problem

[0009] The embodiments of this application provide a method for video encoding. The method includes: obtaining a video image, which is an image of at least one frame of a video; and performing weighted prediction encoding on the video image to generate an image code stream, where the weighted prediction encoding uses at least one set of weighted prediction identification information and parameters.

[0010] The embodiments of this application provide a method for video decoding. The method includes: obtaining an image code stream; analyzing the weighted prediction identification information and parameters in the image code stream; and decoding the image code stream according to the weighted prediction identification information and parameters to generate a reconstructed image.

[0011] The embodiments of this application further provide an apparatus for video encoding. The apparatus includes: an image acquisition module configured to obtain a video image, which is an image of at least one frame of a video; and a video encoding module configured to perform weighted prediction encoding on the video image using at least one set of weighted prediction identification information and parameters to generate an image code stream.

[0012] The embodiments of this application further provide an apparatus for video decoding. The apparatus includes: a code stream acquisition module configured to obtain an image code stream and analyze the weighted prediction identification information and parameters in the image code stream; and an image reconstruction module configured to decode the image code stream according to the weighted prediction identification information and parameters to generate a reconstructed image.

[0013] Embodiments of the present application further provide an electronic device. The electronic device includes one or more processors and a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the embodiments of the present application.

[0014] Embodiments of the present application further provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method described in any one of the embodiments of the present application is implemented.

[0015] According to the embodiments of the present application, by obtaining a video image which is an image of at least one frame in a video, and performing weighted prediction coding on the video image by using at least one set of weighted prediction identification information and parameters, an image code stream is generated, thereby realizing flexible coding of the video image, improving the efficiency of video coding, and reducing the influence of the luminance change of the video image on the coding efficiency.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0017] The embodiments described herein are only used for interpreting the present application and are not used for limiting the present application.

[0018] In the following description, suffixes such as "module", "component", or "unit" for indicating elements are only used to facilitate the description of the present invention and do not have a specific meaning in themselves. Therefore, "module", "component", or "unit" may be used interchangeably.

[0019] FIG. 2 is a flowchart of a video encoding method provided by an embodiment of the present application. The embodiment of the present application may be applied to video encoding in a scene where the luminance changes. The method may be executed by a video encoding device. The video encoding device may be implemented by software and / or hardware, and is usually incorporated into a terminal device. Referring to FIG. 2, the method provided by the embodiment of the present application specifically includes the following steps.

[0020] In step 110, a video image, which is an image of at least one frame of a video, is obtained.

[0021] Here, the video image may be video data that needs to be transmitted. The video image may be data of one frame in a video data sequence, or data of one frame corresponding to a certain time.

[0022] In the embodiment of the present application, the video data may be processed, and the image data of one or more frames therein may be extracted as the video data for video encoding.

[0023] In step 120, weighted prediction encoding using at least one set of weighted prediction identification information and parameters is performed on the video image to generate an image code stream.

[0024] Among them, in the scenes within a video, since the intensity of light gradually changes over time or there is a shadow effect in the video of the same scene, even if the similarity of the background between frames is very high, there is a large difference in brightness and a luminance change exists between the images of adjacent frames. The current frame is equivalent to the one obtained by multiplying the entire previous frame by a weight and adding an offset, and performing video encoding based on the previous frame. Such a process of video encoding using weights and offsets may be referred to as weighted prediction encoding. The process of weighted prediction encoding is mainly related to three prediction parameters, namely weight, offset, and log weight denominator. Among them, the log weight denominator can avoid floating-point operations in the encoding process and expand the weight. The weighted prediction identification information may be the identification information of the parameters used for weighted prediction. The parameters may be the specific parameters of the weighted prediction identification information and may include at least one of weight, offset, and log weight denominator. The image code stream may be the data generated after encoding the video image and may be used for transmission between terminal devices.

[0025] In the embodiments of the present application, weighted prediction encoding may be performed on the video image, and in the process of weighted prediction encoding, one set or multiple sets of weighted prediction identification information and parameters may be utilized. For example, for video images of different frames, different weighted prediction identification information and parameters may be utilized, or for video images of different regions within the same frame, different weighted prediction identification information and parameters may be utilized. In addition, in the process of performing weighted prediction encoding on the video image, different weighted prediction identification information and parameters may be selected according to the luminance change of the video image for video encoding.

[0026] According to this embodiment, by obtaining a video image which is an image of at least one frame in a video, and performing weighted prediction coding on the video image by using at least one set of weighted prediction identification information and parameters to generate an image code stream, flexible coding of the video image can be realized, the efficiency of video coding can be improved, and the influence of the luminance change of the video image on the coding efficiency can be reduced.

[0027] FIG. 3 is a flowchart of another video coding method provided by an embodiment of the present application. The embodiment of the present application is embodied based on the above embodiment. Referring to FIG. 3, the method provided by the embodiment of the present application specifically includes the following steps.

[0028] In step 210, a video image which is an image of at least one frame of a video is obtained.

[0029] In step 220, based on the comparison result between the video image and the reference image, the luminance change situation is determined.

[0030] Here, the reference image may be image data located before or after the currently processed video image in the video data sequence. The reference image may be used for the motion estimation operation of the currently processed video image. The number of reference images may be one frame or multiple frames. The luminance change situation may be the luminance change situation of the video image with respect to the reference image. Specifically, the luminance change situation may be determined by the change of the pixel values of the video image with respect to the reference image.

[0031] In this embodiment, the video image may be compared with the reference image. The comparison method may include calculating the difference value of the pixel values at the corresponding positions, or calculating the difference value of the average pixel values of the video image and the reference image, etc. Based on the comparison result between the video image and the reference image, the luminance change situation may be determined. The luminance change situation may include gradually becoming brighter, gradually becoming darker, not changing, randomly changing, etc.

[0032] Furthermore, based on the above embodiments, the luminance change situation includes at least one of the average luminance change value of the image and the luminance change value of the pixel point.

[0033] Specifically, the luminance change situation of the video image with respect to the reference image may be determined by the average luminance change value of the image and / or the luminance change value of the pixel point. Here, the average luminance change value of the image may refer to the average value change of the luminance value of the current image with respect to the reference image, and the luminance change value of the pixel point may be the change of the luminance value of each pixel point in the video image with respect to the luminance value of the pixel point at the corresponding position in the reference image. Note that the luminance change situation may also be other changes in statistical characteristics of luminance values, such as luminance variance, luminance standard deviation, etc.

[0034] In step 230, based on the luminance change situation of the video image, weighted prediction coding is performed on the video image using at least one set of weighted prediction identification information and parameters.

[0035] In this embodiment, multiple sets of weighted prediction identification information and parameters may be preset in advance, and according to the luminance change situation, the corresponding prediction identification information and parameters may be selected to perform weighted prediction coding on the video image. Note that weighted prediction coding may also be performed on the video image according to the specific content of the luminance change situation. For example, for the luminance changes of video images in different frames, different weighted prediction identification information and parameters may be selected to perform weighted prediction coding, or for different regions in the video image, different weighted prediction identification information and parameters may be selected to perform weighted prediction coding.

[0036] In step 240, the weighted prediction identification information and parameters are written into the image code stream.

[0037] Specifically, after weighted prediction coding for a video image, an image code stream is generated, and in order to facilitate the video decoding process in a subsequent process, the weighted prediction identification information and parameters used in the coding process may be written into the image code stream.

[0038] According to this embodiment, a video image is acquired, the luminance change situation is determined based on the comparison result between the video image and a reference image, weighted prediction coding is performed on the video image based on the luminance change situation, and by writing the weighted prediction identification information and parameters used in the weighted prediction coding into the image code stream, flexible coding of the video image can be realized, the efficiency of video coding can be improved, and the influence of the luminance change of the video image on the coding efficiency can be reduced.

[0039] Furthermore, based on the above embodiment, the step of performing weighted prediction coding on the video image based on the luminance change situation is as follows: If it is a luminance change situation where the image luminance of the entire frame is uniform, the step of performing weighted prediction coding on the video image; If it is a luminance change situation where the image luminance of a partial region is uniform, at least one of the step of determining to perform weighted prediction coding on the image of each partial region in the video image respectively is included.

[0040] Here, that the image luminance of the entire frame is uniform may mean that the luminance change of the entire frame of the image frame of the video image is the same. That the image luminance of a partial region is uniform may mean that there are a plurality of regions in the video image and the luminance changes of each region are not the same.

[0041] In this embodiment, when the luminance change situation is such that the image luminance of the entire frame is uniform, weighted prediction coding may be performed on the entire frame of the video image. When the luminance change situation is such that the image luminance of a partial region is uniform, since various luminance change situations exist in the video image, weighted prediction coding may be performed according to each image region. Note that when the luminance changes of each image region are not the same, the weighted prediction identification information and parameters used in the weighted prediction coding process may be different.

[0042] Furthermore, based on the above embodiment, in the video image, there may be at least one set of the weighted prediction identification information and parameters for the image of the reference image or a partial region of the reference image.

[0043] In this embodiment, the weighted prediction identification information and parameters used in the weighted prediction coding process of the video image are information determined with respect to the reference image. In the video image, there is one set or a plurality of sets of weighted prediction identification information and parameters based on the reference image. When the reference image of the video image is a plurality of frames, in the video image, there may be one set or a plurality of sets of weighted prediction identification information and parameters for each frame of the reference image respectively. Each set of weighted prediction identification information and parameters may have a correlation with the corresponding video image and / or reference image.

[0044] FIG. 4 is a flowchart of another video coding method provided by an embodiment of the present application. The embodiment of the present application is embodied based on the above embodiment. Referring to FIG. 4, the method provided by the embodiment of the present application specifically includes the following steps.

[0045] In step 310, a video image which is an image of at least one frame of the video is acquired.

[0046] In step 320, weighted prediction coding is performed on the video image according to a pre-trained neural network model.

[0047] Here, the neural network model may perform weighted prediction encoding processing on the video image, and may determine the weighted prediction identification information and parameters used for the video image. The neural network model may be generated by training using image samples with weighted prediction identification information and parameters attached. The neural network model may determine the weighted prediction identification information and parameters of the video image, or may directly determine the image code stream of the video image.

[0048] In this embodiment, the video image may be directly or indirectly input into the neural network model, and weighted prediction encoding for the video image may be realized by the neural network model. Note that the input layer of the neural network model may receive the video image or the features of the video image, the neural network model may generate the weighted prediction identification information and parameters used for the weighted prediction encoding of the video image, or may directly perform weighted prediction encoding on the video image.

[0049] In step 330, write the weighted prediction identification information and parameters into the image code stream.

[0050] According to this embodiment, by acquiring a video image, performing weighted prediction encoding on the video image using a pre-trained neural network model, and writing the weighted prediction identification information and parameters used for the weighted prediction encoding into the image code stream generated by the encoding of the video image, flexible encoding of the video image can be realized, the efficiency of video encoding can be improved, and the influence of the luminance change of the video image on the encoding efficiency can be reduced.

[0051] Furthermore, based on the above embodiment, the weighted prediction identification information further includes the neural network model structure and the neural network model parameters.

[0052] Specifically, the weighted prediction identification information may further include a neural network model structure and neural network model parameters. The neural network model structure may be information reflecting the structure of the neural network, such as functions used in fully connected layers, activation functions, loss functions, etc. The neural network model parameters may be specific possible values of the parameters of the neural network model, such as network weight values, the number of hidden layers, etc.

[0053] Furthermore, based on the above embodiments, one set of the weighted prediction identification information and parameters used for the weighted prediction coding corresponds to the video image of one frame or an image of at least one partial region of the video image.

[0054] In this embodiment, one set or multiple sets of weighted prediction identification information and parameters may be used in the weighted prediction coding process. Each set of weighted prediction identification information and parameters corresponds to the video image of one frame or the image of one partial region within the video image in the weighted prediction coding process, and the image of the partial region may be a part of the video image, such as a slice image or a sub-image, etc.

[0055] Furthermore, based on the above embodiments, the standard of the image of the partial region includes at least one of a slice, a tile, a subpicture, a coding tree unit, and a coding unit.

[0056] Specifically, when performing weighted prediction coding on a video image in the form of an image of a partial region, the image of the partial region may be one or more of a slice, a tile, a subpicture, a coding tree unit, and a coding unit.

[0057] Furthermore, based on the above embodiments, the weighted prediction identification information and the parameters are included in at least one parameter set among a sequence layer parameter set, an image layer parameter set, a slice layer parameter set, supplementary enhancement information, video user usability information, image header information, slice header information, network abstraction layer unit header information, a coding tree unit, and a coding unit.

[0058] Specifically, the weighted prediction identification information and the parameters may be written into the image code stream, and the identification information and the parameters are included in all or part of the parameter sets of a sequence layer parameter set, an image layer parameter set, a slice layer parameter set, supplementary enhancement information, video user usability information, image header information, slice header information, and network abstraction layer unit header information, or may be included in a coding tree unit or a coding unit as a new information unit.

[0059] Furthermore, based on the above embodiments, the weighted prediction identification information and the parameters include at least one of reference image index information, weighted prediction enable control information, region-adaptive weighted prediction enable control information, and weighted prediction parameters.

[0060] In this embodiment, the weighted prediction identification information may be reference image index information used to determine a reference image used for luminance change, weighted prediction enable control information used to determine whether to perform weighted prediction coding, region-adaptive weighted prediction enable control information used to determine whether to perform region-weighted prediction coding on a video image, and weighted prediction parameters used in the weighted prediction coding process, which may include weights, offsets, logarithmic weight denominators, etc.

[0061] Furthermore, based on the above embodiments, the image code stream includes a transmission stream or a media file.

[0062] Specifically, the image code stream may be a transmission stream or a media file.

[0063] In one embodiment, FIG. 5 is a schematic diagram of a video encoding method provided by an embodiment of the present application. Referring to FIG. 5, the input of the encoding process is an image included in the video, and the output is an image code stream, or a transmission data stream or a media file including the image code stream. The process of video encoding may include the following steps.

[0064] In step 101, an image is read. The image may be the data of one frame in the video sequence, or the data of one frame corresponding to a certain time.

[0065] In step 102, the luminance change of the image with respect to a certain reference image is detected.

[0066] Here, the reference image may refer to an image before the image in the time line in the video sequence, or an image after the current image in the time line. The reference image is used to perform a motion estimation operation on the current image. The reference image may have one frame or multiple frames.

[0067] The luminance change may refer to the average value change of the luminance values of the current image with respect to the reference image, or the change of the luminance value of each pixel point of the current image with respect to the luminance value of the corresponding position pixel point of the reference image, or the change of other luminance value statistical characteristics.

[0068] When the reference image has multiple frames, the current image needs to detect the luminance change for each frame of the reference image.

[0069] In step 103, based on the detection result of the luminance change in step 102, it is determined whether it is necessary to apply the weighted prediction operation. Here, the basis for the determination includes the trade-off of the weighted prediction operation between the improvement of the coding efficiency and the increase of the coding complexity for a specific luminance change scenario.

[0070] In step 104, when it is necessary to apply the weighted prediction operation, weighted prediction coding is executed, and the index information of the reference image mentioned above, one set or a plurality of sets of weighted prediction parameters, and / or weighted prediction parameter index information are recorded.

[0071] Based on the luminance change statistical information in step 102, the coding unit records the region characteristics of the luminance change. The region characteristics of the luminance change include the characteristic that the luminance change is kept uniform over the entire frame region of the image, or the characteristic that the luminance change is kept uniform in a partial region of the image. The partial region may be one or more slices, one or more tiles, one or more subpictures, one or more coding tree units (CTUs), or one or more coding units (CUs).

[0072] When the luminance change is kept uniform over the entire frame region of the image, it may be regarded that the luminance change of the current image with respect to the reference image is unified, and if not, it may be regarded as not unified.

[0073] When the luminance change situation of the current image is unified, weighted prediction coding is executed, and the index information of the reference image mentioned above and one set of weighted prediction parameters including weights and offsets are recorded. When there are various luminance change situations in each region of the current image, weighted prediction coding is performed according to the image region, and the above-mentioned reference image index information, multiple sets of weighted prediction parameters, and specific weighted prediction parameter index information used for each image region are recorded. Here, one image region may correspond to one set of weighted prediction parameters, or multiple image regions may correspond to one set of weighted prediction parameters.

[0074] In step 105, the weighted prediction identification information and parameters are written into the code stream. The identification information and parameters may be included in all or some of the parameter sets of the sequence layer parameter set, the picture layer parameter set, the slice layer parameter set, the supplementary enhancement information, the video usability information, the picture header information, the slice header information, and the network abstraction layer unit header information, or may be included in the coding tree unit or the coding unit as a new information unit.

[0075] In step 106, if the application of the weighted prediction operation is not required, image coding is directly performed based on the conventional method.

[0076] In step 107, an image code stream, or a transmission stream or a media file including the image code stream is output.

[0077] In another embodiment, FIG. 6 is a schematic diagram of a video coding method provided by an embodiment of the present application. Referring to FIG. 6, according to this embodiment, deep learning technology is applied to region-adaptive weighted prediction coding, and the process of video coding may include the following steps.

[0078] In step 201, an image is read. The image may be the data of one frame in the video sequence, or may be the data of one frame corresponding to a certain time.

[0079] In step 202, it is determined whether to use a weighted prediction scheme based on deep learning technology.

[0080] In step 203, when using deep learning technology, weighted prediction coding of the image is performed using the neural network model and parameters generated by training.

[0081] In some examples, in order to reduce the computational complexity, based on the learning and training of the neural network, it may be simplified to fixed weighted prediction parameters. In the application, for the content of different regions in the image, appropriate weighted prediction parameters are selected to perform weighted prediction coding. The weighted prediction parameters may include weights and offsets. Here, a single image may use all or some of the weights and offsets within a set of numerical values.

[0082] In step 204, when performing weighted prediction coding of the image based on the neural network model, record the weighted prediction parameters required for the deep learning scheme, including but not limited to the neural network model structure and parameters, a set of extracted weighted prediction parameter numerical values, and the parameter index information used for each image region.

[0083] In step 205, when not using deep learning technology, perform region-adaptive weighted prediction coding based on a conventional operation scheme, for example, based on the operation process in Example 1.

[0084] In step 206, write the weighted prediction identification information and parameters into the code stream. The identification information and parameters are included in all or part of the parameter sets of the sequence layer parameter set, the picture layer parameter set, the slice layer parameter set, the supplementary enhancement information, the video user data information, the picture header information, the slice header information, and the network abstraction layer unit header information, or may be included in the coding tree unit or the coding unit as a new information unit.

[0085] In step 207, output the picture code stream, or the transmission stream or media file including the picture code stream.

[0086] FIG. 7 is a flowchart of a video decoding method provided by an embodiment of the present application. The embodiment of the present application may be applied to video decoding in a scene where the luminance changes. This method may be executed by a video decoding device. The video decoding device may be implemented by software and / or hardware methods and is usually incorporated into a terminal device. Referring to FIG. 7, the method provided by the embodiment of the present application specifically includes the following steps.

[0087] In step 410, obtain the picture code stream and analyze the weighted prediction identification information and parameters in the picture code stream.

[0088] In this embodiment, receive the transmitted picture code stream and extract the weighted prediction identification information and parameters from the picture code stream. Here, the picture code stream may include one set or multiple sets of weighted prediction identification information and parameters.

[0089] In step 420, decode the picture code stream according to the weighted prediction identification information and parameters to generate a reconstructed picture.

[0090] Specifically, weighted prediction decoding may be performed on the image code stream using the obtained weighted prediction identification information and parameters, and the image code stream may be processed into a reconstructed image. Here, the reconstructed image may be an image generated based on the transmission code stream.

[0091] According to this embodiment, by obtaining an image code stream, obtaining weighted prediction identification information and parameters in the image code stream, and performing processing on the image code stream according to the obtained weighted prediction identification information and parameters to generate a reconstructed image, dynamic decoding of the image code stream can be realized, the efficiency of video decoding can be improved, and the influence of changes in image luminance on video decoding efficiency can be reduced.

[0092] FIG. 8 is a flowchart of another video decoding method provided by an embodiment of the present application. The embodiment of the present application is a specific implementation based on the above embodiment. Referring to FIG. 8, the method provided by the embodiment of the present application specifically includes the following steps.

[0093] In step 510, an image code stream is obtained, and weighted prediction identification information and parameters in the image code stream are analyzed.

[0094] In step 520, according to a pre-trained neural network model, weighted prediction decoding is performed on the image code stream using the weighted prediction identification information and parameters to generate a reconstructed image.

[0095] Here, the neural network model may be a deep learning model used for decoding the image code stream. This deep learning model may be trained and generated by a sample code stream and a sample image. The neural network model can perform weighted prediction decoding on the image code stream.

[0096] In this embodiment, an image code stream, weighted prediction identification information, and parameters may be input into a pre-trained neural network model, and the neural network model may perform weighted prediction decoding processing on the image code stream to process the image code stream into a reconstructed image.

[0097] According to this embodiment, an image code stream is obtained, weighted prediction identification information and parameters in the image code stream are obtained, and using a pre-trained neural network model, based on the weighted prediction identification information and parameters, by processing the image code stream into a reconstructed image, dynamic decoding of the image code stream can be realized, the efficiency of video decoding can be improved, and the influence on the video decoding efficiency due to changes in image luminance can be reduced.

[0098] Furthermore, based on the above embodiment, the number of weighted prediction identification information and parameters in the image code stream is at least one set.

[0099] Specifically, the image code stream may be information generated by performing weighted prediction coding on a video image. Depending on the weighted prediction coding method, the image code stream may have one set or multiple sets of weighted prediction identification information and parameters. For example, on the encoding side, if weighted prediction coding is performed on different regions within the video image respectively, the image code stream may have multiple sets of weighted prediction identification information and parameters.

[0100] Furthermore, based on the above embodiment, the weighted prediction identification information and the parameters are included in at least one parameter set among a sequence layer parameter set, an image layer parameter set, a slice layer parameter set, supplementary enhancement information, video user quality information, image header information, slice header information, network abstraction layer unit header information, coding tree unit, and coding unit.

[0101] Furthermore, based on the above embodiments, the weighted prediction identification information and the parameters include at least one of reference image index information, weighted prediction enable control information, region adaptation weighted prediction enable control information, and weighted prediction parameters.

[0102] Furthermore, based on the above embodiments, the image code stream includes a transmission stream or a media file.

[0103] Furthermore, based on the above embodiments, the weighted prediction identification information and parameters further include a neural network model structure and neural network model parameters.

[0104] In one embodiment, FIG. 9 is a schematic diagram of a video decoding method provided by an embodiment of the present application. Referring to FIG. 9, the input of the decoding process is an image code stream, or a transmission data stream or a media file including the image code stream, the output is an image constituting a video, and the decoding process of the video image may include the following steps.

[0105] In step 301, the code stream is read.

[0106] In step 302, the code stream is analyzed to obtain weighted prediction identification information.

[0107] The decoding unit analyzes the sequence layer parameter set, the picture layer parameter set, and / or the slice layer parameter set to obtain weighted prediction identification information. Among them, the sequence layer parameter set includes a sequence parameter set (Sequence Parameter Set, SPS), the picture layer parameter set includes a picture parameter set (Picture Parameter Set, PPS) and an adaptation parameter set (Adaptation Parameter Set, APS), and the slice layer parameter set includes APS. The weighted prediction identification information in the sequence layer parameter set may be referenced by the picture layer parameter set and the slice layer parameter set, and the weighted prediction identification information in the picture layer parameter set may be referenced by the slice layer parameter set. The weighted prediction identification information includes, but is not limited to, whether to use multiple weighted predictions of unidirectional weighting prediction, bidirectional weighting prediction, and / or region adaptation for the sequence and / or picture indicated by the current parameter set.

[0108] Here, the determination of whether to adopt the method of multiple weighted predictions of region adaptation includes, but is not limited to, analyzing the binary identifier or the number of sets of weighted prediction parameters (whether it is 1 or more).

[0109] In step 303, based on the weighted prediction identification information, it is determined whether to apply weighted prediction decoding to the current picture.

[0110] In step 304, if it is determined to apply weighted prediction decoding to the current picture, the weighted prediction parameter information is obtained.

[0111] The decoding unit obtains weighted prediction parameter information from the parameter set and / or data header information according to the instruction of the identification information. Here, the parameter set includes SPS, PPS, and APS, and the data header information includes the image header PH and the slice header (SH). The weighted prediction parameter information includes, but is not limited to, whether weighted prediction parameters (weights and offsets) are set for each reference image in the reference image list, the number of sets of weighted prediction parameters set for each reference image, and the specific values of the weighted prediction parameters for each set.

[0112] In step 305, weighted prediction decoding is performed on the current image based on the weighted prediction identification information and parameters.

[0113] Based on the weighted prediction identification information, the decoding unit may perform unified-form weighted prediction decoding on the current complete image, or perform differentiated weighted prediction decoding on each local content within the image.

[0114] In step 306, if it is determined not to apply weighted prediction decoding to the current image, image decoding is performed directly based on the conventional method.

[0115] In step 307, a reconstructed image is generated. Here, the reconstructed image may be used for display or directly saved.

[0116] In another embodiment, FIG. 10 is a schematic diagram of another video decoding method provided according to an embodiment of the present application. Referring to FIG. 10, according to this embodiment, deep learning technology is applied to region-adaptive weighted prediction coding, and the video decoding method may specifically include the following steps.

[0117] In step 401, the code stream is read.

[0118] In step 402, analyze the code stream to obtain weight prediction identification information based on deep learning.

[0119] The decoding unit analyzes the sequence layer parameter set, the picture layer parameter set, and / or the slice layer parameter set to obtain weight prediction identification information based on deep learning. Among them, the sequence layer parameter set includes the sequence parameter set (Sequence Parameter Set, SPS), the picture layer parameter set includes the picture parameter set (Picture Parameter Set, PPS) and the adaptation parameter set (Adaptation Parameter Set, APS), and the slice layer parameter set includes the APS. The weight prediction identification information in the sequence layer parameter set may be referenced by the picture layer parameter set and the slice layer parameter set, and the weight prediction identification information in the picture layer parameter set may be referenced by the slice layer parameter set. The weight prediction identification information based on the deep learning includes, but is not limited to, whether to use weight prediction based on deep learning for the sequence and / or picture indicated by the current parameter set.

[0120] In step 403, determine whether to apply weight prediction decoding based on deep learning to the current picture based on the weight prediction identification information based on deep learning.

[0121] In step 404, if it is determined to apply weight prediction decoding based on deep learning to the current picture, obtain the weight prediction parameters required for the deep learning scheme.

[0122] The decoding unit acquires weighted prediction parameter information from the parameter set and / or data header information according to the instruction of the identification information. Here, the parameter set includes SPS, PPS, and APS, and the data header information includes the image header PH and the slice header (SH). The weighted prediction parameter information includes, but is not limited to, reference image index information, neural network model structure and parameters, all or some of the weighted prediction parameter values, and weighted prediction parameter index information used for the content of each region of the current image.

[0123] In step 405, based on the weighted prediction identification information and parameters, weighted prediction decoding based on deep learning is performed on the current image.

[0124] In step 406, if it is determined not to perform weighted prediction decoding based on deep learning on the current image, image decoding is directly performed based on the conventional method.

[0125] In step 407, a reconstructed image is generated. Here, the reconstructed image may be used for display or directly saved.

[0126] In one embodiment, the identification information of the region-adaptive weighted prediction parameters included in the sequence layer parameter set SPS in the code stream is given in the embodiment. The identification information in the SPS may be referenced by the PPS and APS. In Table 1, the syntax and semantics are defined as follows.

[0127] The sps_weighted_pred_flag is enable control information for applying the weighted prediction technique in the sequence layer to unidirectional prediction slices (P slices). When sps_weighted_pred_flag is equal to 1, it indicates that weighted prediction may be applied to the P slices indicated by the SPS. Conversely, when sps_weighted_pred_flag is equal to 0, it indicates that weighted prediction is not applied to the P slices indicated by the SPS.

[0128] The sps_weighted_bipred_flag is enable control information for applying the weighted prediction technique in the sequence layer to bidirectional prediction slices (B slices). When sps_weighted_bipred_flag is equal to 1, it indicates that explicit weighted prediction may be applied to the B slices indicated by the SPS. Conversely, when sps_weighted_pred_flag is equal to 0, it indicates that explicit weighted prediction is not applied to the B slices indicated by the SPS.

[0129] The sps_wp_multi_weights_flag is enable control information regarding that a single reference picture in the sequence layer has multiple sets of weighted prediction parameters. When sps_wp_multi_weights_flag is equal to 1, it indicates that a single reference picture of the picture indicated by the SPS may have multiple sets of weighted prediction parameters. Conversely, when sps_wp_multi_weights_flag is equal to 0, it indicates that a single reference picture of the picture indicated by the SPS has only a single set of weighted prediction parameters.

[0130] [Table 1]

[0131] In one embodiment, identification information of region-adaptive weighted prediction parameters included in an image layer parameter set PPS in a code stream is given in the examples, and the identification information in the PPS may be referenced by the APS.

[0132] In Table 2, syntax and semantics are defined as follows.

[0133] pps_weighted_pred_flag is enable control information for applying the weighted prediction technique in the image layer to unidirectional prediction slices (P slices). When pps_weighted_pred_flag is equal to 1, it indicates that weighted prediction is applied to the P slices indicated by the PPS. Conversely, when pps_weighted_pred_flag is equal to 0, it indicates that weighted prediction is not applied to the P slices indicated by the PPS. When sps_weighted_pred_flag is equal to 0, the value of pps_weighted_pred_flag should be equal to 0.

[0134] pps_weighted_bipred_flag is enable control information for applying the weighted prediction technique in the image layer to bidirectional prediction slices (B slices). When pps_weighted_bipred_flag is equal to 1, it indicates that explicit weighted prediction is applied to the B slices indicated by the PPS. Conversely, when pps_weighted_pred_flag is equal to 0, it indicates that explicit weighted prediction is not applied to the B slices indicated by the PPS. When sps_weighted_bipred_flag is equal to 0, the value of pps_weighted_bipred_flag should be equal to 0.

[0135] The pps_wp_multi_weights_flag is enable control information regarding that a single reference picture in the picture layer has multiple sets of weighted prediction parameters. When the pps_wp_multi_weights_flag is equal to 1, it indicates that a single reference picture of the picture indicated by the PPS has multiple sets of weighted prediction parameters. Conversely, when the pps_wp_multi_weights_flag is equal to 0, it indicates that a single reference picture of the picture indicated by the PPS has only one set of weighted prediction parameters. When the sps_wp_multi_weights_flag is equal to 0, the value of the pps_wp_multi_weights_flag should be equal to 0.

[0136] When the pps_no_pic_partition_flag is equal to 1, it indicates that picture partitioning is not applied to each picture indicated by the PPS. When the pps_no_pic_partition_flag is equal to 0, it indicates that each picture indicated by the PPS may be divided into multiple tiles or slices.

[0137] When the pps_rpl_info_in_ph_flag is equal to 1, it indicates that the reference picture list (RPL) information exists within the picture header (PH) syntax structure and does not exist within the slice header that does not include the PH syntax structure indicated by the PPS. When the pps_rpl_info_in_ph_flag is equal to 0, it indicates that the RPL information does not exist within the PH syntax structure and may exist within the slice header indicated by the PPS.

[0138] When pps_wp_info_in_ph_flag is equal to 1, it indicates that the weighted prediction information may exist within the PH syntax structure and does not exist within the slice header that does not contain the PH syntax structure indicated by the PPS. When pps_wp_info_in_ph_flag is equal to 0, it indicates that the weighted prediction information does not exist within the PH syntax structure and may exist within the slice header indicated by the PPS.

[0139]

Table 2

[0140] In another embodiment, identification information of the region-adaptive weighted prediction parameters included in the picture header (PH) in the code stream is given in the embodiment. The weighted prediction parameters included in the PH may be referenced by the current picture, a slice within the current picture, and / or a CTU or CU within the current picture.

[0141] In Table 3, the syntax and semantics are defined as follows.

[0142] When ph_inter_slice_allowed_flag is equal to 0, it indicates that all the coded slices of the picture are of the intra prediction type (I slice). When ph_inter_slice_allowed_flag is equal to 1, it indicates that the picture is allowed to include one or more slices of the one-way or bi-directional inter-picture prediction type (P slice or B slice).

[0143] The multi_pred_weights_table( ) is a numerical table containing weighted prediction parameters, where a single reference image may have multiple sets of weighted prediction parameters (weights + offsets). When applying the inter-frame weighted prediction technique and determining that the weighted prediction information may exist within the PH, if pps_wp_multi_weights_flag is equal to 1, the weighted prediction parameters may be obtained from this table.

[0144] The pred_weight_table( ) is a numerical table containing weighted prediction parameters, where a single reference image has only a single set of weighted prediction parameters (weights + offsets). When applying the inter-frame weighted prediction technique and determining that the weighted prediction information may exist within the PH, if pps_wp_multi_weights_flag is equal to 0, the weighted prediction parameters may be obtained from this table.

[0145]

Table 3

[0146] In another embodiment, identification information of the region-adaptive weighted prediction parameters included in the slice header SH in the code stream is provided in the example. The weighted prediction parameters included in the SH may be referenced by the current slice and / or the CTU or CU within the current slice.

[0147] In Table 4, the syntax and semantics are defined as follows.

[0148] When sh_picture_header_in_slice_header_flag is equal to 1, it indicates that the PH syntax structure exists within the slice header. When sh_picture_header_in_slice_header_flag is equal to 0, the PH syntax structure does not exist in the slice header, that is, the slice layer may not fully inherit the identification information of the picture layer, and the encoding tools can be flexibly selected.

[0149] sh_slice_type indicates the encoding type of the slice, and may be an intra-frame encoding type (I slice), a uni-directional inter-frame encoding type (P slice), or a bi-directional inter-frame encoding type (B slice).

[0150] sh_wp_multi_weights_flag is enable control information regarding a single reference picture having multiple sets of weighted prediction parameters in the slice layer. When sh_wp_multi_weights_flag is equal to 1, it indicates that a single reference picture of the picture slice indicated by the slice header SH has multiple sets of weighted prediction parameters. Conversely, when sh_wp_multi_weights_flag is equal to 0, it indicates that a single reference picture of the picture slice indicated by the SH has only one set of weighted prediction parameters. When pps_wp_multi_weights_flag is equal to 0, the value of sh_wp_multi_weights_flag should be equal to 0.

[0151] When it is determined that the inter-frame weighted prediction technique is applied and there is a possibility that the weighted prediction information exists within the SH, if pps_wp_multi_weights_flag is equal to 1, the weighted prediction parameters may be obtained from the numerical table multi_pred_weights_table( ), and if pps_wp_multi_weights_flag is equal to 0, the weighted prediction parameters may be obtained from the numerical table pred_weight_table( ).

[0152]

Table 4

[0153] In another embodiment, the syntax and semantics of the numerical table of the weighted prediction parameters are given in the examples. Both Pred_weight_table( ) and multi_pred_weights_table( ) are numerical tables containing weighted prediction parameters. The difference between them is that the former defines that a single reference image has only a single set of weighted prediction parameters, and the latter defines that a single reference image may have multiple sets of weighted prediction parameters. Specifically, the syntax and semantics of pred_weight_table( ) may refer to the description in the document of the international standard H.266 / VVC version 1. The syntax and semantics of multi_pred_weights_table( ) are given in Table 5 and its description. Here, luma_log2_weight_denom and delta_chroma_log2_weight_denom are amplification factors of the weight factors for luminance and chrominance, respectively, to avoid floating-point operations on the encoding side.

[0154] num_l0_weights represents the number of weight factors that need to be indicated for a number of entries (reference images) in the reference picture list 0 (RPL 0) when pps_wp_info_in_ph_flag is equal to 1. The range of possible values of num_l0_weights is [0, Min(15, num_ref_entries[0][RplsIdx[0]])], where num_ref_entries[listIdx][rplsIdx] represents the number of entries in the reference picture list syntax structure ref_pic_list_struct(listIdx, rplsIdx). When pps_wp_info_in_ph_flag is equal to 1, the variable NumWeightsL0 is set to num_l0_weights. Conversely, when pps_wp_info_in_ph_flag is equal to 0, the variable NumWeightsL0 is set to NumRefIdxActive[0]. Here, the value of NumRefIdxActive[i] - 1 indicates the maximum reference index that may be used for slice decoding in the reference picture list i (RPL i). If the value of NumRefIdxActive[i] is 0, it indicates that there is no reference index used for slice decoding.

[0155] When luma_weight_l0_flag[i] is equal to 1, it indicates that the luma component for which unidirectional prediction is performed using the i-th entry (RefPicList[0][i]) in reference list 0 has a weight factor (weight + offset). When luma_weight_l0_flag[i] is equal to 0, it indicates that the above weight factor does not exist.

[0156] When chroma_weight_l0_flag[i] is equal to 1, it indicates that the chroma prediction value for which unidirectional prediction is performed using the i-th entry (RefPicList[0][i]) in reference list 0 has a weight factor (weight + offset). When chroma_weight_l0_flag[i] is equal to 0, it indicates that the above weight factor does not exist (default situation).

[0157] num_l0_luma_pred_weights[i] indicates the number of weight factors that need to be shown for the luma component of entry i (reference picture i) in reference picture list 0 (RPL 0) when luma_weight_l0_flag[i] is equal to 1, that is, the number of weighted prediction parameters that can be held by the luma component of a single reference picture i in list 0.

[0158] delta_luma_weight_l0[i][k] and luma_offset_l0[i][k] indicate the k-th weight factor and offset value of the i-th reference picture luma component in reference picture list 0, respectively.

[0159] num_l0_chroma_pred_weights[i][j] indicates the number of weight factors that need to be shown for the j-th chroma component of entry i (reference picture i) in reference picture list 0 (RPL 0) when chroma_weight_l0_flag[i] is equal to 1, that is, the number of weighted prediction parameters that can be held by the j-th chroma component of a single reference picture i in list 0.

[0160] delta_chroma_weight_l0[i][j][k] and delta_chroma_offset_l0[i][j][k] indicate the k-th weight factor and offset value of the j-th chroma component of the i-th reference picture in reference picture list 0, respectively.

[0161] For bidirectional prediction, as shown in Table 5, in addition to the above-mentioned weighted prediction parameter identification for reference picture list 0, similar information identification is also required for reference picture list 1 (RPL 1).

[0162]

Table 5

[0163] In another embodiment, the syntax and semantics of the numerical table of another weighted prediction parameter are provided in this example. Both pred_weight_table( ) and multi_pred_weights_table( ) are numerical tables containing weighted prediction parameters. The difference between them is that the former defines that a single reference image has only a single set of weighted prediction parameters, while the latter defines that a single reference image may have multiple sets of weighted prediction parameters. Specifically, the syntax and semantics of pred_weight_table( ) may refer to the description in the document of the international standard H.266 / VVC version 1. The syntax and semantics of multi_pred_weights_table( ) are given in Table 6 and its description. In Embodiment 11, for a large number of entries in the reference image list, that is, for a specified plurality of reference images, it is necessary to determine whether there is a weight factor for each of them. A single reference image with a weight factor may also have a set of weighted prediction parameters (including weights and offsets) in multiple sets. In contrast, this example is a special case of Embodiment 11, considering only one reference image in each reference image list, that is, when it is determined to apply the weighted prediction technique, a set of multiple sets of weighted prediction parameters owned by the reference image is directly specified. The meaning of each field in Table 6 is the same as the semantic interpretation corresponding to each field in Table 5.

[0164]

Table 6

[0165] In another embodiment, the identification information and parameters of region-adaptive weighted prediction are provided in the coding tree unit CTU.

[0166] The weighted prediction parameter information within a CTU may be separately identified, or it may refer to that within other parameter sets (e.g., the sequence layer parameter set SPS, the picture layer parameter set PPS) or header information (e.g., the picture header PH, the slice header SH). It may record the weighted prediction parameter difference value, or it may record the weighted prediction parameter index information and the difference value. When the weighted prediction parameter information within a CTU records the difference value, or records the index information and the difference value, one weighted numerical value is obtained from other parameter sets or header information, and by adding one difference value within the CTU, the weighted prediction parameter finally applied to the CTU or CU can be obtained. By including the weighted prediction parameter information in CTU units, a fine luminance gradation effect at the CTU unit level, such as circular gradation or radial gradation, can be realized.

[0167] When labeling the weighted prediction parameter information within a CTU, the specific bitstream construction method may be as shown in Table 7.

[0168] In Table 7, when a single reference image within the reference picture list RPL has only one set of weighted prediction parameters, for example, when the indication information sh_wp_multi_weights_flag is equal to 0, the weighted prediction parameter difference value may be directly set within the CTU. When a single reference image within the reference picture list has multiple sets of weighted prediction parameters, for example, when the indication information sh_wp_multi_weights_flag is equal to 1, it is necessary to label, within the CTU, the set of weighted prediction parameters with its own index, and the weighted prediction parameter difference value.

[0169] The weighted prediction parameter finally applied to the current CTU is the sum of the weighted prediction parameter of the reference picture and the weighted prediction parameter difference value defined in coding_tree_unit( ).

[0170] The weighted prediction parameter difference value includes, but is not limited to, the weighted prediction parameter difference values for the luminance component in RPL0 (ctu_delta_luma_weight_l0 and ctu_delta_luma_offset_l0), the weighted prediction parameter difference values for the chrominance component in RPL0 (ctu_delta_chroma_weight_l0[ i ] and ctu_delta_chroma_offset_l0[ i ]), the weighted prediction parameter difference values for the luminance component in RPL1 (ctu_delta_luma_weight_l1 and ctu_delta_luma_offset_l1), and the weighted prediction parameter difference values for the chrominance component in RPL1 (ctu_delta_chroma_weight_l1[ i ] and ctu_delta_chroma_offset_l1[ i ]).

[0171] The weighted prediction parameter index value includes, but is not limited to, the weighted prediction parameter index value for the luminance component in RPL0 (ctu_index_l0_luma_pred_weights), the weighted prediction parameter index value for the chrominance component in RPL0 (ctu_index_l0_chroma_pred_weights[ i ]), the weighted prediction parameter index value for the luminance component in RPL1 (ctu_index_l1_luma_pred_weights), and the weighted prediction parameter index value for the chrominance component in RPL1 (ctu_index_l1_chroma_pred_weights[ i ]).

[0172] In short, the weighted prediction parameter finally applied to the current CTU is the sum of the weighted prediction parameter difference value defined in coding_tree_unit( ) and the specific weighted prediction parameter of the reference image.

[0173]

Table 7

[0174] In another embodiment, the identification information and parameters of region - adaptive weighted prediction are given within a coding unit CU.

[0175] The weighted prediction parameter information within the CU may be independently identified, or may refer to other parameter sets (e.g., sequence - layer parameter set SPS, picture - layer parameter set PPS) or header information (e.g., picture header PH, slice header SH) or those within the coding tree unit CTU. The weighted prediction parameter difference value may be recorded, or the weighted prediction parameter index information and the difference value may be recorded. When the weighted prediction parameter information within the CU records the difference value, or records the index information and the difference value, one weighted value is obtained from other parameter sets or header information and added to one difference value within the CU to obtain the weighted prediction parameter finally applied to the CU. By including the weighted prediction parameter information in units of CU, it is possible to achieve a fine - grained luminance gradation effect in units of CU, such as circular gradation or radial gradation.

[0176] When labeling the weighted prediction parameter information within the CU, the specific method of constructing the code stream may be as shown in Table 8.

[0177] Here, cu_pred_weights_adjust_flag indicates whether it is necessary to adjust the weighted prediction parameter value indexed to the current CU. When cu_pred_weights_adjust_flag is equal to 1, it is necessary to adjust the weighted prediction parameter value in the current CU. That is, it indicates that the weighted prediction parameter value finally applied to the current CU is the sum of the weighted prediction parameter value at the CTU level and the difference value labeled at the CU level. When Cu_pred_weights_adjust_flag is equal to 1, it indicates that the weighted prediction parameter value determined at the CTU level is directly used for the current CU.

[0178] In Table 8, the weighted prediction parameter difference values labeled at the CU level include the weighted prediction parameter difference values (cu_delta_luma_weight_l0 and cu_delta_luma_offset_l0) for the luminance components in RPL0.

[0179] When the current CU includes chrominance components, the weighted prediction parameter difference values labeled at the CU level further include the weighted prediction parameter difference values for each chrominance component. When the current CU is bidirectional prediction, the weighted prediction parameter difference values labeled at the CU level further include the weighted prediction parameter difference values for the luminance component and / or each chrominance component in RPL1.

[0180] In short, the weighted prediction parameter finally applied to the current CU is the sum of the specific weighted prediction parameter held by the upper-level data (e.g., CTU level, Slice level, Subpicture level) in the encoding / decoding structure and the weighted prediction parameter difference value defined in coding_unit( ).

[0181] [Table 8]

[0182] In another embodiment, the identification information and parameters of the region-adaptive weighted prediction are provided within the Supplemental Enhancement Information (SEI).

[0183] The NAL unit type in the network abstraction layer unit header information nal_unit_header( ) is set to 23, representing the leading SEI information. ssei_rbsp( ) contains the related coded stream sei_message( ), and the sei_message( ) contains valid data information. The value of payloadType may be different from other SEI information in the current H.266 / VVC version 1 (for example, the value may be 100). In this case, the payload_size_byte contains the coded stream information related to the region-adaptive weighted prediction. The specific method of constructing the coded stream is shown in Table 7.

[0184] When multi_pred_weights_cancel_flag is 1, the SEI information related to the previous picture is cancelled, and the related SEI function is not used for that picture. When multi_pred_weights_cancel_flag is 0, the previous SEI information is continued to be used (in the decoding process, if the current picture does not hold the SEI information, the SEI information of the previous picture held in memory is continued to be used in the decoding process of the current picture), and the related SEI function for that picture is enabled. When multi_pred_weights_persistence_flag is 1, the SEI information is applied to the current picture and the pictures after the current layer. When multi_pred_weights_persistence_flag is 0, the SEI information is applied only to the current picture. The meanings of the other fields in Table 9 are the same as the semantic interpretations corresponding to each field in Table 5.

[0185]

Table 9

[0186] In another embodiment, the identification information and parameters of region-adaptive weighted prediction are given in the media description information. Here, the media description information includes, but is not limited to, media presentation description (MPD) information within the Dynamic Adaptive Streaming over HTTP (DASH) protocol and asset descriptor information within the MPEG Media Transport (MMT) protocol. Taking the media resource descriptor in MMT as an example, Table 10 shows the coding stream configuration method of the weighted prediction information for its region adaptation. The syntax and field meanings in Table 10 are the same as the semantic interpretations corresponding to each field in Table 5.

[0187]

Table 10

[0188] FIG. 11 is a schematic configuration diagram of a video encoding apparatus provided according to an embodiment of the present application. This video encoding apparatus can execute the video encoding method provided according to any embodiment of the present application and has functional modules and beneficial effects corresponding to the executed method. This apparatus may be implemented by software and / or hardware. Specifically, it includes an image acquisition module 610 and a video encoding module 620.

[0189] The image acquisition module 610 is configured to acquire a video image that is an image of at least one frame of a video.

[0190] The video encoding module 620 is configured to perform weighted prediction encoding on the video image by using at least one set of weighted prediction identification information and parameters to generate an image code stream.

[0191] According to this embodiment, the image acquisition module acquires a video image that is an image of at least one frame in the video, and the video encoding module performs weighted prediction encoding on the video image by using at least one set of weighted prediction identification information and parameters to generate an image code stream, thereby realizing flexible encoding of the video image, improving the efficiency of video encoding, and reducing the influence of the luminance change of the video image on the encoding efficiency.

[0192] Furthermore, based on the above embodiment of the present application, the apparatus further includes a change determination module configured to determine the luminance change situation based on the comparison result between the video image and the reference image.

[0193] Furthermore, based on the above embodiment of the present application, the luminance change situation includes at least one of the average luminance change value of the image and the luminance change value of the pixel point.

[0194] Furthermore, based on the above embodiment of the present application, the video encoding module 620 includes an encoding processing unit configured to perform weighted prediction encoding on the video image based on the luminance change situation of the video image.

[0195] Furthermore, based on the above embodiment, specifically, if the luminance change situation is such that the image luminance of the entire frame is uniform, weighted prediction encoding is performed on the video image, and if the luminance change situation is such that the image luminance of each partial region is uniform, weighted prediction encoding is performed on the image of each partial region in the video image respectively.

[0196] Furthermore, based on the above embodiments, the set of the weight prediction identification information and parameters used for weighted prediction coding in the apparatus corresponds to the video image of one frame or the image of at least one partial region of the video image.

[0197] Furthermore, based on the above embodiments, the standard of the image of the partial region in the apparatus includes one of a slice, a tile, a subpicture, a coding tree unit, and a coding unit.

[0198] Furthermore, based on the above embodiments, the apparatus further further includes a code stream writing module configured to write the weight prediction identification information and parameters into the image code stream.

[0199] Furthermore, based on the embodiments of the present application, the weight prediction identification information and the parameters in the apparatus are included in at least one parameter set of a sequence layer parameter set, a picture layer parameter set, a slice layer parameter set, supplementary enhancement information, video usability information, picture header information, slice header information, network abstraction layer unit header information, a coding tree unit, and a coding unit.

[0200] Furthermore, based on the above embodiments, the weight prediction identification information and the parameters in the apparatus include at least one of reference picture index information, weight prediction enable control information, region-adaptive weight prediction enable control information, and weight prediction parameters.

[0201] Furthermore, based on the above embodiments, the image code stream in the apparatus includes a transport stream or a media file.

[0202] Furthermore, based on the above embodiments, the video coding module 620 It further includes a deep learning unit configured to perform weighted prediction encoding on the video image according to a pre-trained neural network model.

[0203] Furthermore, based on the above embodiment, the weighted prediction identification information in the apparatus further includes a neural network model structure and neural network model parameters.

[0204] FIG. 12 is a schematic structural diagram of a video decoding apparatus provided by an embodiment of the present application. This video decoding apparatus can execute the video decoding method provided by any embodiment of the present application, and has a functional module and beneficial effects corresponding to the executed method. This apparatus may be implemented by software and / or hardware. Specifically, it includes a code stream acquisition module 710 and an image reconstruction module 720.

[0205] The code stream acquisition module 710 is configured to acquire an image code stream and analyze the weighted prediction identification information and parameters in the image code stream.

[0206] The image reconstruction module 720 is configured to decode the image code stream according to the weighted prediction identification information and parameters, and generate a reconstructed image.

[0207] According to this embodiment, the code stream acquisition module acquires an image code stream, acquires the weighted prediction identification information and parameters in the image code stream, and the image reconstruction module processes the image code stream according to the acquired weighted prediction identification information and parameters to generate a reconstructed image, thereby realizing dynamic decoding of the image code stream, improving the efficiency of video decoding, and reducing the influence of changes in image luminance on video decoding efficiency.

[0208] Furthermore, based on the above embodiments, the number of weighted prediction identification information and parameters in the image code stream in the device is at least one set.

[0209] Furthermore, based on the above embodiments, the weighted prediction identification information and the parameters in the device are included in at least one parameter set among a sequence layer parameter set, an image layer parameter set, a slice layer parameter set, supplementary enhancement information, video user quality information, image header information, slice header information, network abstraction layer unit header information, a coded tree unit, and a coding unit.

[0210] Furthermore, based on the above embodiments, the weighted prediction identification information and the parameters in the device include at least one of reference image index information, weighted prediction enable control information, region-adaptive weighted prediction enable control information, and weighted prediction parameters.

[0211] Furthermore, based on the above embodiments, the image code stream in the device includes a transport stream or a media file.

[0212] Furthermore, based on the above embodiments, the image reconstruction module 720 includes a deep learning decoding unit configured to perform weighted prediction decoding on the image code stream using the weighted prediction identification information and parameters according to a pre-trained neural network model to generate a reconstructed image.

[0213] Furthermore, based on the above embodiments, the weighted prediction identification information and parameters in the device further include a neural network model structure and neural network model parameters.

[0214] In some examples, FIG. 13 is a schematic structural diagram of an encoding unit provided according to an embodiment of the present application. The encoding unit shown in FIG. 13 is applied to a device that performs encoding processing on a video. The input of the device is an image included in the video, and the output is an image code stream, or a transmission stream or a media file including the image code stream. The encoding unit performs the following steps. In step 501, an image is input. In step 502, the encoding unit processes and encodes the image. The specific operation process is shown in the video encoding method provided by any of the above embodiments. In step 503, a code stream is output.

[0215] In some examples, FIG. 14 is a schematic structural diagram of a decoding unit provided according to an embodiment of the present application. The decoding unit shown in FIG. 14 is applied to a device that performs decoding processing on a video. The input of the device is an image code stream, or a transmission stream or a media file including the image code stream, and the output is an image constituting the video. The decoding unit performs the following steps. In step 601, an image is input. In step 602, the decoding unit analyzes the code stream to obtain an image and decodes the image. An example of the specific operation process is shown in the video decoding method provided by any of the above embodiments. In step 603, the image is output. In step 604, the playback unit plays back the image.

[0216] FIG. 15 is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. This electronic device includes a processor 70, a memory 71, an input device 72, and an output device 73. The number of processors 70 in the electronic device may be one or more, and FIG. 15 takes one processor 70 as an example. The processor 70, the memory 71, the input device 72, and the output device 73 in the electronic device may be connected by a bus or other means, and FIG. 15 shows an example of being connected by a bus.

[0217] The memory 71, as a computer-readable storage medium, can store software programs, computer-executable programs, and modules, for example, modules corresponding to the video encoding device or video decoding device in the embodiments of the present application (image acquisition module 610 and video encoding module 620, or coded stream acquisition module 710 and image reconstruction module 720). The processor 70 executes the software programs, instructions, and modules stored in the memory 71 to execute various functional applications and data processing of the electronic device, that is, to implement the methods described above.

[0218] The memory 71 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by using the electronic device and the like. Further, the memory 71 may include a high-speed random access memory, or may include a non-volatile memory, for example, at least one magnetic disk memory device, a flash memory device, or other non-volatile solid-state memory devices. In some examples, the memory 71 may further include a memory disposed remotely from the processor 70, and these remote memories may be connected to the electronic device via a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0219] The input device 72 may be used to receive input numerical and character information or generate key signal inputs related to user settings and function control of the electronic device. The output device 73 may include a display device such as a display.

[0220] This embodiment further provides a storage medium including computer-executable instructions. When the computer-executable instructions are executed by a computer processor, a video encoding method is executed, and the method includes: obtaining a video image which is an image of at least one frame of the video; performing weighted prediction encoding on the video image to generate an image code stream, where the weighted prediction encoding uses at least one set of weighted prediction identification information and parameters.

[0221] Alternatively, when the computer-executable instructions are executed by a computer processor, a video decoding method is executed, and the method includes: obtaining an image code stream, and analyzing the weighted prediction identification information and parameters in the image code stream; decoding the image code stream according to the weighted prediction identification information and parameters to generate a reconstructed image.

[0222] Through the description of the above embodiments, although the present application can of course be implemented by hardware, in many cases, it may also be implemented by software and the necessary general-purpose hardware methods, which are better embodiments. Based on such an understanding, the essential part of the technical solution of the present application or the part contributing to the prior art may be embodied in the form of a software product. This software product may be stored in a computer-readable storage medium such as a floppy disk (registered trademark), read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disk of a computer, and includes several instructions configured to cause a computer device (which may be a personal computer, server, network device, etc.) to execute the methods described in each embodiment of the present application.

[0223] In addition, in the embodiments of the above device, the units and modules included are only classified based on functional logic, and are not limited to the above classification as long as the corresponding functions can be realized. Also, the specific names of the functional units are only for the purpose of facilitating mutual distinction, and do not limit the protection scope of this application.

[0224] All or some of the steps of the method disclosed above, functional modules / units in a system or device may be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0225] In the hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical assemblies. For example, one physical assembly may have multiple functions, or one function or step may be executed collaboratively by several physical assemblies. Some or all of the physical assemblies may be implemented as software executed by a processor such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). The term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage devices, magnetic cartridges, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer, but is not limited to these. Further, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information transmission medium.

[0226] As described above, some embodiments of the present application have been described with reference to the drawings, but the scope of the rights of the present application is not limited thereby. Any changes, substitutions by equivalents, and improvements made by those skilled in the art without departing from the scope and essence of the present application shall be within the scope of the rights of the present application.

Claims

1. A method for video encoding, comprising: obtaining a video image, which is an image of at least one frame of the video; performing weighted prediction encoding on the video image to generate an image code stream, wherein in the weighted prediction encoding, at least one set of weighted prediction identification information and parameters is used; wherein the step of performing weighted prediction encoding on the video image to generate an image code stream includes: performing weighted prediction encoding on the video image based on the luminance change situation of the video image with respect to a reference image, and the step of performing weighted prediction encoding on the video image based on the luminance change situation of the video image with respect to the reference image includes: if the luminance change situation is such that the image luminance of the entire frame is uniform, performing weighted prediction encoding on the video image; and if the luminance change situation is such that the image luminance of a partial region is uniform, performing weighted prediction encoding on the image of each partial region within the video image; wherein the method includes at least one of the above.

2. The method according to claim 1, further comprising determining a luminance change situation based on a comparison result between the video image and the reference image.

3. The method according to claim 2, wherein the luminance change situation includes at least one of an average luminance change value of the image and a luminance change value of a pixel point.

4. In the video image, there is at least one set of the weighted prediction identification information and parameters for the reference image or an image of a partial region of the reference image. The method according to claim 1.

5. One set of the weighted prediction identification information and parameters used in the weighted prediction encoding corresponds to the video image of one frame or an image of at least one partial region of the video image. The method according to claim 1.

6. The standard of the image of the partial region includes at least one of a slice, a tile, a subpicture, a coding tree unit, and a coding unit. The method according to claim 5.

7. The method according to claim 1, further comprising writing the weighted prediction identification information and parameters into the image code stream.

8. ​ ​ ​ ​ ​ ​ The weighted prediction identification information and the parameters are included in at least one parameter set among a sequence layer parameter set, an image layer parameter set, a slice layer parameter set, supplementary enhancement information, video user utility information, image header information, slice header information, network abstraction layer unit header information, an encoding tree unit, and an encoding unit. The method according to claim 7.

9. The weighted prediction identification information and the parameters include at least one of reference image index information, weighted prediction enable control information, region-adaptive weighted prediction enable control information, and weighted prediction parameters. The method according to claim 1.

10. The step of performing weighted prediction encoding on the video image to generate an image code stream includes: Performing weighted prediction encoding on the video image according to a pre-trained neural network model. The method according to claim 1, including the step described above.

11. The weighted prediction identification information further includes a neural network model structure and neural network model parameters. The method according to claim 10.

12. A method for video decoding, comprising: Obtaining an image code stream and analyzing weighted prediction identification information and parameters in the image code stream; Decoding the image code stream according to the weighted prediction identification information and parameters to generate a reconstructed image; The image code stream is generated by performing weighted prediction encoding on the video image based on the luminance change situation of the video image with respect to a reference image. The step of performing the weighted prediction encoding on the video image based on the luminance change situation of the video image with respect to the reference image includes at least one of the following steps: if the luminance change situation is such that the image luminance of the entire frame is uniform, performing weighted prediction encoding on the video image; if the luminance change situation is such that the image luminance of a partial region is uniform, performing weighted prediction encoding on each image of the partial regions in the video image respectively.

13. The number of the weighted prediction identification information and the parameters in the image code stream is at least one set. The method according to claim 12.

14. The weighted prediction identification information and the parameters are included in at least one parameter set among a sequence layer parameter set, an image layer parameter set, a slice layer parameter set, supplementary enhancement information, video user usability information, image header information, slice header information, network abstraction layer unit header information, an encoding tree unit, and an encoding unit, and / or the weighted prediction identification information and the parameters include at least one of reference image index information, weighted prediction enable control information, region-adaptive weighted prediction enable control information, and weighted prediction parameters The method according to claim 12.

15. The step of decoding the image code stream according to the weighted prediction identification information and the parameters to generate a reconstructed image is performing weighted prediction decoding on the image code stream using the weighted prediction identification information and the parameters according to a pre-trained neural network model to generate a reconstructed image The method according to claim 14, including this.

16. The weighted prediction identification information and the parameters further include a neural network model structure and neural network model parameters The method according to claim 15.

17. An electronic device comprising one or more processors and a memory storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 11 or 12 to 16 Electronic device.

18. A computer-readable storage medium storing one or more programs, wherein the one or more programs can implement the method according to any one of claims 1 to 11 or 12 to 16 when executed by one or more processors Computer-readable storage medium.

Citation Information

Patent Citations

  • Method and apparatus for encoding / decoding video signal

    US20190149836A1

  • Image processing device and method

    WO2012157538A1

  • Encoding method and decoding method

    WO2013057782A1

  • Image decoding device, image coding device, image decoding method, and image coding method

    WO2018037919A1

  • Interframe predictive coding method and device

    WO2018040869A1