Video encoding method, video decoding method, and related apparatus

By adding preset syntax during the video encoding process to control the display and encoding methods of knowledge images, the problems of high encoding latency and large caching overhead are solved, and a more efficient encoding and decoding process is achieved.

CN119342232BActive Publication Date: 2026-05-08ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2023-07-18
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing video coding methods suffer from long coding latency, especially when encoding knowledge images, which leads to increased coding time and significant caching overhead.

Method used

By adding preset syntax to the encoded data, it indicates whether the knowledge image is used for display and/or whether to force the use of skip mode to encode the same content frame, thereby reducing the multiple encoding of the original image content corresponding to the knowledge image, and selectively controlling the display or encoding method of the knowledge image by using skip mode to encode the same content frame.

Benefits of technology

It reduces encoding latency and caching overhead, improves encoding efficiency, and reduces encoding time and hardware resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119342232B_ABST
    Figure CN119342232B_ABST
Patent Text Reader

Abstract

The application discloses a video coding method, a video decoding method and related devices. The video coding method comprises: coding a video to obtain coded data; adding a preset syntax in the coded data to obtain a video bitstream, different values of the preset syntax being used to indicate whether a knowledge image is used for display, and / or whether a same content frame of the knowledge image is encoded by using a skip mode compulsorily, the knowledge image and the same content frame corresponding to the same original image content. The application can reduce coding delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and in particular to a video encoding method, a video decoding method, and related apparatus. Background Technology

[0002] Video image data is relatively large, so it is usually necessary to compress the video pixel data (RGB, YUV, etc.). The compressed data is called a video stream, which is transmitted to the user's end via wired or wireless network for decoding and viewing. The entire video encoding process includes prediction, transformation, quantization, and encoding.

[0003] In video encoding and decoding, to improve compression ratio and reduce the number of codewords to be transmitted, the encoder does not directly encode and transmit pixel values. Instead, it uses intra-frame or inter-frame prediction modes to predict the pixel values ​​of the current block using reconstructed pixels from the coded blocks of the current frame or a reference frame. The pixel value predicted using a certain prediction mode is called the predicted pixel value, and the difference between the predicted pixel value and the original pixel value is called the residual. The encoder only needs to encode a certain prediction mode and the residual generated when using that prediction mode, and the decoder can decode the corresponding pixel value based on this bitstream information. This greatly reduces the number of codewords required for encoding.

[0004] However, during the long-term research and development process, the inventors of this application discovered that current video coding methods still have some shortcomings; for example, there is a relatively long coding delay. Summary of the Invention

[0005] This application provides a video encoding method, a video decoding method, and related apparatus, which can reduce encoding latency.

[0006] To achieve the above objectives, this application provides a video coding method, which includes:

[0007] The video is encoded to obtain encoded data;

[0008] A preset syntax is added to the encoded data to obtain a video stream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display and / or whether to force the use of a skip mode to encode the same content frame of the knowledge image. The knowledge image and the same content frame correspond to the same original image content.

[0009] In one embodiment,

[0010] The process of encoding the video to obtain encoded data includes: encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image;

[0011] The step of adding a preset syntax to the encoded data to obtain a video stream includes: adding the preset syntax to the encoded data of the knowledge image and / or the sequence parameter set corresponding to the knowledge image to indicate that the knowledge image is used for display.

[0012] In one embodiment, the process of encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image includes:

[0013] The playback sequence number of the knowledge image is encoded into the encoded data of the knowledge image.

[0014] In one embodiment, encoding the video includes: encoding the knowledge image to obtain encoded data of the knowledge image; setting a preset mode for all image blocks in the same content frame of the knowledge image to a skip mode, and disabling all filtering tools in the encoding process of the same content frame to obtain encoded data of the same content frame;

[0015] The step of adding a preset syntax to the encoded data to obtain a video stream includes: adding the preset syntax to the encoded data of the knowledge image and / or the sequence parameter set corresponding to the knowledge image to indicate that the same content frames of the knowledge image are encoded using a forced skip mode.

[0016] In one embodiment, the preset syntax has at least a first value, a second value, and / or a third value;

[0017] When the preset syntax value is the first value, it means that the knowledge image is not used for display, and the skip mode is not forced to encode the same content frames of the knowledge image.

[0018] When the preset syntax value is the second value, it means that the knowledge image is used for display;

[0019] When the preset syntax is set to the third value, it means that the knowledge image is not used for display, and the same content frames of the knowledge image are forcibly encoded using the skip mode.

[0020] In one embodiment, the preset syntax includes a first syntax and / or a second syntax;

[0021] When the first syntax takes the fourth value, it means that the knowledge image is not used for display;

[0022] When the first syntax takes the fifth value, it means that the knowledge image is used for display;

[0023] When the second syntax takes the sixth value, it means that the skip mode is not forced to encode the same content frames of the knowledge image;

[0024] When the second syntax takes the seventh value, it means that the skip mode is forced to encode the same content frames of the knowledge image.

[0025] In one embodiment, adding a preset syntax to the encoded data to obtain a video bitstream includes:

[0026] Add the preset syntax to the sequence parameter set and / or image parameter set.

[0027] In one embodiment, the preset syntax includes switch syntax and usage state syntax, and adding the preset syntax to the sequence parameter set and / or image parameter set includes:

[0028] The switch syntax is added to the sequence parameter set. Different values ​​of the switch syntax are used to indicate whether the display scheme of knowledge images is enabled in the image sequence of the sequence parameter set, and / or whether the scheme of forcibly encoding the same content frames of knowledge images using the skip mode is enabled in the image sequence of the sequence parameter set.

[0029] The usage state syntax is added to the image parameter set. Different values ​​of the usage state syntax are used to indicate whether the knowledge image corresponding to the image parameter set is used for display, and / or whether to force the use of skip mode to encode the same content frame of the knowledge image corresponding to the image parameter set.

[0030] In one embodiment, encoding the video to obtain encoded data includes:

[0031] Determine whether to use fragmented transmission for knowledge images;

[0032] If fragmented transmission is used, the knowledge image is divided into at least two image blocks, and then the at least two image blocks of the knowledge image are respectively encoded into different coded data;

[0033] If fragmented transmission is not used, the entire knowledge image frame is directly encoded into a single coded data.

[0034] In one embodiment, determining whether to use fragmented transmission for the knowledge image includes: determining whether to use fragmented transmission for the knowledge image based on the business scenario;

[0035] The business scenario is determined based on the bit overhead and / or image bitrate requirements of the knowledge image.

[0036] In one embodiment, adding a preset syntax to the encoded data to obtain a video bitstream includes:

[0037] If the knowledge image is not transmitted in fragments, the preset syntax in the encoded data indicates that the knowledge image is to be displayed.

[0038] If the knowledge image is transmitted in fragments, the preset syntax in the encoded data indicates that the knowledge image is not used for display.

[0039] To achieve the above objectives, this application also provides a video coding method, the method comprising:

[0040] The knowledge image is forced to use a skip mode to encode the same content frames, and the encoded data of the same content frames is obtained. The knowledge image and its corresponding same content frames correspond to the same original image content.

[0041] To achieve the above objectives, this application also provides a video decoding method, the method comprising:

[0042] A preset syntax is obtained from the video stream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display and / or whether to force the use of a skip mode to encode the same content frame of the knowledge image. The knowledge image and its corresponding same content frame correspond to the same original image content.

[0043] Based on the preset syntax, the knowledge image and / or the same content frames of the knowledge image are decoded.

[0044] In one embodiment, decoding a knowledge image and / or identical content frames of the knowledge image based on the preset syntax includes: forcing the use of a skip mode to encode identical content frames of the knowledge image when the value of the preset syntax indicates that the knowledge image is decoded; and decoding the knowledge image based on the decoded image of the knowledge image, forcing the use of the skip mode to decode to obtain a decoded image of the identical content frames.

[0045] The method further includes: controlling the display unit to display a decoded image of the same content frame.

[0046] To achieve the above objectives, this application also provides a video decoding method, comprising: decoding a video bitstream to obtain a decoded image of a knowledge image;

[0047] The control display unit displays the decoded image of the knowledge image.

[0048] To achieve the above objectives, this application also provides an encoder that includes a processor; the processor is configured to execute instructions to implement the steps of the above method.

[0049] To achieve the above objectives, this application also provides a decoder, which includes a processor; the processor is used to execute instructions to implement the steps of the above method.

[0050] To achieve the above objectives, this application also provides a computer-readable storage medium for storing instruction / program data that can be executed to implement the above methods.

[0051] The video encoding method of this application encodes knowledge images to obtain encoded data; a preset syntax is added to the encoded data to obtain a video bitstream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display, and / or whether to force the use of a skip mode to encode frames with the same content. The frames with the same content and the knowledge image correspond to the same original image content. In this way, the knowledge image can be selectively displayed, thus avoiding multiple encodings of the original image content corresponding to the knowledge image, which can reduce encoding time. Alternatively, the skip mode can be selectively forced to encode frames with the same content of the knowledge image, which can reduce the second encoding time of the original image content corresponding to the knowledge image. Therefore, the method of this application can reduce encoding latency. Attached Figure Description

[0052] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0053] Figure 1 This is a schematic diagram illustrating the encoding order of multiple images in an image sequence in related technologies;

[0054] Figure 2 This is a schematic diagram of one embodiment of the video encoding method of this application;

[0055] Figure 3 This is a schematic diagram of an embodiment of the video encoding method of this application;

[0056] Figure 4 This is a schematic diagram of another embodiment of the video encoding method of this application;

[0057] Figure 5 This is a schematic diagram of one embodiment of the video decoding method of this application;

[0058] Figure 6 This is a schematic diagram of another embodiment of the video decoding method of this application;

[0059] Figure 7 This is a schematic diagram of yet another embodiment of the video decoding method of this application;

[0060] Figure 8 This is a schematic diagram of one embodiment of the encoder of this application;

[0061] Figure 9 This is a schematic diagram of the structure of one embodiment of the decoder of this application;

[0062] Figure 10 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application. In addition, unless otherwise specified (e.g., "or additionally" or "or in alternatives"), the term "or" as used herein refers to a non-exclusive "or" (i.e., "and / or"). Furthermore, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.

[0064] Some video codec standards (such as the SVAC3 standard) introduce the concept of a library picture. In this application, a library picture is represented by an L-frame. In current standards, a library picture is a long-term reference frame encoded using I-frames. The library picture serves only as a reference frame and is not used for display. Library pictures are identified by their Library Picture Index (IDX). Since library pictures are not currently used for display, a playback sequence number (i.e., POC or DOI) is not used to identify them.

[0065] like Figure 1 As shown, these standards also introduce the concept of RL (reference library) frames, which are P-frames or B-frames that reference only the knowledge image. Generally, the original image of an RL frame is the same as the original image of the L-frame it references. Therefore, the original image corresponding to the L-frame is encoded twice. The first encoding is used as the knowledge image (e.g., L0), which is not used for display. The second encoding is used as a P-frame or B-frame that references only that knowledge image (e.g., RL0), which is used for display. Consequently, there is an encoding delay before the regular image is displayed due to the repeated encoding of certain frames.

[0066] Based on this, this application proposes a video encoding method. This method encodes a knowledge image to obtain encoded data. A preset syntax is added to the encoded data to obtain a video stream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display, and / or whether to force the use of a skip mode to encode frames with the same content. The frames with the same content and the knowledge image correspond to the same original image content. Thus, the knowledge image can be selectively displayed, avoiding multiple encodings of the original image content corresponding to the knowledge image, thereby reducing encoding time. Alternatively, the skip mode can be selectively forced to encode frames with the same content of the knowledge image, reducing the time spent on the second encoding of the original image content corresponding to the knowledge image. Therefore, the method of this application can reduce encoding latency.

[0067] The aforementioned frames with identical content can be understood as frames where the knowledge image is encoded first, and then the corresponding original image content is encoded again. Generally, in schemes where the original image content corresponding to the knowledge image is encoded twice, RL frames are usually frames with identical content. However, in schemes where the original image content corresponding to the knowledge image is encoded only once, frames with identical content to the knowledge image generally do not appear, and thus RL frames are generally not frames with identical content.

[0068] Specifically, such as Figure 2 As shown, the video encoding method proposed in this application may include the following steps. It should be noted that the step numbers are for simplification only and are not intended to limit the execution order of the steps. The execution order of each step in this embodiment can be arbitrarily changed without departing from the technical concept of this application.

[0069] S101: Encode the video to obtain encoded data.

[0070] S102: Add preset syntax to the encoded data to obtain the video bitstream.

[0071] Optionally, the knowledge image can be encoded to obtain encoded data. Then, a preset syntax can be added to the encoded data. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display and / or whether to force the use of a skip mode to encode the same content frame. In this way, the knowledge image can be selectively displayed, so that the original image corresponding to the knowledge image does not need to be encoded multiple times, which can reduce the encoding time. Alternatively, the same content frame of the knowledge image can be selectively forced to use a skip mode to encode, which can reduce the second encoding time of the original image content corresponding to the knowledge image. Therefore, the method of this application can reduce the encoding delay.

[0072] The video includes at least one image sequence, and each image sequence includes multiple image frames. Optionally, an image sequence may include at least one knowledge image.

[0073] In one feasible approach, encoding video may include: encoding at least one parameter of the image sequence in the video (e.g., video encoding configuration and compatibility level, video image size and aspect ratio, video frame rate and bit rate, inter-frame prediction and intra-frame prediction settings, quantization parameters, entropy coding mode, reference frame settings, etc.) to obtain encoded data; then, a preset syntax may be added to the encoded data so that the preset syntax and the original parameters of the image sequence in the encoded data form a sequence parameter set of the image sequence. That is, in this feasible approach, a preset syntax may be added to the sequence parameter set (SPS) of the image sequence to uniformly control whether all knowledge image frames in the sequence are used for display and whether the skip mode is forced to encode the same content frames of all knowledge images in the sequence.

[0074] The default syntax can be is_display, and the specific way to add it is shown in the underlined content in Table 1.

[0075] Table 1. A schematic table of sequence parameter sets

[0076] Sequence Parameter Set (RBSP) Definition descriptor sequence_parameter_set_rbsp(){ profile_id u(8) level_id u(8) progressive_sequence u(1) field_coded_sequence u(1) library_picture_enable_flag u(1) if(LibraryPictureEnableFlag){ is_display u(1) } marker_bit f(1) ……

[0077] In another possible implementation, encoding the video may include: encoding at least one parameter of the image sequence in the video (e.g., video encoding configuration and compatibility level, video image size and aspect ratio, video frame rate and bit rate, inter-frame prediction and intra-frame prediction settings, quantization parameters, entropy coding mode, reference frame settings, etc.) to obtain encoded data of the image sequence parameters; then, a preset syntax can be added to the encoded data of the image sequence parameters so that the preset syntax and the original parameters of the image sequence in the encoded data form a sequence parameter set of the image sequence. That is, in this possible implementation, a preset syntax can be added to the sequence parameter set (SPS) of the image sequence to indicate the activation status of the "display scheme of knowledge images" and the "scheme of forcing the use of skip mode to encode the same content frames of knowledge images" in the image sequence.

[0078] Encoding the video may further include: encoding the knowledge images in the aforementioned image sequence to obtain encoded data for the knowledge images; then, predefined syntax may be added to the encoded data of the knowledge images so that the predefined syntax in the encoded data of the knowledge images is used to indicate whether the knowledge image is used for display, and / or whether to force the use of a skip mode to encode the same content frames of the knowledge image. The encoded data of the knowledge images may include an image parameter set and image content data. The predefined syntax may be `is_display`, and as shown in Table 2, predefined syntax can be added to the image parameter set of the knowledge images.

[0079] Table 2 is a schematic table of image parameter sets.

[0080]

[0081]

[0082] Optionally, the encoded data of the image sequence parameters and the encoded data of at least one knowledge image can be combined into a video stream packet, that is, the sequence parameter set and the encoded data of at least one knowledge image can be placed in the same stream packet and transmitted together to the decoding end.

[0083] In a specific example, by using a preset syntax indicator in the sequence parameter set to disable the "knowledge image display scheme" in the image sequence, all knowledge images in the image sequence will not be displayed, regardless of whether the preset syntax exists in the encoded data of the knowledge images or what the value of the preset syntax is. Therefore, it is preferable that by using a preset syntax indicator in the sequence parameter set to disable the "knowledge image display scheme" in the image sequence, the preset syntax can be omitted from the encoded data of the knowledge images in the image sequence, thereby reducing the video compression rate. Alternatively, by using a preset syntax indicator in the sequence parameter set to enable the "knowledge image display scheme" in the image sequence, the preset syntax can be added to the encoded data of each knowledge image in the image sequence, so that the preset syntax in the encoded data of each knowledge image indicates whether each knowledge image is used for display.

[0084] In another specific example, by using a preset syntax instruction in the sequence parameter set to disable the "scheme of forcing the use of skip mode to encode identical content frames of knowledge images" in the image sequence, regardless of whether the preset syntax exists in the encoded data of the knowledge images in the image sequence or what the value of the preset syntax is, all knowledge images in the image sequence will not be forced to use skip mode to encode identical content frames of knowledge images. That is, the original image content corresponding to the knowledge images will be re-encoded using the conventional encoding method of existing technology. Thus, it is preferable that when the preset syntax instruction in the sequence parameter set disables the "scheme of forcing the use of skip mode to encode identical content frames of knowledge images" in the image sequence, the preset syntax does not need to be added to the encoded data of the knowledge images in the image sequence, thereby reducing the video compression rate. Alternatively, by using a preset syntax instruction in the sequence parameter set to enable the "scheme of forcing the use of skip mode to encode identical content frames of knowledge images" in the image sequence, the preset syntax can be added to the encoded data of each knowledge image in the image sequence to indicate whether to force the use of skip mode to encode identical content frames of each knowledge image.

[0085] In another feasible approach, encoding the video may include: encoding a knowledge image to obtain encoded data of the knowledge image; then, predefined syntax may be added to the encoded data of the knowledge image so that the predefined syntax in the encoded data of the knowledge image is used to indicate whether the knowledge image is used for display, and / or whether to force the encoding of the same content frames of the knowledge image using a skip mode. The predefined syntax may be added to the image parameter set of the knowledge image.

[0086] For the original image content of a knowledge image, it can include at least three different processing methods.

[0087] The first approach involves not displaying the decoded image of the knowledge image at the decoding end. In this case, the encoding end encodes the original image content corresponding to the knowledge image twice: once for encoding the undisplayed knowledge image, and once for encoding the same content frame (i.e., the RL image) of the displayed knowledge image. When encoding the same content frame for display, conventional encoding methods can be used to re-encode the original image content corresponding to the knowledge image. For example, the reconstructed image of the knowledge image can be used as a reference image to perform inter-frame encoding on the original image content, resulting in encoded data for the same content frame. However, this second encoding of the original image content corresponding to the knowledge image also incurs significant time overhead. Furthermore, since the decoded image of the knowledge image and the decoded image of the same content frame may differ in this scheme, both the decoded image of the knowledge image and the decoded image of the same content frame need to be stored in a cache at the decoding end, resulting in significant cache overhead. In other words, this approach incurs substantial hardware overhead.

[0088] The second approach is to control the decoding end to display the decoded image of the knowledge image, that is, to set the knowledge image for display. In this way, the original image content corresponding to the knowledge image does not need to be encoded again for display. That is, after encoding the knowledge image, the next frame of the knowledge image can be encoded directly, thus reducing encoding time. Furthermore, for the original image content of the knowledge image, the decoding end will only decode one image, thus avoiding the problem of "needing to simultaneously store the decoded image of the knowledge image and the decoded image of its corresponding content frame in the cache." This approach also saves caching overhead. For example, the knowledge image can be encoded to obtain encoded data; then the encoded data is sent to the decoding end; the decoding end decodes the encoded data to obtain the decoded image of the knowledge image; and the decoding end displays the decoded image of the knowledge image.

[0089] To facilitate video playback by the decoding end based on the video bitstream, when a knowledge image is set for display, its playback sequence number can be encoded into the knowledge image's encoded data. This allows the decoding end to determine the playback order of the knowledge images, thus facilitating video playback. In other words, knowledge images and regular images can be numbered together with their playback sequence numbers (i.e., DOI / POC).

[0090] The third approach is to not display the decoded image of the knowledge image at the decoding end. In this case, the encoding end encodes the original image content corresponding to the knowledge image twice: once for encoding the undisplayed knowledge image, and once for encoding the same content frame (i.e., the RL image) of the displayed knowledge image. When encoding the same content frame for display, a skip mode can be forced to encode the original image content corresponding to the knowledge image. That is, during the second encoding of the original image content of the knowledge image, the prediction mode of all CUs in the image can be set to the skip mode with mv=0 (equivalent to directly copying the reconstructed value after the first encoding), and all filtering tools in this encoding process can be turned off. Since the second encoding is in skip mode and filtering tools are turned off, the time overhead of the second encoding is very small, and the reconstructed image content after encoding is consistent with the first encoding. Therefore, when decoding the same content frame, the decoded image of the knowledge image can be directly copied, thus saving video decoding time.

[0091] Furthermore, for the decoding end, since the decoded image of the knowledge image can be directly copied when decoding the same content frame (i.e., the decoded image of the knowledge image is the same as the decoded image of its same content frame), after obtaining the decoded image of the same content frame of the knowledge image, either the decoded image of the knowledge image or the decoded image of the same content frame of the knowledge image can be deleted. Thus, the cache only retains one of the two, saving cache overhead. More preferably, after obtaining the decoded image of the same content frame of the knowledge image, the decoded image of the knowledge image can be deleted from the cache to facilitate the display of the decoded image of the same content frame of the knowledge image and the subsequent decoding of images using the same content frame as a reference frame.

[0092] Alternatively, any of the above processing methods can be executed by default.

[0093] Alternatively, at least two of the three processing methods mentioned above can be selected as the processing method for the original image content corresponding to the knowledge image.

[0094] In some embodiments, there are two processing methods for the processing of the original image content corresponding to a knowledge image. The encoding end can decide to use one of the two processing methods to process the original image content of the knowledge image. After the decision is made, a preset syntax can be added to the encoding data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs. This allows the decoding end to know which of the two processing methods the encoding end uses to process the original image content of the knowledge image.

[0095] In a specific example, the encoder can decide which of the first and second processing methods described above to use to process the original image content of the knowledge image. After deciding, a preset syntax `is_display` can be added to the encoded data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs. The value of the preset syntax `is_display` informs the decoder which of the first and second processing methods the encoder uses to process the original image content of the knowledge image. Specifically, the preset syntax can include a switch syntax and a usage state syntax. That is, there are two syntax elements controlling the switching between the first and second processing methods: 1. As shown in Table 3, a switch syntax `is_display_sps` is added to the sequence parameter set (SPS); 2. As shown in Table 4, a usage state syntax `is_display_pps` is added to the image parameter set. `is_display_sps` indicates whether the display switching function (i.e., whether the "knowledge image display scheme" is enabled) is enabled for all knowledge image frames in the sequence. When `is_display_sps = 1`, the display switching function is enabled. In this case, `is_display_pps` needs to be transmitted to indicate whether an individual knowledge image frame is displayed. `is_display_pps = 1` indicates that the image frame is displayed, and `is_display_pps = 0` indicates that the image frame is not displayed. When `is_display_sps = 0`, the display switching function is disabled for the entire sequence, and `is_display_pps` does not need to be transmitted.

[0096] Table 3. Another schematic table of sequence parameter sets

[0097]

[0098] Table 4. Another schematic table of image parameter sets

[0099] Intra-frame prediction prediction image parameter set (RBSP) definition descriptor intra_picture_parameter_set_rbsp(){ decode_order_index u(8) intra_picture_start_code f(8) bbv_delay u(32) time_code_flag u(1) if (time_code_flag == '1') time_code u(24) if(LibraryStreamFlag){ library_picture_index ue(v) If (IsDisplaySPS){ u(1) is_display_pps } }

[0100] In another specific example, the encoder can decide which of the second and third processing methods described above to use to process the original image content of the knowledge image. After deciding, the preset syntax is_display can be added to the encoded data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs, so that the decoder knows which of the second and third processing methods described above the encoder uses to process the original image content of the knowledge image through the value of the preset syntax is_display.

[0101] In another specific example, the encoder can decide which of the first and third processing methods described above to use to process the original image content of the knowledge image. After deciding, the preset syntax need_copy_lib can be added to the encoded data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs. The value of the preset syntax need_copy_lib will let the decoder know which of the first and third processing methods described above the encoder uses to process the original image content of the knowledge image.

[0102] In other embodiments, the selection range for the processing method of the original image content corresponding to a knowledge image can include the three processing methods mentioned above. That is, the encoding end can decide which of the three processing methods to use to process the original image content of the knowledge image. After the decision is made, a preset syntax can be added to the encoding data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs, so that the decoding end knows which of the three processing methods the encoding end uses to process the original image content of the knowledge image through the encoding data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs.

[0103] For example, the encoder can decide which of the three processing methods mentioned above to use to process the original image content of the knowledge image. After deciding, the preset syntax is_display can be added to the encoded data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs, so that the decoder knows which of the three processing methods to use through different values ​​of the preset syntax is_display.

[0104] In a specific example, the preset syntax `display_mode` takes a first value (e.g., 0), representing the first processing method described above, i.e., not displaying the knowledge image, and re-encoding the original image content corresponding to the knowledge image using the conventional encoding method of existing technology; the preset syntax `display_mode` takes a second value (e.g., 1), representing the second processing method, i.e., the knowledge image is used for display; the preset syntax `display_mode` takes a third value (e.g., 2), representing the third processing method, i.e., not displaying the knowledge image, and forcibly using the skip mode to encode the same content frames of the knowledge image. As shown in Table 5, a syntax element identifier `display_mode` can be added to the Sequence Parameter Set (SPS) to control the switching of the display mode of the knowledge image, thus uniformly controlling all knowledge image frames in the sequence. `display_mode=0` means all knowledge images in the sequence are displayed; `display_mode=1` means all knowledge images in the sequence are not displayed, but the original image content corresponding to the knowledge images is forcibly encoded using the skip method when encoded a second time; `display_mode=2` means all knowledge images in the sequence are not displayed, and the original image content corresponding to the knowledge images is encoded using the conventional encoding method of existing technology when encoded a second time.

[0105] Table 5 shows another schematic table of sequence parameter sets.

[0106] Sequence Parameter Set (RBSP) Definition descriptor sequence_parameter_set_rbsp(){ profile_id u(8) level_id u(8) progressive_sequence u(1) field_coded_sequence u(1) library_picture_enable_flag u(1) if(LibraryPictureEnableFlag){ display_mode u(2) } marker_bit f(1) ……

[0107] For example, the preset syntax includes a first syntax and a second syntax. Specifically, the encoder can decide which of the three processing methods mentioned above to use to process the original image content of the knowledge image. After deciding, the first syntax `is_display` can be added to the encoded data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs, so that the decoder knows whether the knowledge image needs to be displayed through the value of the first syntax `is_display`. If the first syntax `is_display` indicates that the knowledge image should not be displayed, the second syntax `need_copy_lib` can also be added to the encoded data of the knowledge image and / or the sequence parameter set of the image sequence to which the knowledge image belongs, so that the decoder knows whether to force the use of the skip mode to encode the same content frame of the knowledge image through the value of the second syntax `need_copy_lib`.

[0108] In this context, the first syntax `is_display` takes the fourth value (e.g., 0), representing that the knowledge image is not displayed. The second syntax `is_display` takes the fifth value (e.g., 1), representing the second processing method, i.e., displaying the knowledge image. Further, when the first syntax `is_display` takes the fourth value (e.g., 0), the second syntax can be set. The second syntax `need_copy_lib` takes the sixth value (e.g., 0), representing the first processing method mentioned above, i.e., not displaying the knowledge image and not forcing the use of skip mode to encode the same content frames of the knowledge image; that is, using the conventional encoding method of existing technology to re-encode the original image content corresponding to the knowledge image. The third syntax `need_copy_lib` takes the seventh value (e.g., 1), representing the third processing method, i.e., not displaying the knowledge image and forcing the use of skip mode to encode the same content frames of the knowledge image.

[0109] In a specific example, the on / off syntax (i.e., the first syntax) of the "knowledge image display scheme" is added to the sequence parameter set of the image sequence to indicate the on / off state of the "knowledge image display scheme" in the image sequence through the on / off syntax of the "knowledge image display scheme" in the sequence parameter set; the usage state syntax of the "knowledge image display scheme" and / or the usage state syntax of the "scheme to force the use of skip mode to encode the same content frames of knowledge images" are added to the encoded data of each knowledge image in the image sequence; so that when the "knowledge image display scheme" is not enabled in the image sequence, the "force" in the encoded data of each knowledge image is used to indicate the on / off state of the "knowledge image display scheme". The usage state syntax of "Scheme to encode identical content frames of knowledge images using skip mode" indicates whether to force the use of skip mode to encode identical content frames of each knowledge image; it can also indicate whether each knowledge image is used for display by using the usage state syntax of "Display scheme of knowledge image" in the encoded data of each knowledge image when the "Display scheme of knowledge image" is enabled in the image sequence, and / or indicate whether to force the use of skip mode to encode identical content frames of each knowledge image by using the usage state syntax of "Scheme to force the use of skip mode to encode identical content frames of knowledge images" in the encoded data of each knowledge image.

[0110] Specifically, the first syntax `is_display` with the fourth value can be added to the sequence parameter set of the image sequence. This indicates that the "knowledge image display scheme" is not enabled in the image sequence. Furthermore, the second syntax `need_copy_lib` can be added to the encoded data of each knowledge image in the image sequence. This indicates whether the skip mode is forced for identical content frames in each knowledge image when the knowledge image is not displayed. Encoding is performed; specifically, if the second syntax need_copy_lib in the encoded data of a knowledge image is the sixth value, it means that the first processing method mentioned above is adopted, that is, the knowledge image is not displayed, and the skip mode is not forced to encode the same content frame of the knowledge image, that is, the original image content corresponding to the knowledge image is re-encoded using the conventional encoding method of the existing technology; if the second syntax need_copy_lib in the encoded data of a knowledge image is the seventh value, it means that the third processing method is adopted, that is, the knowledge image is not displayed, and the skip mode is forced to encode the same content frame of the knowledge image. In other embodiments, the first syntax `is_display` being set to a fifth value can be added to the sequence parameter set of the image sequence. This allows the syntax indicating that a "knowledge image display scheme" is enabled in the image sequence through an `is_display` value of 1. Further, the first syntax `is_display` can be added to the encoded data of each knowledge image in the image sequence. This allows the value of the first syntax `is_display` in the encoded data of each knowledge image to indicate whether each knowledge image is used for display. Specifically, if the first syntax `is_display` in the encoded data of a knowledge image is set to a fourth value, it means the knowledge image is not displayed; if the first syntax `is_display` in the encoded data of a knowledge image is set to a fifth value, it means the second processing method is used, i.e., the knowledge image is displayed. Furthermore, when the first syntax `is_display` in the encoded data of a knowledge image is set to a fourth value, a second syntax can be added to the encoded data of that knowledge image to indicate whether to force the use of a skip mode for encoding the same content frame of the knowledge image.

[0111] In another specific example, the switch syntax for "display scheme of knowledge images" and / or the switch syntax for "scheme to force the encoding of the same content frames of knowledge images using skip mode" are added to the sequence parameter set of the image sequence. This is to indicate the on / off state of "display scheme of knowledge images" in the image sequence through the switch syntax for "display scheme of knowledge images" in the sequence parameter set, and / or, to indicate the on / off state of "scheme to force the encoding of the same content frames of knowledge images using skip mode" in the image sequence through the switch syntax for "scheme to force the encoding of the same content frames of knowledge images using skip mode" in the sequence parameter set. The usage state syntax for "display scheme of knowledge images" and / or the usage state syntax for "scheme to force the encoding of the same content frames of knowledge images using skip mode" are added to the encoded data of each knowledge image in the image sequence. Thus, if the sequence parameter set indicates that the image sequence does not enable the "knowledge image display scheme" and does not enable the "scheme to force the use of skip mode to encode the same content frames of knowledge images," it can be assumed that all knowledge images in the image sequence adopt the first processing method described above, that is, knowledge images are not displayed, and the use of skip mode to encode the same content frames of knowledge images is not forced; that is, the original image content corresponding to the knowledge image is re-encoded using the conventional encoding method of the existing technology. If the sequence parameter set indicates that the image sequence does not enable the "knowledge image display scheme" and enables the "scheme to force the use of skip mode to encode the same content frames of knowledge images," the usage status of the "scheme to force the use of skip mode to encode the same content frames of knowledge images" in the encoding data of each knowledge image can be used as a reference. The syntax indicates whether to force the use of skip mode to encode the same content frames of each knowledge image; if the sequence parameter set indicates that the image sequence enables the "knowledge image display scheme" and the "scheme to force the use of skip mode to encode the same content frames of knowledge images" is enabled, the usage state syntax of the "knowledge image display scheme" and / or the usage state syntax of the "scheme to force the use of skip mode to encode the same content frames of knowledge images" can be added to the encoding data of each knowledge image in the image sequence, so as to determine which of the above three processing methods the original image content of the knowledge image is processed by the usage state syntax of the "knowledge image display scheme" and / or the usage state syntax of the "scheme to force the use of skip mode to encode the same content frames of knowledge images" in the encoding data of each knowledge image.

[0112] In another specific example, the usage state syntax for "display scheme of knowledge images" and / or the usage state syntax for "forced skip mode encoding of identical content frames of knowledge images" can be directly added to the sequence parameter set of the image sequence. This allows for unified control over whether all knowledge images in the image sequence use the "display scheme of knowledge images" and / or the "forced skip mode encoding of identical content frames of knowledge images" through the usage state syntax of "display scheme of knowledge images" and / or the "forced skip mode encoding of identical content frames of knowledge images." Specifically, when the usage state syntax of "display scheme of knowledge images" in the sequence parameter set of the image sequence indicates that all knowledge images in the image sequence are displayed, the processing method for the original image content of all knowledge images in the image sequence is the second processing method described above. When the usage state syntax of "display scheme for knowledge images" in the sequence parameter set of an image sequence indicates that all knowledge images in the image sequence are not displayed, the usage state syntax of "force use of skip mode to encode the same content frames of knowledge images" in the sequence parameter set of the image sequence can indicate whether to force the use of skip mode to encode all knowledge images in the image sequence; or, the on / off syntax of "force use of skip mode to encode the same content frames of knowledge images" in the sequence parameter set of the image sequence can also indicate the on / off state of "force use of skip mode to encode the same content frames of knowledge images" in the image sequence. When "force use of skip mode to encode the same content frames of knowledge images" is enabled in an image sequence, the " The scheme of "forcing the use of skip mode to encode identical content frames of knowledge images" is added to the encoding data of each knowledge image in the image sequence using state syntax. This state syntax indicates whether or not the skip mode is forcibly used to encode identical content frames of knowledge images. Alternatively, the scheme can be added to the encoding data of each knowledge image in the image sequence using state syntax. As shown in Table 6, two syntax elements, `is_display` and `need_copy_lib`, can be added to the Sequence Parameter Set (SPS). `is_display` controls whether a knowledge image is displayable, and `need_copy_lib` controls the encoding method used when a knowledge image is not displayed. This provides unified control over whether all knowledge image frames in the sequence are displayed.First, the syntax element `is_display` can be transmitted. `is_display = 1` indicates that all knowledge images in the sequence are displayed; `is_display = 0` indicates that all knowledge images in the sequence are not displayed. Then, when `is_display = 0`, the syntax element `need_copy_lib` needs to be transmitted. `need_copy_lib = 1` forces the second encoding of the original image content corresponding to the knowledge image to use the skip encoding method; `need_copy_lib = 0` encodes the second encoding of the original image content corresponding to the knowledge image using the conventional encoding method of existing technology.

[0113] Table 6 shows another schematic table of sequence parameter sets.

[0114] Sequence Parameter Set (RBSP) Definition descriptor sequence_parameter_set_rbsp(){ profile_id u(8) level_id u(8) progressive_sequence u(1) field_coded_sequence u(1) library_picture_enable_flag u(1) if(LibraryPictureEnableFlag){ is_display u(1) If (IsDisplay == 0) { need_copy_lib u(1) } } marker_bit f(1) ……

[0115] In another specific example, instead of adding the relevant syntax of "display scheme of knowledge image" and / or "scheme to force the encoding of the same content frame of knowledge image using skip mode" to the sequence parameter set of the image sequence, the processing method of the original image content of each knowledge image is directly indicated by the usage state syntax of "display scheme of knowledge image" and / or "scheme to force the encoding of the same content frame of knowledge image using skip mode" in the encoded data of each knowledge image.

[0116] In addition, the decision to use patch transmission (i.e., fragmented transmission) to process knowledge images can be made based on the business scenario.

[0117] The business scenario can be determined based on the bit overhead of the knowledge image and / or the image bitrate requirements.

[0118] For example, if the bitrate requirement for knowledge images is low, the knowledge image can be processed without using patch-based transmission, i.e. Figure 1 As shown, the knowledge image is transmitted directly as a whole frame, that is, the entire knowledge image frame is directly encoded into a bitstream; if the bitrate requirement of the knowledge image is high, such as... Figure 3As shown, knowledge images can be processed using a patch-based transmission method. This involves dividing the knowledge image into at least two image blocks, and then encoding these blocks into different bitstreams. In this way, the decoding end can only decode the knowledge image by obtaining the bitstream containing the encoded data of all image blocks, reducing the impact of large bitstream data on the transmission channel. More preferably, before transmitting the bitstream of the "image encoded from the reference knowledge image," the bitstream packet containing the encoded data of all image blocks of the reference knowledge image is transmitted, so that the decoding end obtains the decoded image of the knowledge image before decoding the "image encoded from the reference knowledge image."

[0119] The bitstream requirements for knowledge images can be related to video resolution requirements, image clarity requirements, and / or the update frequency of knowledge images.

[0120] Optionally, if the update frequency of the knowledge image frame is low (e.g., below the frequency threshold) and a higher bitrate is required to encode the knowledge image, i.e., the bitstream requirement of the knowledge image is relatively high, then the knowledge image can be processed by using a patch transmission method. If the update frequency of the knowledge image frame is high (e.g., above the frequency threshold) and the generated knowledge image does not need to have high image quality, i.e., the bitstream requirement of the knowledge image is relatively low, then the knowledge image can be transmitted as a whole frame, and the impact on the transmission channel is not significant.

[0121] Alternatively, if the video resolution requirement is high (e.g., above the resolution threshold), a higher bitrate is needed to encode the knowledge image, meaning the bitrate requirement for the knowledge image is relatively high. In this case, the knowledge image can be processed using a patch transmission method. If the video resolution requirement is low (e.g., below the resolution threshold), and the generated knowledge image does not need to have high image quality, meaning the bitrate requirement for the knowledge image is relatively low, then transmitting the knowledge image as a whole frame is sufficient, and the impact on the transmission channel is not significant.

[0122] For example, if the bit overhead of the knowledge image is small (e.g., less than the bit threshold), the knowledge image can be processed without using the patch transmission method, that is, the knowledge image can be transmitted as a whole frame, that is, the whole frame of knowledge image is directly encoded into a bitstream. If the bit overhead of the knowledge image is large (e.g., greater than the bit threshold), the knowledge image can be processed using the patch transmission method, that is, the knowledge image is divided into at least two image blocks, and then at least two image blocks of the knowledge image are encoded into different bitstreams, which can reduce the impact of the transmission channel caused by the large bitstream data.

[0123] Optionally, regardless of whether the original image content of the knowledge image is transmitted in patches or in whole frames, the processing method can be any of the three processing methods described above, without limitation. Preferably, the processing method for the original image content of the knowledge image transmitted in patches can be the first processing method described above. The processing method for the original image content of the knowledge image transmitted in whole frames can be either the second or third processing method described above. The choice between the second and third processing methods can be determined based on the encoder's own capabilities and other practical considerations.

[0124] For example, in some actual business scenarios, the frequency of updating knowledge image frames can be higher, and the generated knowledge images do not need to have high image quality. In this case, transmitting the knowledge image as a whole frame is sufficient, and the impact on the transmission channel is not significant. The knowledge image can be displayed directly, and then the image content of the next frame can be encoded. This reduces the latency of displaying one frame of image and eliminates the need for additional buffering. Specifically, such as... Figure 4 As shown, L0 and L1 represent knowledge image frames. The knowledge image with POC=0 is the first frame of the encoded video sequence, and the preset syntax value of the knowledge image with POC=0 is set to the value corresponding to the "knowledge image display scheme". The RL frame with POC=1 is the image content of the second frame of the video sequence. In some other business scenarios, the update frequency of knowledge image frames is low, and a higher bitrate is required to encode the knowledge image. In this case, to avoid impacting the transmission channel, it is necessary to transmit the knowledge image in patches. In this case, the first processing method can be used to process the original image content corresponding to the knowledge image, that is, to be consistent with the patching method of the existing technology.

[0125] This application also provides a video encoding method according to another embodiment. In this embodiment, the video encoding method includes the following steps:

[0126] The knowledge image is forced to use a skip mode to encode the same content frames, and the encoded data of the same content frames is obtained. The knowledge image and the same content frames of the knowledge image correspond to the same original image content.

[0127] Optionally, forcing the use of a skip mode to encode identical content frames of a knowledge image can mean: setting the preset mode of all image blocks in identical image frames of a knowledge image to a skip mode with mv=0, and disabling all filtering tools in the encoding process of identical content frames to obtain encoded data of identical image frames.

[0128] This application also provides a video encoding method according to yet another embodiment. In this embodiment, the video encoding method includes the following steps: encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image.

[0129] Please see Figure 5 , Figure 5 This is a flowchart illustrating one embodiment of the video decoding method of this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily follow that approach. Figure 5 The illustrated process sequence is limited. In this embodiment, the video decoding method includes the following steps:

[0130] S201: Obtain preset syntax from video stream.

[0131] The different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display, and / or whether to force the use of a skip mode to encode the same content frame of the knowledge image, wherein the knowledge image and the same content frame of the knowledge image correspond to the same original image content.

[0132] S202: Decode knowledge images and / or frames containing the same content from knowledge images based on a preset syntax.

[0133] The step of decoding the knowledge image and / or the same content frame of the knowledge image based on the preset syntax includes: decoding only the knowledge image when the value of the preset syntax indicates that the knowledge image is to be displayed; the method further includes: controlling the display unit to display the decoded image of the knowledge image.

[0134] Optionally, decoding the knowledge image and / or identical content frames of the knowledge image based on the preset syntax includes: encoding the identical content frames of the knowledge image by forcing the use of a skip mode under the value indication of the preset syntax, and decoding the knowledge image; using the decoded image of the knowledge image as the predicted image of the identical content frame; obtaining the decoded image of the identical content frame based on the predicted image of the identical content frame; the method further includes: controlling the display unit to display the decoded image of the identical content frame. Wherein, obtaining the decoded image of the identical content frame based on the predicted image of the identical content frame can be: adding the predicted image and the residual image of the identical content frame to obtain the decoded image of the identical content frame.

[0135] Furthermore, based on the preset syntax, decoding the knowledge image and / or the same content frame of the knowledge image includes: forcing the use of a skip mode to encode the same content frame of the knowledge image when the value of the preset syntax indicates that the knowledge image is decoded; directly copying the decoded image of the knowledge image as the decoded image of the same content frame; the method further includes: controlling the display unit to display the decoded image of the same content frame.

[0136] Furthermore, the specific details of the video decoding method correspond to the specific details of the video encoding method in the above embodiments, and can be referred to the relevant content of the video encoding method in the above embodiments, which will not be repeated here.

[0137] Please see Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the video decoding method of this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily follow that approach. Figure 6 The illustrated process sequence is limited. In this embodiment, the video decoding method includes the following steps:

[0138] S301: Decode the video stream to obtain the decoded image of the knowledge image;

[0139] S302: Controls the display unit to display the decoded image of the knowledge image.

[0140] Please see Figure 7 , Figure 7 This is a flowchart illustrating another embodiment of the video decoding method of this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily follow that approach. Figure 7 The illustrated process sequence is limited. In this embodiment, the video decoding method includes the following steps:

[0141] S401: Decode the video stream to obtain the decoded image of the knowledge image;

[0142] S402: Decoding image based on knowledge image, forcing the use of skip mode to decode the same content frame of knowledge image to obtain the decoded image of the same content frame.

[0143] After obtaining the decoded image of the same content frame, the display unit can be controlled to display the decoded image of the same content frame. However, the decoded image of the knowledge image may not be used for display.

[0144] Please see Figure 8 , Figure 8 This is a schematic diagram of one embodiment of the encoder of this application. The encoder 10 includes a processor 12, which executes instructions to implement the video encoding method described above. For details of the implementation process, please refer to the description of the above embodiment, which will not be repeated here.

[0145] Processor 12 can also be referred to as a CPU (Central Processing Unit). Processor 12 may be an integrated circuit chip with signal processing capabilities. Processor 12 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor, or processor 12 can be any conventional processor.

[0146] The encoder 10 may further include a memory 11 for storing instructions and data required for the processor 12 to run.

[0147] The processor 12 is used to execute instructions to implement the methods provided by any embodiment and any non-conflicting combination of the video encoding methods of this application described above.

[0148] Please see Figure 9 , Figure 9 This is a schematic diagram of one embodiment of the encoder of this application. The encoder 20 includes a processor 22, which executes instructions to implement the video decoding method described above. For detailed implementation process, please refer to the description of the above embodiment; it will not be repeated here.

[0149] Processor 22 can also be referred to as CPU (Central Processing Unit). Processor 22 may be an integrated circuit chip with signal processing capabilities. Processor 22 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor can be a microprocessor, or processor 22 can be any conventional processor.

[0150] The encoder 20 may further include a memory 21 for storing instructions and data required for the processor 22 to run.

[0151] The processor 22 is used to execute instructions to implement the methods provided by any embodiment and any non-conflicting combination of the video decoding methods of this application described above.

[0152] Please see Figure 10 , Figure 10This is a schematic diagram of the structure of a computer-readable storage medium in an embodiment of this application. The computer-readable storage medium 30 in this embodiment stores instruction / program data 31. When executed, this instruction / program data 31 implements the methods provided by any embodiment of the video encoding method and video decoding method of this application, as well as any non-conflicting combination thereof. The instruction / program data 31 can be formed into a program file and stored in the storage medium 30 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) or processor can execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium 30 includes various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0154] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0155] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0156] The above are merely embodiments of this application and do not limit the scope of this patent application. Any equivalent structural or procedural changes made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. A video encoding method, characterized in that, The method includes: The video is encoded to obtain encoded data; A preset syntax is added to the encoded data to obtain a video stream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display. The step of adding a preset syntax to the encoded data to obtain a video bitstream includes: If the knowledge image is transmitted in segments, and if the bitstream of other images is inserted between the bitstreams of at least two image blocks of the knowledge image, the preset syntax in the encoded data indicates that the knowledge image is not used for display.

2. The video encoding method according to claim 1, characterized in that, The process of encoding the video to obtain encoded data includes: encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image; The step of adding a preset syntax to the encoded data to obtain a video stream includes: adding the preset syntax to the encoded data of the knowledge image and / or the sequence parameter set corresponding to the knowledge image to indicate that the knowledge image is used for display.

3. The video encoding method according to claim 2, characterized in that, The process of encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image includes: The playback sequence number of the knowledge image is encoded into the encoded data of the knowledge image.

4. The video encoding method according to claim 1, characterized in that, The preset syntax has at least a first value, a second value, and / or a third value; When the preset syntax value is the first value, it means that the knowledge image is not used for display, and the skip mode is not forced to encode the same content frames of the knowledge image. When the preset syntax value is the second value, it means that the knowledge image is used for display; When the preset syntax is set to the third value, it means that the knowledge image is not used for display, and the same content frames of the knowledge image are forcibly encoded using the skip mode.

5. The video encoding method according to claim 1, characterized in that, The preset syntax includes a first syntax and / or a second syntax; When the first syntax takes the fourth value, it means that the knowledge image is not used for display; When the first syntax takes the fifth value, it means that the knowledge image is used for display; When the second syntax takes the sixth value, it means that the skip mode is not forced to encode the same content frames of the knowledge image; When the second syntax takes the seventh value, it means that the skip mode is forced to encode the same content frames of the knowledge image.

6. The video encoding method according to claim 1, characterized in that, The step of adding a preset syntax to the encoded data to obtain a video bitstream includes: Add the preset syntax to the sequence parameter set and / or image parameter set.

7. The video encoding method according to claim 6, characterized in that, The preset syntax includes switch syntax and usage state syntax. Adding the preset syntax to the sequence parameter set and / or image parameter set includes: The switch syntax is added to the sequence parameter set, and different values ​​of the switch syntax are used to indicate whether the display scheme of knowledge images is enabled in the image sequence of the sequence parameter set; The usage state syntax is added to the image parameter set, and different values ​​of the usage state syntax are used to indicate whether the knowledge image corresponding to the image parameter set is used for display.

8. The video encoding method according to claim 1, characterized in that, The process of encoding the video to obtain encoded data includes: Determine whether to use fragmented transmission for knowledge images; If fragmented transmission is used, the knowledge image is divided into at least two image blocks, and then the at least two image blocks of the knowledge image are respectively encoded into different coded data; If fragmented transmission is not used, the entire knowledge image frame is directly encoded into a single coded data.

9. The video encoding method according to claim 8, characterized in that, The determination of whether to use fragmented transmission for knowledge images includes: determining whether to use fragmented transmission for knowledge images based on the business scenario; The business scenario is determined based on the bit overhead and / or image bitrate requirements of the knowledge image.

10. The video encoding method according to claim 8, characterized in that, The step of adding a preset syntax to the encoded data to obtain a video bitstream includes: If the knowledge image is not transmitted in fragments, the preset syntax in the encoded data indicates that the knowledge image is to be displayed. If the knowledge image is transmitted in fragments, the preset syntax in the encoded data indicates that the knowledge image is not used for display.

11. A video encoding method, characterized in that, The method includes: The video is encoded to obtain encoded data; A preset syntax is added to the encoded data to obtain a video bitstream. Different values ​​of the preset syntax are used to indicate whether a knowledge image is used for display and whether to force the use of a skip mode to encode the same content frames of the knowledge image. The knowledge image and the same content frames correspond to the same original image content. The forced use of the skip mode to encode the same content frames of the knowledge image is as follows: set the preset mode of all image blocks in the same content frame to the skip mode, and turn off all filtering tools in the encoding process of the same content frame. The step of adding a preset syntax to the encoded data to obtain a video bitstream includes: If the knowledge image is transmitted in segments, and if the bitstream of other images is inserted between the bitstreams of at least two image blocks of the knowledge image, the preset syntax in the encoded data indicates that the knowledge image is not used for display.

12. The video encoding method according to claim 11, characterized in that, The process of encoding the video to obtain encoded data includes: encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image; The step of adding a preset syntax to the encoded data to obtain a video stream includes: adding the preset syntax to the encoded data of the knowledge image and / or the sequence parameter set corresponding to the knowledge image to indicate that the knowledge image is used for display.

13. The video encoding method according to claim 12, characterized in that, The process of encoding the original image content of the knowledge image only once to obtain the encoded data of the knowledge image includes: The playback sequence number of the knowledge image is encoded into the encoded data of the knowledge image.

14. The video encoding method according to claim 11, characterized in that, The preset syntax has at least a first value, a second value, and / or a third value; When the preset syntax value is the first value, it means that the knowledge image is not used for display, and the skip mode is not forced to encode the same content frames of the knowledge image. When the preset syntax value is the second value, it means that the knowledge image is used for display; When the preset syntax is set to the third value, it means that the knowledge image is not used for display, and the same content frames of the knowledge image are forcibly encoded using the skip mode.

15. The video encoding method according to claim 11, characterized in that, The preset syntax includes a first syntax and / or a second syntax; When the first syntax takes the fourth value, it means that the knowledge image is not used for display; When the first syntax takes the fifth value, it means that the knowledge image is used for display; When the second syntax takes the sixth value, it means that the skip mode is not forced to encode the same content frames of the knowledge image; When the second syntax takes the seventh value, it means that the skip mode is forced to encode the same content frames of the knowledge image.

16. The video encoding method according to claim 11, characterized in that, The step of adding a preset syntax to the encoded data to obtain a video bitstream includes: Add the preset syntax to the sequence parameter set and / or image parameter set.

17. The video encoding method according to claim 16, characterized in that, The preset syntax includes switch syntax and usage state syntax. Adding the preset syntax to the sequence parameter set and / or image parameter set includes: The switch syntax is added to the sequence parameter set. Different values ​​of the switch syntax are used to indicate whether the display scheme of knowledge images is enabled in the image sequence of the sequence parameter set, and whether the scheme of forcibly using the skip mode to encode the same content frames of knowledge images is enabled in the image sequence of the sequence parameter set. The usage state syntax is added to the image parameter set. Different values ​​of the usage state syntax are used to indicate whether the knowledge image corresponding to the image parameter set is used for display and whether to force the use of skip mode to encode the same content frame of the knowledge image corresponding to the image parameter set.

18. The video encoding method according to claim 11, characterized in that, The process of encoding the video to obtain encoded data includes: Determine whether to use fragmented transmission for knowledge images; If fragmented transmission is used, the knowledge image is divided into at least two image blocks, and then the at least two image blocks of the knowledge image are respectively encoded into different coded data; If fragmented transmission is not used, the entire knowledge image frame is directly encoded into a single coded data.

19. The video encoding method according to claim 18, characterized in that, The determination of whether to use fragmented transmission for knowledge images includes: determining whether to use fragmented transmission for knowledge images based on the business scenario; The business scenario is determined based on the bit overhead and / or image bitrate requirements of the knowledge image.

20. The video encoding method according to claim 18, characterized in that, The step of adding a preset syntax to the encoded data to obtain a video bitstream includes: If the knowledge image is not transmitted in fragments, the preset syntax in the encoded data indicates that the knowledge image is to be displayed. If the knowledge image is transmitted in fragments, the preset syntax in the encoded data indicates that the knowledge image is not used for display.

21. A video decoding method, characterized in that, The method includes: A preset syntax is obtained from the video bitstream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display. If the knowledge image is transmitted in segments, and if the bitstream of other images is inserted between the bitstreams of at least two image blocks of the knowledge image, the preset syntax indicates that the knowledge image is not used for display. Based on the preset syntax, the knowledge image and / or the same content frames of the knowledge image are decoded.

22. A video decoding method, characterized in that, The method includes: A preset syntax is obtained from the video stream. Different values ​​of the preset syntax are used to indicate whether the knowledge image is used for display and whether to force the use of a skip mode to encode the same content frame of the knowledge image. The knowledge image and the same content frame correspond to the same original image content. If the knowledge image is transmitted in segments and if the bitstream of other images is inserted between the bitstreams of at least two image blocks of the knowledge image, the preset syntax indicates that the knowledge image is not used for display. Based on the preset syntax, the knowledge image and / or the same content frames of the knowledge image are decoded; Specifically, the forced use of the skip mode to encode identical content frames of the knowledge image means: setting the preset mode of all image blocks in the identical content frame to the skip mode, and disabling all filtering tools in the encoding process of the identical content frame.

23. The video decoding method according to claim 22, characterized in that, The step of decoding the knowledge image and / or the same content frame of the knowledge image based on the preset syntax includes: forcing the use of a skip mode to encode the same content frame of the knowledge image when the value of the preset syntax indicates that the knowledge image is decoded; and forcing the use of a skip mode to decode the same content frame based on the decoded image of the knowledge image to obtain the decoded image of the same content frame. The method further includes: controlling the display unit to display a decoded image of the same content frame.

24. An encoder, characterized in that, The encoder includes a processor; the processor is configured to execute instructions to implement the steps of the method as described in any one of claims 1-20.

25. A decoder, characterized in that, The decoder includes a processor; the processor is configured to execute instructions to implement the steps of the method as described in any one of claims 21-23.

26. A computer-readable storage medium storing instruction / program data thereon, characterized in that, When the instruction / program data is executed by the processor, it implements the steps of the method according to any one of claims 1-23.

Citation Information

Patent Citations

  • Frame reference method and device thereof, and storage medium

    CN111988626A

  • Video Encoder, Video Decoder, and Corresponding Method

    US20210337220A1