Task-oriented video semantic coding and decoding system

Through a multi-level reinforcement learning task-oriented video semantic codec system, artificial design masks are generated and the optimal codec mode is determined using reinforcement learning, which solves the problem of insufficient integration of semantic metrics in inter-frame codec in the prior art, and realizes efficient video compression and support for intelligent visual tasks.

CN120303942APending Publication Date: 2025-07-11DOUYIN VISION CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380083005.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-12-04
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing video encoding and decoding standards are difficult to optimize semantic metrics for intelligent vision tasks, especially in the inter-frame encoding and decoding process, resulting in insufficient video compression efficiency and quality.

Method used

A task-oriented video semantic codec system with multi-level reinforcement learning is adopted to generate manual design masks and task-oriented mode decision components, and gradually use reinforcement learning to determine the optimal codec mode, and combine the codec for video compression and decompression to simplify the complex mode decision space.

Benefits of technology

While maintaining semantic information, it significantly reduces compression costs and is suitable for intelligent visual tasks such as pose estimation, action recognition and object tracking, improving the efficiency and quality of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303942A_ABST
    Figure CN120303942A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding and decoding system for universal semantic compression. The video codec system includes a task-oriented mode decision component configured to receive as inputs an original video and a task-oriented semantic mask, and to stepwise utilize reinforcement learning to determine a task-oriented optimal codec mode. The video codec system also includes a codec configured to compress the original video into a bitstream based on the task-oriented optimal codec mode, or to decompress the bitstream into a reconstructed video based on the task-oriented optimal codec mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2022 / 136217, filed on December 2, 2022, the teachings and disclosures of which are incorporated herein by reference. Background Art

[0004] Digital video occupies the largest bandwidth used on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use may continue to grow. Summary of the Invention

[0005] A first aspect relates to a video codec system for general semantic compression, comprising: a task - oriented mode decision component configured to receive an original video and a task - oriented semantic mask as inputs, and gradually use reinforcement learning to determine an optimal task - oriented codec mode; and a codec configured to compress the original video into a bitstream based on the optimal task - oriented codec mode, or decompress the bitstream into a reconstructed video based on the optimal task - oriented codec mode.

[0006] A second aspect relates to a method implemented on a video codec system, the method comprising: determining an optimal task - oriented codec mode by gradually using reinforcement learning at a task - oriented mode decision component, wherein the task - oriented mode decision component uses the original video and a task - oriented semantic mask as inputs; and performing a conversion between visual media data and a bitstream based on the optimal task - oriented codec mode.

[0007] A third aspect relates to a device for processing video data, comprising: a processor; and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the video codec system according to the first aspect or execute the method according to the second aspect.

[0008] A fourth aspect relates to a non - transitory computer - readable medium, comprising a computer program product for use in a video codec device, the computer program product comprising computer - executable instructions stored on the non - transitory computer - readable medium, such that the computer - executable instructions, when executed by a processor, cause the video codec device to implement the video codec system according to the first aspect or execute the method according to the second aspect.

[0009] A fifth aspect relates to a non - transitory computer - readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining a task - oriented optimal codec mode by gradually utilizing reinforcement learning at a task - oriented mode decision component, wherein the task - oriented mode decision component uses the original video and a task - oriented semantic mask as inputs; and generating a bitstream based on the determination.

[0010] A sixth aspect relates to a method for storing a bitstream of a video, including: determining a task - oriented optimal codec mode by gradually utilizing reinforcement learning at a task - oriented mode decision component, wherein the task - oriented mode decision component uses the original video and a task - oriented semantic mask as inputs; generating a bitstream based on the determination; and storing the bitstream in a non - transitory computer - readable recording medium.

[0011] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0012] These and other features will be more clearly understood from the following detailed description in conjunction with the drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To understand the present disclosure more fully, reference is now made to the following brief description taken in conjunction with the drawings and the detailed description, in which like reference numerals represent like parts.

[0014] Figure 1 An example of a task - oriented video semantic codec system is shown.

[0015] Figure 2 An example of a task - oriented optimal mode decision part is shown.

[0016] Figure 3 A block diagram of an example of a video processing system is shown.

[0017] Figure 4 A block diagram of an example of a video processing device is shown.

[0018] Figure 5 A flowchart of an example of a video processing method is shown.

[0019] Figure 6 A block diagram of an example of a video codec system is shown.

[0020] Figure 7 A block diagram of an example of an encoder is shown.

[0021] Figure 8 A block diagram of an example of a decoder is shown.

[0022] Figure 9 It is a schematic diagram of an example of an encoder.

[0023] Figure 10 It is a flowchart of an example of a video processing method. Detailed implementation manners

[0024] First of all, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or to be developed. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified without departing from the entire scope of the appended claims and their equivalents.

[0025] The use of section headings in this document is for ease of understanding and does not limit the applicability of the technologies and embodiments disclosed in each section only to that section. In addition, the technologies described herein are also applicable to other video codec protocols and designs.

[0026] 1. Preliminary discussion

[0027] Video codec standards [14 - 17] are usually optimized for objective / perceptual metrics, such as the pixel-level metric peak signal-to-noise ratio (PSNR) [4 - 6] or the perceptual-level metric multi-scale SSIM (MS-SSIM) [7, 8]. However, with the development of artificial intelligence, the demand for video codecs that support video coding and decoding for intelligent vision tasks is increasing day by day, such as pose estimation [9], action recognition

[10] , and object tracking

[11] described in machine video coding (VCM) [12, 13]. However, the optimization objective, that is, the task-oriented semantic metric [1], cannot be directly integrated into existing video codec standards.

[0028] To address the above challenges, some pioneering works [1, 2] use reinforcement learning to integrate semantic metrics into traditional codecs. However, they only handle intra-frame coding and decoding, but not inter-frame coding and decoding, which leaves room for task-oriented semantic video coding and decoding. After that, Ref. [3] proposed hierarchical reinforcement learning for semantic video coding and decoding for video segmentation and verified its effectiveness and efficiency in video object segmentation problems. However, the hierarchical mode selection strategy simplifies the complex mode decision space and is therefore not very suitable for dynamic videos that require a more comprehensive design. Based on multi-level reinforcement learning that gradually and comprehensively selects task-oriented optimal modes, the present disclosure describes a task-oriented video semantic coding and decoding system. 2. Detailed implementation manners 2.1 Technical field

[0031] This disclosure belongs to the field of video processing technology. Specifically, a video coding and decoding scheme for task-oriented semantic information maintenance is proposed. While well preserving the task-oriented semantic information, the compression cost is significantly reduced. As described in VCM

[12] , it can be used for many intelligent vision tasks such as pose estimation [9], action recognition

[10] , and object tracking

[11] .

[0032] 2.2 Description of Related Examples

[0033] This disclosure aims at task-oriented video semantic compression. This disclosure describes an optimal mode selection strategy for task orientation based on multi-level reinforcement learning, as Figure 1 shown, this strategy seeks to achieve optimal rate-semantic-distortion optimization in a coarse-to-fine manner. Figure 1 An example of a task-oriented video semantic coding and decoding system is shown. Details are as follows.

[0034] First, an artificially designed mask of the video is generated to indicate semantic importance. Next, the task-oriented mode decision component takes the video and the mask as inputs to generate the best mode. Then, the codec compresses the video in the selected best mode and outputs a bitstream. After storage or transmission, the bitstream is decoded for downstream tasks such as video object segmentation, detection, tracking, or reconstruction. The dashed box represents an optional component.

[0035] Figure 2 An example of the task-oriented optimal mode decision part is shown. Details of the task-oriented mode decision component are as Figure 2 shown. The first obstacle is that the mode space is exponentially proportional to the number of frames and should be simplified. However, due to the introduction of complex temporal dependencies in video coding and decoding by inter-frame prediction, the distortion in the reference region will propagate to all subsequent regions. Therefore, the mode space in video coding and decoding cannot be simplified like in intra-frame coding and decoding. Adopting another method, this disclosure selects the best mode in a coarse-to-fine manner with the help of multi-level reinforcement learning. The higher-level agent generates the optimal mode, and the lower-level agent generates an optimal mode offset centered on the former. Each agent first extracts features and then uses reinforcement learning to represent the distortion propagation phenomenon introduced by inter-frame prediction, which will capture the ideal rate-distortion point that minimizes the rate and distortion costs. Finally, the gradually generated optimal mode offsets are added together and then indicated to the codec for compression.

[0036] As a supplement, to indicate semantic importance to the hierarchical agent, an artificially designed mask is first generated and fed to the agent together with the original video. Additionally, to train the feature extractor and the reinforcement agent, the rate (R) and task-related distortion (D) of the decompressed video that are calculated are assigned as rewards to each level.

[0037] 2.3 Description of Example Embodiments

[0038] As an embodiment of the present disclosure, the specific implementation is presented as follows.

[0039] Figure 1 An example of a task-oriented video semantic codec system is shown. First, an artificially designed mask of the original video is generated to indicate semantic importance. As used herein, the artificially designed mask can be any pre-selected and / or pre-configured mask, such as a task-oriented semantic mask. The task-oriented semantic mask can be an array, where each element can indicate the semantics of the sample points in the corresponding position. For example, if the task is to determine the best mode for a specific foreground object, the mask can indicate the position of the specific foreground object by indicating which sample points include the foreground object and which sample points do not include the foreground object. Thus, the mask can focus the best mode selection process on the sample points relevant to a specific task. The task-oriented semantic mask can be generated by a relevant neural network based on the corresponding task. For example, when the task is object detection, the object detection neural network can determine the position of the relevant object and generate a task-oriented semantic mask that removes all sample points not included in the object. As an example, the video can include an image of a bear in a forest. If the task is to determine the best mode for encoding and decoding the image of the bear, the object detection neural network can generate a task-oriented semantic mask that indicates all the sample points of the bear and indicates all the sample points of the forest. In this way, applying the mask to the original video can effectively remove the forest from consideration, leaving only the bear. In this way, any task can be represented by a task-oriented semantic mask that indicates the sample points that should be considered when determining the best mode for that task.

[0040] Next, the task-oriented mode decision component takes the video and the mask as inputs to generate the best mode. The task-oriented mode decision component is configured to select from one or more codec modes based on the task, which can be input or pre-configured. As an example, the task-oriented semantic mask can be applied to the original video to focus on the process of task-related samples. Then, the masked video can be fed into a convolutional neural network to allow the neural network to perform feature extraction from the masked video in order to select the best mode. Then, the codec compresses the video in the selected best mode and outputs a bitstream. After storage or transmission, the bitstream is decoded for downstream tasks such as video object segmentation, detection, tracking, or reconstruction. The dashed box represents an optional component.

[0041] Figure 2 An example of the task-oriented optimal mode decision part is shown, which can be used to implement Figure 1 the task-oriented mode decision component. Taking the video and the generated task-oriented semantic mask as inputs, the task-oriented semantic mask is applied to the video to obtain the masked video, and the selected task-oriented best mode for the masked video is determined and output. To simplify the decision mode space, the example divides the mode space in a coarse-to-fine manner with the help of multi-level reinforcement learning. The higher-level agent extracts global features from the masked video and outputs the best mode, while the lower-level agent extracts local features from the masked video and outputs the best mode offset centered on the former in a finer manner. Each agent consists of a feature extraction part and a reinforcement learning part, and the reinforcement learning part uses reinforcement learning to capture the best mode offset from the masked video. Then, the codec compresses the video in the selected best mode. After that, the collection rate and task-related distortion are used as rewards to gradually train the multi-level agents in a coarse-to-fine manner.

[0042] 3. Summary of the present disclosure

[0043] The present disclosure describes a general video semantic codec system that minimizes the task-oriented rate-distortion cost. Taking the video and the manually designed mask as inputs, the comprehensive task-oriented mode decision part simplifies the complex decision mode space in a coarse-to-fine manner and then uses hierarchical reinforcement learning agents to gradually select the optimal mode offset.

[0044] 4. List of solutions and embodiments

[0045] 1. A video codec system for general semantic compression, comprising: a manually designed mask generation part; a task-oriented mode decision part, which takes the original video and the corresponding mask as inputs and gradually uses reinforcement learning to determine the optimal codec mode offset for the task; and a codec. The encoder compresses the video into a bitstream under the selected best task-oriented codec mode, and then the decoder decompresses the bitstream into a reconstructed video. The reconstructed video is usually fed into downstream semantic tasks.

[0046] 1.1 The codec according to 1, wherein the codec conforms to a traditional codec standard.

[0047] 1.2 The downstream task according to 1, wherein the task is a semantic task, a reconstruction task, or a combination thereof.

[0048] 1.3 The traditional codec standard according to 1.1, wherein it is High Efficiency Video Coding (HEVC)

[15] , Versatile Video Coding (VVC)

[16] , a to-be-finalized standard (e.g., AVS3)

[17] , or a future video codec standard, etc.

[0049] 1.4 The semantic task according to 1.2, wherein the task is video object segmentation, video object tracking, action recognition, etc. or a combination thereof.

[0050] 1.5 The reconstruction task according to 1.2, wherein compared with the original video, the reconstructed video has minimized pixel-level or perceptual-level metrics.

[0051] 1.6 The pixel-level metric according to 1.4, wherein it is PSNR / Mean Squared Error (MSE), etc.

[0052] 1.7 The perceptual-level metric according to 1.4, wherein it is SSIM, MS-SSIM, Video Multimethod Assessment Fusion (VMAF), Learned Perceptual Image Patch Similarity (LPIPS), etc.

[0053] 2. A task-oriented mode decision part conforming to 1, comprising: N reinforcement learning agents, which output N best mode offsets in a coarse-to-fine manner; and a training strategy, which trains the N agents to output N best mode offsets in a coarse-to-fine manner.

[0054] 2.1 The task-oriented mode decision part according to 2, wherein N = 3.

[0055] 2.2 The task-oriented mode decision part according to 2.1, wherein the first-level agent outputs the best mode for a group of pictures (GOP), the second-level agent outputs the best mode offset for each frame in the GOP, and the third-level agent outputs the best mode offset for the background and foreground of each frame.

[0056] 2.3 The mode according to 2, wherein the mode assigns bitrates to each level.

[0057] 2.4 The mode according to 2, wherein the mode selects quantization parameters for each level.

[0058] 2.5 The mode according to 2, wherein the mode selects Lagrange multipliers for each level.

[0059] 3. A training strategy conforming to 2, which consists of the following parts: The first-level agent takes a video and its mask as inputs, extracts first-level coarse features, and then outputs the first-level best mode; the second-level agent takes a video and its mask as inputs, extracts second-level fine features, and then outputs the second-level best mode offset centered on the first-level best mode; the N-level agent takes a video and its mask as inputs, extracts the finer (N - 1)-level features, and then outputs the N-level best mode offset centered on the (N - 1)-level best mode. The N-level best mode offsets are collected to form the finest best mode for encoding and decoding, and then the codec compresses the original video in the best mode. The decompressed video is used for downstream tasks and the distortion and rate are collected; and the distortion and rate are allocated to agents at different levels to train the agents.

[0060] 3.1 The agent according to 3, wherein it is a reinforcement learning agent, such as a deep Q-network (DQN), an advantage actor-critic (A2C) algorithm, or an asynchronous advantage actor-critic (A3C) algorithm, etc.

[0061] 5. References

[0062] [1]Li X,Shi J,Chen Z.Task-driven semantic coding via reinforcementlearning[J].IEEE Transactions on Image Processing,2021,30:6307-6320.

[0063] [2] Shi J, Chen Z. Reinforced bit allocation under task-driven semantic distortion metrics[C] / / 2020 IEEE international symposium on circuits and systems(ISCAS). IEEE, 2020:1-5.

[0064] [3] Xie G, Li X, Lin S, et al. Hierarchical Reinforcement Learning Based Video Semantic Coding for Segmentation[J]. arXiv preprint arXiv:2208.11529, 2022.

[0065] [4] Hu J H, Peng W H, Chung C H. Reinforcement learning for HEVC / H.265 intra-frame rate control[C] / / 2018 IEEE International Symposium on Circuits and Systems(ISCAS). IEEE, 2018:1-5.

[0066] [5] Chen L C, Hu J H, Peng W H. Reinforcement learning for HEVC / H.265 frame-level bit allocation[C] / / 2018 IEEE 23rd International Conference on Digital Signal Processing(DSP). IEEE, 2018:1-5.

[0067] [6] Zhou M, Wei X, Kwong S, et al. Rate control method based on deep reinforcement learning for dynamic video sequences in HEVC[J]. IEEE Transactions on Multimedia, 2020, 23:1106-1121.

[0068] [7] Wang S, Rehman A, Wang Z, et al. Rate-SSIM optimization for videocoding[C] / / 2011 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2011:833-836.

[0069] [8] Dai W, Au O C, Zhu W, et al. SSIM-based rate-distortion optimization in H.264[C] / / 2014 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP). IEEE, 2014:7343-7347.

[0070] [9] Wang J, Sun K, Cheng T, et al. Deep high-resolution representation learning for visual recognition[J]. IEEE transactions on pattern analysis and machine intelligence, 2020, 43(10):3349-3364.

[0071]

[10] Feichtenhofer C, Fan H, Malik J, et al. Slowfast networks for video recognition[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2019:6202-6211.

[0072]

[11] Wang Z, Zheng L, Liu Y, et al. Towards real-time multi-object tracking[C] / / European Conference on Computer Vision. Springer, Cham, 2020:107-122.

[0073]

[12] Zhang Y. Video Coding for Machines[C] / / ITU Workshop on “The future of media. 2019.

[0074]

[13] Wood D. Task Oriented Video Coding: A Survey[J]. arXiv preprint arXiv:2208.07313, 2022.

[0075]

[14] Wiegand T, Sullivan G J, Bjontegaard G, et al. Overview of the H.264 / AVC video coding standard[J]. IEEE Transactions on circuits and systems for video technology, 2003, 13(7):560 - 576.

[0076]

[15] Sullivan G J, Ohm J R, Han W J, et al. Overview of the high - efficiency video coding(HEVC) standard[J]. IEEE Transactions on circuits and systems for video technology, 2012, 22(12):1649 - 1668.

[0077]

[16] Bross B, Wang Y K, Ye Y, et al. Overview of the versatile video coding(VVC) standard and its applications[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 31(10):3736 - 3764.

[0078]

[17] Fan L,Ma S,Wu F.Overview of AVS video standard[C] / / 2004IEEEInternational Conference on Multimedia and Expo(ICME)(IEEE Cat.No.04TH8763).IEEE,2004,1:423-426.

[0079] Figure 3 FIG. 4000 is a block diagram of an example of a video processing system in which the various techniques of the present disclosure may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0080] System 4000 may include an encoding component 4004 that may implement the various encoding or decoding methods described in this document. Encoding component 4004 may reduce the average bit rate of the video from input 4002 to the output of encoding component 4004 to produce an encoded representation of the video. Encoding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of encoding component 4004 may be stored or transmitted via a communication connection represented by component 4006. A bitstream (or encoded) representation of the video stored or communicated at input 4002 may be used by component 4008 to generate pixel values or a displayable video for transmission to display interface 4010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Additionally, while certain video processing operations are referred to as "encoding" operations or tools, it is understood that encoding tools or operations are used by an encoder, and the corresponding decoding tools or operations that reverse the encoding result will be performed by a decoder.

[0081] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Displayport, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0082] Figure 4 FIG. 1 is a block diagram of an example of a video processing apparatus 4100. The apparatus 4100 can be used to implement one or more methods described herein. The apparatus 4100 can be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The (multiple) processors 4102 can be configured to implement one or more methods described in this document. The (multiple) memories 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the video processing circuitry 4106 can be at least partially included in the processor 4102, for example, a graphics co-processor.

[0083] Figure 5 FIG. 2 is a flowchart of an example of a video processing method 4200. The method 4200 includes step 4202: determining to adopt a video codec system. The video codec system includes a manually designed mask generation component. The video codec system further includes a task-oriented mode decision component configured to receive an original video and a corresponding mask as inputs and progressively use reinforcement learning to determine a task-oriented optimal codec mode offset. The video codec system further includes an encoder configured to compress the original video into a bitstream in the determined task-oriented optimal codec mode. The video codec system further includes a decoder configured to decompress the bitstream into a reconstructed video, where the reconstructed video is fed into a downstream task. In step 4204, a conversion is performed between the visual media data and the bitstream based on the video codec system. The conversion in step 4204 can include encoding at the encoder or decoding at the decoder, depending on the example.

[0084] It should be noted that the method 4200 can be implemented in an apparatus for processing video data including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to execute the method 4200. Additionally, the method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium that, when executed by a processor, cause the video codec device to execute the method 4200.

[0085] Figure 6FIG. 0 shows a block diagram of an example of a video codec system 4300 that can utilize the techniques of the present disclosure. The video codec system 4300 can include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, where the source device 4310 can be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, where the destination device 4320 can be referred to as a video decoding device.

[0086] The source device 4310 can include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 can include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data can include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 4316 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be directly transmitted to the destination device 4320 via the I / O interface 4316 over a network 4330. The encoded video data can also be stored on a storage medium / server 4340 for access by the destination device 4320.

[0087] The destination device 4320 can include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 can include a receiver and / or a modem. The I / O interface 4326 can obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 can decode the encoded video data. The display device 4322 can display the decoded video data to a user. The display device 4322 can be integrated with the destination device 4320, or can be external to the destination device 4320, which can be configured to interface with an external display device.

[0088] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or further standards.

[0089] Figure 7 FIG. 13 shows a block diagram of an example of a video encoder 4400, which can be Figure 6Video encoder 4314 in the system 4300 shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0090] The functional components of video encoder 4400 can include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy encoding unit 4414. The prediction unit 4402 can include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406.

[0091] In other examples, video encoder 4400 can include more, fewer, or different functional components. In one example, the prediction unit 4402 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0092] In addition, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but are shown separately in the example of video encoder 4400 for purposes of explanation.

[0093] The segmentation unit 4401 can segment a picture into one or more video blocks. Video encoder 4400 and video decoder 4500 can support various video block sizes.

[0094] The mode selection unit 4403 can select, for example, one of a plurality of codec modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 4407 to generate residual block data, and provide it to the reconstruction unit 4412 to reconstruct the coded block to be used as a reference picture. In some examples, the mode selection unit 4403 can select an intra-inter combined prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 can also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0095] To perform inter prediction on a current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from the cache 4413 with the current video block. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and the decoded samples of pictures from the cache 4413 other than the picture associated with the current video block.

[0096] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.

[0097] In some examples, the motion estimation unit 4404 may perform uni-directional prediction on the current video block, and the motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 4404 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0098] In other examples, the motion estimation unit 4404 may perform bi-directional prediction on the current video block, the motion estimation unit 4404 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 4404 may then generate a reference index and a motion vector, the reference index indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and the motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0099] In some examples, the motion estimation unit 4404 may output a complete set of motion information for use in the decoding process of the decoder. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is similar enough to the motion information of neighboring video blocks.

[0100] In one example, the motion estimation unit 4404 may indicate a value in the syntax structure associated with the current video block to the video decoder 4500, where the value indicates that the current video block has the same motion information as another video block.

[0101] In another example, the motion estimation unit 4404 may indicate another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0102] As discussed above, the video encoder 4400 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0103] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0104] The residual generation unit 4407 may generate residual data for the current video block by subtracting the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0105] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform a subtraction operation.

[0106] The transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0107] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0108] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.

[0109] After the reconstruction unit 4412 reconstructs the video block, a loop filter operation may be performed to reduce blockiness artifacts in the video block.

[0110] The entropy encoding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0111] Figure 8 A block diagram showing an example of a video decoder 4500, the video decoder 4500 may be Figure 6 the video decoder 4324 in the system 4300 shown. The video decoder 4500 may be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0112] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 may perform a decoding process generally opposite to the encoding process described for the video encoder 4400.

[0113] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream can include entropy-coded video data (e.g., coded blocks of video data). The entropy decoding unit 4501 can decode the entropy-coded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 4502 can determine this information, for example, by performing AMVP and Merge modes.

[0114] The motion compensation unit 4502 can generate motion-compensated blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used with sub-pixel precision can be included in the syntax element.

[0115] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of video blocks to calculate the interpolation of sub-integer pixels for reference blocks. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 according to the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.

[0116] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information that describes how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0117] The intra prediction unit 4503 can form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0118] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. Then, the decoded video block is stored in the cache 4507, and the cache 4507 provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.

[0119] Figure 9It is a schematic diagram of an example of an encoder 4600. The encoder 4600 is applicable to the technology for implementing VVC. The encoder 4600 includes three loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses predefined filters, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets respectively and by applying a finite impulse response (FIR) filter, where the transcoded side information signals the offsets and the filter coefficients. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool for attempting to capture and repair the artifacts created by the previous stages.

[0120] The encoder 4600 further includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive the input video. The intra prediction component 4608 is configured to perform intra prediction, and the ME / MC component 4610 is configured to perform inter prediction using the reference pictures obtained from a reference picture buffer 4612. The residual blocks from the inter prediction or the intra prediction are fed to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, and the quantized residual transform coefficients are fed to an entropy coding component 4618. The entropy coding component 4618 performs entropy coding on the prediction results and the quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting the images to the DF 4602, the SAO 4604, and the ALF 4606 for filtering before these images are stored in the reference picture buffer 4612.

[0121] Figure 10It is a flowchart of an example of video processing method 4700. Method 4700 includes step 4702: selecting a task-oriented semantic mask based on a task. The mask can be selected to focus on task-related samples. The task-oriented semantic mask can be selected and / or generated by a relevant neural network. At step 4704, the task-oriented semantic mask and the original video are received at a first-level agent, a second-level agent, and an N-level agent in a task-oriented mode decision component. At step 4706, the task-oriented semantic mask is applied to the original video to obtain a masked video. Then, the agents extract coarse features at the first-level agent, fine features at the second-level agent, and N-1 level finer features at the N-level agent from the masked video. At step 4708, based on the offsets of the extracted features, reinforcement learning is gradually utilized at the task-oriented mode decision component to determine the task-oriented optimal codec mode. At step 4710, a conversion between visual media data and a bitstream is performed based on the task-oriented optimal codec mode. The conversion at step 4710 can include encoding at an encoder or decoding at a decoder, depending on the example.

[0122] It should be noted that method 4700 can be implemented in a device for processing video data including a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to execute method 4200. Additionally, method 4700 can be executed by a non-transitory computer-readable medium including a computer program product for use in a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, and the computer-executable instructions, when executed by a processor, cause the video codec device to execute method 4700.

[0123] Next, a list of some example preferred solutions is provided.

[0124] The following solutions illustrate examples of the techniques discussed herein.

[0125] 1. A video codec system for general semantic compression, including: a manually designed mask generation component; a task-oriented mode decision component configured to receive an original video and a corresponding mask as inputs and gradually utilize reinforcement learning to determine a task-oriented optimal codec mode offset; a codec for encoding configured to compress the original video into a bitstream in the determined task-oriented optimal codec mode; and a codec for decoding configured to decompress the bitstream into a reconstructed video, where the reconstructed video is fed into a downstream task.

[0126] 2. The video codec system according to Solution 1, wherein the codec complies with a codec standard.

[0127] 3. The video codec system according to Solution 1 or 2, wherein the downstream tasks include semantic tasks, reconstruction tasks, or a combination thereof.

[0128] 4. The video codec system according to any one of Solutions 1-3, wherein the codec standard is High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Audio Video Standard 3 (AVS3), or a combination thereof.

[0129] 5. The video codec system according to any one of Solutions 1-4, wherein the downstream tasks include video object segmentation, video object tracking, action recognition, or a combination thereof.

[0130] 6. The video codec system according to any one of Solutions 1-5, wherein the downstream tasks include a reconstruction task, and compared with the original video, the reconstructed video has minimized pixel-level or perceptual-level metrics.

[0131] 7. The video codec system according to any one of Solutions 1-6, wherein a pixel-level metric is adopted, and the pixel-level metric includes Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), or a combination thereof.

[0132] 8. The video codec system according to any one of Solutions 1-7, wherein a perceptual-level metric is adopted, and the perceptual-level metric includes Structural Similarity Index Measure (SSIM), Multi-Scale SSIM (MS-SSIM), Video Multimethod Assessment Fusion (VMAF), Learned Perceptual Image Patch Similarity (LPIPS), or a combination thereof.

[0133] 9. The video codec system according to any one of Solutions 1-8 further includes a task-oriented mode decision component, and the task-oriented mode decision component includes: N reinforcement learning agents configured to output N best mode offsets in a coarse-to-fine manner; and a training policy component configured to train the N agents to output N best mode offsets in a coarse-to-fine manner.

[0134] 10. The video codec system according to any one of Solutions 1-9, wherein N = 3.

[0135] 11. The video codec system according to any one of Solutions 1-10, wherein the first-level agent outputs the best mode for a group of pictures (GOP), the second-level agent outputs the best mode offset for each frame in the GOP, and the third-level agent outputs the best mode offset for the background and foreground of each frame.

[0136] 12. The video codec system according to any one of Solutions 1-11, wherein the mode is allocated based on the bit rate of each level.

[0137] 13. The video codec system according to any one of Solutions 1-12, wherein the mode is selected based on the quantization parameter of each level.

[0138] 14. The video codec system according to any one of Solutions 1-13, wherein the mode is selected based on the Lagrange multiplier of each level.

[0139] 15. The video codec system according to any one of Solutions 1-14, wherein the first-level agent is configured to receive a video and a mask as inputs, extract first-level coarse features, and then output the first-level best mode; the second-level agent is configured to receive the video and the mask as inputs, extract second-level fine features, and output the second-level best mode offset centered on the first-level best mode; the N-level agent is configured to receive the video and the mask as inputs, extract the finer (N-1)-level features, and output the N-level best mode offset centered on the (N-1)-level best mode; the N-level best mode offsets are collected to form the finest best mode for encoding and decoding; the codec is configured to use the finest best mode to compress the original video; the decompressed video is used for downstream tasks, distortion acquisition, and rate acquisition; and the distortion and rate are allocated to agents at different levels to train agents at different levels.

[0140] 16. The video codec system according to any one of Solutions 1-15, wherein the agent includes a reinforcement learning agent, and the reinforcement learning agent includes a deep Q-network (DQN), an advantage actor-critic (A2C) algorithm, or an asynchronous advantage actor-critic (A3C) algorithm.

[0141] 17. A method includes: determining to adopt the video codec system according to any one of Solutions 1-16; and performing a conversion between visual media data and a bitstream based on the video codec system.

[0142] 18. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the video codec system according to any one of Solutions 1-16.

[0143] 19. A non-transitory computer-readable medium, comprising a computer program product for use in a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, the computer-executable instructions, when executed by a processor, causing the video codec device to implement the video codec system according to any one of Solutions 1-16.

[0144] 20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining to adopt the video codec system according to any one of Solutions 1-16; and generating the bitstream based on the determination.

[0145] 21. A method for storing a bitstream of a video, comprising: determining to adopt the video codec system according to any one of Solutions 1-16; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0146] 22. A method, apparatus, or system described in this document.

[0147] The following solutions show other examples of the technology discussed herein.

[0148] 1. A video codec system for general semantic compression, comprising: a task-oriented mode decision component configured to receive an original video and a task-oriented semantic mask as inputs and gradually use reinforcement learning to determine an optimal task-oriented codec mode; and a codec configured to compress the original video into a bitstream based on the optimal task-oriented codec mode, or decompress the bitstream into a reconstructed video based on the optimal task-oriented codec mode.

[0149] 2. The video codec system according to Solution 1, further comprising a mask generation component configured to select a task-oriented semantic mask based on a task.

[0150] 3. The video codec system according to Solution 1 or 2, wherein the codec complies with a codec standard.

[0151] 4. The video codec system according to any one of Solutions 1-3, wherein the codec standard is High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Audio Video Standard 3 (AVS3), or a combination thereof.

[0152] 5. The video codec system according to any one of Solutions 1-4, wherein the reconstructed video is fed into a downstream task including a semantic task, a reconstruction task, or a combination thereof.

[0153] 6. The video codec system according to any one of Solutions 1-5, wherein the reconstructed video is fed into a downstream task including video object segmentation, video object tracking, action recognition, or a combination thereof.

[0154] 7. The video codec system according to any one of Solutions 1-6, wherein the reconstructed video is fed into a downstream task including a reconstruction task, and compared with the original video, the reconstructed video has a minimized pixel-level metric or perceptual-level metric.

[0155] 8. The video codec system according to any one of Solutions 1-7, wherein the pixel-level metric includes Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), or a combination thereof.

[0156] 9. The video codec system according to any one of Solutions 1-8, wherein the perceptual-level metric includes Structural Similarity Index Measure (SSIM), Multi-Scale SSIM (MS-SSIM), Video Multimethod Assessment Fusion (VMAF), Learned Perceptual Image Patch Similarity (LPIPS), or a combination thereof.

[0157] 10. The video codec system according to any one of Solutions 1-9, wherein the task-oriented mode decision component includes N reinforcement learning agents, and the N reinforcement learning agents are configured to output N best mode offsets for the original video in a coarse-to-fine manner based on the task-oriented semantic mask.

[0158] 11. The video codec system according to any one of Solutions 1-10, wherein the task-oriented mode decision component further includes a training policy component, and the training policy component is configured to train the N agents to output the N best mode offsets in a coarse-to-fine manner.

[0159] 12. The video codec system according to any one of Solutions 1-11, wherein N = 3.

[0160] 13. The video codec system according to any one of Solutions 1-12, wherein the first-level agent outputs the best mode for a group of pictures (GOP), the second-level agent outputs the best mode offset for each frame in the GOP, and the third-level agent outputs the best mode offset for the background and foreground of each frame.

[0161] 14. The video codec system according to any one of Solutions 1-13, wherein the best mode is allocated based on the bit rate of each level of agent, the best mode is selected based on the quantization parameter of each level of agent, or the best mode is selected based on the Lagrange multiplier of each level of agent.

[0162] 15. The video codec system according to any one of Solutions 1-14, wherein the first-level agent is configured to receive the original video and the task-oriented semantic mask as inputs, extract first-level coarse features from the original video based on the task-oriented semantic mask, and then output the first-level best mode based on the first-level coarse features; the second-level agent is configured to receive the original video and the task-oriented semantic mask as inputs, extract second-level fine features from the original video based on the task-oriented semantic mask, and output a second-level best mode offset based on the second-level fine features centered on the first-level best mode; the Nth-level agent is configured to receive the original video and the task-oriented semantic mask as inputs, extract (N-1)th-level finer features from the original video based on the task-oriented semantic mask, and output an Nth-level best mode offset based on the (N-1)th-level finer features centered on the (N-1)th-level best mode; the Nth-level best mode offsets are collected to form the finest best mode for encoding and decoding; the codec is configured to use the finest best mode to compress the original video; the decompressed video is used for downstream tasks, distortion acquisition, and rate acquisition; and the distortion and the rate are allocated to different levels of agents to train the different levels of agents.

[0163] 16. The video codec system according to any one of Solutions 1-15, wherein the learning agent is a reinforcement learning agent, and the reinforcement learning agent includes a deep Q-network (DQN), an advantage actor-critic (A2C) algorithm, or an asynchronous advantage actor-critic (A3C) algorithm.

[0164] 17. A method implemented on a video codec system, the method comprising: progressively determining an optimal task-oriented codec mode by leveraging reinforcement learning at a task-oriented mode decision component, wherein the task-oriented mode decision component uses the original video and a task-oriented semantic mask as inputs; and performing a conversion between visual media data and a bitstream based on the task-oriented optimal codec mode.

[0165] 18. The method according to claim 17, further comprising: compressing the original video into a bitstream based on the task-oriented optimal codec mode, or decompressing the bitstream into a reconstructed video based on the task-oriented optimal codec mode.

[0166] 19. The method according to solution 17 or 18, further comprising: selecting a task-oriented semantic mask based on a task.

[0167] 20. The method according to any one of solutions 17-19, wherein the codec complies with a codec standard, and the codec standard is High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Audio Video Standard 3 (AVS3), or a combination thereof.

[0168] 21. The method according to any one of solutions 17-20, wherein the reconstructed video is fed into a downstream task including a semantic task, a reconstruction task, video object segmentation, video object tracking, action recognition, or a combination thereof.

[0169] 22. The method according to any one of solutions 17-21, wherein the reconstructed video is fed into a downstream task including a reconstruction task, and compared with the original video, the reconstructed video has a minimized pixel-level metric or perceptual-level metric.

[0170] 23. The method according to any one of solutions 17-22, wherein the pixel-level metric includes Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), or a combination thereof.

[0171] 24. The method according to any one of solutions 17-23, wherein the perceptual-level metric includes Structural Similarity Index Measure (SSIM), Multi-Scale SSIM (MS-SSIM), Video Multimethod Assessment Fusion (VMAF), Learned Perceptual Image Patch Similarity (LPIPS), or a combination thereof.

[0172] 25. The method according to any one of Solutions 17-24, wherein the task-oriented mode decision component includes N reinforcement learning agents, and the N reinforcement learning agents are configured to output N best mode offsets for the original video in a coarse-to-fine manner based on the task-oriented semantic mask.

[0173] 26. The method according to any one of Solutions 17-25, wherein the task-oriented mode decision component further includes a training policy component, and the training policy component is configured to train the N agents to output the N best mode offsets in a coarse-to-fine manner.

[0174] 27. The method according to any one of Solutions 17-26, wherein N = 3.

[0175] 28. The method according to any one of Solutions 17-27, wherein the first-level agent outputs the best mode for a group of pictures (GOP), the second-level agent outputs the best mode offset for each frame in the GOP, and the third-level agent outputs the best mode offset for the background and foreground of each frame.

[0176] 29. The method according to any one of Solutions 17-28, wherein the best mode is allocated based on the bit rate of each level of agent, the best mode is selected based on the quantization parameter of each level of agent, or the best mode is selected based on the Lagrange multiplier of each level of agent.

[0177] 30. The method according to any one of Solutions 17 - 29 further includes: receiving the original video and the task - oriented semantic mask as inputs at a first - level agent, extracting first - level coarse features from the original video based on the task - oriented semantic mask, and then outputting a first - level best mode based on the first - level coarse features; receiving the original video and the task - oriented semantic mask as inputs at a second - level agent, extracting second - level fine features from the original video based on the task - oriented semantic mask, and outputting a second - level best mode offset based on the second - level fine features centered on the first - level best mode; and receiving the original video and the task - oriented semantic mask as inputs at an N - th level agent, extracting (N - 1) - th level finer features from the original video based on the task - oriented semantic mask, and outputting an N - th level best mode offset based on the (N - 1) - th level finer features centered on the (N - 1) - th level best mode; wherein, the N - th level best mode offsets are collected to form the finest best mode for encoding and decoding, the codec uses the finest best mode to compress the original video, the decompressed video is used for downstream tasks, distortion acquisition, and rate acquisition, and the distortion and the rate are assigned to different levels of agents to train the different levels of agents.

[0178] 31. The method according to any one of Solutions 17 - 30, wherein the learning agent is a reinforcement - learning agent, and the reinforcement - learning agent includes a Deep Q - Network (DQN), an Advantage Actor - Critic (A2C) algorithm, or an Asynchronous Advantage Actor - Critic (A3C) algorithm.

[0179] 32. The method according to any one of Solutions 1 - 31, wherein the conversion includes encoding the visual media data into the bitstream.

[0180] 33. The method according to any one of Solutions 1 - 31, wherein the conversion includes decoding the visual media data from the bitstream.

[0181] 34. A device for processing video data includes: a processor; and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the video codec system according to any one of Solutions 1 - 16 or execute the method according to any one of Solutions 17 - 33.

[0182] 35. A non-temporary computer-readable medium, comprising a computer program product for use by a video codec device, wherein the computer program product comprises computer executable instructions stored on the non-temporary computer-readable medium, and the computer executable instructions, when executed by a processor, enable the video codec device to implement a video codec system according to any one of Solutions 1-16 or execute a method according to any one of Solutions 17-33.

[0183] 36. A non-temporary computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method comprises: determining a task-oriented optimal encoding and decoding mode by gradually utilizing reinforcement learning at a task-oriented mode decision component, wherein the task-oriented mode decision component uses an original video and a task-oriented semantic mask as input; and generating the bitstream based on the determination.

[0184] 37. A method for storing a bitstream of a video, comprising: determining a task-oriented optimal encoding and decoding mode by gradually utilizing reinforcement learning at a task-oriented mode decision component, wherein the task-oriented mode decision component uses an original video and a task-oriented semantic mask as input; generating the bitstream based on the determination; and storing the bitstream in a non-temporary computer-readable recording medium.

[0185] In the described solution, an encoder can comply with the format rules by generating a coded representation according to the format rules. In the described solution, a decoder can parse syntax elements in the coded representation according to the format rules using known information of the presence and absence of syntax elements to generate decoded video.

[0186] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may correspond, for example, to bits propagated at the same or different positions within the bitstream defined by the syntax. For example, a macroblock may be encoded based on a transformed and coded error residual value and may also use bits in the header and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on the determination, using known information that some fields may or may not exist, as described in the above solution. Similarly, the encoder may determine whether to include or not include a specific syntax field, and generate the coded representation accordingly by including or excluding the syntax field from the coded representation.

[0187] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. A computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” includes all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the relevant computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information to be transmitted to a suitable receiver apparatus.

[0188] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including a compiled or interpreted language, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program need not correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple co-related files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0189] The processing and logical flows described in this document can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0190] A processor suitable for executing a computer program includes, for example, general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks such as internal hard disks or removable hard disks; magneto-optical disks; and compact disc read-only memory (CD ROM) and digital versatile disc read-only memory (DVD-ROM) discs. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0191] Although this patent document contains many details, these details should not be construed as limitations on any subject or the scope of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Moreover, although features may operate in certain combinations as described above, and even were initially claimed in such a manner, in some cases, one or more features in a claimed combination may be excluded from that combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.

[0192] Similarly, although operations are depicted in the figures in a particular order, this should not be understood to require that such operations be performed in the particular order or sequence shown, or that all of the illustrated operations be performed to achieve a desired result. Additionally, the partitioning of various system components in the embodiments described in this patent document should not be understood to require such partitioning in all embodiments.

[0193] Only a few implementations and examples have been described, and other implementations, improvements, and variations may be made based on what is described and illustrated in this patent document.

[0194] When there is no intermediate component other than the line, trace, or other medium between the first component and the second component, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component other than the line, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupled" and its variants include both direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" means a range of ± 10% of the subsequent number.

[0195] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples should be considered illustrative rather than restrictive and are not intended to be limited to the details given herein. For example, the various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.

[0196] In addition, the techniques, systems, subsystems, and methods described and shown as discrete or separate in the various embodiments may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the disclosure. Other items shown or discussed as being coupled may be directly connected or may also be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Those skilled in the art can identify other examples of changes, substitutions, and alterations and can make them without departing from the spirit and scope disclosed herein.

Claims

1. A video codec system for general semantic compression, comprising: A task-oriented mode decision component, configured to receive an original video and a task-oriented semantic mask as inputs, and gradually use reinforcement learning to determine an optimal task-oriented codec mode; And A codec, configured to compress the original video into a bitstream based on the optimal task-oriented codec mode, or decompress the bitstream into a reconstructed video based on the optimal task-oriented codec mode.

2. The video codec system according to claim 1, further comprising a mask generation component configured to select a task-oriented semantic mask based on a task.

3. The video codec system according to claim 1 or 2, wherein The codec complies with a codec standard.

4. The video codec system according to any one of claims 1-3, wherein, The codec standard is High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Audio Video Standard 3 (AVS3), or a combination thereof.

5. The video codec system according to any one of claims 1-4, wherein, The reconstructed video is fed into a downstream task including a semantic task, a reconstruction task, or a combination thereof.

6. The video codec system according to any one of claims 1-5, wherein, The reconstructed video is fed into a downstream task including video object segmentation, video object tracking, action recognition, or a combination thereof.

7. The video codec system according to any one of claims 1-6, wherein, The reconstructed video is fed into a downstream task including a reconstruction task, and compared with the original video, the reconstructed video has a minimized pixel-level metric or perceptual-level metric.

8. The video codec system according to any one of claims 1-7, wherein, The pixel-level metric includes Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), or a combination thereof.

9. The video codec system according to any one of claims 1-8, wherein, The perceptual-level metric includes Structural Similarity Index Measure (SSIM), Multi-Scale SSIM (MS-SSIM), Video Multimethod Assessment Fusion (VMAF), Learned Perceptual Image Patch Similarity (LPIPS), or a combination thereof.

10. The video codec system according to any one of claims 1-9, wherein, The task-oriented mode decision component includes N reinforcement learning agents, and the N reinforcement learning agents are configured to output N best mode offsets for the original video in a coarse-to-fine manner based on the task-oriented semantic mask.

11. The video codec system according to any one of claims 1-10, wherein, The task-oriented mode decision component further includes a training strategy component, and the training strategy component is configured to train the N agents to output the N best mode offsets in a coarse-to-fine manner.

12. The video codec system according to any one of claims 1-11, wherein, N=3。 13. The video codec system according to any one of claims 1-12, wherein, The first-level agent outputs the best mode for a Group of Pictures (GOP), the second-level agent outputs the best mode offset for each frame in the GOP, and the third-level agent outputs the best mode offset for the background and foreground of each frame.

14. The video codec system according to any one of claims 1-13, wherein, The best mode is assigned based on the bitrate of each level of agent, the best mode is selected based on the quantization parameter of each level of agent, or the best mode is selected based on the Lagrange multiplier of each level of agent.

15. The video codec system according to any one of claims 1-14, wherein, The first-level agent is configured to receive the original video and the task-oriented semantic mask as inputs, extract first-level coarse features from the original video based on the task-oriented semantic mask, and then output a first-level best mode based on the first-level coarse features. Among them, the second-level agent is configured to receive the original video and the task-oriented semantic mask as inputs, extract second-level fine features from the original video based on the task-oriented semantic mask, and output a second-level best-mode offset based on the second-level fine features centered on the first-level best mode. Among them, the N-level agent is configured to receive the original video and the task-oriented semantic mask as inputs, extract finer (N - 1)-level features from the original video based on the task-oriented semantic mask, and output an N-level best-mode offset based on the (N - 1)-level finer features centered on the (N - 1)-level best mode. Among them, the N-level best-mode offsets are collected to form the finest best mode for encoding and decoding. Among them, the codec is configured to use the finest best mode to compress the original video. Among them, the decompressed video is used for downstream tasks, distortion acquisition, and rate acquisition, and Among them, the distortion and the rate are allocated to agents at different levels to train the agents at different levels.

16. The video codec system according to any one of claims 1-15, wherein, The learning agent is a reinforcement learning agent, and the reinforcement learning agent includes a Deep Q-Network (DQN), an Advantage Actor-Critic (A2C) algorithm, or an Asynchronous Advantage Actor-Critic (A3C) algorithm.

17. A method implemented on a video codec system, comprising: Determining a task-oriented optimal codec mode by gradually leveraging reinforcement learning at a task-oriented mode decision component, where the task-oriented mode decision component uses the original video and a task-oriented semantic mask as inputs; and Performing a conversion between visual media data and a bitstream based on the task-oriented optimal codec mode.

18. The method according to claim 17, further comprising: Compressing the original video into a bitstream based on the task-oriented optimal codec mode, or decompressing the bitstream into a reconstructed video based on the task-oriented optimal codec mode.

19. The method according to claim 17 or 18, further comprising: Selecting a task-oriented semantic mask based on the task.

20. The method according to any one of claims 17-19, wherein The codec complies with a codec standard, and the codec standard is High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Audio Video Standard 3 (AVS3), or a combination thereof.

21. The method according to any one of claims 17 - 20, wherein, The reconstructed video is fed into downstream tasks including semantic tasks, reconstruction tasks, video object segmentation, video object tracking, action recognition, or a combination thereof.

22. The method according to any one of claims 17 - 21, wherein, The reconstructed video is fed into downstream tasks including a reconstruction task, and compared with the original video, the reconstructed video has minimized pixel-level metrics or perceptual-level metrics.

23. The method according to any one of claims 17-22, wherein, The pixel-level metrics include Peak Signal-to-Noise Ratio (PSNR), Mean Squared Error (MSE), or a combination thereof.

24. The method according to any one of claims 17-23, wherein, The perceptual-level metrics include Structural Similarity Index Measure (SSIM), Multi-Scale SSIM (MS-SSIM), Video Multimethod Assessment Fusion (VMAF), Learned Perceptual Image Patch Similarity (LPIPS), or a combination thereof.

25. The method according to any one of claims 17-24, wherein The task-oriented mode decision component includes N reinforcement learning agents, and the N reinforcement learning agents are configured to output N optimal mode offsets for the original video in a coarse-to-fine manner based on the task-oriented semantic mask.

26. The method according to any one of claims 17-25, wherein, The task-oriented mode decision component further includes a training policy component, and the training policy component is configured to train the N agents to output the N optimal mode offsets in a coarse-to-fine manner.

27. The method according to any one of claims 17 - 26, wherein, N=3。 28. The method according to any one of claims 17 to 27, wherein The first-level agent outputs the optimal mode for a group of pictures (GOP), the second-level agent outputs the optimal mode offset for each frame in the GOP, and the third-level agent outputs the optimal mode offset for the background and foreground of each frame.

29. The method according to any one of claims 17-28, wherein, The optimal mode is allocated based on the bit rate of each level of agent, the optimal mode is selected based on the quantization parameter of each level of agent, or the optimal mode is selected based on the Lagrange multiplier of each level of agent.

30. The method according to any one of claims 17-29, further comprising: Receiving the original video and the task-oriented semantic mask as inputs at the first-level agent, extracting first-level coarse features from the original video based on the task-oriented semantic mask, and then outputting a first-level optimal mode based on the first-level coarse features; Receiving the original video and the task-oriented semantic mask as inputs at the second-level agent, extracting second-level fine features from the original video based on the task-oriented semantic mask, and outputting a second-level optimal mode offset based on the second-level fine features centered on the first-level optimal mode; And Receiving the original video and the task-oriented semantic mask as inputs at the Nth-level agent, extracting (N-1)th-level finer features from the original video based on the task-oriented semantic mask, and outputting an Nth-level optimal mode offset based on the (N-1)th-level finer features centered on the (N-1)th-level optimal mode, wherein the Nth-level optimal mode offsets are collected to form the finest optimal mode for encoding and decoding, wherein the codec uses the finest optimal mode to compress the original video, wherein the decompressed video is used for downstream tasks, distortion collection, and rate collection, and wherein the distortion and the rate are allocated to different levels of agents to train the different levels of agents.

31. The method according to any one of claims 17 - 30, wherein, The learning agent is a reinforcement learning agent, and the reinforcement learning agent includes a deep Q network (DQN), an advantage actor-critic (A2C) algorithm, or an asynchronous advantage actor-critic (A3C) algorithm.

32. The method according to any one of claims 1-31, wherein, The conversion includes encoding the visual media data into the bitstream.

33. The method according to any one of claims 1-31, wherein, The conversion includes decoding the visual media data from the bitstream.

34. An apparatus for processing video data, comprising: A processor; And A non-transitory memory having instructions thereon, wherein, When executed by the processor, the instructions cause the processor to implement the video codec system according to any one of claims 1-16 or execute the method according to any one of claims 17-33.

35. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, wherein the computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, and when executed by a processor, the computer-executable instructions cause the video codec device to implement the video codec system according to any one of claims 1-16 or execute the method according to any one of claims 17-33.

36. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a video processing device, wherein, The method includes: Determining an optimal task-oriented codec mode by progressively leveraging reinforcement learning at a task-oriented mode decision component, wherein the task-oriented mode decision component uses the original video and a task-oriented semantic mask as inputs; and Generating the bitstream based on the determination.

37. A method for storing a bitstream of a video, comprising: Determining an optimal task-oriented codec mode by progressively leveraging reinforcement learning at a task-oriented mode decision component, wherein the task-oriented mode decision component uses the original video and a task-oriented semantic mask as inputs; Generating the bitstream based on the determination; and Storing the bitstream in a non-transitory computer-readable recording medium.

Citation Information

Cited By

  • Self-adaptive decoding method and system for compressed image file

    CN121567875A