Video compression processing method and apparatus, and electronic device
By using the multi-functional target neural network model, the unification of multiple video compression processing functions is achieved, the complexity problem caused by the single function of the neural network model is solved, and the efficiency of video compression processing is improved.
Patent Information
- Application Number
- PCT/CN2024/141281
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the neural network model used for video compression processing has a relatively single function, resulting in a complex video compression processing process.
The target neural network model with multiple video compression processing functions is adopted, and a variety of video compression processing functions are realized through a unified neural network model, including loop filtering, super resolution and post-processing filtering.
It improves the functional diversity of the neural network model of video compression processing and reduces the complexity of video compression processing.
Smart Images

Figure CN2024141281_03072025_PF_FP_ABST
Abstract
Description
Video compression processing method, device and electronic equipment
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311843781.6 filed in China on December 28, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application belongs to the field of video compression technology, and specifically relates to a video compression processing method, device and electronic equipment. Background Art
[0004] In related art, different neural network models are required to achieve different video compression processing functions. For example, loop filtering requires a neural network-based loop filter, while super-resolution requires a neural network-based super-resolution processor. Therefore, the neural network models used for video compression processing in related art are relatively simple in function. Summary of the Invention
[0005] The embodiments of the present application provide a video compression processing method, device, and electronic device, which can solve the problem in related technologies that the neural network model used for video compression processing has relatively single functions.
[0006] In a first aspect, a video compression processing method is provided, which is performed by an encoding end, and the method includes:
[0007] Inputting the first image into a target neural network model, wherein the target neural network model is a neural network model having N video compression processing functions, where N is an integer greater than 1;
[0008] Performing a first processing on the first image based on the target neural network model to obtain a second image;
[0009] The video compression processing method corresponding to the first processing is implemented by one or more of N video compression processing functions.
[0010] In a second aspect, a video compression processing method is provided, which is performed by a decoding end, and the method includes:
[0011] Inputting the first image into a target neural network model, wherein the target neural network model is a neural network model having N video compression processing functions, where N is an integer greater than 1;
[0012] Performing a first processing on the first image based on the target neural network model to obtain a second image;
[0013] The video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0014] In a third aspect, a video compression processing device is provided, which is applied to an encoding end. The device includes:
[0015] A first input module is configured to input a first image into a target neural network model, wherein the target neural network model is a neural network model having N video compression processing functions, where N is an integer greater than 1;
[0016] A first processing module, configured to perform a first processing on the first image based on the target neural network model to obtain a second image;
[0017] The video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0018] In a fourth aspect, a video compression processing device is provided, which is applied to a decoding end, and includes:
[0019] A first input module is configured to input a first image into a target neural network model, wherein the target neural network model is a neural network model having N video compression processing functions, where N is an integer greater than 1;
[0020] A first processing module, configured to perform a first processing on the first image based on the target neural network model to obtain a second image;
[0021] The video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0022] In a fifth aspect, an electronic device is provided, which terminal includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0023] In a sixth aspect, an electronic device is provided, comprising a processor and a communication interface, wherein the processor is used to: input a first image into a target neural network model, wherein the target neural network model is a neural network model with N types of video compression processing functions, where N is an integer greater than 1; perform a first processing on the first image based on the target neural network model to obtain a second image; wherein the video compression processing method corresponding to the first processing is determined by the target neural network model.
[0024] In a seventh aspect, an electronic device is provided, comprising: a memory configured to store video data, and a processing circuit configured to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0025] In an eighth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0026] In the ninth aspect, a coding and decoding system is provided, comprising: a coding end device and a decoding end device, wherein the coding end device can be used to execute the steps of the method described in the first aspect, and the decoding end device can be used to execute the steps of the method described in the second aspect.
[0027] In the tenth aspect, a chip is provided, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0028] In the eleventh aspect, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0029] In an embodiment of the present application, a first image is input into a target neural network model, and a first process is performed on the first image based on the target neural network model to obtain a second image. The target neural network model is a neural network model having N types of video compression processing functions, and the video compression processing method corresponding to the first process is implemented by one or more of the N types of video compression processing functions. In this way, the functional diversity of the video compression processing neural network model can be increased, thereby enabling multiple video compression processing functions to be implemented through a unified neural network model, which helps reduce the complexity of video compression processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1 is a schematic diagram of a coding and decoding system provided in an embodiment of the present application;
[0031] FIG2 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;
[0032] FIG3 is a schematic structural diagram of a decoder provided in an embodiment of the present application;
[0033] FIG4 is a flow chart of a video compression processing method provided by an embodiment of the present application;
[0034] FIG5 is a schematic diagram of Example 1 provided in the embodiments of the present application;
[0035] FIG6 is a schematic diagram of Example 2 provided in the embodiments of the present application;
[0036] FIG7 is a schematic diagram of Example 3 provided in the embodiments of the present application;
[0037] FIG8 is a structural diagram of a video compression processing device provided in an embodiment of the present application;
[0038] FIG9 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0039] FIG10 is a structural diagram of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0041] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0042] FIG1 is a schematic diagram of a codec system 10 provided in an embodiment of the present application. The technical solution of the embodiment of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) of video data. The video data includes original unencoded video, encoded video, decoded (e.g., reconstructed) video, or syntax elements.
[0043] As shown in FIG1 , a codec system 10 includes a source device 100 that provides encoded video data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides the video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (e.g., a smartwatch or a wearable camera), a television, a camera, a display device, an in-vehicle device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an aircraft, a robot, a satellite, and the like.
[0044] In the example of Figure 1, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, while the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 may not include some of the components in Figure 1, or may also include other components outside of Figure 1. For example, the source device 100 can receive video data from an external data source (such as an external camera). Similarly, the destination device 110 can be connected to an external display device interface without including an integrated display device. For another example, the memory 102 and the memory 113 can be external memories.
[0045] Although FIG1 illustrates source device 100 and destination device 110 as separate devices, in some examples, the two may be integrated into a single device. In such embodiments, the functions corresponding to source device 100 and the functions corresponding to destination device 110 may be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.
[0046] In some examples, source device 100 and destination device 110 can perform one-way video transmission or two-way video transmission. If it is two-way video transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, that is, each of source device 100 and destination device 110 includes an encoder and a decoder.
[0047] Data source 101 represents a source of video data (i.e., raw, unencoded video data) and provides a series of pictures containing the video data to encoder 200, which encodes the picture data. Data source 101 of source device 100 may include a video capture device (such as a video camera), a video archive containing previously captured raw video, or a video feed interface for receiving video from a video content provider. Alternatively, data source 101 may generate computer graphics-based data as the source video, or combine real-time video, archived video, and computer-generated video. In these cases, encoder 200 encodes the captured, pre-captured, or computer-generated video data. Encoder 200 may rearrange the pictures from the order in which they were received (sometimes referred to as "display order") into an encoding order. Encoder 200 may generate a bitstream comprising the encoded video data. Source device 100 may then output the encoded video data to communication medium 120 via output interface 104 for receipt or retrieval, for example, by input interface 111 of destination device 110.
[0048] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general purpose memories. In some examples, the memory 102 may store raw video data from the data source 101, and the memory 113 may store decoded video data from the decoder 300. Additionally or alternatively, the memories 102 and 113 may store software instructions executable by, for example, the encoder 200 and the decoder 300, respectively. Although the memory 102 and the memory 113 are shown separately from the encoder 200 and the decoder 300 in this example, it should be understood that the encoder 200 and the decoder 300 may also include internal memory for functionally similar or equivalent purposes. If the encoder 200 and the decoder 300 are arranged on the same hardware device, the memory 102 and the memory 113 may be the same memory. Furthermore, the memories 102 and 113 may store, for example, encoded video data output from the encoder 200 and input to the decoder 300. In some examples, portions of memory 102 , 113 may be allocated as one or more video buffers, eg, for storing raw, decoded, or encoded video data.
[0049] In some examples, source device 100 can output the encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access the encoded data from memory 113 via input interface 111. Memory 113 or storage 102 can include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a Digital Versatile Disc (DVD), a Compact Disc Read-Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0050] Output interface 104 may include any type of medium or device capable of transmitting encoded video data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded video data directly from source device 100 to destination device 110 in real time. The encoded video data may be modulated according to a communication standard of a wireless communication protocol and transmitted to destination device 110.
[0051] The communication medium 120 may include a transient medium such as a wireless broadcast or a wired network transmission. For example, the communication medium 120 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium) such as a hard disk, a flash drive, a compact disk, a digital video disk, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0052] In some embodiments, the communication medium 120 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 100 to the destination device 110. For example, a server (not shown) can receive encoded video from the source device 100 and provide the encoded video data to the destination device 110, for example, via a network transmission. The server may include a web server (e.g., for a website), a server configured to provide a file transfer protocol service (such as the File Transfer Protocol (FTP) or the File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or an evolved Multimedia Broadcast Multicast Service (eMBMS) server, a Network-attached storage (NAS) device, and the like. The server can implement one or more HTTP streaming protocols, such as the Moving Pictures Experts Group (MPEG) Media Transport (MMT) protocol, the Dynamic Adaptive Streaming over HTTP (DASH) protocol, the HTTP Live Streaming (HLS) protocol, or the Real Time Streaming Protocol (RTSP).
[0053] The destination device 110 can access the encoded video data from the server, for example, via a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection) or a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.) for accessing the encoded video data stored on the server.
[0054] The output interface 104 and the input interface 111 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard or the IEEE 802.15 standard (e.g., ZigBee™), the Bluetooth standard, or other physical components. In examples where the output interface 104 and the input interface 111 include wireless components, the output interface 104 and the input interface 111 may be configured to transmit data, such as encoded video data, according to WIFI, Ethernet, a cellular network (such as the 4th Generation Mobile Communication Technology (4G), Long Term Evolution (LTE), LTE-Advanced, the 5th Generation Mobile Communication Technology (5G), the 6th Generation Mobile Networks (6G), etc.).
[0055] The technology provided in the embodiments of the present application can be applied to support video encoding and decoding in one or more multimedia applications such as: video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0056] Input interface 111 of destination device 110 receives an encoded video bitstream from communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), wherein the syntax elements are used to decode the encoded data units to obtain decoded video data. Display device 114 displays the decoded video data to a user. Display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0057] The encoder 200 and the decoder 300 may be implemented as one or more of a variety of processing circuits, which may include a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the technology is implemented in whole or in part in software, the device may store instructions for the software in an appropriate non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of the present application.
[0058] The encoder 200 and the decoder 300 can be processed based on the following video coding and decoding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding (HEVC)), H.266 (also known as Versatile Video Coding (VVC), the second generation Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, the first generation of open media alliance video (Alliance for Open Media Video 1, AV1), the first generation audio and video coding standard (Audio Video coding Standard 1, AVS1), AVS2, AVS3 or the next generation video standard protocol, which is not specifically limited in the embodiments of the present application.
[0059] Generally, the encoder 200 and the decoder 300 can perform block-based encoding and decoding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in the encoding or decoding process). For example, a block can include a two-dimensional matrix of samples of luminance or chrominance data. For example, the encoder 200 and the decoder 300 can encode and decode video data represented in the YUV format, where "Y" represents brightness (Luminance or Luma), that is, grayscale values, and "U" and "V" represent chrominance (Chrominance or Chroma).
[0060] 2 , which is a schematic diagram of the structure of an encoder 200 provided in an embodiment of the present application, and the encoder 200 may be the encoder 200 in FIG1 . In the example of FIG2 , the encoder 200 includes a memory 201, a coding parameter determination unit 210, a residual generation unit 202, a transform processing unit 203, a quantization unit 204, an inverse quantization unit 205, an inverse transform processing unit 206, a reconstruction unit 207, a filter unit 208, a decoded picture buffer (DPB) 209, and an entropy coding unit 220.
[0061] The memory 201 can store video data to be encoded. For example, the encoder 200 can receive and store video data from the data source 101 shown in FIG1 . In some examples, the memory 201 can be on the same chip as other components of the encoder 200 (as shown in FIG2 ), or can be independent of the chip where these components are located.
[0062] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is used to obtain a first prediction block for the current block using the inter-frame prediction mode, and the intra-frame prediction unit 213 is used to obtain a second prediction block for the current block using the intra-frame prediction mode. The mode selection unit 211 is used to obtain a target prediction block based on the first prediction block and the second prediction block, and to determine a final prediction mode. In addition, the coding parameter determination unit 210 may also include other functional units, such as a functional unit for determining the division method of the coding unit (CU), a functional unit for determining the transform type of the residual data of the CU or the quantization parameter of the residual data of the CU, etc.
[0063] For ease of description and understanding, the embodiment of the present application refers to the CU to be processed in the current image as the current CU, and the image block to be processed in the current CU as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0064] The inter prediction unit 212 may include a motion estimation unit and a motion compensation unit. For inter prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 209).
[0065] The motion estimation unit can generate one or more motion vectors (MVs) that represent the position of a reference block in a reference picture relative to the position of a current block in a current picture. The motion compensation unit can interpolate to obtain a prediction value with the accuracy indicated by the motion vectors.
[0066] The encoding parameter determination unit 210 may provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the original unencoded video data of the current block from the memory 201 and calculates the residual between the current block and the target prediction block to obtain a residual block. In some examples, the functionality of the residual generation unit 202 may be implemented using one or more subtractor circuits that perform binary subtraction.
[0067] As an example, the coding parameter determination unit 210 may provide a syntax element representing the coding parameter to the entropy coding unit 220 for encoding. The coding parameter includes one or more of the CU partitioning method, the final prediction mode, the transform type of the CU residual data, or the quantization parameter of the CU residual data.
[0068] The transform processing unit 203 transforms the residual block output by the residual generation unit 202 to obtain a transform coefficient block. The transform may include discrete cosine transform (DCT), integer transform, directional transform, or KL transform (Karhunen-Loeve Transform). In some examples, the encoder 200 may not include the transform processing unit 203.
[0069] The quantization unit 204 may quantize the transform coefficients in the transform coefficient block according to a quantization parameter (QP) value associated with the current block to generate a quantized transform coefficient block.
[0070] The inverse quantization unit 205 and the inverse transform processing unit 206 can respectively perform inverse quantization and inverse transform on the transform coefficient block to obtain a reconstructed residual block. The reconstruction unit 207 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0071] The filter unit 208 may perform one or more filter operations on the reconstructed block. For example, the filter unit 208 may be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, the encoder 200 may not include the filter unit 208.
[0072] The encoder 200 stores the reconstructed picture obtained from the reconstructed block in the DPB 209. For example, in examples where the operation of the filter unit 208 is not required, the reconstruction unit 207 may store the reconstructed block in the DPB 209. In examples where the operation of the filter unit 208 is required, the filter unit 208 may store the filtered reconstructed block in the DPB 209. The inter-frame prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter-frame prediction on blocks of subsequent pictures to be encoded. In some examples, the DPB 209 may be replaced with other types of memory.
[0073] The entropy coding unit 220 may entropy encode syntax elements of other components in the encoder 200 and output encoded video data. For example, the entropy coding unit 220 may entropy encode the quantized transform coefficient blocks from the quantization unit 204. As another example, the entropy coding unit 220 may entropy encode syntax elements from the coding parameter determination unit 210 (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction).
[0074] It is understandable that the composition of the encoder 200 shown in FIG2 is merely an illustration and does not constitute a limitation to the embodiments of the present application.
[0075] FIG3 is a schematic diagram of the structure of a decoder 300 provided in an embodiment of the present application, and the decoder 300 may be the decoder 300 described in FIG1 . In the example of FIG3 , the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0076] The entropy decoding unit 302 can receive the encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements, where the syntax elements indicate encoding parameters, and the encoding parameters include one or more of the CU partitioning method, the final prediction mode, the transform type of the CU's residual data, or the quantization parameter of the CU's residual data.
[0077] When the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter-frame prediction mode, the prediction block of the current CU can be obtained by the inter-frame prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra-frame prediction mode, the prediction block of the current CU can be obtained by the intra-frame prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may also include a unit for performing prediction functions according to other prediction modes.
[0078] The CPB 301 can obtain and store encoded video data from the communication medium 120 shown in Figure 1. The DPB 307 is used to store decoded pictures. Optionally, the CPB 301 and DPB 307 can also be replaced with other types of memory, which is not specifically limited in this application. In some examples, the CPB 301 can be on the same chip as other components of the decoder 300 (as shown in the figure), or it can be independent of the chip where these components are located.
[0079] The decoder 300 can perform reconstruction operations on each block separately. The entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients and transform information (such as QP or transform mode indication) to obtain quantized transform coefficients. The quantized transform coefficients are dequantized by the dequantization unit 303 to obtain a transform coefficient block including the transform coefficients. The transform coefficient block is inversely transformed by the inverse transform processing unit 304 to generate a residual block corresponding to the current block. This inverse transform is the inverse operation of the above-mentioned transform.
[0080] The reconstruction unit 305 may reconstruct the current block according to the prediction block and the residual block. For example, the reconstruction unit 305 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.
[0081] The filter unit 306 may perform one or more filter operations on the reconstructed block. For example, the type of the filter unit 306 may refer to the type of the filter unit 208 and will not be described in detail here. In some examples, the operation of the filter unit 306 may be skipped.
[0082] The decoder 300 may store the reconstructed picture obtained from the reconstructed block in the DPB 307. For example, in an example in which the operation of the filter unit 306 is not performed, the reconstruction unit 305 may store the reconstructed block in the DPB 307. In an example in which the operation of the filter unit 306 is performed, the filter unit 306 may store the filtered reconstructed block in the DPB 307. The decoder 300 may output a decoded picture (e.g., decoded video) from the DPB 307 for subsequent presentation to a display device (such as the display device 114 of FIG. 1 ).
[0083] Before describing the embodiments of the present application, the following briefly introduces the relevant technologies:
[0084] 1. Neural Network-Based Loop Filter Solution
[0085] To reduce the impact of distortion effects such as blocking artifacts and ringing on video quality, video compression standards incorporate loop filters, such as deblocking filters and sample adaptive compensation (SAB). Deblocking filters reduce blocking artifacts, while SAB improves ringing artifacts. These loop filters can be placed within the codec loop and effectively improve both subjective and objective video quality.
[0086] With the widespread application of neural network-based methods in image processing, some neural network-based loop filter solutions have also shown good performance in video compression. Related technologies attempt to use neural network filters to partially replace traditional filters, or even completely replace traditional filters.
[0087] 2. Neural Network-Based Super-Resolution
[0088] Neural network-based super-resolution approaches utilize neural networks to reconstruct images for super-resolution, enhancing clarity and detail. This approach achieves super-resolution by training neural networks to learn image features. Common neural networks include convolutional neural networks and recurrent neural networks. Within image super-resolution approaches, neural networks can be used for single-image super-resolution reconstruction and for the joint super-resolution reconstruction of multiple images. This approach has been widely applied in fields such as image processing, medical image analysis, and security monitoring, and continues to evolve and improve.
[0089] In the related art, loop filtering requires the use of a neural network-based loop filter, while super-resolution requires a neural network-based super-resolution processor. In other words, different neural network models are required to achieve different video compression processing functions. As can be seen, the neural network models used for video compression processing in the related art have relatively simple functions, which leads to a relatively complex video compression process.
[0090] The following describes the video compression processing method provided by the embodiments of the present application in conjunction with the accompanying drawings. The video compression processing method provided by the embodiments of the present application can be performed by an encoding end, such as the encoder 200 shown in Figure 1 or Figure 2, or by a decoding end, such as the decoder 300 described in Figure 1 or Figure 3. The encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a video decoding device.
[0091] FIG4 is a flowchart of a video compression processing method provided by an embodiment of the present application. As shown in FIG4 , the video compression processing method includes the following steps:
[0092] Step 401: Inputting a first image into a target neural network model, wherein the target neural network model is a neural network model having N video compression processing functions, where N is an integer greater than 1;
[0093] Step 401: Perform a first processing on the first image based on the target neural network model to obtain a second image; wherein the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0094] The first image includes at least one of the following: a reconstructed block, a predicted block, a boundary strength, a residual block, and a transform coefficient block.
[0095] The reconstructed block refers to a reconstructed sample block of a signal obtained by decoding the bit stream by the decoder and constituting a decoded image, which may be a sample after filtering or a sample before filtering.
[0096] The prediction block refers to a prediction sample block of a signal obtained by using reconstructed samples through prediction technology during the decoding process.
[0097] The boundary strength refers to the boundary strength value of the signal obtained by the deblocking filtering technology according to the pattern information and sample changes on both sides of the boundary.
[0098] The residual block refers to a block composed of the difference between the original samples and the predicted samples of the image block of the signal.
[0099] The transform coefficient block refers to a block composed of transform coefficients in the frequency domain obtained by transforming the residual block through a transform technique.
[0100] The first image is used as an input to the target neural network model. The target neural network model may perform a first processing on the first image based on the determined video compression processing mode to obtain a second image. The second image may include a filtered reconstructed block or a luminance reconstructed block at a first resolution, where the first resolution is higher than a resolution of the reconstructed block corresponding to the first image.
[0101] The reason for determining the video compression processing method corresponding to the first processing is that the target neural network model has multiple (including two) video compression processing functions. Since the target neural network model has multiple video compression processing functions, the target neural network model can be used as a unified neural network model to implement multiple video compression processing functions. In this way, when multiple video compression processing is required for an image, it can be achieved through a unified neural network model, without the need to input the image multiple times into multiple single-function neural network models for processing, which is conducive to reducing the complexity of video compression processing.
[0102] In an embodiment of the present application, a first image is input into a target neural network model, and a first process is performed on the first image based on the target neural network model to obtain a second image. The target neural network model is a neural network model having N types of video compression processing functions, and the video compression processing method corresponding to the first process is implemented by one or more of the N types of video compression processing functions. In this way, the functional diversity of the video compression processing neural network model can be increased, thereby enabling multiple video compression processing functions to be implemented through a unified neural network model, which helps reduce the complexity of video compression processing.
[0103] It should be noted that the video compression processing method provided in the embodiment of the present application is applicable to both the model application stage (or model inference stage) and the model training stage, and the embodiment of the present application does not limit this.
[0104] The target neural network model provided in the embodiment of the present application can be set at either the encoding end or the decoding end, and the embodiment of the present application does not limit this.
[0105] Optionally, the N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function;
[0106] The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method and a post-processing filtering processing method.
[0107] In the loop filtering process, the sizes (width and height) of the input image and the output image are the same.
[0108] Exemplarily, when performing loop filtering processing as a loop filtering model, the sizes (size in the width direction and size in the height direction) of the source image, the input image, and the output image are the same.
[0109] Exemplarily, when loop filtering is performed as a loop filtering 2x model, the source image is twice the size of the input image (the size in the width direction and the size in the height direction), and the size of the input image and the output image (the size in the width direction and the size in the height direction) are the same.
[0110] It should be noted that the input image and the output image are respectively the input and output of the video compression processing method.
[0111] The process of loop filtering based on the target neural network model can be: first, the first luminance data and the first chrominance data are extracted through the convolution layer and activation layer of the target neural network model (it should be noted that each convolution layer can correspond to an activation layer, such as the activation layer can be a parametric rectified linear unit (PReLU)) to generate K feature maps, where K is a positive integer. Then, these feature maps are subjected to feature fusion through the convolution layer and the activation layer to generate S fused feature maps, where S is a positive integer. Then, these fused feature maps are feature enhanced through multiple residual blocks to generate L enhanced feature maps, where L is a positive integer. Finally, these enhanced feature maps are subjected to operations such as convolution layers and activation layers to generate filtered reconstructed image blocks.
[0112] There may be multiple super-resolution processing multiples, and different super-resolution multiples may be used as different super-resolution processing modes.
[0113] For example, when super-resolution processing is performed as a super-resolution 1.5x model, the source image is 1.5 times the size of the input image (the size in the width direction and the size in the height direction), and the output image is 1.5 times the size of the input image (the size in the width direction and the size in the height direction).
[0114] For example, when super-resolution processing is performed as a super-resolution 2x model, the source image is twice the size of the input image (the size in the width direction and the size in the height direction), and the output image is twice the size of the input image (the size in the width direction and the size in the height direction).
[0115] Super-resolution processing can also be called super-resolution reconstruction processing. The process of super-resolution reconstruction processing based on the target neural network model can be: first, the first luminance data and the first chrominance data are subjected to feature extraction through the convolution layer and activation layer of the target neural network model to generate T feature maps, where T is a positive integer. Next, these feature maps are subjected to feature fusion through the convolution layer and activation layer to generate M fused feature maps, where M is a positive integer. Then, these fused feature maps are subjected to feature enhancement through multiple residual blocks to generate P enhanced feature maps, where P is a positive integer. Finally, these enhanced feature maps are subjected to operations such as convolution layers, activation layers, and upsampling to generate high-resolution image blocks.
[0116] The post-processing filtering method is roughly the same as the loop filtering method. The process of post-processing filtering based on the target neural network model can refer to the above-mentioned loop filtering process. To avoid repetition, this will not be described in detail. It should be noted that the above-mentioned video compression processing functions and various video compression processing methods are only examples. When there are other video compression processing functions, the target neural network model can continue to expand other video compression processing functions, and the embodiments of this application do not limit this.
[0117] In an embodiment of the present application, the target neural network model itself can determine the video compression processing method corresponding to the first processing, or other models or other modules can determine the video compression processing method corresponding to the first processing, and inform the target neural network model of the determined video compression processing method.
[0118] In the embodiment of the present application, determining the video compression processing method corresponding to the first processing may include multiple implementation methods, which are described below respectively.
[0119] In some embodiments, the video compression processing method corresponding to the first processing is determined based on target information, and the target information is information used to indicate the video compression processing method.
[0120] In some embodiments, the method further comprises:
[0121] The target information is input into the target neural network model.
[0122] Here, the video compression processing method corresponding to the first processing can be determined by the target neural network model based on the target information.
[0123] In this embodiment, in addition to inputting the first image into the target neural network model, target information is also input into the target neural network model. In other words, the input data of the target neural network model includes the target information and the first image.
[0124] The above-mentioned target information can be understood as function indication information, which is used to indicate the video compression processing function.
[0125] Exemplarily, the target information is a function identifier.
[0126] Optionally, the target information includes at least one of the following:
[0127] A first identifier, wherein an index value of the first identifier is mapped to a corresponding video compression processing mode;
[0128] The second identifier includes a multiple of the image width direction and a multiple of the image height direction.
[0129] Exemplarily, the first identifier may be an index value such as 0, 1, 2, or 3.
[0130] The meaning of each index value may be, for example:
[0131] Mark 0: loop filter model;
[0132] Identifier 1: Super-resolution 2x model;
[0133] Mark 2: Super resolution 1.5x model;
[0134] Identification 3: Post-processing filter model.
[0135] If the first flag is 0, the target neural network model is used as a loop filtering model for filtering processing.
[0136] If the first flag is 1, the target neural network model is super-resolution processed as a super-resolution 2x model.
[0137] If the first flag is 2, the target neural network model is super-resolution processed as a super-resolution 1.5x model.
[0138] If the first flag is 3, the target neural network model is used as a post-processing filter model.
[0139] The meaning of each index value identifier may also be, for example:
[0140] Mark 0: loop filter model;
[0141] Identification 1: Loop filter 2x model;
[0142] Mark 2: Super-resolution 2x model;
[0143] Identification 3: Post-processing filter model.
[0144] If the first flag is 0, the target neural network model is used as a loop filtering model for filtering processing.
[0145] If the first flag is 1, the target neural network model is filtered as a loop filter 2x model.
[0146] If the first identifier is 2, the target neural network model is super-resolution processed as a super-resolution 2x model, and the magnification is 2 times.
[0147] If the first flag is 3, the target neural network model is used as a post-processing filter model.
[0148] The target information adopts the above-mentioned first identification method which is relatively concise and can directly determine the video compression processing method based on the target information.
[0149] During training, the target neural network model can learn the functions of the model by itself to improve the generalization of the model.
[0150] Exemplarily, the second identifier may be in a data format such as (1, 1), (2, 2), or (1.5, 1.5).
[0151] If the second identifier is (1, 1), indicating that the multiple in the image width direction and the multiple in the image height direction are both 1, the target neural network model can be used as a loop filtering model for filtering processing.
[0152] If the second identifier is (2, 2), indicating that the multiples in the image width direction and the image height direction are both 2, the target neural network model can be used as a super-resolution 2x model for super-resolution processing.
[0153] If the second identifier is (1.5, 1.5), indicating that the multiples in the image width direction and the image height direction are both 1.5, the target neural network model can be used as a super-resolution 1.5x model for super-resolution processing.
[0154] The target information adopts the above-mentioned second identification method which is more flexible and can provide effective input information for the target neural network model.
[0155] In this embodiment, target information is used to assist the target neural network model in determining the video compression processing method. This method is simple and efficient and can improve the performance of the target neural network model.
[0156] It should be noted that in this embodiment, the first image can be an image obtained through pre-processing. For example, the third image can be up-sampled in width and height according to the super-resolution factor to obtain the first image. The first image can also be an image that has not been pre-processed, and this embodiment of the present application is not limited to this. When target information is input, regardless of whether the first image has been pre-processed, the target neural network model can determine the video compression processing method based on the target information.
[0157] In some embodiments, before inputting the first image into the target neural network model, the method further includes:
[0158] performing a second process on the third image to obtain the first image;
[0159] The video compression processing method corresponding to the first processing is determined based on the first image.
[0160] For some video compression processing functions, the input image needs to be pre-processed.
[0161] For example, for a super-resolution 2x model, the width and height of the input image need to be upsampled to 2 times the original size. In this case, the pre-processing is to upsample the width and height of the input image to 2 times the original size.
[0162] For example, for a 1.5x super-resolution model, the width and height of the input image need to be upsampled to 1.5 times the original size. In this case, the pre-processing is to upsample the width and height of the input image to 1.5 times the original size.
[0163] In this embodiment, the process of performing the second processing on the third image to obtain the first image can be understood as the pre-processing process. In other words, the first image is the image obtained after the pre-processing of the third image. Here, the third image can be understood as the input image.
[0164] A pre-processing module can be set at the encoding end or the decoding end to implement pre-processing of the input image.
[0165] In this embodiment, the input image (i.e., the third image) is pre-processed to obtain a first image, which is then input into the target neural network model for the first processing. Since the first image is a pre-processed image, it contains certain specific image features. Therefore, the corresponding video compression processing method can be determined based on the image features of the first image.
[0166] For example, if the first image corresponds to an image obtained by sampling the input image to twice its original size in the width and height directions, the target neural network model can determine that the corresponding video compression processing method is super-resolution 2x.
[0167] For example, if the first image corresponds to an image obtained by sampling the input image to 1.5 times the original image in the width and height directions, the target neural network model can determine that the corresponding video compression processing method is super-resolution 1.5x.
[0168] In this embodiment, pre-processing the input image before inputting it into the target neural network model allows the model's input data to better adapt to the network, improving the performance of the target neural network model in some aspects. Furthermore, the target neural network model can acquire the image features of the pre-processed first image through self-learning. This eliminates the need to input additional auxiliary information or instructions to the target neural network model to determine the corresponding video compression processing method, improving the performance of the target neural network model in some aspects.
[0169] Optionally, performing the second processing on the third image to obtain the first image includes:
[0170] performing a second processing on the third image according to target information to obtain the first image, wherein the target information is information for indicating a video compression processing mode;
[0171] The video compression processing mode corresponding to the first processing is determined based on at least one of the first image and the target information.
[0172] In this embodiment, by inputting target information into the pre-processing module, the pre-processing module can determine the processing method of the second processing based on the target information.
[0173] Exemplarily, the video compression processing mode corresponding to the first processing is determined based on the first image.
[0174] Exemplarily, the video compression processing mode corresponding to the first processing is determined based on the target information.
[0175] Exemplarily, the video compression processing mode corresponding to the first processing is determined based on the first image and the target information.
[0176] Here, the target information can refer to the relevant description of the aforementioned target information. To avoid repetition, it will not be described in detail.
[0177] Optionally, the target information is information for indicating a super-resolution processing mode;
[0178] The performing a second processing on the third image according to the target information to obtain the first image includes at least one of the following:
[0179] performing upsampling processing on the third image in the width direction and the height direction according to the target information to obtain the first image;
[0180] A super-resolution multiple is determined according to the target information, and up-sampling processing is performed on the width direction and the height direction of the third image according to the super-resolution multiple to obtain the first image.
[0181] It should be noted that upsampling can be done using a resampling encoding method, namely Reference Picture Resampling (RPR). With the increase in high-resolution videos, video transmission in bandwidth-constrained environments has become a huge challenge. In VVC, RPR can adaptively adjust the resolution based on network conditions. When the network bandwidth is low, it can downsample and encode low-resolution (LR) frames. When the network bandwidth improves, it can upsample and encode high-resolution (HR) original frames.
[0182] In some embodiments, the method further comprises:
[0183] performing a second processing on the third image according to target information to obtain the first image, wherein the target information is information for indicating a video compression processing mode;
[0184] Inputting the target information into the target neural network model;
[0185] Among them, the video compression processing method corresponding to the first processing is determined by the target neural network model based on at least one of the first image and the target information.
[0186] As previously described, the process of performing the second processing on the third image to obtain the first image can be understood as a pre-processing process. In this embodiment, target information is input to the pre-processing module so that the pre-processing module can determine the processing method for the second processing based on the target information. Target information is also input to the target neural network model so that the target neural network model directly determines the video compression processing method based on the target information.
[0187] Here, the target information can refer to the relevant description of the aforementioned target information. To avoid repetition, it will not be described in detail.
[0188] In this implementation, pre-processing the input image before inputting it into the target neural network model allows the model's input data to better adapt to the network, improving the performance of the target neural network model in some aspects. The target information can assist the target neural network model in training or reasoning, improving its performance in some aspects.
[0189] The following provides a specific embodiment by taking the target neural network model to determine the video compression processing method as an example:
[0190] Example 1: Function Identification Assisted Target Neural Network Model Training and Inference
[0191] As shown in Figure 5, the following steps are included:
[0192] Step 11: Acquire a first image;
[0193] Step 12: Inputting the first image and the function identifier into a target neural network model, and performing video compression processing on the first image based on the target neural network model;
[0194] Step 13: Obtain a second image output by the target neural network model.
[0195] In this embodiment, the target neural network model is used to implement at least two functions in video compression processing, and the target neural network model determines the video compression processing method based on the function identifier.
[0196] Example 2: Function Identification Assisted Pre-Processing, and Function Identification Assisted Target Neural Network Model Training and Inference
[0197] As shown in Figure 6, the following steps are included:
[0198] Step 21: Acquire a third image;
[0199] Step 22: Pre-process the third image according to the function identifier to obtain a first image;
[0200] Step 23: Inputting the first image and the function identifier into a target neural network model, and performing video compression processing on the first image based on the target neural network model;
[0201] Step 24: Obtain a second image output by the target neural network model.
[0202] In this embodiment, the target neural network model is used to implement at least two functions in video compression processing, and the target neural network model determines the video compression processing method based on the function identifier.
[0203] It should be noted that once the structure of the target neural network model is fixed, it usually cannot process different super-resolution functions at the same time.
[0204] Compared with Example 1, Example 2 adds a pre-processing operation. Therefore, the pre-processing can make the size of the input image of the target neural network model the same, so that the target neural network model can process different super-resolution functions.
[0205] Example 3: Function identification assists pre-processing, and pre-processing assists the training and reasoning of the target neural network model.
[0206] As shown in Figure 7, the following steps are included:
[0207] Step 31: Acquire a third image;
[0208] Step 32: Pre-process the third image according to the function identifier to obtain a first image;
[0209] Step 33: Inputting the first image into a target neural network model, and performing video compression processing on the first image based on the target neural network model;
[0210] Step 34: Obtain a second image output by the target neural network model.
[0211] In this embodiment, the target neural network model is used to implement at least two functions in video compression processing. The function identifier is not used as an input to the target neural network model. The target neural network model determines the video compression processing method based on the image features of the first image.
[0212] Compared with Example 2, Example 3 can moderately reduce the computational complexity and parameter amount of the neural network model.
[0213] In summary, in the embodiments of the present application, the functional diversity of the video compression processing neural network model can be improved, so that multiple video compression processing functions can be realized through a unified neural network model, which is conducive to reducing the complexity of video compression processing.
[0214] The video compression processing method provided in the embodiment of the present application can be executed by a video compression processing device. In the embodiment of the present application, the video compression processing device provided in the embodiment of the present application is described by taking the video compression processing method performed by the video compression processing device as an example.
[0215] 8 , the embodiment of the present application further provides a video compression processing device, which can be applied to an encoding end or a decoding end. As shown in FIG8 , the video compression processing device 800 includes:
[0216] A first input module 801 is configured to input a first image into a target neural network model, wherein the target neural network model is a neural network model having N types of video compression processing functions, where N is an integer greater than 1;
[0217] The first processing module 802 is used to perform a first processing on the first image based on the target neural network model to obtain a second image.
[0218] The video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0219] Optionally, the video compression processing mode corresponding to the first processing is determined based on target information, and the target information is information used to indicate the video compression processing mode.
[0220] Optionally, the video compression processing device 800 further includes:
[0221] A second input module is used to input target information into the target neural network model, wherein the target information is information for indicating a video compression processing method;
[0222] Among them, the video compression processing method corresponding to the first processing is determined by the target neural network model based on the target information.
[0223] Optionally, the video compression processing device 800 further includes:
[0224] a second processing module, configured to perform a second processing on the third image to obtain the first image;
[0225] The video compression processing method corresponding to the first processing is determined based on the first image.
[0226] Optionally, the video compression processing method corresponding to the first processing is determined by the target neural network model based on the first image.
[0227] Optionally, the video compression processing device 800 further includes:
[0228] a third processing module, configured to perform a second processing on the third image according to target information to obtain the first image, wherein the target information is information indicating a video compression processing mode;
[0229] The video compression processing mode corresponding to the first processing is determined based on at least one of the first image and the target information.
[0230] Optionally, the target information is information for indicating a super-resolution processing mode;
[0231] The third processing module is specifically configured to:
[0232] performing upsampling processing on the third image in the width direction and the height direction according to the target information to obtain the first image;
[0233] A super-resolution multiple is determined according to the target information, and up-sampling processing is performed on the width direction and the height direction of the third image according to the super-resolution multiple to obtain the first image.
[0234] Optionally, the video compression processing device 800 further includes:
[0235] A third input module, configured to input the target information into the target neural network model;
[0236] Among them, the video compression processing method corresponding to the first processing is determined by the target neural network model based on at least one of the first image and the target information.
[0237] Optionally, the target information includes at least one of the following:
[0238] A first identifier, wherein an index value of the first identifier is mapped to a corresponding video compression processing mode;
[0239] The second identifier includes a multiple of the image width direction and a multiple of the image height direction.
[0240] Optionally, the N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function;
[0241] The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method and a post-processing filtering processing method.
[0242] In summary, in the embodiments of the present application, the functional diversity of the video compression processing neural network model can be improved, so that multiple video compression processing functions can be realized through a unified neural network model, which is conducive to reducing the complexity of video compression processing.
[0243] The video compression processing device provided in the embodiment of the present application can implement the various processes implemented in the method embodiments of Figures 4 to 7 and achieve the same technical effects. To avoid repetition, they will not be described here.
[0244] As shown in Figure 9, an embodiment of the present application further provides an electronic device 900, comprising a processor 901 and a memory 902, wherein the memory 902 stores a program or instruction that can be run on the processor 901. For example, when the electronic device 900 is an encoding end device, the program or instruction is executed by the processor 901 to implement the various steps of the above-mentioned video compression processing method embodiment, and can achieve the same technical effect. When the electronic device 900 is a decoding end device, the program or instruction is executed by the processor 901 to implement the various steps of the above-mentioned video compression processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here. Optionally, the memory 902 can be the memory 102 or the memory 113 in the embodiment shown in Figure 1, and the processor 901 can implement the functions of the encoder 200 or the decoder 300 in the embodiments shown in Figures 1 to 3.
[0245] The present application also provides an electronic device including: a memory configured to store video data; and a processing circuit configured to implement the steps of the above-described video compression processing method embodiment. Optionally, the memory may be memory 102 or memory 113 in the embodiment shown in FIG1 , and the processing circuit may implement the functions of encoder 200 or decoder 300 in the embodiments shown in FIG1 through FIG3 .
[0246] The present application also provides an electronic device including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute a program or instruction to implement the steps in the method embodiment shown in FIG4 . This device embodiment corresponds to the aforementioned method embodiment, and each implementation process and implementation method of the aforementioned method embodiment are applicable to this terminal embodiment and can achieve the same technical effects.
[0247] The electronic device may be a terminal, or may be other devices other than a terminal, such as a server, a network attached storage (NAS), etc.
[0248] Among them, the terminal can be a mobile phone, tablet personal computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile Internet device (MID), augmented reality (AR), virtual reality (VR) equipment, mixed reality (MR) equipment, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipborne equipment, pedestrian user equipment (PUE), smart home (home appliances with wireless communication function, such as refrigerator, TV, washing machine or furniture, etc.), game console, personal computer (PC), ATM or self-service machine and other terminal-side devices. Wearable devices include: smart watches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among them, vehicle-mounted devices can also be called vehicle-mounted terminals, vehicle-mounted controllers, vehicle-mounted modules, vehicle-mounted components, vehicle-mounted chips, or vehicle-mounted units, etc. It should be noted that the specific type of terminal is not limited in the embodiments of this application.
[0249] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0250] For example, the electronic device may include but is not limited to the source device 100 or the destination device 110 shown in FIG. 1 .
[0251] Taking an electronic device as a terminal as an example, FIG10 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of the present application.
[0252] The terminal 1000 includes but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009 and at least some of the components of the processor 1010.
[0253] Those skilled in the art will appreciate that the terminal 1000 may also include a power supply (such as a battery) to power various components. The power supply may be logically connected to the processor 1010 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The terminal structure shown in FIG10 does not limit the terminal. The terminal may include more or fewer components than shown, or may combine certain components, or have different component arrangements, which will not be described in detail here.
[0254] It should be understood that in an embodiment of the present application, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes image data of a static picture or video obtained by an image acquisition device (such as a camera) in a video acquisition mode or an image acquisition mode, or may process the obtained point cloud data. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0255] In the embodiment of the present application, after receiving downlink data from a network-side device, the RF unit 1001 may transmit the data to the processor 1010 for processing. Furthermore, the RF unit 1001 may send uplink data to the network-side device. Typically, the RF unit 1001 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like.
[0256] The memory 1009 can be used to store software programs or instructions and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 1009 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0257] Processor 1010 may include one or more processing units. Optionally, processor 1010 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1010.
[0258] The processor 1010 is configured to:
[0259] Inputting the first image into a target neural network model, wherein the target neural network model is a neural network model having N video compression processing functions, where N is an integer greater than 1;
[0260] Performing a first processing on the first image based on the target neural network model to obtain a second image;
[0261] The video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0262] Optionally, the video compression processing mode corresponding to the first processing is determined based on target information, and the target information is information used to indicate the video compression processing mode.
[0263] Optionally, the processor 1010 is further configured to:
[0264] Inputting target information into the target neural network model, wherein the target information is information for indicating a video compression processing method;
[0265] Among them, the video compression processing method corresponding to the first processing is determined by the target neural network model based on the target information.
[0266] Optionally, the processor 1010 is further configured to:
[0267] performing a second process on the third image to obtain the first image;
[0268] The video compression processing method corresponding to the first processing is determined based on the first image.
[0269] Optionally, the processor 1010 is further configured to:
[0270] performing a second processing on the third image according to target information to obtain the first image, wherein the target information is information for indicating a video compression processing mode;
[0271] The video compression processing mode corresponding to the first processing is determined based on at least one of the first image and the target information.
[0272] Optionally, the target information is information for indicating a super-resolution processing mode;
[0273] Optionally, the processor 1010 is further configured to perform at least one of the following:
[0274] performing upsampling processing on the third image in the width direction and the height direction according to the target information to obtain the first image;
[0275] A super-resolution multiple is determined according to the target information, and up-sampling processing is performed on the width direction and the height direction of the third image according to the super-resolution multiple to obtain the first image.
[0276] Optionally, the processor 1010 is further configured to:
[0277] performing a second processing on the third image according to target information to obtain the first image, wherein the target information is information for indicating a video compression processing mode;
[0278] Inputting the target information into the target neural network model;
[0279] Among them, the video compression processing method corresponding to the first processing is determined by the target neural network model based on at least one of the first image and the target information.
[0280] Optionally, the target information includes at least one of the following:
[0281] A first identifier, wherein an index value of the first identifier is mapped to a corresponding video compression processing mode;
[0282] The second identifier includes a multiple of the image width direction and a multiple of the image height direction.
[0283] Optionally, the N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function;
[0284] The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method and a post-processing filtering processing method.
[0285] In summary, in the embodiments of the present application, the functional diversity of the video compression processing neural network model can be improved, so that multiple video compression processing functions can be realized through a unified neural network model, which is conducive to reducing the complexity of video compression processing.
[0286] It can be understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described here.
[0287] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned video compression processing method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0288] The processor is the processor in the terminal described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as ROM, RAM, a magnetic disk, or an optical disk. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0289] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned video compression processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0290] It should be understood that the chip mentioned in the embodiments of the present application may include a system-level chip (also referred to as a system chip, a chip system or a system-on-chip chip), and may also include an independent display chip, etc.
[0291] An embodiment of the present application further provides a computer program / program product, which is stored in a storage medium. The computer program / program product is executed by at least one processor to implement the various processes of the above-mentioned video compression processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0292] An embodiment of the present application also provides a video compression processing system, including: an encoding end device and a decoding end device, wherein the encoding end device can be used to execute the steps of the video compression processing method described above, and the decoding end device can be used to execute the steps of the video compression processing method described above.
[0293] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0294] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of a computer software product plus a necessary general-purpose hardware platform, or of course, by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes a number of instructions for enabling a terminal or network-side device to execute the methods described in each embodiment of the present application.
[0295] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms of implementation methods without departing from the purpose of this application and the scope of protection of the claims. These implementation methods are all within the protection of this application.
Claims
1. A video compression processing method, comprising: Inputting a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1; Performing a first processing on the first image based on the target neural network model to obtain a second image; Wherein, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
2. The method according to claim 1, wherein, The video compression processing method corresponding to the first processing is determined based on target information, and the target information is information for indicating a video compression processing method.
3. The method according to claim 1, wherein Before the step of inputting the first image into the target neural network model, the method further comprises: Performing a second processing on a third image to obtain the first image; Wherein, the video compression processing method corresponding to the first processing is determined based on the first image.
4. The method according to claim 1, wherein Before the step of inputting the first image into the target neural network model, the method further comprises: Performing a second processing on a third image according to target information to obtain the first image, where the target information is information for indicating a video compression processing method; Wherein, the video compression processing method corresponding to the first processing is determined based on at least one of the first image and the target information.
5. The method according to claim 4, wherein, The target information is information for indicating a super-resolution processing method; The step of performing a second processing on a third image according to target information to obtain the first image includes at least one of the following: According to the target information, performing upsampling processing on the width direction and height direction of the third image to obtain the first image; Determining a super-resolution multiple according to the target information, and performing upsampling processing on the width direction and height direction of the third image according to the super-resolution multiple to obtain the first image.
6. The method according to any one of claims 2, 4, and 5, wherein The target information includes at least one of the following: A first identifier, where the index value of the first identifier is mapped to the corresponding video compression processing method; A second identifier, where the second identifier includes the multiple of the image width direction and the multiple of the image height direction.
7. The method according to any one of claims 1 to 6, wherein The N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function; The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
8. A video compression processing apparatus, comprising: A first input module, configured to input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1; A first processing module, configured to perform a first processing on the first image based on the target neural network model to obtain a second image; Wherein, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
9. The device according to claim 8, wherein The video compression processing method corresponding to the first processing is determined based on target information, and the target information is information for indicating a video compression processing method.
10. The apparatus according to claim 8, further comprising: A second processing module, configured to perform a second processing on a third image to obtain the first image; Among them, the video compression processing method corresponding to the first processing is determined based on the first image.
11. The apparatus according to claim 8, further comprising: A third processing module, configured to perform a second processing on a third image according to target information to obtain the first image, where the target information is information for indicating a video compression processing method; Among them, the video compression processing method corresponding to the first processing is determined based on at least one of the first image and the target information.
12. The device according to claim 11, wherein, The target information is information for indicating a super-resolution processing method; The third processing module is specifically configured to perform at least one of the following: According to the target information, perform upsampling processing on the width direction and height direction of the third image to obtain the first image; Determine a super-resolution multiple according to the target information, and perform upsampling processing on the width direction and height direction of the third image according to the super-resolution multiple to obtain the first image.
13. The device according to any one of claims 9, 11 and 12, wherein The target information includes at least one of the following: A first identifier, where the index value of the first identifier is mapped to the corresponding video compression processing method; A second identifier, where the second identifier includes a multiple in the image width direction and a multiple in the image height direction.
14. The apparatus according to any one of claims 8 to 13, wherein, The N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function; The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
15. An electronic device, comprising a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
16. A readable storage medium, where a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
17. A chip, where the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
GAN-based security video compression method, device and equipment, and storage medium
CN111541900A
Video encoding method, decoding method, encoder, decoder, and AI accelerator
CN114868390A
Video processing method, device, equipment, decoder, system and storage medium
CN115943422A
Multi-neural network model for filtering during video coding and decoding
CN116349226A
Cross-domain emergency command and dispatch method and system
CN118764649A