Video compression processing method and device and electronic equipment
By using a multi-functional target neural network model to uniformly process the image, the problem of single video compression processing function in the prior art is solved, and a more complex video compression processing is achieved.
Patent Information
- Application Number
- CN202311843781.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the neural network model used for video compression processing has a relatively single function, resulting in a complex video compression processing process.
The target neural network model with multiple video compression processing functions is adopted to uniformly process the images and realize multiple video compression processing functions.
It improves the functional diversity of the neural network model of video compression processing and reduces the complexity of video compression processing.
Smart Images

Figure CN120238656A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video compression, and specifically relates to a video compression processing method, device, and electronic device. Background Art
[0002] In the related art, in order to implement different video compression processing functions, different neural network models need to be used. For example, for the loop filtering function, a loop filter based on a neural network needs to be used, and for the super-resolution function, a super-resolution processor based on a neural network needs to be used. It can be seen that in the related art, the functions of the neural network models for video compression processing are relatively single. Summary of the Invention
[0003] Embodiments of this application provide a video compression processing method, device, and electronic device, which can solve the problem that the functions of the neural network models for video compression processing in the related art are relatively single.
[0004] In a first aspect, a video compression processing method is provided, which is executed by an encoding end. The method includes:
[0005] Input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0006] Perform a first processing on the first image based on the target neural network model to obtain a second image;
[0007] Among them, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0008] In a second aspect, a video compression processing method is provided, which is executed by a decoding end. The method includes:
[0009] Input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0010] Perform a first processing on the first image based on the target neural network model to obtain a second image;
[0011] Among them, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0012] In a third aspect, a video compression processing device is provided, which is applied to an encoding end. The device includes:
[0013] A first input module, configured to input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0014] A first processing module, configured to perform a first processing on the first image based on the target neural network model to obtain a second image;
[0015] Wherein, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0016] In a fourth aspect, a video compression processing device is provided, which is applied to a decoding end. The device includes:
[0017] A first input module, configured to input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0018] A first processing module, configured to perform a first processing on the first image based on the target neural network model to obtain a second image;
[0019] Wherein, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0020] In a fifth aspect, an electronic device is provided. The terminal includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0021] In a sixth aspect, an electronic device is provided, including a processor and a communication interface. Wherein, the processor is configured to: input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1; perform a first processing on the first image based on the target neural network model to obtain a second image; wherein, the video compression processing method corresponding to the first processing is determined by the target neural network model.
[0022] In a seventh aspect, an electronic device is provided, including: a memory configured to store video data, and a processing circuit configured to implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.
[0023] In an eighth aspect, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0024] In a ninth aspect, a coding and decoding system is provided, including: an encoding end device and a decoding end device, where the encoding end device can be used to execute the steps of the method described in the first aspect, and the decoding end device can be used to execute the steps of the method described in the second aspect.
[0025] In a tenth aspect, a chip is provided, where the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0026] In an eleventh aspect, a computer program / program product is provided, where the computer program / program product is stored in a storage medium, and the program / program product is executed by at least one processor to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0027] In an embodiment of the present application, a first image is input into a target neural network model, and based on the target neural network model, the first image is subjected to a first process to obtain a second image. The target neural network model is a neural network model with N video compression processing functions, and the video compression processing method corresponding to the first process is implemented by one or more of the N video compression processing functions. In this way, the functional diversity of the video compression processing neural network model can be improved, so that multiple video compression processing functions can be implemented through a unified neural network model, which is beneficial to reducing the complexity of video compression processing. Description of the Drawings
[0028] Figure 1 is a schematic diagram of the coding and decoding system provided by an embodiment of the present application;
[0029] Figure 2 is a schematic structural diagram of the encoder provided by an embodiment of the present application;
[0030] Figure 3 is a schematic structural diagram of the decoder provided by an embodiment of the present application;
[0031] Figure 4 is a flowchart of a video compression processing method provided by an embodiment of the present application;
[0032] Figure 5 is a schematic diagram of Embodiment 1 provided by an embodiment of the present application;
[0033] Figure 6 It is a schematic diagram of Embodiment 2 provided by an embodiment of the present application;
[0034] Figure 7 It is a schematic diagram of Embodiment 3 provided by an embodiment of the present application;
[0035] Figure 8 It is a structural diagram of a video compression processing device provided by an embodiment of the present application;
[0036] Figure 9 It is a structural diagram of an electronic device provided by an embodiment of the present application;
[0037] Figure 10 It is a structural diagram of a terminal provided by an embodiment of the present application. Specific implementation manners
[0038] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0039] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.
[0040] Figure 1 It is a schematic diagram of the codec system 10 provided by an embodiment of the present application. The technical solution of the embodiment of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) video data. Among them, the video data includes original unencoded video, encoded video, decoded (for example, reconstructed) video, or syntax elements, etc.
[0041] Such as Figure 1As shown, the encoding and decoding system 10 includes a source device 100 that provides encoded video data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (such as a smart watch or a wearable camera), a television, a camera, a display device, a vehicle-mounted device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an aircraft, a robot, a satellite, etc.
[0042] In Figure 1 the example of, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, while the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 may not include Figure 1 some components in Figure 1 or may also include other components outside of. For example, the source device 100 may receive video data from an external data source (such as an external camera). Similarly, the destination device 110 may interface with an external display device and does not include an integrated display device. For another example, the memory 102 and the memory 113 may be external memories.
[0043] Although Figure 1 the source device 100 and the destination device 110 are depicted as separate devices, in some examples, the two may also be integrated into one device. In such embodiments, the functions corresponding to the source device 100 and the functions corresponding to the destination device 110 may be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.
[0044] In some examples, the source device 100 and the destination device 110 may perform unidirectional video transmission or bidirectional video transmission. If it is bidirectional video transmission, the source device 100 and the destination device 110 may operate in a substantially symmetric manner, that is, each of the source device 100 and the destination device 110 includes an encoder and a decoder.
[0045] The data source 101 represents a source of video data (i.e., raw, unencoded video data) and provides successive pictures containing the video data to the encoder 200, which encodes the data of the pictures. The data source 101 of the source device 100 may include a video capture device (such as a video camera), a video archive containing previously captured raw video, or a video feed interface for receiving video from a video content provider. Alternatively, the data source 101 may generate computer graphics-based data as the source video, or combine live video, archived video, and computer-generated video. In these cases, the encoder 200 encodes the captured, pre-captured, or computer-generated video data. The encoder 200 may rearrange the pictures from the received order (sometimes referred to as the "display order") into the encoding order. The encoder 200 may generate a bitstream including the encoded video data. The source device 100 may then output the encoded video data via the output interface 104 onto the communication medium 120 for reception or retrieval by, for example, the input interface 111 of the destination device 110.
[0046] The memories 102 of the source device 100 and 113 of the destination device 110 represent general memories. In some examples, the memory 102 may store the raw video data from the data source 101, and the memory 113 may store the decoded video data from the decoder 300. Additionally or alternatively, the memories 102, 113 may store software instructions executable by, for example, the encoder 200 and the decoder 300. Although the memory 102 and the memory 113 are shown separately from the encoder 200 and the decoder 300 in this example, it should be understood that the encoder 200 and the decoder 300 may also include internal memories for functionally similar or equivalent purposes. If the encoder 200 and the decoder 300 are arranged on the same hardware device, the memory 102 and the memory 113 may be the same memory. Furthermore, the memories 102, 113 may store, for example, the encoded video data output from the encoder 200 and input to the decoder 300. In some examples, portions of the memories 102, 113 may be allocated as one or more video buffers, for example, for storing raw, decoded, or encoded video data.
[0047] In some examples, the source device 100 may output the encoded data from the output interface 104 to the memory 113. Similarly, the destination device 110 may access the encoded data from the memory 113 via the input interface 111. The memory 113 or the memory 102 may include any of a variety of distributed or local access data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memories (CD-ROMs), flash memories, volatile or non-volatile memories, or any other suitable digital storage media for storing encoded video data.
[0048] The output interface 104 may include any type of medium or device capable of sending the encoded video data from the source device 100 to the destination device 110. For example, the output interface 104 may include a transmitter or transceiver, such as an antenna, configured to send the encoded video data from the source device 100 directly in real time to the destination device 110. The encoded video data may be modulated according to the communication standards of a wireless communication protocol and sent to the destination device 110.
[0049] The communication medium 120 may include transient media, such as wireless broadcasts or wired network transmissions. For example, the communication medium 120 may include the radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0050] In some embodiments, the communication medium 120 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 100 to the destination device 110. For example, a server (not shown) may receive the encoded video from the source device 100 and provide the encoded video data to the destination device 110, e.g., via network transmission to the destination device 110. The server may include, for example, a web server (e.g., for a website), a server configured to provide file transfer protocol services such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol, a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server may implement one or more HTTP streaming protocols such as MPEG Media Transport (MMT) protocol, Dynamic Adaptive Streaming over HTTP (DASH) protocol, HTTP Live Streaming (HLS) protocol, or Real Time Streaming Protocol (RTSP), etc.
[0051] The destination device 110 may access the encoded video data from the server, e.g., via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing the encoded video data stored on the server.
[0052] The output interface 104 and the input interface 111 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 standard or the IEEE 802.15 standard (e.g., ZigBeeTM), the Bluetooth standard, etc., or other physical components. In an example where the output interface 104 and the input interface 111 include wireless components, the output interface 104 and the input interface 111 may be configured to transmit data, such as encoded video data, according to WIFI, Ethernet, a cellular network (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).
[0053] The technology provided by the embodiments of this application can be applied to support video encoding and decoding in one or more of the following multimedia applications: video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0054] The input interface 111 of the destination device 110 receives an encoded video bitstream from the communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), where the syntax elements are used to decode the encoded data units to obtain decoded video data. The display device 114 displays the decoded video data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid-crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0055] The encoder 200 and the decoder 300 may be implemented as one or more of various processing circuits, which may include a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the technology is implemented in whole or in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided by the embodiments of this application.
[0056] The encoder 200 and the decoder 300 can perform processing based on the following video coding and decoding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding, HEVC), H.266 (also known as Versatile Video Coding, VVC), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, Alliance for Open Media Video 1 (AV1), Audio Video Coding Standard 1 (AVS1), AVS2, AVS3, or the next-generation video standard protocol. The embodiments of the present application do not make specific limitations.
[0057] Generally, the encoder 200 and the decoder 300 can perform block-based coding and decoding of pictures. The term "block" generally refers to a structure including data to be processed (e.g., encoded, decoded, or otherwise used during the encoding or decoding process). For example, a block can include a two-dimensional matrix of samples of luminance or chrominance data. For example, the encoder 200 and the decoder 300 can perform coding and decoding on video data represented in the YUV format.
[0058] See Figure 2 This figure is a schematic structural diagram of the encoder 200 provided by the embodiments of the present application. The encoder 200 can be the Figure 1 encoder 200 in. In the Figure 2 example, the encoder 200 includes a memory 201, an encoding parameter determination unit 210, a residual generation unit 202, a transformation processing unit 203, a quantization unit 204, an inverse quantization unit 205, an inverse transformation processing unit 206, a reconstruction unit 207, a filter unit 208, a decoded picture buffer (DPB) 209, and an entropy encoding unit 220.
[0059] The memory 201 can store the video data to be encoded. For example, the encoder 200 can receive the video data from the Figure 1 data source 101 shown and store it. In some examples, the memory 201 can be on the same chip as other components of the encoder 200 (as shown in Figure 2 ), or can be independent of the chip where these components are located.
[0060] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is configured to obtain a first prediction block of a current block by using an inter-frame prediction mode, the intra-frame prediction unit 213 is configured to obtain a second prediction block of the current block by using an intra-frame prediction mode, and the mode selection unit 211 is configured to obtain a target prediction block based on the first prediction block and the second prediction block, and determine a final prediction mode. In addition, the coding parameter determination unit 210 may further include other functional units, such as a functional unit for determining a partitioning manner of a coding unit (CU), a functional unit for determining a transform type of residual data of the CU, or a functional unit for determining a quantization parameter of the residual data of the CU, etc.
[0061] For the convenience of description and understanding, in the embodiments of the present application, the CU to be processed in the current image is referred to as the current CU, and the image block to be processed in the current CU is referred to as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0062] The inter-frame prediction unit 212 may include a motion estimation unit and a motion compensation unit. For the inter-frame prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (for example, one or more previously encoded and decoded pictures stored in the DPB 209).
[0063] The motion estimation unit may form one or more motion vectors (MVs) of the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit may obtain a predicted value with the accuracy indicated by the motion vector through interpolation.
[0064] The coding parameter determination unit 210 may provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the original unencoded video data of the current block from the memory 201, and calculates the residual between the current block and the target prediction block to obtain a residual block. In some examples, the function of the residual generation unit 202 may be implemented by using one or more subtractor circuits that perform binary subtraction.
[0065] As an example, the coding parameter determination unit 210 may provide syntax elements representing coding parameters to the entropy coding unit 220 for encoding. The coding parameters include one or more of a partitioning manner of the CU, a final prediction mode, a transform type of the residual data of the CU, or a quantization parameter of the residual data of the CU, etc.
[0066] The transform processing unit 203 performs a transform on the residual block output by the residual generation unit 202 to obtain a transform coefficient block. This transform may include a Discrete Cosine Transform (DCT), an integer transform, a directional transform, a Karhunen-Loeve transform, etc. In some examples, the encoder 200 may not include the transform processing unit 203.
[0067] The quantization unit 204 may quantize the transform coefficients in the transform coefficient block according to the quantization parameter (QP) value associated with the current block to generate a quantized transform coefficient block.
[0068] The inverse quantization unit 205 and the inverse transform processing unit 206 may perform inverse quantization and inverse transform on the transform coefficient block respectively to obtain a reconstructed residual block. The reconstruction unit 207 may generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0069] The filter unit 208 may perform one or more filter operations on the reconstructed block. For example, the filter unit 208 may be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, the encoder 200 may not include the filter unit 208.
[0070] The encoder 200 stores the reconstructed picture obtained from the reconstructed block in the DPB 209. For example, in an example where the operation of the filter unit 208 is not required, the reconstruction unit 207 may store the reconstructed block in the DPB 209. In an example where the operation of the filter unit 208 is required, the filter unit 208 may store the filtered reconstructed block in the DPB 209. The inter prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter prediction on the blocks of the subsequent picture to be encoded. In some examples, the DPB 209 may be replaced by other types of memories.
[0071] The entropy coding unit 220 may perform entropy coding on the syntax elements of other components in the encoder 200 and output the encoded video data. For example, the entropy coding unit 220 may perform entropy coding on the quantized transform coefficient block from the quantization unit 204. As another example, the entropy coding unit 220 may perform entropy coding on the syntax elements (such as motion information for inter prediction or intra mode information for intra prediction) from the coding parameter determination unit 210.
[0072] It can be understood that Figure 2 the composition of the encoder 200 described above is only illustrative and does not constitute a limitation to the embodiments of the present application.
[0073] Figure 3 FIG. is a schematic structural diagram of a decoder 300 provided by an embodiment of the present application. The decoder 300 may be Figure 1 the decoder 300 described above. In Figure 3 the example of, the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0074] The entropy decoding unit 302 may receive the encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements, where the syntax elements indicate encoding parameters, and the encoding parameters include one or more of the partitioning method of the CU, the final prediction mode, the transform type of the residual data of the CU, or the quantization parameter of the residual data of the CU, etc.
[0075] In the case where the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter prediction mode, the prediction block of the current CU may be obtained through the inter prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra prediction mode, the prediction block of the current CU may be obtained through the intra prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may further include a unit for performing a prediction function according to other prediction modes.
[0076] The CPB 301 may obtain and store the encoded video data from the communication medium 120 as shown in Figure 1 FIG. The DPB 307 is used to store the decoded pictures. Optionally, the CPB 301 and the DPB 307 may also be replaced with other types of memories, which are not specifically limited in the present application. In some examples, the CPB 301 may be on the same chip as other components of the decoder 300 (as shown in the figure), or may be independent of the chip where these components are located.
[0077] The decoder 300 can perform reconstruction operations on each block individually. The entropy decoding unit 302 can perform entropy decoding on the syntax elements of the quantized transform coefficients and the transform information (such as QP or transform mode indication) to obtain the quantized transform coefficients. The inverse quantization unit 303 performs inverse quantization on the quantized transform coefficients to obtain a transform coefficient block including transform coefficients. The inverse transform processing unit 304 performs an inverse transform on the transform coefficient block to generate a residual block corresponding to the current block, and the inverse transform is the reverse operation of the above-mentioned transform.
[0078] The reconstruction unit 305 can reconstruct the current block based on the prediction block and the residual block. For example, the reconstruction unit 305 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.
[0079] The filter unit 306 can perform one or more filter operations on the reconstructed block. For example, the type of the filter unit 306 can refer to the type of the filter unit 208, which will not be elaborated here. In some examples, the operation of the filter unit 306 can be skipped.
[0080] The decoder 300 can store the reconstructed picture obtained from the reconstructed block in the DPB 307. For example, in an example where the operation of the filter unit 306 is not performed, the reconstruction unit 305 can store the reconstructed block in the DPB 307. In an example where the operation of the filter unit 306 is performed, the filter unit 306 can store the filtered reconstructed block in the DPB 307. The decoder 300 can output the decoded picture (such as decoded video) from the DPB 307 for subsequent presentation to a display device (such as Figure 1 display device 114).
[0081] Before describing the embodiments of the present application, the related technologies will be briefly introduced as follows:
[0082] I. Neural Network-Based Loop Filter Scheme
[0083] To reduce the impact of distortion effects such as blocking effect and ringing effect on video quality, some loop filters are introduced in video compression standards, such as deblocking filter, sample adaptive compensation and other modules. The deblocking filter is used to reduce the blocking effect, and the sample adaptive compensation is used to improve the ringing effect. These loop filters can be located in the codec loop and can effectively improve the subjective and objective quality of the video.
[0084] With the wide application of neural network-based methods in image processing, some neural network-based loop filter schemes also have good performance in the direction of video compression. In the related technologies, attempts are made to replace some traditional filters with neural network filters or completely replace traditional filters.
[0085] II. Super-Resolution Scheme Based on Neural Network
[0086] The super-resolution scheme based on neural network refers to using a neural network to perform super-resolution reconstruction on an image to improve the clarity and details of the image. This method can train a neural network to learn the features of the image, thereby achieving super-resolution reconstruction of the image. Common neural networks include convolutional neural networks and recurrent neural networks, etc. In the image super-resolution scheme, a neural network can be used for single-image super-resolution reconstruction, joint super-resolution reconstruction of multiple images, etc. This method has been widely applied in the fields of image processing, medical image analysis, security monitoring, etc., and is constantly developing and improving.
[0087] In the related art, the loop filtering function needs to use a loop filter based on a neural network, and the super-resolution function needs to use a super-resolution processor based on a neural network. That is to say, in order to implement different video compression processing functions, different neural network models need to be used. It can be seen that in the related art, the functions of the neural network models used for video compression processing are relatively single, which leads to a relatively complex video compression processing process.
[0088] The following introduces the video compression processing method provided by the embodiments of the present application in conjunction with the accompanying drawings. The video compression processing method provided by the embodiments of the present application can be executed by the encoding end, such as Figure 1 or Figure 2 the encoder 200 shown, or can be executed by the decoding end, such as Figure 1 or Figure 3 the decoder 300 described. Among them, the encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a video decoding device.
[0089] Figure 4 shows a flowchart of a video compression processing method provided by an embodiment of the present application. As Figure 4 shown, the video compression processing method includes the following steps:
[0090] Step 401: Input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0091] Step 401: Perform a first processing on the first image based on the target neural network model to obtain a second image; wherein, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0092] The first image includes at least one of the following: a reconstructed block, a predicted block, a boundary strength, a residual block, and a transform coefficient block.
[0093] The reconstructed block refers to: a block of reconstructed samples of a signal that is decoded by a decoder from a bitstream and forms a decoded image, which can be filtered samples or unfiltered samples before filtering.
[0094] The predicted block refers to: a block of predicted samples of a signal calculated by a prediction technique using already reconstructed samples during the decoding process.
[0095] The boundary strength refers to: a boundary strength value of a signal obtained by a deblocking filtering technique based on the pattern information on both sides of a boundary and sample variations.
[0096] The residual block refers to: a block composed of the difference between the original samples and the predicted samples of an image block of a signal.
[0097] The transform coefficient block refers to: a block composed of transform coefficients in the frequency domain obtained by applying a transform technique to a residual block.
[0098] The first image is input into a target neural network model. The target neural network model can perform a first processing on the first image based on the determined video compression processing method to obtain a second image. The second image may include a filtered reconstructed block or a luminance reconstructed block of a first resolution, where the first resolution is higher than the resolution of the reconstructed block corresponding to the first image.
[0099] The reason for determining the video compression processing method corresponding to the first processing is that the target neural network model has multiple (including two) video compression processing functions. Since the target neural network model has multiple video compression processing functions, the target neural network model can be implemented as a unified neural network model to achieve multiple video compression processing functions. In this way, when multiple video compression processes need to be performed on an image, it can be achieved through a unified neural network model without the need to input the image multiple times into multiple single-function neural network models for processing, which helps to reduce the complexity of video compression processing.
[0100] In an embodiment of the present application, the first image is input into the target neural network model, and the target neural network model performs a first processing on the first image to obtain a second image. The target neural network model is a neural network model with N video compression processing functions, and the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions. In this way, the functional diversity of the video compression processing neural network model can be improved, and thus, multiple video compression processing functions can be achieved through a unified neural network model, which is beneficial to reducing the complexity of video compression processing.
[0101] It should be noted that the video compression processing method provided by the embodiments of the present application is applicable to both the model application stage (or model inference stage) and the model training stage, and the embodiments of the present application do not limit this.
[0102] The target neural network model provided by the embodiments of the present application can be set at the encoding end or the decoding end, and the embodiments of the present application do not limit this.
[0103] Optionally, the N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function;
[0104] The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
[0105] For the loop filtering processing method, the sizes (width-wise size and height-wise size) of the input image and the output image are the same.
[0106] Exemplarily, when performing loop filtering as a loop filtering model, the sizes (width-wise size and height-wise size) of the source image, the input image, and the output image are the same.
[0107] Exemplarily, when performing loop filtering as a loop filtering 2x model, the source image is 2 times the size (width-wise size and height-wise size) of the input image, and the sizes (width-wise size and height-wise size) of the input image and the output image are the same.
[0108] It should be noted that the input image and the output image are the input and output of the video compression processing method, respectively.
[0109] The process of performing loop filtering based on the target neural network model can be as follows: First, the first luminance data and the first chrominance data are subjected to feature extraction through the convolutional layer and the activation layer of the target neural network model (it should be noted that each convolutional layer can correspond to an activation layer, such as the activation layer can be a Parametric Rectified Linear Unit (PReLU)) to generate K feature maps, where K is a positive integer. Then, these feature maps are subjected to feature fusion through the convolutional layer and the activation layer to generate S fused feature maps, where S is a positive integer. Next, these fused feature maps are subjected to feature enhancement through multiple residual blocks to generate L enhanced feature maps, where L is a positive integer. Finally, these enhanced feature maps are subjected to operations such as convolutional layer and activation layer to generate the filtered reconstructed image block.
[0110] The multiple of the super-resolution processing can include various ones, and different super-resolution multiples can be used as different super-resolution processing methods.
[0111] Exemplarily, when performing super-resolution processing as a super-resolution 1.5x model, the source image is 1.5 times the size of the input image (the size in the width direction and the size in the height direction), and the output image is 1.5 times the size of the input image (the size in the width direction and the size in the height direction).
[0112] Exemplarily, when performing super-resolution processing as a super-resolution 2x model, the source image is 2 times the size of the input image (the size in the width direction and the size in the height direction), and the output image is 2 times the size of the input image (the size in the width direction and the size in the height direction).
[0113] Super-resolution processing can also be referred to as super-resolution reconstruction processing. The process of performing super-resolution reconstruction processing based on the target neural network model can be as follows: First, the first luminance data and the first chrominance data are subjected to feature extraction through the convolutional layer and the activation layer of the target neural network model to generate T feature maps, where T is a positive integer. Then, these feature maps are subjected to feature fusion through the convolutional layer and the activation layer to generate M fused feature maps, where M is a positive integer. Next, these fused feature maps are subjected to feature enhancement through multiple residual blocks to generate P enhanced feature maps, where P is a positive integer. Finally, these enhanced feature maps are subjected to operations such as convolutional layer, activation layer, and upsampling to generate high-resolution image blocks.
[0114] The post-processing filtering processing method is roughly the same as the loop filtering processing method. The process of performing post-processing filtering processing based on the target neural network model can refer to the above loop filtering processing process. To avoid repetition, it will not be elaborated here. It should be noted that the above various video compression processing functions and various video compression processing methods are only examples. When there are other video compression processing functions, the target neural network model can continue to expand other video compression processing functions, which are not limited in the embodiments of the present application.
[0115] In the embodiments of the present application, the video compression processing method corresponding to the first processing can be determined by the target neural network model itself, or can be determined by other models or other modules, and the determined video compression processing method is notified to the target neural network model.
[0116] In the embodiments of the present application, determining the video compression processing method corresponding to the first processing may include various implementation manners, which will be described separately below.
[0117] In some embodiments, the video compression processing method corresponding to the first processing is determined based on target information, and the target information is information for indicating the video compression processing method.
[0118] In some embodiments, the method further includes:
[0119] Input the target information into the target neural network model.
[0120] Here, the video compression processing method corresponding to the first processing can be determined by the target neural network model based on the target information.
[0121] In this embodiment, in addition to inputting the first image into the target neural network model, the target information is also input into the target neural network model. That is to say, the input data of the target neural network model includes the target information and the first image.
[0122] The above-mentioned target information can be understood as function indication information for indicating the video compression processing function.
[0123] Exemplarily, the target information is a function identifier.
[0124] Optionally, the target information includes at least one of the following:
[0125] A first identifier, the index value of the first identifier is mapped to the corresponding video compression processing method;
[0126] A second identifier, the second identifier includes a multiple in the image width direction and a multiple in the image height direction.
[0127] Exemplarily, the first identifier can be index values such as 0, 1, 2, 3, etc.
[0128] The meanings represented by each index value can be, for example:
[0129] Identifier 0: Loop filter model;
[0130] Identifier 1: Super-resolution 2x model;
[0131] Identifier 2: Super-resolution 1.5x model;
[0132] Identifier 3: Post-processing filter model.
[0133] If the first identifier is 0, the target neural network model performs filtering processing as a loop filter model.
[0134] If the first identifier is 1, the target neural network model performs super-resolution processing as a super-resolution 2x model.
[0135] If the first identifier is 2, the target neural network model performs super-resolution processing as a super-resolution 1.5x model.
[0136] If the first identifier is 3, the target neural network model is a post-processing filter model.
[0137] The meanings represented by each index value can also be, for example:
[0138] Identification 0: Loop filter model;
[0139] Identification 1: Loop filter 2x model;
[0140] Identification 2: Super-resolution 2x model;
[0141] Identification 3: Post-processing filter model.
[0142] If the first identification is 0, the target neural network model performs filtering as a loop filter model.
[0143] If the first identification is 1, the target neural network model performs filtering as a loop filter 2x model.
[0144] If the first identification is 2, the target neural network model performs super-resolution as a super-resolution 2x model, and the magnification factor is 2 times.
[0145] If the first identification is 3, the target neural network model is a post-processing filter model.
[0146] The target information is relatively concise in the above-described manner of the first identification, and the video compression processing method can be directly determined based on the target information.
[0147] During training, the target neural network model can learn the functions of the model by itself to improve the generalization of the model.
[0148] Exemplarily, the second identification can be data formats such as (1, 1), (2, 2), (1.5, 1.5), etc.
[0149] If the second identification is (1, 1), indicating that the magnification factors in both the image width direction and the image height direction are 1, the target neural network model can perform filtering as a loop filter model.
[0150] If the second identification is (2, 2), indicating that the magnification factors in both the image width direction and the image height direction are 2, the target neural network model can perform super-resolution as a super-resolution 2x model.
[0151] If the second identification is (1.5, 1.5), indicating that the magnification factors in both the image width direction and the image height direction are 1.5, the target neural network model can perform super-resolution as a super-resolution 1.5x model.
[0152] The target information is relatively flexible in the above-described manner of the second identification, and can provide effective input information for the target neural network model.
[0153] In this embodiment, the target neural network model is assisted by target information to determine the video compression processing method. This method is simple and efficient and can improve the performance of the target neural network model.
[0154] It should be noted that in this embodiment, the first image can be an image obtained through preprocessing. For example, according to the super-resolution multiple, the width and height directions of the third image are upsampled to obtain the first image. The first image can also be an image without preprocessing, and the embodiments of the present application do not limit this. In the case where target information is input, regardless of whether the first image has been preprocessed, the target neural network model can determine the video compression processing method based on the target information.
[0155] In some embodiments, before inputting the first image into the target neural network model, the method further includes:
[0156] Performing a second process on the third image to obtain the first image;
[0157] Wherein, the video compression processing method corresponding to the first process is determined based on the first image.
[0158] For some video compression processing functions, preprocessing of the input image is required.
[0159] Exemplarily, for a super-resolution 2x model, the width and height directions of the input image need to be upsampled to 2 times the original. In this case, the preprocessing is to upsample the width and height directions of the input image to 2 times the original.
[0160] Exemplarily, for a super-resolution 1.5x model, the width and height directions of the input image need to be upsampled to 1.5 times the original. In this case, the preprocessing is to upsample the width and height directions of the input image to 1.5 times the original.
[0161] In this embodiment, the process of performing a second process on the third image to obtain the first image can be understood as the preprocessing process. That is to say, the first image is an image obtained after preprocessing of the third image. Here, the third image can be understood as the input image.
[0162] A preprocessing module can be set at the encoding end or the decoding end to implement preprocessing of the input image.
[0163] In this embodiment, the input image (i.e., the third image) is first preprocessed to obtain the first image, and then the first image is input into the target neural network model for the first process. Since the first image is an image obtained through preprocessing, the first image contains some specific image features. Therefore, the corresponding video compression processing method can be determined based on the image features of the first image.
[0164] Exemplarily, if the first image corresponds to an image obtained by sampling the input image in both the width and height directions by a factor of 2, the target neural network model may determine that the corresponding video compression processing method is super-resolution 2x.
[0165] Exemplarily, if the first image corresponds to an image obtained by sampling the input image in both the width and height directions by a factor of 1.5, the target neural network model may determine that the corresponding video compression processing method is super-resolution 1.5x.
[0166] In this embodiment, the input image is preprocessed and then input into the target neural network model, which can make the input data of the model better adapt to the network and improve the performance of the target neural network model in some aspects. Moreover, the target neural network model can obtain the image features of the preprocessed first image through self-learning. In this way, the corresponding video compression processing method can be determined without inputting other auxiliary information or indication information into the target neural network model, which improves the performance of the target neural network model in some aspects.
[0167] Optionally, the second processing of the third image to obtain the first image includes:
[0168] Performing second processing on the third image according to the target information to obtain the first image, where the target information is information for indicating a video compression processing method;
[0169] Wherein, the video compression processing method corresponding to the first processing is determined based on at least one of the first image and the target information.
[0170] In this embodiment, by inputting the target information into the preprocessing module, the preprocessing module can determine the processing method of the second processing based on the target information.
[0171] Exemplarily, the video compression processing method corresponding to the first processing is determined based on the first image.
[0172] Exemplarily, the video compression processing method corresponding to the first processing is determined based on the target information.
[0173] Exemplarily, the video compression processing method corresponding to the first processing is determined based on the first image and the target information.
[0174] Here, the target information can refer to the relevant description of the foregoing target information. To avoid repetition, it will not be elaborated herein.
[0175] Optionally, the target information is information for indicating a super-resolution processing method;
[0176] Performing a second processing on the third image according to the target information to obtain the first image includes at least one of the following:
[0177] Performing upsampling processing on the width direction and height direction of the third image according to the target information to obtain the first image;
[0178] Determining a super-resolution multiple according to the target information, and performing upsampling processing on the width direction and height direction of the third image according to the super-resolution multiple to obtain the first image.
[0179] It should be noted that the upsampling processing can adopt a resampling coding method, that is, reference picture resampling (RPR). With the increase of high-resolution videos, it has brought huge challenges to video transmission under bandwidth-limited conditions. In VVC, RPR can adaptively adjust the resolution according to the network situation. When the network bandwidth is low, it can downsample and encode low-resolution (LR) frames, and when the network bandwidth improves, it can upsample and encode the original high-resolution (HR) frames.
[0180] In some embodiments, the method further includes:
[0181] Performing a second processing on the third image according to the target information to obtain the first image, where the target information is information for indicating a video compression processing method;
[0182] Inputting the target information into the target neural network model;
[0183] Wherein, the video compression processing method corresponding to the first processing is determined by the target neural network model based on at least one of the first image and the target information.
[0184] As described above, the process of performing a second processing on the third image to obtain the first image can be understood as a preprocessing process. In this embodiment, the target information is input to the preprocessing module so that the preprocessing module can determine the processing method of the second processing based on the target information. The target information is also input to the target neural network model so that the target neural network model can directly determine the video compression processing method based on the target information.
[0185] Here, the target information can refer to the relevant description of the foregoing target information. To avoid repetition, it will not be elaborated here.
[0186] In this embodiment, the input image is pre-processed before being input into the target neural network model, enabling the input data of the model to better adapt to the network and improving the performance of the target neural network model in some aspects. The target information can assist the target neural network model in training or inference, improving the performance of the target neural network model in some aspects.
[0187] The following takes the determination of the video compression processing method by the target neural network model as an example to provide specific embodiments:
[0188] Embodiment 1: Functional identification assisting the training and inference of the target neural network model
[0189] As Figure 5 shown, it includes the following steps:
[0190] Step 11: Obtain the first image;
[0191] Step 12: Input the first image and the functional identification into the target neural network model, and perform video compression processing on the first image based on the target neural network model;
[0192] Step 13: Obtain the second image output by the target neural network model.
[0193] In this embodiment, the target neural network model is used to implement at least two functions in video compression processing, and the target neural network model determines the video compression processing method based on the functional identification.
[0194] Embodiment 2: Functional identification assisting pre-processing, and functional identification assisting the training and inference of the target neural network model As Figure 6 shown, it includes the following steps:
[0195] Step 21: Obtain the third image;
[0196] Step 22: According to the functional identification, perform pre-processing on the third image to obtain the first image;
[0197] Step 23: Input the first image and the functional identification into the target neural network model, and perform video compression processing on the first image based on the target neural network model;
[0198] Step 24: Obtain the second image output by the target neural network model.
[0199] In this embodiment, the target neural network model is used to implement at least two functions in video compression processing, and the target neural network model determines the video compression processing method based on the functional identification.
[0200] It should be noted that once the structure of the target neural network model is fixed, it usually cannot handle different super-resolution functions simultaneously.
[0201] In Example 2 compared with Example 1, due to the addition of preprocessing operations, the sizes of the input images of the target neural network model can be made the same through preprocessing, so that the target neural network model can handle different super-resolution functions.
[0202] Example 3: Function identification assisted preprocessing, and preprocessing assisted training and inference of the target neural network model are as Figure 7 shown, and include the following steps:
[0203] Step 31: Obtain a third image;
[0204] Step 32: According to the function identification, perform preprocessing on the third image to obtain a first image;
[0205] Step 33: Input the first image into the target neural network model, and perform video compression processing on the first image based on the target neural network model;
[0206] Step 34: Obtain a second image output by the target neural network model.
[0207] In this embodiment, the target neural network model is used to implement at least two functions in video compression processing. The function identification is not used as an input to the target neural network model, and the target neural network model determines the video compression processing method based on the image features of the first image.
[0208] Example 3 compared with Example 2 can moderately reduce the computational complexity and the number of parameters of the neural network model.
[0209] In summary, in the embodiments of the present application, the functional diversity of the neural network model for video compression processing can be improved, so that multiple video compression processing functions can be implemented through a unified neural network model, which is beneficial to reducing the complexity of video compression processing.
[0210] For the video compression processing method provided by the embodiments of the present application, the execution subject may be a video compression processing device. In the embodiments of the present application, taking the video compression processing device executing the video compression processing method as an example, the video compression processing device provided by the embodiments of the present application is described.
[0211] Referring to Figure 8 , the embodiments of the present application also provide a video compression processing device, which can be applied to the encoding end or the decoding end. As Figure 8 shown, the video compression processing device 800 includes:
[0212] A first input module 801 for inputting a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0213] A first processing module 802 for performing a first processing on the first image based on the target neural network model to obtain a second image.
[0214] Wherein, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0215] Optionally, the video compression processing method corresponding to the first processing is determined based on target information, where the target information is information for indicating a video compression processing method.
[0216] Optionally, the video compression processing device 800 further includes:
[0217] A second input module for inputting target information into the target neural network model, where the target information is information for indicating a video compression processing method;
[0218] Wherein, the video compression processing method corresponding to the first processing is determined by the target neural network model based on the target information.
[0219] Optionally, the video compression processing device 800 further includes:
[0220] A second processing module for performing a second processing on a third image to obtain the first image;
[0221] Wherein, the video compression processing method corresponding to the first processing is determined based on the first image.
[0222] Optionally, the video compression processing method corresponding to the first processing is determined by the target neural network model based on the first image.
[0223] Optionally, the video compression processing device 800 further includes:
[0224] A third processing module for performing a second processing on a third image according to target information to obtain the first image, where the target information is information for indicating a video compression processing method;
[0225] Wherein, the video compression processing method corresponding to the first processing is determined based on at least one of the first image and the target information.
[0226] Optionally, the target information is information for indicating a super-resolution processing method;
[0227] The third processing module is specifically configured to perform at least one of the following:
[0228] Perform upsampling processing on the width direction and height direction of the third image according to the target information to obtain the first image;
[0229] Determine the super-resolution multiple according to the target information, and perform upsampling processing on the width direction and height direction of the third image according to the super-resolution multiple to obtain the first image.
[0230] Optionally, the video compression processing device 800 further includes:
[0231] A third input module, configured to input the target information into the target neural network model;
[0232] Wherein, the video compression processing method corresponding to the first processing is determined by the target neural network model based on at least one of the first image and the target information.
[0233] Optionally, the target information includes at least one of the following:
[0234] A first identifier, the index value of the first identifier is mapped to the corresponding video compression processing method;
[0235] A second identifier, the second identifier includes the multiple in the image width direction and the multiple in the image height direction.
[0236] Optionally, the N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function;
[0237] The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
[0238] In summary, in the embodiments of the present application, the functional diversity of the video compression processing neural network model can be improved, so that multiple video compression processing functions can be realized through a unified neural network model, which is beneficial to reducing the complexity of video compression processing.
[0239] The video compression processing device provided by the embodiments of the present application can implement Figures 4 to 7 each process implemented by the method embodiments, and achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0240] As Figure 9As shown in the figure, an embodiment of the present application further provides an electronic device 900, including a processor 901 and a memory 902. A program or instruction that can run on the processor 901 is stored on the memory 902. For example, when the electronic device 900 is an encoding-end device, when the program or instruction is executed by the processor 901, each step of the above-described video compression processing method embodiment is implemented, and the same technical effect can be achieved. When the electronic device 900 is a decoding-end device, when the program or instruction is executed by the processor 901, each step of the above-described video compression processing method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, details are not described here again. Optionally, the memory 902 may be Figure 1 the memory 102 or the memory 113 in the embodiment shown, and the processor 901 may implement Figures 1 to 3 the functions of the encoder 200 or the decoder 300 in the embodiment shown.
[0241] An embodiment of the present application further provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement each step of the above-described video compression processing method embodiment. Optionally, the memory may be Figure 1 the memory 102 or the memory 113 in the embodiment shown, and the processing circuit may implement Figures 1 to 3 the functions of the encoder 200 or the decoder 300 in the embodiment shown.
[0242] An embodiment of the present application further provides an electronic device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps in the Figure 4 method embodiment shown. This device embodiment corresponds to the above method embodiment. Each implementation process and implementation manner of the above method embodiment can be applied to this terminal embodiment, and the same technical effect can be achieved.
[0243] The above electronic device may be a terminal or other devices other than a terminal, such as a server, a Network Attached Storage (NAS), etc.
[0244] Among them, the terminal can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile internet device (MID), an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipborne device, a pedestrian user equipment (PUE), a smart home (home devices with wireless communication functions, such as refrigerators, TVs, washing machines, or furniture, etc.), a game console, a personal computer (PC), a teller machine, or a self-service machine, etc., which are terminal-side devices. Wearable devices include: smart watches, smart bracelets, smart earphones, smart glasses, smart jewelry (such as smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among them, the vehicle user equipment can also be referred to as a vehicle terminal, a vehicle controller, a vehicle module, a vehicle component, a vehicle chip, or a vehicle unit, etc. It should be noted that the specific type of the terminal is not limited in the embodiments of the present application.
[0245] The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server. The cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), or cloud computing services based on big data and artificial intelligence platforms, etc.
[0246] Exemplarily, the above-mentioned electronic devices can include, but are not limited to Figure 1 the types of the source device 100 or the destination device 110 shown.
[0247] Taking the electronic device as a terminal as an example, Figure 10 is a schematic diagram of the hardware structure of a terminal for implementing an embodiment of the present application.
[0248] The terminal 1000 includes, but is not limited to, at least some components such as a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010.
[0249] Those skilled in the art can understand that the terminal 1000 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 10 The terminal structure shown in the figure does not limit the terminal. The terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements, which will not be elaborated here.
[0250] It should be understood that in the embodiments of the present application, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes the image data of a static picture or video obtained by an image acquisition device (such as a camera) in a video acquisition mode or an image acquisition mode, or may process the obtained point cloud data. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0251] In the embodiments of the present application, after the radio frequency unit 1001 receives downlink data from a network side device, it can be transmitted to the processor 1010 for processing; in addition, the radio frequency unit 1001 can send uplink data to the network side device. Generally, the radio frequency unit 1001 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc.
[0252] The memory 1009 can be used to store software programs or instructions as well as various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0253] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010.
[0254] Among them, the processor 1010 is used for:
[0255] Input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1;
[0256] Perform a first process on the first image based on the target neural network model to obtain a second image;
[0257] Among them, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
[0258] Optionally, the video compression processing method corresponding to the first processing is determined based on target information, and the target information is information for indicating a video compression processing method.
[0259] Optionally, the processor 1010 is further configured to:
[0260] Input the target information into the target neural network model, where the target information is information for indicating a video compression processing method;
[0261] Among them, the video compression processing method corresponding to the first processing is determined by the target neural network model based on the target information.
[0262] Optionally, the processor 1010 is further configured to:
[0263] Perform a second processing on the third image to obtain the first image;
[0264] Among them, the video compression processing method corresponding to the first processing is determined based on the first image.
[0265] Optionally, the processor 1010 is further configured to:
[0266] Perform a second processing on the third image according to the target information to obtain the first image, where the target information is information for indicating a video compression processing method;
[0267] Among them, the video compression processing method corresponding to the first processing is determined based on at least one of the first image and the target information.
[0268] Optionally, the target information is information for indicating a super-resolution processing method;
[0269] Optionally, the processor 1010 is further configured to perform at least one of the following:
[0270] Perform upsampling processing on the width direction and height direction of the third image according to the target information to obtain the first image;
[0271] Determine the super-resolution multiple according to the target information, and perform upsampling processing on the width direction and height direction of the third image according to the super-resolution multiple to obtain the first image.
[0272] Optionally, the processor 1010 is further configured to:
[0273] Perform a second processing on the third image according to the target information to obtain the first image, where the target information is information for indicating a video compression processing method;
[0274] Input the target information into the target neural network model;
[0275] Wherein, the video compression processing method corresponding to the first processing is determined by the target neural network model based on at least one of the first image and the target information.
[0276] Optionally, the target information includes at least one of the following:
[0277] A first identifier, the index value of the first identifier is mapped to the corresponding video compression processing method;
[0278] A second identifier, the second identifier includes a multiple in the image width direction and a multiple in the image height direction.
[0279] Optionally, at least two of the N video compression processing functions include a loop filtering function, a super-resolution function, and a post-processing filtering function;
[0280] The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
[0281] In summary, in the embodiments of the present application, the functional diversity of the video compression processing neural network model can be improved, so that multiple video compression processing functions can be realized through a unified neural network model, which is beneficial to reducing the complexity of video compression processing.
[0282] It can be understood that the implementation processes of the various implementation manners mentioned in this embodiment can refer to the relevant descriptions of the method embodiments and achieve the same or corresponding technical effects. To avoid repetition, they will not be elaborated here.
[0283] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned video compression processing method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.
[0284] Wherein, the processor is the processor in the terminal described in the above embodiment. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disks, or optical disks, etc. In some examples, the readable storage medium can be a non-transitory readable storage medium.
[0285] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement each process of the above-mentioned video compression processing method embodiment, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.
[0286] It should be understood that the chips mentioned in the embodiments of the present application may include system-on-chip (also referred to as system chip, chip system or system-on-chip), and may also include independent display chips, etc.
[0287] The embodiments of the present application further provide a computer program / program product, which is stored in a storage medium and is executed by at least one processor to implement the various processes of the above-mentioned video compression processing method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0288] The embodiments of the present application further provide a video compression processing system, including: an encoding end device and a decoding end device. The encoding end device can be used to execute the steps of the above-mentioned video compression processing method, and the decoding end device can be used to execute the steps of the above-mentioned video compression processing method.
[0289] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0290] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of a computer software product plus a necessary general hardware platform, and of course, can also be implemented by hardware. This computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disc, etc.) and includes several instructions for causing a terminal or a network-side device to execute the methods described in the various embodiments of the present application.
[0291] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms of implementation manners without departing from the purpose of the present application and the scope protected by the claims. These implementation manners are all within the protection scope of the present application.
Claims
1. A video compression processing method, characterized in that, The method includes: Input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1; Perform a first processing on the first image based on the target neural network model to obtain a second image; Among them, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
2. The method according to claim 1, wherein The video compression processing method corresponding to the first processing is determined based on target information, and the target information is information for indicating a video compression processing method.
3. The method according to claim 1, wherein Before the step of inputting the first image into the target neural network model, the method further includes: Perform a second processing on a third image to obtain the first image; Among them, the video compression processing method corresponding to the first processing is determined based on the first image.
4. The method according to claim 1, characterized in that, Before the step of inputting the first image into the target neural network model, the method further includes: Perform a second processing on the third image according to the target information to obtain the first image, where the target information is information for indicating a video compression processing method; Among them, the video compression processing method corresponding to the first processing is determined based on at least one of the first image and the target information.
5. The method according to claim 4, wherein The target information is information for indicating a super-resolution processing method; The step of performing a second processing on the third image according to the target information to obtain the first image includes at least one of the following: According to the target information, perform upsampling processing on the width direction and height direction of the third image to obtain the first image; Determine a super-resolution multiple according to the target information, and perform upsampling processing on the width direction and height direction of the third image according to the super-resolution multiple to obtain the first image.
6. The method according to any one of claims 2, 4 and 5, characterized in that The target information includes at least one of the following: A first identifier, where the index value of the first identifier is mapped to the corresponding video compression processing method; A second identifier, where the second identifier includes the multiple of the image width direction and the multiple of the image height direction.
7. The method according to any one of claims 1 to 6, characterized in that, The N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function; The video compression processing method corresponding to the first processing includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
8. A video compression processing device, characterized in that, It includes: A first input module, configured to input a first image into a target neural network model, where the target neural network model is a neural network model with N video compression processing functions, and N is an integer greater than 1; A first processing module, configured to perform a first processing on the first image based on the target neural network model to obtain a second image; Among them, the video compression processing method corresponding to the first processing is implemented by one or more of the N video compression processing functions.
9. The device according to claim 8, characterized in that, The video compression processing method corresponding to the first processing is determined based on target information, and the target information is information for indicating a video compression processing method.
10. The device according to claim 8, characterized in that, It further includes: A second processing module, configured to perform a second processing on a third image to obtain the first image; Among them, the video compression processing method corresponding to the first processing is determined based on the first image.
11. The device according to claim 8, characterized in that, It further includes: A third processing module, configured to perform a second process on a third image according to target information, to obtain the first image, where the target information is information for indicating a video compression processing method; Among them, the video compression processing method corresponding to the first process is determined based on at least one of the first image and the target information.
12. The device according to claim 11, wherein, The target information is information for indicating a super-resolution processing method; The third processing module is specifically configured to perform at least one of the following: According to the target information, perform upsampling processing on the width direction and the height direction of the third image to obtain the first image; Determine a super-resolution multiple according to the target information, and perform upsampling processing on the width direction and the height direction of the third image according to the super-resolution multiple to obtain the first image.
13. The device according to any one of claims 9, 11 and 12, characterized in that, The target information includes at least one of the following: A first identifier, where an index value of the first identifier is mapped to a corresponding video compression processing method; A second identifier, where the second identifier includes a multiple in the image width direction and a multiple in the image height direction.
14. The device according to any one of claims 8 to 13, characterized in that, The N video compression processing functions include at least two of a loop filtering function, a super-resolution function, and a post-processing filtering function; The video compression processing method corresponding to the first process includes one of a loop filtering processing method, a super-resolution processing method, and a post-processing filtering processing method.
15. An electronic device, characterized in that, Comprising a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
16. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
17. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the steps of the method according to any one of claims 1 to 7.