Video coding method and related equipment
By performing RPR downsampling and feature fusion processing on high-resolution video frames, low-resolution video frames are generated, solving the problem of high-frequency information loss and achieving more efficient encoding and decoding effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2024-10-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video coding techniques lose a lot of high-frequency information during the resampling and downsampling of the reference image, especially in areas with complex textures, making it difficult to recover high-quality, high-resolution images from video frames.
By performing reference image resampling (RPR) downsampling and feature fusion on the high-resolution first video frame, a low-resolution third video frame is generated, which retains more high-frequency information and is used in the encoding process to improve encoding efficiency.
It is easier to recover high-quality images in subsequent encoding and decoding processes, thereby improving encoding performance and encoding gain.
Smart Images

Figure CN121967744A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio and video technology, specifically relating to a video encoding method and related equipment. Background Technology
[0002] The increasing use of high-resolution and high-frame-rate videos provides a better visual experience, but it also leads to a dramatic increase in video data volume, posing significant challenges to transmission and storage in universities. Video encoding and decoding compress video signals by reducing information redundancy. Image super-resolution technology can convert low-resolution images into higher-resolution images. To meet the needs of real-time video communication, image super-resolution is combined with video coding. For example, in video coding standards, video frames are encoded at low resolution and upsampled to the original resolution at the decoding end using bilinear or bicubic interpolation methods. Building on this, Versatile Video Coding (VVC) proposes a Reference Picture Resampling (RPR) technique, which can adaptively change the resolution within the bitstream without inserting Instantaneous Decoder Refresh (IDR) frames or Intra Random Access Pictures (IRAP) frames. This avoids network congestion caused by excessively large IDR or IRAP frames. When network bandwidth is limited, the image to be encoded is converted to a low-resolution image during the encoding stage, and then restored to a high-resolution image at the decoding end, thereby reducing the amount of data transmitted and saving bandwidth. However, when network conditions are good or bandwidth is sufficient, the original resolution of the video frames is transmitted to ensure video quality.
[0003] RPR technology has made significant progress in improving coding efficiency. However, current technologies generally optimize the upsampling part of RPR, ignoring the problem that the downsampling part of RPR loses a lot of high-frequency information. This problem is more serious in areas with complex textures and makes it difficult to restore video frames to high-quality, high-resolution images. Summary of the Invention
[0004] This application provides a video encoding method and related equipment that can solve the problem of a large amount of high-frequency information loss caused by RPR-based downsampling encoding.
[0005] Firstly, a video encoding method is provided, the method comprising:
[0006] The encoding end performs downsampling on the first video frame based on reference image resampling (RPR) to obtain the second video frame;
[0007] The encoding end performs feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame;
[0008] The encoding end determines the target bitstream based on the third video frame.
[0009] Secondly, a video encoding apparatus is provided, the apparatus comprising:
[0010] The downsampling module is used to downsample the first video frame based on reference image resampling (RPR) to obtain the second video frame;
[0011] A fusion module is used to perform feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame;
[0012] The encoding module is used to determine the target bitstream based on the third video frame.
[0013] Thirdly, a video encoding apparatus is provided, the apparatus being configured to perform the steps of the method described in the first aspect.
[0014] Fourthly, an electronic device is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.
[0015] Fifthly, an electronic device is provided, including a processor and a communication interface, wherein the processor is configured to downsample a first video frame based on reference image resampling (RPR) to obtain a second video frame; perform feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame; and determine a target bitstream based on the third video frame.
[0016] A sixth aspect provides an electronic device comprising: a memory configured to store video data, and processing circuitry configured to implement the steps of the method described in the first aspect.
[0017] In a seventh aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0018] Eighthly, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method as described in the first aspect.
[0019] In a ninth aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to perform the steps of the method as described in the first aspect.
[0020] In this embodiment, the encoding end performs RPR downsampling and feature fusion processing on the high-resolution first video frame to obtain a low-resolution third video frame. This allows the third video frame to retain richer high-frequency information. The third video frame is used in the subsequent encoding process, making it easier to restore the low-resolution third video frame to a high-quality image in the subsequent encoding and decoding process. This can improve the encoding gain and achieve better encoding performance. Attached Figure Description
[0021] Figure 1 This diagram illustrates the encoding / decoding system provided in an embodiment of this application.
[0022] Figure 2 This is a schematic diagram of the encoder provided in an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of the decoder provided in an embodiment of this application;
[0024] Figure 4 This diagram illustrates the steps of the video encoding method provided in the embodiments of this application.
[0025] Figure 5 This diagram illustrates the video coding framework applied to the video coding method provided in the embodiments of this application.
[0026] Figure 6 This diagram illustrates the structure of the downsampling unit within the video coding framework used in the embodiments of this application.
[0027] Figure 7 This diagram illustrates the composition of residual blocks in the video coding method provided in this application embodiment.
[0028] Figure 8 This diagram illustrates the downsampling neural network model in the video coding method provided in this application embodiment.
[0029] Figure 9 This is a schematic diagram illustrating the structure of the video encoding apparatus provided in an embodiment of this application;
[0030] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0031] Figure 11 This is a schematic diagram showing the structure of the terminal provided in the embodiments of this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0033] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0034] Figure 1 This is a schematic diagram of the encoding / decoding system 10 provided in an embodiment of this application. The technical solution of this application embodiment relates to encoding and decoding (CODEC) video data (including encoding or decoding). The video data includes original unencoded video, encoded video, decoded (e.g., reconstructed) video, or syntax elements, etc.
[0035] like Figure 1 As shown, the encoding / decoding system 10 includes a source device 100, which provides encoded video data to be decoded and displayed by the destination device 110. Specifically, the source device 100 provides video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of the following: desktop computer, laptop computer, tablet computer, set-top box, mobile phone, wearable device (e.g., smartwatch or wearable camera), television, camera, display device, in-vehicle device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, digital media player, video game console, video conferencing equipment, video streaming equipment, broadcast receiver equipment, broadcast transmitter equipment, spacecraft, aircraft, robot, satellite, etc.
[0036] exist Figure 1 In this example, source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. Destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. Source device 100 represents an example of a video encoding device, while destination device 110 represents an example of a video decoding device. In other examples, source device 100 and destination device 110 may not include... Figure 1 Some components, or may include Figure 1 Other components besides the source device 100. For example, the source device 100 can receive video data from an external data source (such as an external camera). Similarly, the destination device 110 can interface with an external display device, without including an integrated display device. Furthermore, the memory 102 and memory 113 can be external memories.
[0037] Although Figure 1 Source device 100 and destination device 110 are illustrated as separate devices, but in some examples, they may also be integrated into a single device. In such embodiments, the same hardware or software, or separate hardware or software, or any combination thereof, may be used to implement the functionality corresponding to source device 100 and destination device 110.
[0038] In some examples, source device 100 and destination device 110 can perform unidirectional or bidirectional video transmission. If it is bidirectional video transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, that is, each of source device 100 and destination device 110 includes an encoder and a decoder.
[0039] Data source 101 represents the source of video data (i.e., raw, unencoded video data) and provides encoder 200 with a series of images containing video data, which encoder 200 encodes. Data source 101 of source device 100 may include video acquisition devices (such as video cameras), video archives containing previously acquired raw video, or video feed interfaces for receiving video from video content providers. Alternatively, data source 101 may generate computer graphics-based data as source video, or combine live video, archived video, and computer-generated video. In these cases, encoder 200 encodes the acquired, pre-acquired, or computer-generated video data. Encoder 200 may rearrange the images from the received order (sometimes referred to as the "display order") according to the encoded order. Encoder 200 may generate a bitstream including the encoded video data. Source device 100 may then output the encoded video data to communication medium 120 via output interface 104 for reception or retrieval, for example, by input interface 111 of destination device 110.
[0040] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memory. In some examples, memory 102 may store raw video data from data source 101, and memory 113 may store decoded video data from decoder 300. Additionally or alternatively, memories 102 and 113 may respectively store software instructions executable by, for example, encoder 200 and decoder 300. Although memories 102 and 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memories 102 and 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded video data output from encoder 200 and input to decoder 300. In some examples, portions of memories 102 and 113 may be allocated as one or more video buffers, for example, to store raw, decoded, or encoded video data.
[0041] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0042] Output interface 104 may include any type of medium or device capable of transmitting encoded video data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded video data directly from source device 100 to destination device 110 in real time. The encoded video data may be modulated according to the communication standards of a wireless communication protocol and transmitted to destination device 110.
[0043] Communication medium 120 may include transient media, such as wireless broadcasting or wired network transmission. For example, communication medium 120 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 120 may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). Communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0044] In some implementations, the communication medium 120 may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive encoded video from source device 100 and provide the encoded video data to destination device 110, for example, via network transmission. The server may include (e.g., a web server for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or Evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT), Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), or Real Time Streaming Protocol (RTSP).
[0045] Destination device 110 can access encoded video data from a server, for example, via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.) for accessing encoded video data stored on the server.
[0046] Output interface 104 and input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 or IEEE 802.15 standard (e.g., ZigBee™), Bluetooth standard, or other physical components. In an example where output interface 104 and input interface 111 include wireless components, output interface 104 and input interface 111 can be configured to transmit data, such as encoded video data, via Wi-Fi, Ethernet, or cellular networks (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).
[0047] The technology provided in this application can be applied to support video encoding and decoding in one or more multimedia applications such as video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission, digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0048] The input interface 111 of the destination device 110 receives an encoded video bitstream from the communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, image groups, images, slices, blocks, etc.), where the syntax elements are used to decode the encoded data units to obtain decoded video data. The display device 114 displays the decoded video data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0049] The encoder 200 and decoder 300 can be implemented as one or more of various processing circuits, which may include microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented wholly or partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of this application.
[0050] The encoder 200 and decoder 300 can process based on the following video codec standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding, HEVC), H.266 (also known as Versatile Video Coding, VVC), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, Alliance for Open Media Video 1 (AV1), Audio Video Coding Standard 1 (AVS1), AVS2, AVS3, or next-generation video standard protocols. This application does not specifically limit the implementation of these protocols.
[0051] Typically, encoder 200 and decoder 300 can perform block-based encoding and decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding or decoding). For example, a block can include a two-dimensional matrix of samples of luminance or chrominance data. For example, encoder 200 and decoder 300 can encode and decode video data represented in YUV format.
[0052] See Figure 2 The figure is a schematic diagram of the structure of the encoder 200 provided in an embodiment of this application. The encoder 200 can be... Figure 1 Encoder 200 in. In Figure 2 In the example, encoder 200 includes memory 201, encoding parameter determination unit 210, residual generation unit 202, transform processing unit 203, quantization unit 204, inverse quantization unit 205, inverse transform processing unit 206, reconstruction unit 207, filter unit 208, decoded picture buffer (DPB) 209, and entropy encoding unit 220.
[0053] Memory 201 can store video data to be encoded; for example, encoder 200 can retrieve the data from the encoder. Figure 1 The data source 101 shown receives and stores video data. In some examples, the memory 201 may be on the same chip as other components of the encoder 200 (e.g., ...). Figure 2 (As shown), they can also be independent of the chip in which these components are located.
[0054] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is used to obtain a first prediction block for the current block using an inter-frame prediction mode. The intra-frame prediction unit 213 is used to obtain a second prediction block for the current block using an intra-frame prediction mode. The mode selection unit 211 is used to obtain a target prediction block based on the first and second prediction blocks and determine the final prediction mode. Furthermore, the coding parameter determination unit 210 may also include other functional units, such as functional units for determining the partitioning method of coding units (CUs), functional units for determining the transformation type of the residual data of the CUs, or functional units for determining the quantization parameters of the residual data of the CUs.
[0055] For ease of description and understanding, in the embodiments of this application, the CU to be processed in the current image is referred to as the current CU, and the image block to be processed in the current CU is referred to as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0056] Inter-frame prediction unit 212 may include a motion estimation unit and a motion compensation unit. For inter-frame prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (e.g., one or more previously encoded / decoded pictures stored in DPB 209).
[0057] The motion estimation unit can generate one or more motion vectors (MVs) representing the position of a reference block in a reference image relative to the position of the current block in the current image. The motion compensation unit can then use interpolation to obtain a predicted value with the precision indicated by the motion vectors.
[0058] The encoding parameter determination unit 210 can provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the raw uncoded video data of the current block from the memory 201 and calculates the residual between the current block and the target prediction block to obtain the residual block. In some examples, the function of the residual generation unit 202 can be implemented using one or more subtractor circuits that perform binary subtraction.
[0059] As an example, the encoding parameter determination unit 210 can provide the entropy encoding unit 220 with syntax elements representing encoding parameters for encoding. The encoding parameters include one or more of the following: the partitioning method of the CU, the final prediction mode, the transformation type of the residual data of the CU, or the quantization parameters of the residual data of the CU.
[0060] The transformation processing unit 203 transforms the residual block output by the residual generation unit 202 to obtain a transform coefficient block. This transformation may include Discrete Cosine Transform (DCT), integer transformation, direction transformation, or Karhunen-Loeve transformation, etc. In some examples, the encoder 200 may not include the transformation processing unit 203.
[0061] Quantization unit 204 can quantize the transform coefficients in the transform coefficient block according to the quantization parameter (QP) value associated with the current block to generate a quantized transform coefficient block.
[0062] The inverse quantization unit 205 and the inverse transform processing unit 206 can perform inverse quantization and inverse transform on the transform coefficient block, respectively, to obtain the reconstructed residual block. The reconstruction unit 207 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0063] Filter unit 208 can perform one or more filter operations on the reconstructed block. For example, filter unit 208 can be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, encoder 200 may not include filter unit 208.
[0064] Encoder 200 stores the reconstructed image obtained from the reconstructed blocks in DPB 209. For example, in an example where the operation of filter unit 208 is not required, reconstruction unit 207 can store the reconstructed blocks in DPB 209. In an example where the operation of filter unit 208 is required, filter unit 208 can store the filtered reconstructed blocks in DPB 209. Inter-frame prediction unit 212 retrieves the reconstructed image from DPB 209 to perform inter-frame prediction on blocks of subsequent images to be encoded. In some examples, DPB 209 can be replaced with other types of memory.
[0065] Entropy coding unit 220 can entropy code the syntax elements of other components in encoder 200 to output encoded video data. For example, entropy coding unit 220 can entropy code the quantized transform coefficient block from quantization unit 204. As another example, entropy coding unit 220 can entropy code the syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from coding parameter determination unit 210.
[0066] Understandable Figure 2 The composition of the encoder 200 described is merely illustrative and does not constitute a limitation on the embodiments of this application.
[0067] Figure 3 This is a schematic diagram of the structure of the decoder 300 provided in the embodiments of this application. The decoder 300 can be... Figure 1 The decoder 300 is described. Figure 3 In the example, the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0068] The entropy decoding unit 302 can receive encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements. The syntax elements indicate encoding parameters, including one or more of the following: CU partitioning method, final prediction mode, transformation type of CU residual data, or quantization parameters of CU residual data.
[0069] When the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter-frame prediction mode, the prediction block of the current CU can be obtained through the inter-frame prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra-frame prediction mode, the prediction block of the current CU can be obtained through the intra-frame prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may also include a unit for performing prediction functions according to other prediction modes.
[0070] CPB 301 can be obtained from, for example Figure 1 The communication medium 120 shown acquires and stores encoded video data. DPB 307 is used to store decoded images. Optionally, CPB 301 and DPB 307 can be replaced with other types of memory; this application does not impose specific limitations. In some examples, CPB 301 can be on the same chip as other components of the decoder 300 (as shown), or it can be on a separate chip from those components.
[0071] Decoder 300 can perform reconstruction operations on each block individually. Entropy decoding unit 302 can entropy decode the syntax elements and transform information (e.g., QP or transform mode indication) of the quantized transform coefficients to obtain the quantized transform coefficients. Dequantization unit 303 dequantizes the quantized transform coefficients to obtain a transform coefficient block including the transform coefficients. Inverse transform processing unit 304 performs an inverse transform on the transform coefficient block to generate a residual block corresponding to the current block; this inverse transform is the reverse operation of the above transform.
[0072] Reconstruction unit 305 can reconstruct the current block based on the prediction block and the residual block. For example, reconstruction unit 305 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.
[0073] Filter unit 306 can perform one or more filter operations on the reconstructed block. For example, the type of filter unit 306 can be referenced to the type of filter unit 208, and will not be described again here. In some examples, the operations of filter unit 306 can be skipped.
[0074] Decoder 300 can store the reconstructed image obtained from the reconstructed blocks in DPB 307. For example, in an example where filter unit 306 is not operated, reconstruction unit 305 can store the reconstructed blocks in DPB 307. In an example where filter unit 306 is operated, filter unit 306 can store the filtered reconstructed blocks in DPB 307. Decoder 300 can output the decoded image (e.g., decoded video) from DPB 307 for use with a display device (such as...). Figure 1 The subsequent presentation of the display device 114).
[0075] The video encoding method provided in this application embodiment is described below with reference to the accompanying drawings. The video encoding method provided in this application embodiment can be executed by an encoding end, for example... Figure 1 or Figure 2 The encoder 200 is shown. The encoding end and the decoding end can be implemented by software, hardware or a combination thereof. When implemented by hardware, the encoding end can be called an encoding end device or a video encoding device, and the decoding end can be called a decoding end device or a video decoding device.
[0076] like Figure 4 As shown in the figure, this application provides a video encoding method, the method including:
[0077] Step 401: The encoding end performs downsampling on the first video frame based on reference image resampling (RPR) to obtain the second video frame;
[0078] Step 402: The encoding end performs feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame.
[0079] Step 402: The encoding end determines the target bitstream based on the third video frame.
[0080] Optionally, the encoding end performs inter-frame prediction or intra-frame prediction on the third video frame to obtain a predicted frame; the encoding end calculates the residual between the third video frame and the predicted frame, and performs transformation, quantization, and entropy coding on the residual to obtain the target bitstream.
[0081] In this embodiment, the encoding end performs RPR downsampling and feature fusion processing on the high-resolution first video frame to obtain a third video frame, which retains richer high-frequency information. This third video frame is used in the subsequent encoding process, making it easier to restore the low-resolution third video frame into a high-quality image in the subsequent encoding and decoding process, thereby improving the encoding gain and achieving better encoding performance.
[0082] Optionally, such as Figure 5 The diagram shown illustrates the video encoding framework applied to the video encoding method provided in this embodiment. Figure 5 As shown, video data (including multiple first video frames) is input to the downsampling unit. The downsampling unit performs RPR downsampling and feature fusion on the first video frames to obtain a third video frame. This third video frame is input to the inter-frame / intra-frame prediction module. The prediction result is compared with the third video frame output by the downsampling unit to calculate the residual. After conversion and quantization, the residual is input to the entropy encoding / decoding unit to obtain the bitstream, which is then transmitted to the decoder. Simultaneously, the residual result is processed by the subsequent inverse quantization unit and inverse transform unit, and added to the prediction result to obtain the second prediction frame. This second prediction frame is then input to the loop filter to obtain the reconstructed frame. Finally, the reconstructed frame is restored to its original resolution by the upsampling unit as the final encoding result and stored in the Decoded Picture Buffer (DPB).
[0083] Optionally, such as Figure 6 As shown, the downsampling unit includes an RPR downsampling filter and a downsampling neural network module. The RPR downsampling filter is used to downsample the high-resolution first video frame based on reference image resampling (RPR) to obtain a low-resolution second video frame. The downsampling neural network model is used to perform feature fusion processing on the high-resolution first video frame and the low-resolution second video frame to obtain a low-resolution third video frame.
[0084] In at least one optional embodiment of this application, step 402 includes:
[0085] The first and second video frames are fused using a downsampling neural network model to obtain the third video frame.
[0086] Optionally, the downsampling neural network model is a deep learning-based neural network (NN) tool. This embodiment uses a deep neural downsampling network to process video frames. Specifically, the intermediate variables after RPR processing are used as input to the downsampling neural network model, and the original high-resolution video frames from the encoding process are used to guide the downsampling process. This allows the low-resolution video frames output by the downsampling neural network model to retain richer high-frequency information, thereby enabling the downsampling neural network model to better adapt to the encoding environment, achieving the goal of reducing bitrate and improving encoding efficiency; and it also makes it easier to recover high-quality images in subsequent encoding and decoding processes, improving encoding and decoding gain.
[0087] The embodiments of this application, by replacing the downsampling filter to generate high-quality low-resolution video frames, can reduce residuals and lower the bitrate during the encoding process. Simultaneously, these high-quality low-resolution video frames are more easily restored to high-quality high-resolution frames at the decoding end, thereby improving the encoding quality. This video encoding method can be used for video and image encoding, decoding, streaming, and storage implementations, and is not specifically limited herein.
[0088] In at least one embodiment of this application, step 402 includes:
[0089] Determine the target feature image based on the first video frame, or based on the first video frame and the second video frame;
[0090] The third video frame is determined based on the luminance Y component and chrominance UV component of the target feature image, and the Y component and UV component of the second video frame.
[0091] Optionally, step 402 can be understood as: the processing of the first video frame and the second video frame by the downsampling neural network model.
[0092] In one implementation, based on the first video frame, the method includes:
[0093] The first video frame is convolved to obtain the target feature image.
[0094] In another implementation, determining the target feature image based on the first video frame and the second video frame includes:
[0095] The first video frame and the second video frame are convolved respectively to obtain a first feature image and a second feature image; the first feature image and the second feature image are coupled to obtain the target feature image.
[0096] In this method, the coupling of the first feature image and the second feature image can further improve the peak signal-to-noise ratio (PSNR), thereby making the quality of the target feature higher and closer to the video frames in the original video.
[0097] Optionally, the size of the first feature image obtained by convolution is the same as the size of the second feature image. For example, the first video frame is convolved using a first convolutional layer to obtain the first feature image; the second video frame is convolved using a second convolutional layer to obtain the second feature image; wherein the convolution kernel of the first convolutional layer is the same as the convolution kernel of the second convolutional layer, and the stride of the first convolutional layer is greater than the stride of the second convolutional layer, so that the size of the first feature image is the same as the size of the second feature image.
[0098] In one example, the stride of the first convolutional layer is 2, and the stride of the second convolutional layer is 1; the convolutional kernels of the first and second convolutional layers are both 3×3.
[0099] In an optional embodiment of this application, determining the third video frame based on the luminance Y component and chrominance UV components of the target feature image, and the Y component and UV components of the second video frame, includes:
[0100] The Y and UV components of the target feature image are downsampled using a convolutional layer to obtain the fourth Y component and the fourth UV component.
[0101] The fourth Y component and the fourth UV component are subjected to residual processing to obtain the fifth Y component and the fifth UV component;
[0102] The sixth Y component is determined based on the sum of the Y component of the second video frame and the fifth Y component.
[0103] The sixth UV component is determined based on the sum of the UV components of the second video frame and the fifth UV component.
[0104] The third video frame is determined based on the sixth Y component and the sixth UV component.
[0105] In another optional embodiment of this application, determining the third video frame based on the luminance Y component and chrominance UV components of the target feature image, and the Y component and UV components of the second video frame, includes:
[0106] The Y component and UV component of the target feature image are subjected to residual processing to obtain the first Y component and the first UV component.
[0107] The third video frame is determined based on the first Y component and the first UV component, as well as the Y component and UV component of the second video frame.
[0108] The step of performing residual processing on the luminance Y component and chrominance UV component of the target feature image to obtain the first Y component and the first UV component includes:
[0109] The Y component of the target feature image is input into m residual blocks for residual processing to obtain the first Y component;
[0110] The UV components of the target feature image are input into n residual blocks for residual processing to obtain the first UV component;
[0111] Where m and n are integers greater than or equal to 1.
[0112] For example, such as Figure 7 As shown, the residual block consists of two 1×1 convolutional layers, a 3×3 separable convolutional layer, and an activation function layer. The result of the convolution is added to the input of the residual block to form a residual connection (i.e., the output of the residual block).
[0113] In one implementation, the value of m is different from the value of n.
[0114] In this implementation, using different numbers of residual blocks to process the Y and UV components respectively can more effectively extract the features of each component. This allows for optimization based on the characteristics of each component in complex image content, thereby improving the overall network performance.
[0115] Optionally, determining the third video frame based on the first Y component and the first UV component, and the Y component and UV component of the second video frame, includes:
[0116] The first Y component and the first UV component are downsampled using a convolutional layer to obtain the second Y component and the second UV component.
[0117] The third Y component is determined based on the sum of the Y component of the second video frame and the second Y component.
[0118] The third UV component is determined based on the UV component of the second video frame and the sum of the second UV components;
[0119] The third video frame is determined based on the third Y component and the third UV component.
[0120] It should be noted that in the step of determining the third Y component based on the sum of the Y component of the second video frame and the second Y component, the sum can be directly used as the third Y component, or the sum can be further convolved and the convolution result can be used as the third Y component, thereby improving PSNR.
[0121] Alternatively, in the step of determining the third UV component based on the UV component of the second video frame and the sum of the second UV components, the sum can be directly used as the third UV component, or the sum can be further convolved and the convolution result can be used as the third UV component, thereby improving PSNR.
[0122] For example, the convolutional layer that downsamples the first Y component consists of a 3×3 convolutional layer, a 3×3 separable convolutional layer, and an activation function layer with a stride of 1; the convolutional layer that downsamples the first UV component consists of a 3×3 convolutional layer, a 3×3 separable convolutional layer, and an activation function layer with a stride of 2.
[0123] Alternatively, in one example, a schematic diagram of a deep learning-based downsampling neural network model is shown below. Figure 8 As shown, Figure 8 In this model, the downsampling neural network takes high-resolution raw frames (UV and Y components) and low-resolution RPR frames (UV and Y components) as input. First, the input signal is passed through a 3×3 convolutional layer to extract initial features. To better fuse the features, the stride of the high-resolution raw frame convolutional layer is 2. The width and height of the convolutional feature map are the same as the low-resolution RPR feature map. The two feature maps are then coupled together (e.g., through...). Figure 8 The image is coupled with a connection layer, a 1×1 convolutional layer, and an activation function layer, and then divided into Y and UV components for separate processing. The Y and UV components are processed through different numbers of residual blocks. Using different numbers of residual blocks to process the Y and UV components separately can more effectively extract the features of each component, optimizing for the characteristics of each channel in complex image content, thereby improving the overall network performance. Next, the Y and UV components are residually connected to the corresponding components of the RPR low-resolution image, and the UV component is downsampled through a convolutional layer with a stride of 2. Finally, low-resolution Y and UV component results are output; these low-resolution Y and UV component results form the third video frame.
[0124] In at least one embodiment of this application, the method further includes:
[0125] A training dataset is obtained, which includes multiple fourth video frames of a video file and a fifth video frame obtained by sampling the fourth video frames multiple times; the resolution of the fifth video frame is higher than that of the fourth video frame.
[0126] The downsampling neural network is trained based on the fourth and fifth video frames to determine the downsampling neural network model.
[0127] Optionally, the fourth video frame can be understood as the original video frame of the video file, while the fifth video frame can be understood as an ultra-high resolution video frame obtained by sampling the original video frame multiple times. In this embodiment of the application, the high resolution video frame and the ultra-high resolution video frame are used as inputs for model training, which enables the video frame output by the model to be closer to the high resolution ground real frame.
[0128] In other words, the goal of training a downsampled neural network model is to optimize parameters such as weights and biases. First, a large number of high-resolution video sequence frames are used as the dataset to improve the model's generalization ability. Then, these video frames, along with ultra-high-resolution video frames (obtained through bicubic upsampling), are input into the neural network model. After downsampling by the network, the loss between the network output and the high-resolution ground truth frames (or the original frames) is calculated to evaluate the model's performance. The loss functions used include absolute difference (SAD) and mean squared error (MSE). Next, the backpropagation algorithm is used to calculate the gradients of each parameter, and these gradients are used to update the parameter values. This process is repeated until a predetermined convergence criterion is met. After training, the optimized parameters are saved for use in the subsequent inference phase.
[0129] In summary, the downsampling unit in this embodiment outputs high-quality low-resolution video frames, effectively enhancing the quality of low-resolution video frames and preserving more high-frequency information. These high-quality low-resolution video frames reduce residuals in the encoding process and facilitate the restoration of high-resolution frames, thereby optimizing the overall efficiency and quality of video encoding. Furthermore, experimental results show that when applying the video encoding method provided in this embodiment, the YUV components of the video frame have gains of (-0.50%, -0.52%, -0.48%) compared to existing technical solutions under the average conditions of all test sequences.
[0130] The video encoding method provided in this application can be executed by a video encoding device. As an example, the device can be an electronic device or a component within an electronic device, such as a chip or circuit. This application uses a video encoding device executing the video encoding method as an example to illustrate the video encoding device provided in this application.
[0131] like Figure 9As shown in the figure, this application embodiment also provides a video encoding apparatus, the apparatus comprising:
[0132] The downsampling module 901 is used to downsample the first video frame based on reference image resampling (RPR) to obtain the second video frame;
[0133] The fusion module 902 is used to perform feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame.
[0134] Encoding module 903 is used to determine the target bitstream based on the third video frame.
[0135] As an optional embodiment, the fusion module includes:
[0136] The fusion module includes:
[0137] The determination submodule is used to determine the target feature image based on the first video frame, or based on the first video frame and the second video frame;
[0138] The first fusion submodule is used to determine the third video frame based on the luminance Y component and chrominance UV component of the target feature image, and the Y component and UV component of the second video frame.
[0139] As an optional embodiment, the determining submodule includes:
[0140] The first convolutional unit is used to convolve the first video frame to obtain the target feature image;
[0141] or,
[0142] The second convolutional unit is used to convolve the first video frame and the second video frame respectively to obtain a first feature image and a second feature image; and to couple the first feature image and the second feature image to obtain the target feature image.
[0143] As an optional embodiment, the size of the first feature image obtained by convolution is the same as the size of the second feature image.
[0144] As an optional embodiment, the first fusion submodule includes:
[0145] The first residual unit is used to perform residual processing on the Y component and UV component of the target feature image respectively to obtain the first Y component and the first UV component.
[0146] The first determining unit is configured to determine the third video frame based on the first Y component and the first UV component, as well as the Y component and UV component of the second video frame.
[0147] As an optional embodiment, the first fusion submodule includes:
[0148] The first downsampling unit is used to downsample the Y component and UV component of the target feature image using a convolutional layer to obtain the fourth Y component and the fourth UV component.
[0149] The second residual unit is used to perform residual processing on the fourth Y component and the fourth UV component respectively to obtain the fifth Y component and the fifth UV component.
[0150] The second determining unit is used to determine the sixth Y component based on the sum of the Y component of the second video frame and the fifth Y component.
[0151] The third determining unit is used to determine the sixth UV component based on the sum of the UV components of the second video frame and the fifth UV component.
[0152] The fourth determining unit is used to determine the third video frame based on the sixth Y component and the sixth UV component.
[0153] As an optional embodiment, the first residual unit includes:
[0154] The first residual subunit is used to input the Y component of the target feature image into m residual blocks for residual processing to obtain the first Y component;
[0155] The second residual subunit is used to input the UV components of the target feature image into n residual blocks for residual processing to obtain the first UV component;
[0156] Where m and n are integers greater than or equal to 1.
[0157] As an optional embodiment, the value of m is different from the value of n.
[0158] As an optional embodiment, the first determining unit includes:
[0159] The downsampling subunit is used to downsample the first Y component and the first UV component using a convolutional layer to obtain the second Y component and the second UV component.
[0160] The first determining subunit is used to determine the third Y component based on the sum of the Y component of the second video frame and the second Y component.
[0161] The second determining subunit is used to determine the third UV component based on the sum of the UV components of the second video frame and the second UV components.
[0162] The third determining subunit is used to determine the third video frame based on the third Y component and the third UV component.
[0163] As an optional embodiment, the fusion module includes:
[0164] The second fusion submodule is used to perform feature fusion processing on the first video frame and the second video frame using a downsampling neural network model to obtain the third video frame.
[0165] As an optional embodiment, the apparatus further includes:
[0166] An acquisition module is used to acquire a training dataset, which includes multiple fourth video frames of a video file, and a fifth video frame obtained by sampling the fourth video frames multiple times; the resolution of the fifth video frame is higher than the resolution of the fourth video frame.
[0167] The training module is used to train the downsampling neural network based on the fourth video frame and the fifth video frame to determine the downsampling neural network model.
[0168] In this embodiment, the encoding end performs RPR downsampling and feature fusion processing on the high-resolution first video frame to obtain a low-resolution third video frame. This allows the third video frame to retain richer high-frequency information. The third video frame is used in the subsequent encoding process, making it easier to restore the low-resolution third video frame into a high-quality image in the subsequent encoding and decoding process. This can improve the encoding gain and achieve better encoding performance.
[0169] It should be noted that the video encoding device provided in this application embodiment is a device capable of executing the above video encoding method. Therefore, all embodiments of the above video encoding method are applicable to this device and can achieve the same or similar beneficial effects, which will not be repeated here.
[0170] The video encoding device provided in this application embodiment can achieve... Figures 4 to 8 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.
[0171] like Figure 10As shown, this application embodiment also provides an electronic device 1000, including a processor 1001 and a memory 1002. The memory 1002 stores a program or instructions that can run on the processor 1001. For example, when the electronic device 1000 is an encoding device, when the program or instructions are executed by the processor 1001, they implement the various steps of the above-described video encoding method embodiment and achieve the same technical effect. To avoid repetition, further details are omitted here. Optionally, the memory 1002 may be... Figure 1 The processor 1001 can implement the memory 102 or memory 113 in the illustrated embodiment. Figure 1-3 The functions of the encoder 200 or decoder 300 in the illustrated embodiment.
[0172] This application also provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the steps of the video encoding method embodiments described above. Optionally, the memory may be... Figure 1 The processing circuitry of memory 102 or memory 113 in the illustrated embodiment can implement... Figure 1-3 The functions of the encoder 200 or decoder 300 in the illustrated embodiment.
[0173] This application embodiment also provides an electronic device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement, for example... Figure 4 The steps in the method embodiment shown are illustrated. This device embodiment corresponds to the above method embodiment, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect.
[0174] The processor or processing circuit in this application embodiment may include general-purpose processors, special-purpose processors, etc., such as central processing units (CPUs), microprocessors, digital signal processors (DSPs), artificial intelligence (AI) processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), network processors (NPs), field-programmable gate arrays (FPGAs), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The communication interface in this application embodiment may include transceivers, pins, circuits, buses, etc.
[0175] The aforementioned electronic devices can be terminals or other devices besides terminals, such as servers, network attached storage (NAS), etc.
[0176] The terminal can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, mixed reality (MR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0177] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server. A cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0178] For example, the aforementioned electronic devices may include, but are not limited to, those described above. Figure 1 The type of source device 100 or destination device 110 shown.
[0179] Taking electronic devices as terminals as an example, Figure 11 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.
[0180] The terminal 1100 includes, but is not limited to, at least some of the following components: radio frequency unit 1101, network module 1102, audio output unit 1103, input unit 1104, sensor 1105, display unit 1106, user input unit 1107, interface unit 1108, memory 1109, and processor 1110.
[0181] Those skilled in the art will understand that the terminal 1100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 11 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0182] It should be understood that, in this embodiment, the input unit 1104 may include a graphics processor 11041 and a microphone 11042. The graphics processor 11041 processes image data of still images or videos obtained by an image acquisition device (such as a camera) in video acquisition mode or image acquisition mode, or it may process the obtained point cloud data. The display unit 1106 may include a display panel 11061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0183] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1101 can transmit it to the processor 1110 for processing; in addition, the radio frequency unit 1101 can send uplink data to the network-side device. Typically, the radio frequency unit 1101 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0184] The memory 1109 can be used to store software programs or instructions, as well as various data. The memory 1109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1109 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1109 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0185] Processor 1110 may include one or more processing units; optionally, processor 1110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1110.
[0186] The processor 1110 is configured to downsample the first video frame using Reference Image Resampling (RPR) to obtain a second video frame; perform feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame; and determine the target bitstream based on the third video frame.
[0187] In this embodiment, the encoding end performs RPR downsampling and feature fusion processing on the high-resolution first video frame to obtain a low-resolution third video frame. This allows the third video frame to retain richer high-frequency information. The third video frame is used in the subsequent encoding process, making it easier to restore the low-resolution third video frame into a high-quality image in the subsequent encoding and decoding process. This can improve the encoding gain and achieve better encoding performance.
[0188] It should be noted that the terminal provided in this application embodiment is a terminal capable of executing the above video encoding method. Therefore, all embodiments of the above video encoding method are applicable to this terminal and can achieve the same or similar beneficial effects, which will not be repeated here.
[0189] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the method embodiment and achieve the same or corresponding technical effect. To avoid repetition, it will not be described again here.
[0190] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video encoding method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0191] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0192] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described video encoding method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0193] It should be understood that the chips mentioned in the embodiments of this application may include system-on-a-chip (also known as system chip, chip system, or system-on-a-chip) or discrete display chips, etc.
[0194] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described video encoding method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0195] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0196] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0197] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. A video encoding method, characterized in that, The method includes: The encoding end performs downsampling on the first video frame based on reference image resampling (RPR) to obtain the second video frame; The encoding end performs feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame; The encoding end determines the target bitstream based on the third video frame.
2. The method according to claim 1, characterized in that, The encoding end performs feature fusion processing on the first video frame and the second video frame to obtain a third video frame, including: Determine the target feature image based on the first video frame, or based on the first video frame and the second video frame; The third video frame is determined based on the luminance Y component and chrominance UV component of the target feature image, and the Y component and UV component of the second video frame.
3. The method according to claim 2, characterized in that, Determining a target feature image based on the first video frame, or based on the first video frame and the second video frame, includes: The first video frame is convolved to obtain the target feature image; or, The first video frame and the second video frame are convolved respectively to obtain a first feature image and a second feature image; the first feature image and the second feature image are coupled to obtain the target feature image.
4. The method according to claim 3, characterized in that, The size of the first feature image obtained by convolution is the same as the size of the second feature image.
5. The method according to claim 2, characterized in that, The third video frame is determined based on the luminance Y component and chrominance UV components of the target feature image, and the Y component and UV components of the second video frame, including: The Y component and UV component of the target feature image are subjected to residual processing to obtain the first Y component and the first UV component. The third video frame is determined based on the first Y component and the first UV component, as well as the Y component and UV component of the second video frame.
6. The method according to claim 2, characterized in that, The third video frame is determined based on the luminance Y component and chrominance UV components of the target feature image, and the Y component and UV components of the second video frame, including: The Y and UV components of the target feature image are downsampled using a convolutional layer to obtain the fourth Y component and the fourth UV component. The fourth Y component and the fourth UV component are subjected to residual processing to obtain the fifth Y component and the fifth UV component; The sixth Y component is determined based on the sum of the Y component of the second video frame and the fifth Y component. The sixth UV component is determined based on the sum of the UV components of the second video frame and the fifth UV component. The third video frame is determined based on the sixth Y component and the sixth UV component.
7. The method according to claim 5, characterized in that, The step of performing residual processing on the luminance Y component and chrominance UV component of the target feature image to obtain the first Y component and the first UV component includes: The Y component of the target feature image is input into m residual blocks for residual processing to obtain the first Y component; The UV components of the target feature image are input into n residual blocks for residual processing to obtain the first UV component; Where m and n are integers greater than or equal to 1.
8. The method according to claim 7, characterized in that, The value of m is different from the value of n.
9. The method according to claim 5, characterized in that, Determining the third video frame based on the first Y component and the first UV component, and the Y component and UV component of the second video frame, includes: The first Y component and the first UV component are downsampled using a convolutional layer to obtain the second Y component and the second UV component. The third Y component is determined based on the sum of the Y component of the second video frame and the second Y component. The third UV component is determined based on the UV component of the second video frame and the sum of the second UV components; The third video frame is determined based on the third Y component and the third UV component.
10. The method according to any one of claims 1-9, characterized in that, The encoding end performs feature fusion processing on the first video frame and the second video frame to obtain a third video frame, including: The first and second video frames are fused using a downsampling neural network model to obtain the third video frame.
11. The method according to claim 10, characterized in that, The method further includes: A training dataset is obtained, which includes multiple fourth video frames of a video file and a fifth video frame obtained by sampling the fourth video frames multiple times; the resolution of the fifth video frame is higher than that of the fourth video frame. The downsampling neural network is trained based on the fourth and fifth video frames to determine the downsampling neural network model.
12. A video encoding device, characterized in that, The device includes: The downsampling module is used to downsample the first video frame based on reference image resampling (RPR) to obtain the second video frame. A fusion module is used to perform feature fusion processing on the first video frame and the second video frame to obtain a third video frame; wherein the resolution of the first video frame is higher than the resolution of the third video frame; The encoding module is used to determine the target bitstream based on the third video frame.
13. The apparatus according to claim 12, characterized in that, The fusion module includes: The determination submodule is used to determine the target feature image based on the first video frame, or based on the first video frame and the second video frame; The first fusion submodule is used to determine the third video frame based on the luminance Y component and chrominance UV component of the target feature image, and the Y component and UV component of the second video frame.
14. The apparatus according to claim 13, characterized in that, The determining submodule includes: The first convolutional unit is used to convolve the first video frame to obtain the target feature image; or, The second convolutional unit is used to convolve the first video frame and the second video frame respectively to obtain a first feature image and a second feature image; and to couple the first feature image and the second feature image to obtain the target feature image.
15. The apparatus according to claim 14, characterized in that, The size of the first feature image obtained by convolution is the same as the size of the second feature image.
16. The apparatus according to claim 13, characterized in that, The first fusion submodule includes: The first residual unit is used to perform residual processing on the Y component and UV component of the target feature image respectively to obtain the first Y component and the first UV component. The first determining unit is configured to determine the third video frame based on the first Y component and the first UV component, as well as the Y component and UV component of the second video frame.
17. The apparatus according to claim 13, characterized in that, The first fusion submodule includes: The first downsampling unit is used to downsample the Y component and UV component of the target feature image using a convolutional layer to obtain the fourth Y component and the fourth UV component. The second residual unit is used to perform residual processing on the fourth Y component and the fourth UV component respectively to obtain the fifth Y component and the fifth UV component. The second determining unit is used to determine the sixth Y component based on the sum of the Y component of the second video frame and the fifth Y component. The third determining unit is used to determine the sixth UV component based on the sum of the UV components of the second video frame and the fifth UV component. The fourth determining unit is used to determine the third video frame based on the sixth Y component and the sixth UV component.
18. The apparatus according to claim 16, characterized in that, The first residual unit includes: The first residual subunit is used to input the Y component of the target feature image into m residual blocks for residual processing to obtain the first Y component; The second residual subunit is used to input the UV components of the target feature image into n residual blocks for residual processing to obtain the first UV component; Where m and n are integers greater than or equal to 1.
19. The apparatus according to claim 18, characterized in that, The value of m is different from the value of n.
20. The apparatus according to claim 16, characterized in that, The first determining unit includes: The downsampling subunit is used to downsample the first Y component and the first UV component using a convolutional layer to obtain the second Y component and the second UV component. The first determining subunit is used to determine the third Y component based on the sum of the Y component of the second video frame and the second Y component. The second determining subunit is used to determine the third UV component based on the sum of the UV components of the second video frame and the second UV components. The third determining subunit is used to determine the third video frame based on the third Y component and the third UV component.
21. The apparatus according to any one of claims 12-20, characterized in that, The fusion module includes: The second fusion submodule is used to perform feature fusion processing on the first video frame and the second video frame using a downsampling neural network model to obtain the third video frame.
22. The apparatus according to claim 21, characterized in that, The device further includes: An acquisition module is used to acquire a training dataset, which includes multiple fourth video frames of a video file, and a fifth video frame obtained by sampling the fourth video frames multiple times; the resolution of the fifth video frame is higher than the resolution of the fourth video frame. The training module is used to train the downsampling neural network based on the fourth video frame and the fifth video frame to determine the downsampling neural network model.
23. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the video encoding method as described in any one of claims 1 to 11.
24. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the video encoding method as described in any one of claims 1 to 11.
25. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method as described in any one of claims 1 to 11.