Image block compression processing method and apparatus, and electronic device
By inputting the brightness data and chrominance data of the image blocks into different convolutional layers of the neural network model, the problem of large calculation complexity and large parameters in the unified neural network model is solved, and efficient video compression processing is achieved.
Patent Information
- Application Number
- PCT/CN2024/138698
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-26
AI Technical Summary
In the unified neural network model, the brightness data and chrominance data need to be bound together to input, resulting in increased computational complexity and larger parameters.
By inputting the brightness data and chrominance data of the image blocks to different convolutional layers of the neural network model for video compression processing, the input data amount and calculation complexity of the convolutional layer are reduced.
It effectively reduces the computational complexity and parameter quantity of neural networks and improves the efficiency of video compression processing.
Smart Images

Figure CN2024138698_26062025_PF_FP_ABST
Abstract
Description
Image block compression processing method, device and electronic equipment
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311752618.9 filed in China on December 19, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application belongs to the field of video processing technology, and specifically relates to an image block compression processing method, device and electronic equipment. Background Art
[0004] In the unified neural network model, luminance data and chrominance data need to be bundled together and input into the same convolutional layer of the neural network. Because the width and height of the input data of the same convolutional layer must be the same, when the width and height of the luminance data and chrominance data use the chrominance subsampling format and the width and height of the chrominance data are different from those of the luminance data, the chrominance data needs to be upsampled. This method will increase the computational complexity of the neural network. Bundling the luminance data and chrominance data together and inputting the same convolutional layer of the neural network will cause the luminance data and chrominance data to have the same number of parameters on this convolutional layer, which will result in a larger number of parameters in the neural network. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, and electronic device for image block compression processing to reduce the computational complexity and parameter quantity of a neural network.
[0006] In a first aspect, a method for image block compression processing is provided, the method comprising:
[0007] Acquire first luminance data and first chrominance data of the image block;
[0008] Inputting the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0009] Acquire second luminance data and second chrominance data output by the neural network model.
[0010] In a second aspect, an image block compression processing device is provided, comprising:
[0011] A first acquisition module, configured to acquire first luminance data and first chrominance data of an image block;
[0012] a processing module, configured to input the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0013] The second acquisition module is used to obtain the second brightness data and the second chromaticity data output by the neural network model.
[0014] In a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0015] In a fourth aspect, an electronic device is provided, comprising a processor and a communication interface, wherein the processor is configured to obtain first luminance data and first chrominance data of an image block;
[0016] Inputting the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0017] Acquire second luminance data and second chrominance data output by the neural network model.
[0018] In a fifth aspect, a coding and decoding system is provided, comprising: an image block compression processing device, wherein the image block compression processing device can be used to perform the steps of the method described in the first aspect.
[0019] In a sixth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0020] In a seventh aspect, a chip is provided, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the steps of the method described in the first aspect.
[0021] In an eighth aspect, a computer program / program product is provided, wherein the computer program / program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect.
[0022] In an embodiment of the present application, first luminance data and first chrominance data of an image block are obtained respectively, and then the first luminance data and the first chrominance data are respectively input into different convolutional layers of a neural network model for video compression processing, thereby obtaining second luminance data and second chrominance data output by the neural network model. When inputting data, the embodiment of the present application splits the luminance data and chrominance data and inputs them into different convolutional layers of the neural network model respectively. Each convolutional layer only processes part of the data it receives during processing, thereby reducing the amount of input data to the convolutional layer, and further reducing the data processing complexity of the convolutional layer, thereby achieving the purpose of reducing the computational complexity and parameter quantity of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] FIG1 is a schematic diagram of a coding and decoding system provided in an embodiment of the present application;
[0024] FIG2 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;
[0025] FIG3 is a schematic structural diagram of a decoder provided in an embodiment of the present application;
[0026] FIG4 is a schematic flow chart of an image block compression processing method according to an embodiment of the present application;
[0027] FIG5 is a schematic diagram of a video compression process in application scenario 1;
[0028] FIG6 is a schematic diagram of the residual processing process;
[0029] FIG7 is a schematic diagram of the video compression process of application scenario 2;
[0030] FIG8 is a schematic diagram of the video compression process of application case 3;
[0031] FIG9 is a schematic diagram of the video compression process of application scenario 4;
[0032] FIG10 is a schematic diagram of the video compression process of application case five;
[0033] FIG11 is a schematic diagram of the video compression process for application case 6;
[0034] FIG12 is a module diagram of an image block compression processing device according to an embodiment of the present application;
[0035] FIG13 is a schematic structural diagram of an electronic device according to an embodiment of the present application;
[0036] FIG14 is a schematic structural diagram of a communication device according to an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0038] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0039] FIG1 is a schematic diagram of a codec system provided in an embodiment of the present application. The technical solution of the embodiment of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) of video data. The video data includes original unencoded video, encoded video, decoded (e.g., reconstructed) video, or syntax elements.
[0040] As shown in FIG1 , the codec system includes a source device 100, which provides encoded video data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides the video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (e.g., a smart watch or a wearable camera), a television, a camera, a display device, an in-vehicle device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an aircraft, a robot, a satellite, and the like.
[0041] In the example of Figure 1, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, while the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 may not include some of the components in Figure 1, or may also include other components outside of Figure 1. For example, the source device 100 can receive video data from an external data source (such as an external camera). Similarly, the destination device 110 can be connected to an external display device interface without including an integrated display device. For another example, the memory 102 and the memory 113 can be external memories.
[0042] Although FIG1 illustrates source device 100 and destination device 110 as separate devices, in some examples, the two may be integrated into a single device. In such embodiments, the functions corresponding to source device 100 and the functions corresponding to destination device 110 may be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.
[0043] In some examples, source device 100 and destination device 110 can perform one-way video transmission or two-way video transmission. If it is two-way video transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, that is, each of source device 100 and destination device 110 includes an encoder and a decoder.
[0044] Data source 101 represents a source of video data (i.e., raw, unencoded video data) and provides a series of pictures containing the video data to encoder 200, which encodes the picture data. Data source 101 of source device 100 may include a video capture device (such as a video camera), a video archive containing previously captured raw video, or a video feed interface for receiving video from a video content provider. Alternatively, data source 101 may generate computer graphics-based data as the source video, or combine real-time video, archived video, and computer-generated video. In these cases, encoder 200 encodes the captured, pre-captured, or computer-generated video data. Encoder 200 may rearrange the pictures from the order in which they were received (sometimes referred to as "display order") into an encoding order. Encoder 200 may generate a bitstream comprising the encoded video data. Source device 100 may then output the encoded video data to communication medium 120 via output interface 104 for receipt or retrieval, for example, by input interface 111 of destination device 110.
[0045] Memory 102 of source device 100 and memory 113 of destination device 110 represent general purpose memories. In some examples, memory 102 may store raw video data from data source 101, and memory 113 may store decoded video data from decoder 300. Additionally or alternatively, memories 102 and 113 may store software instructions executable by, for example, encoder 200 and decoder 300, respectively. Although memory 102 and memory 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memory 102 and memory 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded video data output from encoder 200 and input to decoder 300. In some examples, portions of memory 102 , 113 may be allocated as one or more video buffers, eg, for storing raw, decoded, or encoded video data.
[0046] In some examples, source device 100 can output the encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access the encoded data from memory 113 via input interface 111. Memory 113 or storage 102 can include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a Digital Versatile Disc (DVD), a Compact Disc Read-Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0047] Output interface 104 may include any type of medium or device capable of transmitting encoded video data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded video data directly from source device 100 to destination device 110 in real time. The encoded video data may be modulated according to a communication standard of a wireless communication protocol and transmitted to destination device 110.
[0048] The communication medium 120 may include a transient medium such as a wireless broadcast or a wired network transmission. For example, the communication medium 120 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium) such as a hard disk, a flash drive, a compact disk, a digital video disk, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0049] In some embodiments, the communication medium 120 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 100 to the destination device 110. For example, a server (not shown) can receive encoded video from the source device 100 and provide the encoded video data to the destination device 110, for example, via a network transmission. The server may include a web server (e.g., for a website), a server configured to provide a file transfer protocol service (such as the File Transfer Protocol (FTP) or the File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or an evolved Multimedia Broadcast Multicast Service (eMBMS) server, a Network-attached storage (NAS) device, and the like. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT) protocol, Dynamic Adaptive Streaming over HTTP (DASH) protocol, HTTP Live Streaming (HLS) protocol or Real Time Streaming Protocol (RTSP).
[0050] Destination device 110 can access the encoded video data from the server, for example, via a wireless channel (e.g., a Wi-Fi connection) or a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.) for accessing the encoded video data stored on the server.
[0051] The output interface 104 and the input interface 111 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to the IEEE 802.11 standard or the IEEE 802.15 standard (e.g., ZigBee™), the Bluetooth standard, or other physical components. In examples where the output interface 104 and the input interface 111 include wireless components, the output interface 104 and the input interface 111 may be configured to communicate data, such as encoded video data, according to WIFI, Ethernet, a cellular network (such as 4G, Long Term Evolution (LTE), LTE-Advanced, Fifth Generation Mobile Communication Technology (5G), Sixth Generation Mobile Communication Technology (6G), etc.).
[0052] The technology provided in the embodiments of the present application can be applied to support video encoding and decoding in one or more multimedia applications such as: video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0053] The input interface 111 of the destination device 110 receives an encoded video bitstream from the communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), wherein the syntax elements are used to decode the encoded data units to obtain decoded video data. The display device 114 displays the decoded video data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0054] The encoder 200 and the decoder 300 may be implemented as one or more of a variety of processing circuits, which may include a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the technology is implemented in whole or in part in software, the device may store instructions for the software in an appropriate non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of the present application.
[0055] The encoder 200 and the decoder 300 can be processed based on the following video coding and decoding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding (HEVC)), H.266 (also known as Versatile Video Coding (VVC), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, Video Compression Format Version 8 (VP8), VP9, Alliance for Open Media Video 1 (AV1), Audio Video coding Standard 1 (AVS1), AVS2, AVS3 or next-generation video standard protocols, which are not specifically limited in the embodiments of the present application.
[0056] Generally, the encoder 200 and decoder 300 can perform block-based encoding and decoding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in the encoding or decoding process). For example, a block can include a two-dimensional matrix of samples of luma or chroma data. For example, the encoder 200 and decoder 300 can encode and decode video data represented in YUV format.
[0057] 2 , which is a schematic diagram of the structure of an encoder 200 provided in an embodiment of the present application, and the encoder 200 may be the encoder 200 in FIG1 . In the example of FIG2 , the encoder 200 includes a memory 201, a coding parameter determination unit 210, a residual generation unit 202, a transform processing unit 203, a quantization unit 204, an inverse quantization unit 205, an inverse transform processing unit 206, a reconstruction unit 207, a filter unit 208, a decoded picture buffer (DPB) 209, and an entropy coding unit 220.
[0058] The memory 201 can store video data to be encoded. For example, the encoder 200 can receive and store video data from the data source 101 shown in FIG1 . In some examples, the memory 201 can be on the same chip as other components of the encoder 200 (as shown in FIG2 ), or can be independent of the chip where these components are located.
[0059] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is used to obtain a first prediction block of the current block using the inter-frame prediction mode, and the intra-frame prediction unit 213 is used to obtain a second prediction block of the current block using the intra-frame prediction mode. The mode selection unit 211 is used to obtain a target prediction block based on the first prediction block and the second prediction block, and to determine a final prediction mode. In addition, the coding parameter determination unit 210 may also include other functional units, such as a functional unit for determining the division method of the coding unit (CU), a functional unit for determining the transform type of the residual data of the CU or the quantization parameter of the residual data of the CU, etc.
[0060] For ease of description and understanding, the embodiment of the present application refers to the CU to be processed in the current image as the current CU, and the image block to be processed in the current CU as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0061] The inter prediction unit 212 may include a motion estimation unit and a motion compensation unit. For inter prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in the DPB 209).
[0062] The motion estimation unit may form one or more motion vectors (MVs) that represent the position of a reference block in a reference picture relative to the position of a current block in a current picture, and the motion compensation unit may obtain a prediction value of the accuracy indicated by the motion vectors through interpolation.
[0063] The encoding parameter determination unit 210 may provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the original unencoded video data of the current block from the memory 201 and calculates the residual between the current block and the target prediction block to obtain a residual block. In some examples, the functionality of the residual generation unit 202 may be implemented using one or more subtractor circuits that perform binary subtraction.
[0064] As an example, the coding parameter determination unit 210 may provide a syntax element representing the coding parameter to the entropy coding unit 220 for encoding. The coding parameter includes one or more of the CU partitioning method, the final prediction mode, the transform type of the CU residual data, or the quantization parameter of the CU residual data.
[0065] The transform processing unit 203 transforms the residual block output by the residual generating unit 202 to obtain a transform coefficient block. The transform may include discrete cosine transform (DCT), integer transform, directional transform, or Karhunen-Loeve transform. In some examples, the encoder 200 may not include the transform processing unit 203.
[0066] The quantization unit 204 may quantize the transform coefficients in the transform coefficient block according to a quantization parameter (QP) value associated with the current block to generate a quantized transform coefficient block.
[0067] The inverse quantization unit 205 and the inverse transform processing unit 206 can respectively perform inverse quantization and inverse transform on the transform coefficient block to obtain a reconstructed residual block. The reconstruction unit 207 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0068] The filter unit 208 may perform one or more filter operations on the reconstructed block. For example, the filter unit 208 may be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, the encoder 200 may not include the filter unit 208.
[0069] The encoder 200 stores the reconstructed picture obtained from the reconstructed block in the DPB 209. For example, in examples where the operation of the filter unit 208 is not required, the reconstruction unit 207 may store the reconstructed block in the DPB 209. In examples where the operation of the filter unit 208 is required, the filter unit 208 may store the filtered reconstructed block in the DPB 209. The inter-frame prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter-frame prediction on blocks of subsequent pictures to be encoded. In some examples, the DPB 209 may be replaced with other types of memory.
[0070] The entropy coding unit 220 may entropy encode syntax elements of other components in the encoder 200 and output encoded video data. For example, the entropy coding unit 220 may entropy encode the quantized transform coefficient blocks from the quantization unit 204. As another example, the entropy coding unit 220 may entropy encode syntax elements from the coding parameter determination unit 210 (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction).
[0071] It is understandable that the composition of the encoder 200 shown in FIG2 is merely an illustration and does not constitute a limitation to the embodiments of the present application.
[0072] FIG3 is a schematic diagram of the structure of a decoder 300 provided in an embodiment of the present application, and the decoder 300 may be the decoder 300 described in FIG1 . In the example of FIG3 , the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0073] The entropy decoding unit 302 can receive the encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements, where the syntax elements indicate encoding parameters, and the encoding parameters include one or more of the CU partitioning method, the final prediction mode, the transform type of the CU's residual data, or the quantization parameter of the CU's residual data.
[0074] When the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter-frame prediction mode, the prediction block of the current CU can be obtained by the inter-frame prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra-frame prediction mode, the prediction block of the current CU can be obtained by the intra-frame prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may also include a unit for performing prediction functions according to other prediction modes.
[0075] The CPB 301 can obtain and store encoded video data from the communication medium 120 shown in Figure 1. The DPB 307 is used to store decoded pictures. Optionally, the CPB 301 and DPB 307 can also be replaced with other types of memory, which is not specifically limited in this application. In some examples, the CPB 301 can be on the same chip as other components of the decoder 300 (as shown in the figure), or it can be independent of the chip where these components are located.
[0076] The decoder 300 can perform reconstruction operations on each block separately. The entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients and transform information (such as QP or transform mode indication) to obtain quantized transform coefficients. The quantized transform coefficients are dequantized by the dequantization unit 303 to obtain a transform coefficient block including the transform coefficients. The transform coefficient block is inversely transformed by the inverse transform processing unit 304 to generate a residual block corresponding to the current block. This inverse transform is the inverse operation of the above-mentioned transform.
[0077] The reconstruction unit 305 may reconstruct the current block according to the prediction block and the residual block. For example, the reconstruction unit 305 may add samples of the residual block to corresponding samples of the prediction block to reconstruct the current block.
[0078] The filter unit 306 may perform one or more filter operations on the reconstructed block. For example, the type of the filter unit 306 may refer to the type of the filter unit 208 and will not be described in detail here. In some examples, the operation of the filter unit 306 may be skipped.
[0079] The decoder 300 may store the reconstructed picture obtained from the reconstructed block in the DPB 307. For example, in an example in which the operation of the filter unit 306 is not performed, the reconstruction unit 305 may store the reconstructed block in the DPB 307. In an example in which the operation of the filter unit 306 is performed, the filter unit 306 may store the filtered reconstructed block in the DPB 307. The decoder 300 may output a decoded picture (e.g., decoded video) from the DPB 307 for subsequent presentation to a display device (such as the display device 114 of FIG. 1 ).
[0080] The following describes the image block compression processing method, device, and electronic device provided by the embodiments of the present application in conjunction with the accompanying drawings. The image block prediction method provided by the embodiments of the present application can be executed by an encoding end, such as the encoder 200 shown in Figure 1 or Figure 2. The image block prediction method provided by the embodiments of the present application can be executed by a decoding end, such as the decoder 300 described in Figure 1 or Figure 3. The encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a video decoding device.
[0081] The following first describes the technologies related to the embodiments of the present application.
[0082] 1. Neural network-based loop filter solution
[0083] To reduce the impact of distortion effects like blocking and ringing on video quality, video compression standards incorporate loop filters, such as deblocking and adaptive sample compensation. Deblocking filters reduce blocking artifacts, while adaptive sample compensation improves ringing artifacts. These loop filters effectively improve both subjective and objective video quality within the codec loop.
[0084] With the widespread application of neural network-based methods in image processing, many neural network-based loop filter solutions have also achieved good performance in video compression. Some attempts have been made to completely replace traditional filters with neural network filters, while others have been made to partially replace traditional filters with neural network filters.
[0085] However, neural network filters have the problems of high performance, high complexity, large number of parameters and large amount of calculation.
[0086] 2. Super-resolution solution based on neural network
[0087] Neural network-based super-resolution approaches utilize neural networks to reconstruct images for super-resolution, enhancing clarity and detail. This approach achieves super-resolution by training neural networks to learn image features. Common neural networks include convolutional neural networks and recurrent neural networks. Within image super-resolution approaches, neural networks can be used for single-image super-resolution reconstruction and for the joint super-resolution reconstruction of multiple images. This approach has been widely applied in fields such as image processing, medical image analysis, and security monitoring, and continues to evolve and improve.
[0088] 3. Chroma Subsampling
[0089] The human eye is less sensitive to chrominance signals than luminance signals. Therefore, the chrominance signal in an image can be downsampled so that adjacent pixels share the same chrominance value. Chroma subsampling formats include 4:4:4, 4:2:2, 4:1:1, and 4:2:0.
[0090] Among them, 4:4:4 means that the horizontal resolution and vertical resolution of brightness and chrominance are the same; 4:2:2 means that the horizontal resolution of chrominance is 1 / 2 of the brightness, and the vertical resolution is equal to the brightness; 4:1:1 means that the horizontal resolution of chrominance is 1 / 4 of the brightness, and the vertical resolution is equal to the brightness; 4:2:0 means that the horizontal resolution and vertical resolution of chrominance are 1 / 2 of the brightness.
[0091] In the unified network model, brightness and chromaticity are bundled together as the input of the neural network. The data size of brightness and chromaticity must be the same, which will cause the following problems:
[0092] 1. For the chroma subsampling sampling format, the chroma needs to be upsampled before it can be used as the input of the neural network; and chroma upsampling will increase the computational complexity of the model;
[0093] 2. Bright color texture processing has different levels of difficulty, and the human eye is relatively insensitive to chroma. Binding input will increase the number of model parameters.
[0094] The image block compression processing method, device, and electronic device provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings through some embodiments and their application scenarios.
[0095] As shown in FIG4 , an embodiment of the present application provides an image block compression processing method, including:
[0096] Step 401, obtaining first luminance data and first chrominance data of an image block;
[0097] Step 402: Input the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0098] Step 403: Acquire the second luminance data and the second chrominance data output by the neural network model.
[0099] It should be noted that, when performing video compression processing, it is usually necessary to divide each video frame into multiple image blocks based on the video image frame, and process each image block separately. In the embodiments of the present application, the image block refers to a portion of the image in the video image frame. By processing each image block, the video image frame is processed and the video compression processing is achieved.
[0100] It should be noted that the embodiment of the present application obtains the first luminance data and the first chrominance data of the image block respectively, and then inputs the first luminance data and the first chrominance data into different convolutional layers of the neural network model for video compression processing, thereby obtaining the second luminance data and the second chrominance data output by the neural network model. When inputting data, the embodiment of the present application splits the luminance data and the chrominance data and inputs them into different convolutional layers of the neural network model respectively. Each convolutional layer only processes the part of the data it receives during processing, thereby reducing the amount of input data of the convolutional layer, and further reducing the data processing complexity of the convolutional layer, thereby achieving the purpose of reducing the computational complexity and parameter quantity of the neural network.
[0101] Optionally, in one implementation, the component data of the first brightness data includes, but is not limited to, at least one of the following:
[0102] Luma reconstruction block, luma prediction block, luma boundary intensity, luma residual block, luma transform coefficient block.
[0103] Optionally, the luminance reconstruction block refers to a reconstructed sample block of the luminance signal obtained by the decoder according to the bit stream decoding and constituting the decoded image, which can be a filtered sample or a pre-filtered sample; the luminance prediction block refers to a predicted sample block of the luminance signal obtained by using the reconstructed samples through the prediction technology during the decoding process; the luminance boundary intensity refers to the boundary intensity value of the luminance signal obtained by the deblocking filtering technology based on the pattern information and sample changes on both sides of the boundary; the luminance residual block refers to the difference between the original sample and the predicted sample of the image block of the luminance signal; the luminance transform coefficient block refers to the transform coefficient in the frequency domain obtained by the transform technology of the luminance residual block.
[0104] Optionally, in one implementation, the component data of the first chrominance data includes at least one of the following:
[0105] Blue component data;
[0106] Red component data.
[0107] Optionally, in one implementation, the blue component data and / or the red component data include but are not limited to at least one of the following component type data:
[0108] Chroma reconstruction block, chroma prediction block, chroma boundary intensity, chroma residual block, chroma transform coefficient block.
[0109] Optionally, the chroma reconstruction block refers to a reconstructed sample block of the chroma signal obtained by the decoder according to the bit stream decoding and constituting the decoded image, which can be a filtered sample or a pre-filtered sample; the chroma prediction block refers to a predicted sample block of the chroma signal obtained by using the reconstructed samples through the prediction technology during the decoding process; the chroma boundary strength refers to the boundary strength value of the chroma signal obtained by the deblocking filtering technology based on the pattern information and sample changes on both sides of the boundary; the chroma residual block refers to the difference between the original sample and the predicted sample of the image block of the chroma signal, and the chroma transform coefficient block refers to the transform coefficient in the frequency domain obtained by the chroma residual block through the transform technology.
[0110] Optionally, the second luminance data includes a filtered luminance reconstruction block or a luminance reconstruction block of a first resolution, and the first resolution is higher than a resolution of the luminance reconstruction block corresponding to the first luminance data; or
[0111] The second chroma data includes a filtered chroma reconstruction block or a chroma reconstruction block with a second resolution, and the second resolution is higher than a resolution of the chroma reconstruction block corresponding to the first chroma data.
[0112] It should be noted that a neural network-based loop filter scheme can be used to compress the video, and a neural network-based super-resolution scheme can also be used to compress the video; optionally, when the video is compressed by filtering through a neural network model, the second luminance data obtained at this time includes the filtered luminance reconstruction block, and the second chrominance data obtained includes the filtered chrominance reconstruction block; optionally, when the video is compressed by super-resolution processing through a neural network model, the second luminance data obtained at this time includes the luminance reconstruction block of the first resolution, and the second chrominance data obtained includes the chrominance reconstruction block of the second resolution.
[0113] It should be noted that, in order to enable the convolution layer to process the first luminance data and the second chrominance data, it is necessary to first obtain the step size of the convolution layer before inputting the first luminance data and the second chrominance data into the convolution layer; optionally, in one implementation, before step 402, the method further includes:
[0114] Determining, according to a target chroma subsampling format of the image block, strides of convolutional layers corresponding to the first luminance data and the second chroma data, respectively;
[0115] The target chroma sub-sampling format is the chroma sub-sampling format of the image block output by the neural network model.
[0116] It should be noted here that the step size of the convolution layer corresponding to the first luminance data and the step size of the convolution layer corresponding to the second chrominance data can be the same or different; whether the two are the same can be determined by the size of the first luminance data and the first chrominance data.
[0117] Optionally, for example, if the target chroma subsampling format is 4:2:0, it indicates that the horizontal resolution and vertical resolution of chroma are 1 / 2 of the luminance, based on which the stride of the convolution layer corresponding to the first luminance data is determined to be 2, and the stride of the convolution layer corresponding to the first chroma data is determined to be 1. For example, if the target chroma subsampling format is 4:4:4, it indicates that the horizontal resolution and vertical resolution of luminance and chroma are the same, based on which the stride of the convolution layer corresponding to the first luminance data is determined to be 1, and the stride of the convolution layer corresponding to the first chroma data is determined to be 1; alternatively, the stride of the convolution layer corresponding to the first luminance data is determined to be 2, and the stride of the convolution layer corresponding to the first chroma data is determined to be 2.
[0118] It should be noted that, in order to ensure that the output size of the first luminance data and the first chrominance data after passing through the convolution layer meets the requirements, for different chrominance subsampling formats, the first chrominance data needs to be sampled before being input into the convolution layer; optionally, in one implementation, before inputting the first luminance data and the first chrominance data into different convolution layers of the neural network model for video compression processing, the method further includes:
[0119] If the target chroma subsampling format of the image block is different from the initial chroma subsampling format, adjusting the width or height of the image block corresponding to the first chroma data according to a relationship between the target chroma subsampling format and the initial chroma subsampling format;
[0120] The initial chroma sub-sampling format is the chroma sub-sampling format of the image block before inputting into the neural network model.
[0121] It should be noted that by adjusting the width or height of the image block corresponding to the first chroma data based on the target chroma subsampling format and the initial chroma subsampling format before inputting the neural network convolution layer, the first chroma data is adjusted to the width and height corresponding to the target chroma subsampling format, thereby ensuring that the chroma data that meets the requirements of the target chroma subsampling format can be output.
[0122] It should be noted that the first chroma data needs to be processed only when the target chroma subsampling format is different from the initial chroma subsampling format. If the target chroma subsampling format is the same as the initial chroma subsampling format, the first chroma data does not need to be processed.
[0123] Optionally, for example, when the target chroma sub-sampling format is 4:2:0, if the initial chroma sub-sampling format is 4:4:4, the width and height of the image block related to the first chroma data need to be reduced to half of the original; if the initial chroma sub-sampling format is 4:2:2, it means that the horizontal resolution of the chroma is 1 / 2 of the luminance, based on which it is determined that the height of the image block related to the first chroma data needs to be reduced to half of the original; if the initial chroma sub-sampling format is 4:1:1, it means that the horizontal resolution of the chroma is 1 / 4 of the luminance, and the vertical resolution is equal to the luminance, based on which it is determined that the height of the image block related to the first chroma data needs to be reduced to half of the original and the width needs to be enlarged to twice the original. For example, when the target chroma sub-sampling format is 4:4:4, if the initial chroma sub-sampling format is 4:2:0, the width and height of the image block related to the first chroma data need to be enlarged to twice the original size; if the initial chroma sub-sampling format is 4:2:2, the width of the image block related to the first chroma data needs to be enlarged to twice the original size; if the initial chroma sub-sampling format is 4:1:1, the width of the image block related to the first chroma data needs to be enlarged to four times the original size.
[0124] Optionally, in one implementation, inputting the first brightness data into a convolutional layer of a neural network model includes one of the following:
[0125] B11, inputting all component data of the first brightness data into the first convolutional layer;
[0126] It should be noted that since the first luminance data may include multiple different component data, in this case, the different component data in the first luminance data are all input into the same convolutional layer and processed uniformly. In other words, in this case, the first luminance data is not differentiated by component data.
[0127] For example, if the first luminance data includes a luminance reconstruction block, a luminance prediction block, and a luminance boundary intensity, the luminance reconstruction block, the luminance prediction block, and the luminance boundary intensity all need to be input into the same first-step convolutional layer.
[0128] It should be noted that by inputting all component data of the first brightness data into the same convolutional layer, the number of convolutional layers used for the first brightness data can be reduced, thereby improving the processing rate of feature extraction.
[0129] B12. Input different component data in the first brightness data into the first step-length convolution layer corresponding to the component data respectively;
[0130] It should be noted that, in this case, the number of convolution layers required is equal to the number of component data included in the first brightness data, that is, different component data are input into different convolution layers.
[0131] For example, the first luminance data includes a luminance reconstruction block, a luminance prediction block, and luminance boundary intensity. In this case, the luminance reconstruction block needs to be input into the first step length convolution layer 1, the luminance prediction block needs to be input into the first step length convolution layer 2, and the luminance boundary intensity needs to be input into the first step length convolution layer 3.
[0132] It should be noted that, by inputting different component data of the first brightness data into different convolutional layers, feature extraction can be performed based on the different component data respectively, thereby improving the accuracy of feature extraction.
[0133] Optionally, in one implementation, inputting the first chromaticity data into a convolutional layer of a neural network model includes one of the following:
[0134] B21. Input all component data of the first chrominance data into a convolutional layer of a second step length;
[0135] It should be noted that because the first luminance data may include multiple different component data, in this case, the different component data of the first chrominance data are all input into the same convolutional layer and processed uniformly. In other words, in this case, the first chrominance data is not differentiated by component data.
[0136] For example, if the first chrominance data includes a chrominance reconstruction block, a chrominance prediction block, and a chrominance boundary intensity, the chrominance reconstruction block, the chrominance prediction block, and the chrominance boundary intensity all need to be input into the same convolutional layer with a second step length.
[0137] It should be noted that by inputting all component data of the first chrominance data into the same convolutional layer, the number of convolutional layers used for the first chrominance data can be reduced, thereby improving the processing rate of feature extraction.
[0138] B22. Input the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the blue component data, and input the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the red component data;
[0139] It should be noted that, in this case, the number of convolution layers required is equal to the number of component data included in the first chrominance data, that is, different component data are input into different convolution layers.
[0140] For example, the red component data in the first chromaticity data is input to the convolution layer 1 of the second step length, and the blue component data in the first chromaticity data is input to the convolution layer 2 of the first step length.
[0141] It should be noted that, by inputting different component data of the first chromaticity data into different convolutional layers, feature extraction can be performed based on the different component data respectively, thereby improving the accuracy of feature extraction.
[0142] B23. Inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step length corresponding to the component type data respectively;
[0143] It should be noted that, since the red component data and the blue component data may include multiple component type data, in this case, the same component type data is input into the same convolution layer, and different component type data are input into different convolution layers.
[0144] For example, the red component data includes the chroma reconstruction block and the chroma prediction block corresponding to the red component, and the blue component data includes the chroma reconstruction block and the chroma prediction block corresponding to the blue component. The chroma reconstruction blocks in the red component data and the blue component data are input into the convolution layer 1 of the second step length, and the chroma prediction blocks in the red component data and the blue component data are input into the convolution layer 2 of the second step length.
[0145] It should be noted that, by inputting different component types of the first chromaticity data into different convolutional layers, feature extraction can be performed based on the different component types of the data, thereby improving the accuracy of feature extraction.
[0146] Optionally, in one implementation, the inputting the different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step size corresponding to the component type data respectively includes at least one of the following:
[0147] B231. When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance reconstruction block, add a luma reconstruction block to the chrominance reconstruction block, and input the chrominance reconstruction block and the luma reconstruction block into a convolutional layer of a second step size corresponding to the chrominance reconstruction block.
[0148] Optionally, in one implementation, the specific implementation of adding a luma reconstruction block to the chroma reconstruction block and inputting the chroma reconstruction block and the luma reconstruction block into a convolutional layer with a second step size corresponding to the chroma reconstruction block includes:
[0149] determining whether to downsample the luma reconstruction block according to the initial chroma subsampling format of the image block;
[0150] When it is determined that the luminance reconstruction block needs to be downsampled, the luminance reconstruction block after the downsampling process is added to the chrominance reconstruction block, and inputted into the convolution layer of the second step length corresponding to the chrominance reconstruction block;
[0151] When it is determined that the luminance reconstruction block does not need to be downsampled, the luminance reconstruction block is added to the chrominance reconstruction block and input into the convolution layer of the second step size corresponding to the chrominance reconstruction block.
[0152] It should be noted that when adding luminance reconstruction data features to the chrominance-based convolutional layer for luminance reconstruction feature extraction, the relationship between luminance and chrominance components can be used to obtain more useful feature maps, thereby improving the accuracy of chrominance reconstruction feature extraction.
[0153] B232. When the component type data corresponding to the blue component data and the red component data in the first chromaticity data include a chromaticity prediction block, a luminance prediction block is added to the chromaticity prediction block, and the chromaticity prediction block and the luminance prediction block are input into the convolution layer of the second step length corresponding to the chromaticity prediction block.
[0154] It should be noted that when adding luminance prediction data features to the chrominance-based luminance prediction feature extraction using the convolutional layer, the relationship between luminance and chrominance components can be used to obtain more useful feature maps, thereby improving the accuracy of chrominance prediction feature extraction.
[0155] B24. Inputting the different component type data in the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the component type data, and inputting the different component type data in the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the component type data;
[0156] It should be noted that in this case, not only the component data is distinguished, but also the component type data is distinguished, that is, different component type data under the same component data are input into different convolutional layers, and the same component type data under different component data are also input into different convolutional layers.
[0157] For example, the red component data in the first chrominance data includes the chrominance reconstruction block and the chrominance prediction block corresponding to the red component, and the blue component data includes the chrominance reconstruction block and the chrominance prediction block corresponding to the blue component. Then, the chrominance reconstruction block in the red component data is input into the convolution layer 1 of the second step length, and the chrominance prediction block in the red component data is input into the convolution layer 2 of the second step length. Then, the chrominance reconstruction block in the blue component data is input into the convolution layer 3 of the second step length, and the chrominance prediction block in the blue component data is input into the convolution layer 4 of the second step length.
[0158] It should be noted that by inputting different component type data under different component data in the first chromaticity data into different convolutional layers, feature extraction can be performed based on different component type data under different component data, thereby improving the accuracy of feature extraction.
[0159] It should be noted that the above-mentioned B11 or B12 can be combined with any implementation of B21-B24.
[0160] Optionally, in order to further reduce the amount of brightness data, in one implementation, the method further includes:
[0161] performing downsampling processing on the auxiliary data included in the first brightness data;
[0162] The auxiliary data is component data of the first luminance data other than the luminance reconstruction block; and the first step length of the convolution layer corresponding to the auxiliary data is different from the first step length of the convolution layer corresponding to the luminance reconstruction block.
[0163] It should be noted that, since the final result of video compression processing is the luminance reconstruction block, the luminance reconstruction block is mainly considered for video compression processing. Other luminance data can be used as auxiliary data in a smaller amount. Therefore, the auxiliary luminance data can be downsampled to reduce the resolution, thereby reducing the data amount and reducing the computational complexity of the neural network model.
[0164] Optionally, in one implementation, the number of channels of the convolution layer corresponding to the first luminance data is different from the number of channels of the convolution layer corresponding to the first chrominance data;
[0165] Optionally, the number of channels refers to the number of channels of the feature map output by the convolutional layer of the neural network model.
[0166] It should be noted that since the first luminance data and the first chrominance data are input into different convolutional layers, different numbers of channels can be set for different data based on the importance of the first luminance data and the first chrominance data. Since the number of channels represents the number of parameters of the model, the method of the present application can flexibly set the number of parameters of different types of data, thereby achieving the purpose of reducing the number of parameters of the model.
[0167] Optionally, the process of filtering based on the neural network model in the embodiment of the present application may be: first, the first luminance data and the first chrominance data are subjected to feature extraction through a convolution layer and an activation layer (it should be noted that each convolution layer may correspond to an activation layer, such as the activation layer may be a parametric rectified linear unit (PReLU)) to generate N feature maps, where N is a positive integer. Next, these feature maps are subjected to feature fusion through a convolution layer and an activation layer to generate M fused feature maps, where M is a positive integer. Then, these fused feature maps are subjected to feature enhancement through multiple residual blocks to generate K enhanced feature maps, where K is a positive integer. Finally, these enhanced feature maps are subjected to operations such as convolution layers and activation layers to generate filtered reconstructed image blocks.
[0168] Optionally, the process of super-resolution reconstruction based on the neural network model in the embodiment of the present application may be: first, the first luminance data and the first chrominance data are subjected to feature extraction through a convolution layer and an activation layer to generate N feature maps, where N is a positive integer. Next, these feature maps are subjected to feature fusion through a convolution layer and an activation layer to generate M fused feature maps, where M is a positive integer. Then, these fused feature maps are subjected to feature enhancement through multiple residual blocks to generate K enhanced feature maps. Finally, these enhanced feature maps are subjected to operations such as convolution layers, activation layers, and upsampling to generate high-resolution image blocks.
[0169] The following is an example of a video compression process using filtering based on a neural network model.
[0170] Application 1: All component data of the first luminance data are input to the convolution layer with a step size of 2, and all component data of the first chrominance data are input to the convolution layer with a step size of 1
[0171] As shown in Figure 5, Y[3] is the first luminance data with an input channel of 3, which is input into the convolution layer (the output channel is d1), and then passes through the activation layer for feature extraction; UV[6] is the first chrominance data with an input channel of 6, which is input into the convolution layer (the output channel is d2), and then passes through the activation layer for feature extraction; QPbase[1] is the quantization parameter of the video sequence with an input channel of 1, which is input into the convolution layer (the output channel is d4), and then passes through the activation layer for feature extraction; QPslice[1] is the quantization parameter of each frame with an input channel of 1, which is input into the convolution layer (the output channel is d4), and then passes through the activation layer for feature extraction; IPB[1] is the mode information (intra-frame prediction, inter-frame unidirectional prediction, inter-frame bidirectional prediction) of the coding block with an input channel of 1, which is input into the convolution layer (the output channel is d5), and then passes through the activation layer for feature extraction. The extracted features are input into the convolution layer (output channel is d6), and then through the activation layer for feature fusion. The fused features are then input into multiple residual blocks for feature enhancement. The enhanced features are pixel-mixed through the convolution layer and the activation layer, and cropped to obtain the reconstructed luminance block. The enhanced features are cropped through the convolution layer and the activation layer to obtain the reconstructed chrominance block.
[0172] Optionally, as shown in FIG6 , one implementation process of using the residual block to perform feature enhancement is as follows:
[0173] The input data passes through a convolution layer to obtain the output feature map, and then passes through a convolution layer again to obtain the output feature map; the output feature map is added to the input data to obtain the residual feature.
[0174] Application scenario 2: All component data of the first luminance data are input to the convolution layer with a step size of 2, the blue component data in the first chrominance data is input to the convolution layer with a step size of 1 corresponding to the blue component data, and the red component data in the first chrominance data is input to the convolution layer with a step size of 1 corresponding to the red component data.
[0175] As shown in Figure 7, Y[3] is the first luminance data with an input channel number of 3, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; U[3] is the blue component data with an input channel number of 3, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; V[3] is the red component data with an input channel number of 3, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; QPbase[1] is the quantization parameter of the video sequence with an input channel number of 1, which is input into the convolution layer (output channel is d4), and then passes through the activation layer for feature extraction; QPslice[1] is the quantization parameter of each frame with an input channel number of 1, which is input into the convolution layer (output channel is d4), and then passes through the activation layer for feature extraction; IPB[1] is the mode information (intra-frame prediction, inter-frame unidirectional prediction, inter-frame bidirectional prediction) of the coding block with an input channel number of 1, which is input into the convolution layer (output channel is d5), and then passes through the activation layer for feature extraction. The extracted features are input into the convolution layer (output channel is d6), and then through the activation layer for feature fusion. The fused features are then input into multiple residual blocks for feature enhancement. The enhanced features are pixel-mixed through the convolution layer and the activation layer, and cropped to obtain the reconstructed luminance block. The enhanced features are cropped through the convolution layer and the activation layer to obtain the reconstructed chrominance block.
[0176] Application scenario 3: Different component data in the first luminance data are respectively input into the convolution layer with a step size of 2 corresponding to the component data, and different component type data corresponding to the blue component data and the red component data in the first chrominance data are respectively input into the convolution layer with a step size of 1 corresponding to the component type data.
[0177] As shown in Figure 8, REC EXT Y is the brightness reconstruction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; Pred EXT Y is the brightness prediction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; BS EXT Y is the brightness boundary intensity, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; REC EXT UV is the chroma reconstruction block, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; Pred EXTUV is the chroma prediction block, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; BS EXT UV is the chroma boundary intensity, which is input into the convolution layer (output channel is d2) and then passes through the activation layer for feature extraction; QPbase[1] is the quantization parameter of the video sequence with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passes through the activation layer for feature extraction; QPslice[1] is the quantization parameter of each frame with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passes through the activation layer for feature extraction; IPB[1] is the mode information (intra-frame prediction, inter-frame unidirectional prediction, inter-frame bidirectional prediction) of the coding block with an input channel of 1, which is input into the convolution layer (output channel is d5) and then passes through the activation layer for feature extraction. The extracted features are input into the convolution layer (output channel is d6) and then passed through the activation layer for feature fusion. The fused features are then input into multiple residual blocks for feature enhancement. The enhanced features are pixel-mixed through the convolution layer and activation layer, and cropped to obtain the reconstructed luminance block. The enhanced features are cropped through the convolution layer and activation layer to obtain the reconstructed chroma block.
[0178] Application scenario 4: Different component data in the first luminance data are respectively input to the convolution layer with a step size of 2 corresponding to the component data, different component type data in the blue component data in the first chromaticity data are respectively input to the convolution layer with a step size of 1 corresponding to the component type data, and different component type data in the red component data in the first chromaticity data are respectively input to the convolution layer with a step size of 1 corresponding to the component type data.
[0179] As shown in Figure 9, REC EXT Y is the brightness reconstruction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; Pred EXT Y is the brightness prediction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; BS EXT Y is the brightness boundary intensity, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; REC EXT U is the chroma reconstruction block corresponding to the blue component data, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; Pred EXT U is the chroma prediction block corresponding to the blue component data, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; BS EXT U is the chromaticity boundary intensity corresponding to the blue component data, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; REC EXTV is the chroma reconstruction block corresponding to the red component data, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; Pred EXT V is the chroma prediction block corresponding to the red component data, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; BS EXT V is the chroma boundary intensity corresponding to the red component data, which is input into the convolution layer (output channel is d2) and then passed through the activation layer for feature extraction; QPbase[1] is the quantization parameter of the video sequence with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passed through the activation layer for feature extraction; QPslice[1] is the quantization parameter of each frame with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passed through the activation layer for feature extraction; IPB[1] is the mode information (intra-frame prediction, inter-frame unidirectional prediction, inter-frame bidirectional prediction) of the coding block with an input channel of 1, which is input into the convolution layer (output channel is d5) and then passed through the activation layer for feature extraction. The extracted features are input into the convolution layer (output channel is d6) and then passed through the activation layer for feature fusion. The fused features are then input into multiple residual blocks for feature enhancement. The enhanced features are pixel-mixed through the convolution layer and activation layer, and cropped to obtain the reconstructed luminance block. The enhanced features are cropped through the convolution layer and activation layer to obtain the reconstructed chrominance block.
[0180] Application scenario 5: Different component data in the first luminance data are respectively input to the convolution layer with a step size of 2 corresponding to the component data, and different component type data corresponding to the blue component data and the red component data in the first chrominance data are respectively input to the convolution layer with a step size of 1 corresponding to the component type data, and the luminance reconstruction block is added to the chrominance reconstruction block.
[0181] As shown in Figure 10, REC EXT Y is the brightness reconstruction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; Pred EXT Y is the brightness prediction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; BS EXT Y is the brightness boundary intensity, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; REC EXT UV&Y is a chroma reconstruction block that is added to the brightness reconstruction block, input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; Pred EXT UV is the chroma prediction block, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; BS EXTUV is the chroma boundary intensity, which is input into the convolution layer (output channel is d2) and then passes through the activation layer for feature extraction; QPbase[1] is the quantization parameter of the video sequence with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passes through the activation layer for feature extraction; QPslice[1] is the quantization parameter of each frame with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passes through the activation layer for feature extraction; IPB[1] is the mode information (intra-frame prediction, inter-frame unidirectional prediction, inter-frame bidirectional prediction) of the coding block with an input channel of 1, which is input into the convolution layer (output channel is d5) and then passes through the activation layer for feature extraction. The extracted features are input into the convolution layer (output channel is d6) and then passed through the activation layer for feature fusion. The fused features are then input into multiple residual blocks for feature enhancement. The enhanced features are pixel-mixed through the convolution layer and activation layer, and cropped to obtain the reconstructed luminance block. The enhanced features are cropped through the convolution layer and activation layer to obtain the reconstructed chroma block.
[0182] Application scenario six: Different component data in the first luminance data are input to the convolution layer with a step size of 2 corresponding to the component data, and different component type data corresponding to the blue component data and the red component data in the first chrominance data are input to the convolution layer with a step size of 1 corresponding to the component type data, and the luminance prediction block and the luminance boundary intensity are downsampled before input.
[0183] As shown in Figure 11, REC EXT Y is the brightness reconstruction block, which is input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; Pred EXT Y is the brightness prediction block, which is first downsampled and then input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; BS EXT Y is the brightness boundary intensity, which is first downsampled and then input into the convolution layer (output channel is d1), and then passes through the activation layer for feature extraction; REC EXT UV is the chroma reconstruction block, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; Pred EXT UV is the chroma prediction block, which is input into the convolution layer (output channel is d2), and then passes through the activation layer for feature extraction; BS EXTUV is the chroma boundary intensity, which is input into the convolution layer (output channel is d2) and then passes through the activation layer for feature extraction; QPbase[1] is the quantization parameter of the video sequence with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passes through the activation layer for feature extraction; QPslice[1] is the quantization parameter of each frame with an input channel of 1, which is input into the convolution layer (output channel is d4) and then passes through the activation layer for feature extraction; IPB[1] is the mode information (intra-frame prediction, inter-frame unidirectional prediction, inter-frame bidirectional prediction) of the coding block with an input channel of 1, which is input into the convolution layer (output channel is d5) and then passes through the activation layer for feature extraction. The extracted features are input into the convolution layer (output channel is d6) and then passed through the activation layer for feature fusion. The fused features are then input into multiple residual blocks for feature enhancement. The enhanced features are pixel-mixed through the convolution layer and activation layer, and cropped to obtain the reconstructed luminance block. The enhanced features are cropped through the convolution layer and activation layer to obtain the reconstructed chroma block.
[0184] In summary, the embodiments of the present application adjust the input data of the neural network model to reduce the amount of data input into the neural network model as much as possible, thereby reducing the data processing complexity and parameter amount of the neural network model.
[0185] The image block compression processing method provided in the embodiment of the present application can be executed by an image block compression processing device. In the embodiment of the present application, the image block compression processing device is used as an example to illustrate the image block compression processing method provided in the embodiment of the present application.
[0186] As shown in FIG12 , an image block compression processing apparatus 1200 according to an embodiment of the present application includes:
[0187] A first acquisition module 1201 is configured to acquire first luminance data and first chrominance data of an image block;
[0188] A processing module 1202 is configured to input the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0189] The second acquisition module 1203 is used to obtain the second brightness data and the second chromaticity data output by the neural network model.
[0190] Optionally, the device further includes:
[0191] a determination module, configured to determine, according to a target chroma subsampling format of the image block, the step sizes of the convolution layers corresponding to the first luminance data and the second chroma data, respectively;
[0192] The target chroma sub-sampling format is the chroma sub-sampling format of the image block output by the neural network model.
[0193] Optionally, the device further includes:
[0194] an adjusting module, configured to adjust a width or a height of the image block corresponding to the first chroma data according to a relationship between the target chroma subsampling format and the initial chroma subsampling format if a target chroma subsampling format of the image block is different from an initial chroma subsampling format;
[0195] The initial chroma sub-sampling format is the chroma sub-sampling format of the image block before inputting into the neural network model.
[0196] Optionally, the processing module 1202 is configured to implement one of the following:
[0197] Input all component data of the first brightness data into the first convolutional layer;
[0198] Different component data in the first brightness data are respectively input into the first step-length convolution layer corresponding to the component data.
[0199] Optionally, the processing module 1202 is configured to implement one of the following:
[0200] Inputting all component data of the first chrominance data into a convolutional layer of a second step length;
[0201] Inputting the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the blue component data, and inputting the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the red component data;
[0202] Inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step length corresponding to the component type data respectively;
[0203] The different component type data in the blue component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data, and the different component type data in the red component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data.
[0204] Optionally, the processing module 1202 is configured to implement at least one of the following:
[0205] When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance reconstruction block, adding a luma reconstruction block to the chrominance reconstruction block, and inputting the chrominance reconstruction block and the luma reconstruction block into a convolution layer of a second step size corresponding to the chrominance reconstruction block;
[0206] When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance prediction block, a luminance prediction block is added to the chrominance prediction block, and the chrominance prediction block and the luminance prediction block are input into the convolution layer of the second step size corresponding to the chrominance prediction block.
[0207] Optionally, the processing module 1202 is configured to:
[0208] determining whether to downsample the luma reconstruction block according to the initial chroma subsampling format of the image block;
[0209] When it is determined that the luminance reconstruction block needs to be downsampled, the luminance reconstruction block after the downsampling process is added to the chrominance reconstruction block, and inputted into the convolution layer of the second step length corresponding to the chrominance reconstruction block;
[0210] When it is determined that the luminance reconstruction block does not need to be downsampled, the luminance reconstruction block is added to the chrominance reconstruction block and input into the convolution layer of the second step size corresponding to the chrominance reconstruction block.
[0211] Optionally, the device further includes:
[0212] a sampling module, configured to perform down-sampling processing on the auxiliary data included in the first brightness data;
[0213] The auxiliary data is component data of the first luminance data other than the luminance reconstruction block; and the first step length of the convolution layer corresponding to the auxiliary data is different from the first step length of the convolution layer corresponding to the luminance reconstruction block.
[0214] Optionally, the component data of the first brightness data includes at least one of the following:
[0215] Luminance reconstruction block, luma prediction block, luma boundary intensity, luma residual block, luma transform coefficient block.
[0216] Optionally, the component data of the first chrominance data includes at least one of the following:
[0217] Blue component data;
[0218] Red component data.
[0219] Optionally, the blue component data and / or the red component data include at least one of the following component type data:
[0220] Chroma reconstruction block, chroma prediction block, chroma boundary intensity, chroma residual block, chroma transform coefficient block.
[0221] Optionally, the second luminance data includes a filtered luminance reconstruction block or a luminance reconstruction block of a first resolution, and the first resolution is higher than a resolution of the luminance reconstruction block corresponding to the first luminance data; or
[0222] The second chroma data includes a filtered chroma reconstruction block or a chroma reconstruction block with a second resolution, and the second resolution is higher than a resolution of the chroma reconstruction block corresponding to the first chroma data.
[0223] Optionally, the number of channels of the convolution layer corresponding to the first luminance data is different from the number of channels of the convolution layer corresponding to the first chrominance data;
[0224] The number of channels refers to the number of channels of the feature map output by the convolutional layer of the neural network model.
[0225] It should be noted that the device embodiment corresponds to the above method, and all implementation methods in the above method embodiment are applicable to the device embodiment and can achieve the same technical effects.
[0226] The image block prediction device in the embodiments of the present application can be an electronic device, such as an electronic device with an operating system, or a component of an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the terminal can include, but is not limited to, the types of terminal 11 listed above. Other devices can include servers, network attached storage (NAS), etc., and are not specifically limited in the embodiments of the present application.
[0227] The image block compression processing device provided in the embodiment of the present application can implement the various processes implemented in the method embodiment of Figure 4 and achieve the same technical effects. To avoid repetition, it will not be described here.
[0228] An embodiment of the present application further provides an electronic device, comprising a processor and a communication interface, wherein the processor is configured to obtain first luminance data and first chrominance data of an image block;
[0229] Inputting the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0230] Acquire second luminance data and second chrominance data output by the neural network model.
[0231] Optionally, the processor is further configured to:
[0232] Determining, according to a target chroma subsampling format of the image block, strides of convolutional layers corresponding to the first luminance data and the second chroma data, respectively;
[0233] The target chroma sub-sampling format is the chroma sub-sampling format of the image block output by the neural network model.
[0234] Optionally, the processor is further configured to:
[0235] If the target chroma subsampling format of the image block is different from the initial chroma subsampling format, adjusting the width or height of the image block corresponding to the first chroma data according to a relationship between the target chroma subsampling format and the initial chroma subsampling format;
[0236] The initial chroma sub-sampling format is the chroma sub-sampling format of the image block before inputting into the neural network model.
[0237] Optionally, the processor is configured to implement one of the following:
[0238] Input all component data of the first brightness data into the first convolutional layer;
[0239] Different component data in the first brightness data are respectively input into the first step-length convolution layer corresponding to the component data.
[0240] Optionally, the processor is configured to implement one of the following:
[0241] Inputting all component data of the first chrominance data into a convolutional layer of a second step length;
[0242] Inputting the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the blue component data, and inputting the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the red component data;
[0243] Inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step length corresponding to the component type data respectively;
[0244] The different component type data in the blue component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data, and the different component type data in the red component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data.
[0245] Optionally, the processor is configured to implement at least one of the following:
[0246] When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance reconstruction block, adding a luma reconstruction block to the chrominance reconstruction block, and inputting the chrominance reconstruction block and the luma reconstruction block into a convolution layer of a second step size corresponding to the chrominance reconstruction block;
[0247] When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance prediction block, a luminance prediction block is added to the chrominance prediction block, and the chrominance prediction block and the luminance prediction block are input into the convolution layer of the second step size corresponding to the chrominance prediction block.
[0248] Optionally, the processor is configured to:
[0249] determining whether to downsample the luma reconstruction block according to the initial chroma subsampling format of the image block;
[0250] When it is determined that the luminance reconstruction block needs to be downsampled, the luminance reconstruction block after the downsampling process is added to the chrominance reconstruction block, and inputted into the convolution layer of the second step length corresponding to the chrominance reconstruction block;
[0251] When it is determined that the luminance reconstruction block does not need to be downsampled, the luminance reconstruction block is added to the chrominance reconstruction block and input into the convolution layer of the second step size corresponding to the chrominance reconstruction block.
[0252] Optionally, the processor is further configured to:
[0253] performing downsampling processing on the auxiliary data included in the first brightness data;
[0254] The auxiliary data is component data of the first luminance data other than the luminance reconstruction block; and the first step length of the convolution layer corresponding to the auxiliary data is different from the first step length of the convolution layer corresponding to the luminance reconstruction block.
[0255] Optionally, the component data of the first brightness data includes at least one of the following:
[0256] Luma reconstruction block, luma prediction block, luma boundary intensity, luma residual block, luma transform coefficient block.
[0257] Optionally, the component data of the first chrominance data includes at least one of the following:
[0258] Blue component data;
[0259] Red component data.
[0260] Optionally, the blue component data and / or the red component data include at least one of the following component type data:
[0261] Chroma reconstruction block, chroma prediction block, chroma boundary intensity, chroma residual block, chroma transform coefficient block.
[0262] Optionally, the second luminance data includes a filtered luminance reconstruction block or a luminance reconstruction block of a first resolution, and the first resolution is higher than a resolution of the luminance reconstruction block corresponding to the first luminance data; or
[0263] The second chroma data includes a filtered chroma reconstruction block or a chroma reconstruction block with a second resolution, and the second resolution is higher than a resolution of the chroma reconstruction block corresponding to the first chroma data.
[0264] Optionally, the number of channels of the convolution layer corresponding to the first luminance data is different from the number of channels of the convolution layer corresponding to the first chrominance data;
[0265] The number of channels refers to the number of channels of the feature map output by the convolutional layer of the neural network model.
[0266] The electronic device embodiment corresponds to the above method embodiment, and each implementation process and implementation method of the above method embodiment are applicable to the electronic device embodiment and can achieve the same technical effect. Specifically, Figure 12 is a schematic diagram of the hardware structure of an electronic device implementing the embodiment of the present application.
[0267] The electronic device 1300 includes but is not limited to: a radio frequency unit 1301, a network module 1302, an audio output unit 1303, an input unit 1304, a sensor 1305, a display unit 1306, a user input unit 1307, an interface unit 1308, a memory 1309 and at least some of the components of the processor 1310.
[0268] Those skilled in the art will appreciate that the electronic device 1300 may further include a power source (e.g., a battery) to power various components. The power source may be logically connected to the processor 1310 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The electronic device structure shown in FIG13 does not limit the electronic device. The electronic device may include more or fewer components than shown, or may combine certain components, or have different component arrangements, which will not be described in detail here.
[0269] It should be understood that in an embodiment of the present application, the input unit 1304 may include a graphics processing unit (GPU) 13041 and a microphone 13042, and the graphics processor 13041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1306 may include a display panel 13061, and the display panel 13061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1307 includes a touch panel 13071 and at least one of the other input devices 13072. The touch panel 13071 is also called a touch screen. The touch panel 13071 may include two parts: a touch detection device and a touch controller. Other input devices 13072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0270] In the embodiment of the present application, after receiving downlink data from the access network device, the radio frequency unit 1301 can transmit the data to the processor 1310 for processing. In addition, the radio frequency unit 1301 can send uplink data to the network-side device. Generally, the radio frequency unit 1301 includes but is not limited to an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc.
[0271] The memory 1309 can be used to store software programs or instructions and various data. The memory 1309 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1309 may include a volatile memory or a non-volatile memory, or the memory 1309 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1309 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0272] Processor 1310 may include one or more processing units. Optionally, processor 1310 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1310.
[0273] The processor 1310 is configured to:
[0274] Acquire first luminance data and first chrominance data of the image block;
[0275] Inputting the first luminance data and the first chrominance data into different convolutional layers of a neural network model for video compression processing;
[0276] Acquire second luminance data and second chrominance data output by the neural network model.
[0277] Optionally, the processor 1310 is further configured to:
[0278] Determining, according to a target chroma subsampling format of the image block, strides of convolutional layers corresponding to the first luminance data and the second chroma data, respectively;
[0279] The target chroma sub-sampling format is the chroma sub-sampling format of the image block output by the neural network model.
[0280] Optionally, the processor 1310 is further configured to:
[0281] If the target chroma subsampling format of the image block is different from the initial chroma subsampling format, adjusting the width or height of the image block corresponding to the first chroma data according to a relationship between the target chroma subsampling format and the initial chroma subsampling format;
[0282] The initial chroma sub-sampling format is the chroma sub-sampling format of the image block before inputting into the neural network model.
[0283] Optionally, the processor 1310 is configured to implement one of the following:
[0284] Input all component data of the first brightness data into the first convolutional layer;
[0285] Different component data in the first brightness data are respectively input into the first step-length convolution layer corresponding to the component data.
[0286] Optionally, the processor 1310 is configured to implement one of the following:
[0287] Inputting all component data of the first chrominance data into a convolutional layer of a second step length;
[0288] Inputting the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the blue component data, and inputting the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the red component data;
[0289] Inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step length corresponding to the component type data respectively;
[0290] The different component type data in the blue component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data, and the different component type data in the red component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data.
[0291] Optionally, the processor 1310 is configured to implement at least one of the following:
[0292] When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance reconstruction block, adding a luma reconstruction block to the chrominance reconstruction block, and inputting the chrominance reconstruction block and the luma reconstruction block into a convolution layer of a second step size corresponding to the chrominance reconstruction block;
[0293] When the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance prediction block, a luminance prediction block is added to the chrominance prediction block, and the chrominance prediction block and the luminance prediction block are input into the convolution layer of the second step size corresponding to the chrominance prediction block.
[0294] Optionally, the processor 1310 is configured to:
[0295] determining whether to downsample the luma reconstruction block according to the initial chroma subsampling format of the image block;
[0296] When it is determined that the luminance reconstruction block needs to be downsampled, the luminance reconstruction block after the downsampling process is added to the chrominance reconstruction block, and inputted into the convolution layer of the second step length corresponding to the chrominance reconstruction block;
[0297] When it is determined that the luminance reconstruction block does not need to be downsampled, the luminance reconstruction block is added to the chrominance reconstruction block and input into the convolution layer of the second step size corresponding to the chrominance reconstruction block.
[0298] Optionally, the processor 1310 is further configured to:
[0299] performing downsampling processing on the auxiliary data included in the first brightness data;
[0300] The auxiliary data is component data of the first luminance data other than the luminance reconstruction block; and the first step length of the convolution layer corresponding to the auxiliary data is different from the first step length of the convolution layer corresponding to the luminance reconstruction block.
[0301] Optionally, the component data of the first brightness data includes at least one of the following:
[0302] Luminance reconstruction block, luma prediction block, luma boundary intensity, luma residual block, luma transform coefficient block.
[0303] Optionally, the component data of the first chrominance data includes at least one of the following:
[0304] Blue component data;
[0305] Red component data.
[0306] Optionally, the blue component data and / or the red component data include at least one of the following component type data:
[0307] Chroma reconstruction block, chroma prediction block, chroma boundary intensity, chroma residual block, chroma transform coefficient block.
[0308] Optionally, the second luminance data includes a filtered luminance reconstruction block or a luminance reconstruction block of a first resolution, and the first resolution is higher than a resolution of the luminance reconstruction block corresponding to the first luminance data; or
[0309] The second chroma data includes a filtered chroma reconstruction block or a chroma reconstruction block with a second resolution, and the second resolution is higher than a resolution of the chroma reconstruction block corresponding to the first chroma data.
[0310] Optionally, the number of channels of the convolution layer corresponding to the first luminance data is different from the number of channels of the convolution layer corresponding to the first chrominance data;
[0311] The number of channels refers to the number of channels of the feature map output by the convolutional layer of the neural network model.
[0312] Preferably, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the various processes of the above-mentioned image block compression processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here.
[0313] An embodiment of the present application also provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned image block compression processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0314] The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0315] Optionally, as shown in Figure 14, an embodiment of the present application also provides a communication device 1400, including a processor 1401 and a memory 1402, and the memory 1402 stores a program or instruction that can be run on the processor 1401. When the program or instruction is executed by the processor 1401, the various steps of the above-mentioned image block compression processing method embodiment are implemented and the same technical effect can be achieved.
[0316] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned image block compression processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0317] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0318] An embodiment of the present application further provides a computer program / program product, which is stored in a storage medium. The computer program / program product is executed by at least one processor to implement the various processes of the above-mentioned image block compression processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0319] An embodiment of the present application further provides a coding and decoding system, comprising: an image block compression processing device, which can be used to execute the steps of the above-mentioned image block compression processing method.
[0320] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0321] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0322] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A method for image block compression processing, comprising: Acquire first brightness data and first chrominance data of the image block; Inputting the first brightness data and the first chrominance data into different convolutional layers of a neural network model respectively for video compression processing; Obtain second brightness data and second chromaticity data output by the neural network model.
2. The method according to claim 1, wherein: Also includes: Determining, according to a target chroma subsampling format of the image block, the step sizes of the convolutional layers corresponding to the first luminance data and the second chroma data respectively; The target chroma sub-sampling format is the chroma sub-sampling format of the image block output by the neural network model.
3. The method according to claim 1 or 2, wherein: Before the first brightness data and the first chrominance data are respectively input into different convolutional layers of a neural network model for video compression processing, the method further includes: If the target chroma subsampling format of the image block is different from the initial chroma subsampling format, adjusting the width or height of the image block corresponding to the first chroma data according to the relationship between the target chroma subsampling format and the initial chroma subsampling format; The initial chroma sub-sampling format is the chroma sub-sampling format of the image block before being input into the neural network model.
4. The method according to any one of claims 1 to 3, wherein: The first brightness data is input to the convolutional layer of the neural network model, including one of the following: Input all component data of the first brightness data into the first convolutional layer; Different component data in the first brightness data are respectively input into the first step-length convolution layer corresponding to the component data.
5. The method according to any one of claims 1 to 4, wherein: Inputting the first chromaticity data into a convolutional layer of a neural network model includes one of the following: Input all component data of the first chrominance data into a convolutional layer of a second step length; Inputting the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the blue component data, and inputting the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the red component data; Inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step length corresponding to the component type data respectively; The different component type data in the blue component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data, and the different component type data in the red component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data.
6. The method according to claim 5, wherein: The step of inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the component type data respectively includes at least one of the following: In a case where the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance reconstruction block, adding a luminance reconstruction block to the chrominance reconstruction block, and inputting the chrominance reconstruction block and the luminance reconstruction block to a convolutional layer of a second step length corresponding to the chrominance reconstruction block; When the component type data corresponding to the blue component data and the red component data in the first chromaticity data include a chromaticity prediction block, a luminance prediction block is added to the chromaticity prediction block, and the chromaticity prediction block and the luminance prediction block are input into the convolution layer of the second step length corresponding to the chromaticity prediction block.
7. The method according to claim 6, wherein: The step of adding a luminance reconstruction block to the chrominance reconstruction block and inputting the chrominance reconstruction block and the luminance reconstruction block into a convolutional layer of a second step length corresponding to the chrominance reconstruction block comprises: Determining whether to downsample the luminance reconstruction block according to the initial chroma subsampling format of the image block; When it is determined that the luminance reconstruction block needs to be downsampled, the luminance reconstruction block after the downsampling process is added to the chrominance reconstruction block, and inputted into the convolution layer of the second step length corresponding to the chrominance reconstruction block; When it is determined that the luminance reconstruction block does not need to be downsampled, the luminance reconstruction block is added to the chrominance reconstruction block and input into the convolution layer of the second step length corresponding to the chrominance reconstruction block.
8. The method according to any one of claims 1 to 7, wherein: Also includes: performing down-sampling processing on the auxiliary data included in the first brightness data; The auxiliary data is component data of the first brightness data except for the brightness reconstruction block; the first step length of the convolution layer corresponding to the auxiliary data is different from the first step length of the convolution layer corresponding to the brightness reconstruction block.
9. The method according to any one of claims 1 to 8, wherein: The component data of the first brightness data includes at least one of the following: Luma reconstruction block, luma prediction block, luma boundary strength, luma residual block, luma transform coefficient block.
10. The method according to any one of claims 1 to 8, wherein: The component data of the first chrominance data includes at least one of the following: Blue component data; Red component data.
11. The method according to claim 10, wherein: The blue component data and / or the red component data include at least one of the following component type data: Chroma reconstruction block, chroma prediction block, chroma boundary intensity, chroma residual block, chroma transform coefficient block.
12. The method according to claim 1, wherein: The second brightness data includes a filtered brightness reconstruction block or a brightness reconstruction block of a first resolution, and the first resolution is higher than a resolution of the brightness reconstruction block corresponding to the first brightness data; or The second chrominance data includes a filtered chrominance reconstruction block or a chrominance reconstruction block of a second resolution, and the second resolution is higher than a resolution of the chrominance reconstruction block corresponding to the first chrominance data.
13. The method according to claim 1, wherein: The number of channels of the convolution layer corresponding to the first luminance data is different from the number of channels of the convolution layer corresponding to the first chrominance data.
14. An image block compression processing device, comprising: A first acquisition module, used to acquire first brightness data and first chrominance data of an image block; A processing module, used for inputting the first brightness data and the first chrominance data into different convolutional layers of a neural network model for video compression processing; The second acquisition module is used to acquire second brightness data and second chromaticity data output by the neural network model.
15. The device according to claim 14, wherein: Also includes: A determination module, configured to determine the step sizes of the convolutional layers corresponding to the first luminance data and the second chrominance data, respectively, according to a target chrominance subsampling format of the image block; The target chroma sub-sampling format is the chroma sub-sampling format of the image block output by the neural network model.
16. The device according to claim 14 or 15, wherein: Also includes: an adjustment module, configured to adjust the width or height of the image block corresponding to the first chroma data according to a relationship between the target chroma subsampling format and the initial chroma subsampling format if the target chroma subsampling format of the image block is different from the initial chroma subsampling format; The initial chroma sub-sampling format is the chroma sub-sampling format of the image block before being input into the neural network model.
17. The device according to any one of claims 14 to 16, wherein: The processing module is used to implement one of the following: Input all component data of the first brightness data into the first convolutional layer; Different component data in the first brightness data are respectively input into the first step-length convolution layer corresponding to the component data.
18. The device according to any one of claims 14 to 17, wherein: The processing module is used to implement one of the following: Input all component data of the first chrominance data into a convolutional layer of a second step length; Inputting the blue component data in the first chromaticity data to the convolution layer of the second step length corresponding to the blue component data, and inputting the red component data in the first chromaticity data to the convolution layer of the second step length corresponding to the red component data; Inputting different component type data corresponding to the blue component data and the red component data in the first chromaticity data into the convolution layer of the second step length corresponding to the component type data respectively; The different component type data in the blue component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data, and the different component type data in the red component data in the first chromaticity data are respectively input into the convolution layer of the second step length corresponding to the component type data.
19. The device according to claim 18, wherein: The processing module is used to implement at least one of the following: In a case where the component type data corresponding to the blue component data and the red component data in the first chrominance data include a chrominance reconstruction block, adding a luminance reconstruction block to the chrominance reconstruction block, and inputting the chrominance reconstruction block and the luminance reconstruction block to a convolutional layer of a second step length corresponding to the chrominance reconstruction block; When the component type data corresponding to the blue component data and the red component data in the first chromaticity data include a chromaticity prediction block, a luminance prediction block is added to the chromaticity prediction block, and the chromaticity prediction block and the luminance prediction block are input into the convolution layer of the second step length corresponding to the chromaticity prediction block.
20. The device according to claim 19, wherein The processing module is used for: Determining whether to downsample the luminance reconstruction block according to the initial chroma subsampling format of the image block; When it is determined that the luminance reconstruction block needs to be downsampled, the luminance reconstruction block after the downsampling process is added to the chrominance reconstruction block, and inputted into the convolution layer of the second step length corresponding to the chrominance reconstruction block; When it is determined that the luminance reconstruction block does not need to be downsampled, the luminance reconstruction block is added to the chrominance reconstruction block and input into the convolution layer of the second step length corresponding to the chrominance reconstruction block.
21. The device according to any one of claims 14 to 20, wherein: Also includes: A sampling module, used for down-sampling the auxiliary data included in the first brightness data; The auxiliary data is component data of the first brightness data except for the brightness reconstruction block; the first step length of the convolution layer corresponding to the auxiliary data is different from the first step length of the convolution layer corresponding to the brightness reconstruction block.
22. An electronic device comprising a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the image block compression processing method according to any one of claims 1 to 13 are implemented.
23. A readable storage medium storing a program or instruction, wherein the program or instruction, when executed by a processor, implements the steps of the image block compression processing method according to any one of claims 1 to 13.
24. A chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps of the image block compression processing method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Image processing method, image processing device, terminal and readable storage medium
CN115187679A
End-to-end neural network based video coding
CN116114247A
YUV Image Processing Method and System Using A Neural Network Composed Of Dual-Path Blocks
KR102599753B1
Image signal processor for processing images
US20190108618A1