Neural network-based video processing method and apparatus, and related device
By acquiring and determining the offset values of the brightness quantization parameters and chromaticity quantization parameters, the problem that neural network models are difficult to control brightness and chromaticity output quality is solved, and independent quality control of the output image and better brightness balance are achieved.
Patent Information
- Application Number
- PCT/CN2024/138697
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-26
AI Technical Summary
The neural network model based on quantization parameters is difficult to control the brightness output quality and the chromaticity output quality respectively, which makes it difficult to balance the inference results of brightness chromaticity.
By obtaining the offset values of the brightness quantization parameters and the chromaticity quantization parameters, the brightness quantization parameters and chromaticity quantization parameters are determined respectively as input data of the neural network so that the neural network can independently control the output quality of brightness and chromaticity.
The independent control of the brightness and chromaticity output quality of the output image is realized, thereby better balancing the inference results of the neural network for brightness.
Smart Images

Figure CN2024138697_26062025_PF_FP_ABST
Abstract
Description
Video processing method, device and related equipment based on neural network
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese Patent Application No. 202311760147.6 filed on December 19, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application belongs to the field of coding and decoding technology, and specifically relates to a video processing method, apparatus and related equipment based on a neural network. Background Art
[0004] In unified network models, quantization parameters (such as sequence quantization parameters or frame quantization parameters) are typically used as neural network inputs to provide auxiliary information and guide the network training to achieve better performance. However, based on these quantization parameters, the network model has difficulty controlling the quality of luminance output and chrominance output separately, and as a result, the network model has difficulty in properly balancing the inference results of luminance and chrominance. Summary of the Invention
[0005] The embodiments of the present application provide a video processing method, apparatus and related equipment based on a neural network, which can solve the problem that the quantization parameter network model in the related technology is difficult to control the brightness output quality and the chrominance output quality separately, and thus the network model is difficult to balance the inference results of brightness and chrominance well.
[0006] In a first aspect, a neural network-based video processing method is provided, which is performed by an electronic device. The method includes:
[0007] Obtaining a first offset value and a second offset value, wherein the first offset value refers to an offset value of a parameter value of a luma quantization parameter relative to a first quantization parameter value, and the second offset value refers to an offset value of a parameter value of a chroma quantization parameter relative to a second quantization parameter value;
[0008] determining a luminance quantization parameter according to the first offset value, and determining a chrominance quantization parameter according to the second offset value;
[0009] According to the brightness quantization parameter and the chrominance quantization parameter, input data of the neural network is obtained.
[0010] In a second aspect, a video processing device based on a neural network is provided, comprising:
[0011] a first acquiring module, configured to acquire a first offset value and a second offset value, wherein the first offset value refers to an offset value of a parameter value of a luma quantization parameter relative to a first quantization parameter value, and the second offset value refers to an offset value of a parameter value of a chroma quantization parameter relative to a second quantization parameter value;
[0012] a determination module, configured to determine a luma quantization parameter according to the first offset value, and determine a chroma quantization parameter according to the second offset value;
[0013] The second acquisition module is used to obtain input data of the neural network according to the brightness quantization parameter and the chrominance quantization parameter.
[0014] In a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0015] In a fourth aspect, an electronic device is provided, comprising a processor and a communication interface, wherein the processor is used to obtain a first offset value and a second offset value, the first offset value referring to the offset value of the parameter value of the luminance quantization parameter relative to the first quantization parameter value, and the second offset value referring to the offset value of the parameter value of the chrominance quantization parameter relative to the second quantization parameter value; determine the luminance quantization parameter and the chrominance quantization parameter based on the first offset value and the second offset value; and obtain input data of a neural network based on the luminance quantization parameter and the chrominance quantization parameter.
[0016] In a fifth aspect, an electronic device is provided, comprising: a memory configured to store video data, and a processing circuit configured to implement the steps of the method described in the first aspect.
[0017] In a sixth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0018] In a seventh aspect, a chip is provided, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the steps of the method described in the first aspect.
[0019] In an eighth aspect, a computer program / program product is provided, wherein the computer program / program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect.
[0020] In an embodiment of the present application, a first offset value and a second offset value are obtained, wherein the first offset value refers to the offset value of the parameter value of the luminance quantization parameter relative to the first quantization parameter value, and the second offset value refers to the offset value of the parameter value of the chrominance quantization parameter relative to the second quantization parameter value; the luminance quantization parameter is determined based on the first offset value, and the chrominance quantization parameter is determined based on the second offset value; and the input data of the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter. In this solution, the input data of the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter brightness, that is, luminance and chrominance no longer share the same quantization parameter in the neural network, but each corresponds to a quantization parameter. The luminance quantization parameter corresponding to luminance enables the neural network to effectively control the luminance output quality of the output image, and the chrominance quantization parameter corresponding to chrominance enables the neural network to effectively control the chrominance output quality of the output image, thereby better balancing the neural network's reasoning results on brightness and chrominance. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG1 is a schematic diagram showing a flow chart of a video processing method based on a neural network according to an embodiment of the present application;
[0022] FIG2 is a block diagram of a neural network-based video processing device according to an embodiment of the present application;
[0023] FIG3 is a block diagram showing the structure of an electronic device according to an embodiment of the present application;
[0024] FIG4 shows a structural block diagram of a terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0026] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0027] The term "indication" in this application can be either a direct indication (or explicit indication) or an indirect indication (or implicit indication). A direct indication can be understood as the sender explicitly informing the receiver of specific information, the operation to be performed, or the requested result, etc. in the instruction sent; an indirect indication can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the operation to be performed or the requested result, etc. based on the judgment result.
[0028] In order to enable those skilled in the art to better understand the embodiments of the present application, the following description is first given.
[0029] 1. Neural network-based loop filter solution;
[0030] To reduce the impact of distortion effects like blocking and ringing on video quality, video compression standards incorporate loop filters, such as deblocking and adaptive sample compensation. Deblocking filters reduce blocking artifacts, while adaptive sample compensation improves ringing artifacts. These loop filters effectively improve both subjective and objective video quality within the codec loop.
[0031] With the widespread application of neural network-based methods in image processing, many neural network-based loop filter solutions also have good performance in the direction of video compression. Neural network filters can completely replace traditional filters, or they can partially replace traditional filters.
[0032] 2. Super-resolution solution based on neural network;
[0033] Neural network-based super-resolution approaches utilize neural networks to reconstruct images for super-resolution, enhancing clarity and detail. This approach achieves super-resolution by training neural networks to learn image features. Common neural networks include convolutional neural networks and recurrent neural networks. Within image super-resolution approaches, neural networks can be used for single-image super-resolution reconstruction and for the joint super-resolution reconstruction of multiple images. This approach has been widely applied in fields such as image processing, medical image analysis, and security monitoring, and continues to evolve and improve.
[0034] The neural network-based video processing method provided by the embodiment of the present application is described in detail below with reference to some embodiments and their application scenarios in conjunction with the accompanying drawings.
[0035] As shown in FIG1 , an embodiment of the present application provides a video processing method based on a neural network, which is executed by an electronic device. The method includes:
[0036] Step 101: Obtain a first offset value and a second offset value, wherein the first offset value refers to an offset value of a luminance quantization parameter relative to a first quantization parameter value, and the second offset value refers to an offset value of a chrominance quantization parameter relative to a second quantization parameter value.
[0037] Optionally, the first quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a chrominance quantization parameter.
[0038] Optionally, the second quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a brightness quantization parameter.
[0039] The above-mentioned sequence-level quantization parameter refers to a quantization parameter for an image frame sequence, and the above-mentioned frame-level quantization parameter refers to a quantization parameter for an image frame.
[0040] Optionally, the first quantization parameter and the second quantization parameter are obtained from an encoded bitstream.
[0041] Of course, the first quantization parameter value and the second quantization parameter value may also be pre-set quantization parameter values. The first quantization parameter value and the second quantization parameter value may be the same or different.
[0042] Step 102: Determine a luma quantization parameter according to the first offset value, and determine a chroma quantization parameter according to the second offset value.
[0043] Optionally, the quantization parameter is used to dequantize the quantized coefficients during the decoding process. The quantization parameter controls the compression ratio and video quality. A smaller quantization parameter results in lower coefficient precision, but a higher compression ratio and lower video quality. A larger quantization parameter results in higher coefficient precision, but a lower compression ratio and better video quality.
[0044] Optionally, the luminance quantization parameter is used to control the compression rate of a video luminance signal and the quality of the video luminance signal, and the chrominance quantization parameter is used to control the compression rate of a video chrominance signal and the quality of the video chrominance signal.
[0045] As an implementation manner, the first offset value is added to the first quantization parameter value to obtain the luminance quantization parameter; the second offset value is added to the second quantization parameter value to obtain the chrominance quantization parameter.
[0046] Step 103: Obtain input data of a neural network according to the brightness quantization parameter and the chrominance quantization parameter.
[0047] The brightness quantization parameter and the chromaticity quantization parameter are respectively input into a convolutional layer of the neural network. Based on the two convolutional layers, the feature map of the corresponding brightness quantization parameter and the feature map of the chromaticity quantization parameter can be extracted, that is, the feature map of the brightness quantization parameter and the feature map of the chromaticity quantization parameter can be provided to the neural network, so that the neural network can subsequently improve the brightness output quality of the output image based on the feature map of the brightness quantization parameter, and improve the chromaticity output quality of the output image based on the feature map of the chromaticity quantization parameter, thereby better balancing the inference results of the neural network on brightness and chromaticity.
[0048] Optionally, the neural network includes a filtering network or a super-resolution network.
[0049] In an embodiment of the present application, a first offset value and a second offset value are obtained, wherein the first offset value refers to the offset value of the parameter value of the luminance quantization parameter relative to the first quantization parameter value, and the second offset value refers to the offset value of the parameter value of the chrominance quantization parameter relative to the second quantization parameter value; the luminance quantization parameter is determined based on the first offset value, and the chrominance quantization parameter is determined based on the second offset value; and the input data of the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter. In this solution, the input data of the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter brightness, that is, luminance and chrominance no longer share the same quantization parameter in the neural network, but each corresponds to a quantization parameter. The luminance quantization parameter corresponding to luminance enables the neural network to effectively control the luminance output quality of the output image, and the chrominance quantization parameter corresponding to chrominance enables the neural network to effectively control the chrominance output quality of the output image, thereby better balancing the neural network's reasoning results on brightness and chrominance.
[0050] Optionally, obtaining the first offset value and the second offset value includes:
[0051] Get the brightness quantization parameter index and the chrominance quantization parameter index;
[0052] According to the correspondence between the quantization parameter index and the offset value, a first offset value corresponding to the luma quantization parameter index and a second offset value corresponding to the chroma quantization parameter index are obtained.
[0053] Optionally, the above correspondence is pre-set, for example, a correspondence between a quantization parameter index in a quantization parameter index set and an offset value in a preset offset value set is pre-set. It should be noted that the preset offset value set corresponding to the luma quantization parameter index and the preset offset value set corresponding to the chroma quantization parameter index may be the same or different.
[0054] Optionally, the luma quantization parameter index and the chroma luma parameter index may be the same or different.
[0055] Exemplarily, assuming that the preset offset value set corresponding to the luma quantization parameter index is the same as the preset offset value set corresponding to the chroma quantization parameter index, for example, the preset offset value set is {-5, 0, 5}, if the luma quantization parameter index or the chroma quantization parameter index is 0, the corresponding offset value is 0; if the luma quantization parameter index or the chroma quantization parameter index is 1, the corresponding offset value is -5; if the quantization parameter index luma quantization parameter index or the chroma quantization parameter index is 2, the corresponding offset value is 5; if the first quantization parameter value or the second quantization parameter value is the parameter value of the sequence-level quantization parameter or the parameter value of the frame-level quantization parameter, the final luma quantization parameter or the chroma quantization parameter is the offset value plus the parameter value of the corresponding sequence-level quantization parameter or the parameter value of the frame-level quantization parameter; if the first quantization parameter value is the parameter value of the sequence-level quantization parameter or the parameter value of the frame-level quantization parameter, and the second quantization parameter value is the parameter value of the luma quantization parameter, the final luma quantization parameter is the offset value plus the parameter value of the corresponding sequence-level quantization parameter or the parameter value of the frame-level quantization parameter, and the chroma quantization parameter is the offset value plus the parameter value of the chroma quantization parameter.
[0056] In an embodiment of the present application, a luminance quantization parameter index and a chrominance quantization parameter index are parsed from the encoded bitstream, and then a corresponding offset value is obtained based on the correspondence between the quantization parameter index and the offset value, and then the corresponding quantization parameter and chrominance quantization parameter can be obtained based on the offset value.
[0057] As an optional implementation manner, obtaining, according to the correspondence between the quantization parameter index and the offset value, the first offset value corresponding to the luma quantization parameter index and the second offset value corresponding to the chroma quantization parameter index includes:
[0058] According to the correspondence between the brightness quantization parameter index and the brightness offset value, obtaining a first offset value corresponding to the brightness quantization parameter index;
[0059] According to the correspondence between the chromaticity quantization parameter index and the chromaticity offset value, a second offset value corresponding to the chromaticity quantization parameter index is obtained.
[0060] In this implementation, the preset offset value set corresponding to the luma quantization parameter index is different from the preset offset value set corresponding to the chroma quantization parameter index. For example, the luma quantization parameter index corresponds to a first preset offset value set, which includes at least one luma offset value, and the chroma quantization parameter index corresponds to a second preset offset value set, which includes at least one chroma offset value.
[0061] In this implementation, the luminance quantization parameter index and the chrominance quantization parameter index correspond to different preset offset value sets, which can more flexibly set the offset values corresponding to the luminance quantization parameter index and the chrominance quantization parameter index, so that the neural network can better control the luminance output quality and the chrominance output quality.
[0062] Optionally, the brightness quantization parameter is a sequence-level quantization parameter, or a frame-level quantization parameter, or a block-level quantization parameter;
[0063] Alternatively, the chrominance quantization parameter is a sequence-level quantization parameter, a frame-level quantization parameter, or a block-level quantization parameter.
[0064] The above-mentioned sequence-level quantization parameter refers to a quantization parameter for an image frame sequence, the above-mentioned frame-level quantization parameter refers to a quantization parameter for an image frame, and the above-mentioned block-level quantization parameter refers to a quantization parameter for an image block.
[0065] Through the above-mentioned luminance quantization parameters, the neural network can control the luminance output quality of the image frame sequence, or control the luminance output quality of the image frame, or control the luminance output quality of the image block; through the above-mentioned chrominance quantization parameters, the neural network can control the chrominance output quality of the image frame sequence, or control the chrominance output quality of the image frame, or control the chrominance output quality of the image block.
[0066] Optionally, the input data further includes: a reconstructed image block; or the input data further includes: a predicted image block, a boundary strength, coding block mode information, at least one of a sequence-level quantization parameter and a frame-level quantization parameter, and the reconstructed image block; the method further includes:
[0067] Performing filtering based on the neural network and the input data to obtain a filtered image block;
[0068] Alternatively, super-resolution reconstruction processing is performed based on the neural network and the input data to obtain a reconstructed image block, and the resolution of the reconstructed image block is higher than the resolution of the reconstructed image block.
[0069] In an embodiment of the present application, each parameter in the input data is input into a convolutional layer of the neural network, and each of the convolutional layers may correspond to an activation layer, such as the activation layer may be a Parametric Rectified Linear Unit (PReLU).
[0070] The filtering process based on a neural network and input data can be as follows: First, the input data passes through convolutional layers and activation layers for feature extraction, generating N feature maps, where N is a positive integer. Next, these feature maps are fused through convolutional layers and activation layers to generate M fused feature maps, where M is a positive integer. These fused feature maps are then enhanced through multiple residual blocks to generate K enhanced feature maps, where K is a positive integer. Finally, these enhanced feature maps pass through convolutional layers and activation layers to generate filtered reconstructed image blocks.
[0071] The process of performing super-resolution reconstruction based on the neural network and the input data may be as follows: first, the input data is subjected to feature extraction through a convolutional layer and an activation layer to generate N feature maps, where N is a positive integer. Next, these feature maps are subjected to feature fusion through a convolutional layer and an activation layer to generate M fused feature maps, where M is a positive integer. Then, these fused feature maps are subjected to feature enhancement through multiple residual blocks to generate K enhanced feature maps. Finally, these enhanced feature maps are subjected to operations such as convolutional layers, activation layers, and upsampling to generate high-resolution image blocks.
[0072] The above-mentioned scheme of the embodiment of the present application obtains the input data of the neural network based on the brightness quantization parameter and the chromaticity quantization parameter brightness, that is, brightness and chromaticity no longer share the same quantization parameter in the neural network, but correspond to a quantization parameter respectively. The brightness quantization parameter corresponding to the brightness enables the neural network to effectively control the brightness output quality of the output image, and the chromaticity quantization parameter corresponding to the chromaticity enables the neural network to effectively control the chromaticity output quality of the output image, thereby better balancing the inference results of the neural network on brightness and chromaticity.
[0073] The neural network video processing method provided in the embodiment of the present application can be executed by a neural network video processing device. In the embodiment of the present application, the neural network video processing device provided in the embodiment of the present application is described by taking the neural network video processing device executing the neural network video processing method as an example.
[0074] As shown in FIG2 , the embodiment of the present application further provides a video processing device 200 based on a neural network, comprising:
[0075] A first acquisition module 201 is configured to acquire a first offset value and a second offset value, wherein the first offset value refers to an offset value of a luminance quantization parameter relative to a first quantization parameter value, and the second offset value refers to an offset value of a chrominance quantization parameter relative to a second quantization parameter value;
[0076] a determination module 202, configured to determine a luma quantization parameter according to the first offset value, and to determine a chroma quantization parameter according to the second offset value;
[0077] The second acquisition module 203 is configured to obtain input data of a neural network according to the brightness quantization parameter and the chrominance quantization parameter.
[0078] Optionally, the first acquisition module includes:
[0079] A first acquisition submodule is used to acquire a luminance quantization parameter index and a chrominance quantization parameter index;
[0080] The second acquisition submodule is configured to acquire a first offset value corresponding to the luma quantization parameter index and a second offset value corresponding to the chroma quantization parameter index according to a correspondence between the quantization parameter index and the offset value.
[0081] Optionally, the second acquisition submodule includes:
[0082] A first acquiring unit, configured to acquire a first offset value corresponding to the luma quantization parameter index according to a correspondence between the luma quantization parameter index and the luma offset value;
[0083] The second acquiring unit is configured to acquire a second offset value corresponding to the chromaticity quantization parameter index according to a correspondence between the chromaticity quantization parameter index and the chromaticity offset value.
[0084] Optionally, the first quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a chrominance quantization parameter;
[0085] Alternatively, the second quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a brightness quantization parameter.
[0086] Optionally, the brightness quantization parameter is a sequence-level quantization parameter, or a frame-level quantization parameter, or a block-level quantization parameter;
[0087] Alternatively, the chrominance quantization parameter is a sequence-level quantization parameter, a frame-level quantization parameter, or a block-level quantization parameter.
[0088] Optionally, the input data further includes: a reconstructed image block; or the input data further includes a predicted image block, a boundary strength, coding block mode information, at least one of a sequence-level quantization parameter and a frame-level quantization parameter, and the reconstructed image block; the apparatus further includes:
[0089] A third acquisition module is used to perform filtering processing based on the neural network and the input data to obtain a filtered image block; or to perform super-resolution reconstruction processing based on the neural network and the input data to obtain a reconstructed image block, and the resolution of the reconstructed image block is higher than the resolution of the reconstructed image block.
[0090] The device of the embodiment of the present application obtains a first offset value and a second offset value, wherein the first offset value refers to the offset value of the parameter value of the luminance quantization parameter relative to the first quantization parameter value, and the second offset value refers to the offset value of the parameter value of the chrominance quantization parameter relative to the second quantization parameter value; determines the luminance quantization parameter based on the first offset value, and determines the chrominance quantization parameter based on the second offset value; and obtains input data for the neural network based on the luminance quantization parameter and the chrominance quantization parameter. In this scheme, the input data for the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter brightness, that is, luminance and chrominance no longer share the same quantization parameter in the neural network, but each corresponds to a quantization parameter. The luminance quantization parameter corresponding to luminance enables the neural network to effectively control the luminance output quality of the output image, and the chrominance quantization parameter corresponding to chrominance enables the neural network to effectively control the chrominance output quality of the output image, thereby better balancing the neural network's inference results on brightness and chrominance.
[0091] The neural network-based video processing device provided in the embodiment of the present application can implement the various processes implemented in the method embodiment of Figure 1 and achieve the same technical effects. To avoid repetition, it will not be described here.
[0092] As shown in Figure 3, an embodiment of the present application also provides an electronic device 300, including a processor 301 and a memory 302. The memory 302 stores a program or instruction that can be run on the processor 301. When the program or instruction is executed by the processor 301, the various steps of the above-mentioned neural network-based video processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0093] An embodiment of the present application also provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the various steps of the above-mentioned neural network-based video processing method embodiment.
[0094] The present application also provides an electronic device including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is configured to execute a program or instruction to implement the steps in the method embodiment shown in FIG1 . This device embodiment corresponds to the above-described method embodiment, and each implementation process and implementation method of the above-described method embodiment are applicable to this electronic device embodiment and can achieve the same technical effects.
[0095] The electronic device may be a terminal, or may be other devices other than a terminal, such as a server, a network attached storage (NAS), etc.
[0096] Among them, the terminal can be a mobile phone, tablet personal computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile Internet device (MID), augmented reality (AR), virtual reality (VR) equipment, mixed reality (MR) equipment, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipborne equipment, pedestrian user equipment (PUE), smart home (home appliances with wireless communication function, such as refrigerator, TV, washing machine or furniture, etc.), game console, personal computer (PC), ATM or self-service machine and other terminal-side devices. Wearable devices include: smart watches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among them, vehicle-mounted devices can also be called vehicle-mounted terminals, vehicle-mounted controllers, vehicle-mounted modules, vehicle-mounted components, vehicle-mounted chips, or vehicle-mounted units, etc. It should be noted that the specific type of terminal is not limited in the embodiments of this application.
[0097] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0098] Taking an electronic device as a terminal as an example, FIG4 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of the present application.
[0099] The terminal 400 includes but is not limited to: a radio frequency unit 401, a network module 402, an audio output unit 403, an input unit 404, a sensor 405, a display unit 406, a user input unit 407, an interface unit 408, a memory 409 and at least some of the components of the processor 410.
[0100] Those skilled in the art will appreciate that terminal 400 may further include a power source (e.g., a battery) to power various components. The power source may be logically connected to processor 410 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The terminal structure shown in FIG4 does not limit the terminal. The terminal may include more or fewer components than shown, or may combine certain components or arrange the components differently, which will not be described in detail here.
[0101] It should be understood that in the embodiment of the present application, the input unit 404 may include a graphics processing unit (GPU) 4041 and a microphone 4042. The graphics processor 4041 processes the image data of a static picture or video obtained by an image acquisition device (such as a camera) in a video acquisition mode or an image acquisition mode, or may process the obtained point cloud data. The display unit 406 may include a display panel 4061, which may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 407 includes a touch panel 4071 and at least one of other input devices 4072. The touch panel 4071 is also called a touch screen. The touch panel 4071 may include two parts: a touch detection device and a touch controller. Other input devices 4072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0102] In the embodiment of the present application, after receiving downlink data from a network-side device, the radio frequency unit 401 may transmit the data to the processor 410 for processing. Furthermore, the radio frequency unit 401 may send uplink data to the network-side device. Typically, the radio frequency unit 401 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and the like.
[0103] The memory 409 can be used to store software programs or instructions and various data. The memory 409 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 409 may include a volatile memory or a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 409 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0104] Processor 410 may include one or more processing units. Optionally, processor 410 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 410.
[0105] Among them, the processor 410 is used to obtain a first offset value and a second offset value, where the first offset value refers to the offset value of the parameter value of the luminance quantization parameter relative to the first quantization parameter value, and the second offset value refers to the offset value of the parameter value of the chrominance quantization parameter relative to the second quantization parameter value; determine the luminance quantization parameter according to the first offset value, and determine the chrominance quantization parameter according to the second offset value; and obtain input data of the neural network according to the luminance quantization parameter and the chrominance quantization parameter.
[0106] In an embodiment of the present application, a first offset value and a second offset value are obtained, wherein the first offset value refers to the offset value of the parameter value of the luminance quantization parameter relative to the first quantization parameter value, and the second offset value refers to the offset value of the parameter value of the chrominance quantization parameter relative to the second quantization parameter value; the luminance quantization parameter is determined based on the first offset value, and the chrominance quantization parameter is determined based on the second offset value; and the input data of the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter. In this solution, the input data of the neural network is obtained based on the luminance quantization parameter and the chrominance quantization parameter brightness, that is, luminance and chrominance no longer share the same quantization parameter in the neural network, but each corresponds to a quantization parameter. The luminance quantization parameter corresponding to luminance enables the neural network to effectively control the luminance output quality of the output image, and the chrominance quantization parameter corresponding to chrominance enables the neural network to effectively control the chrominance output quality of the output image, thereby better balancing the neural network's reasoning results on brightness and chrominance.
[0107] It can be understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described here.
[0108] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned neural network-based video processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0109] The processor is the processor in the terminal described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as ROM, RAM, a magnetic disk, or an optical disk. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0110] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned neural network-based video processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0111] It should be understood that the chip mentioned in the embodiments of the present application may include a system-level chip (also referred to as a system chip, a chip system or a system-on-chip chip), and may also include an independent display chip, etc.
[0112] An embodiment of the present application further provides a computer program / program product, which is stored in a storage medium and is executed by at least one processor to implement the various processes of the above-mentioned neural network-based video processing method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0113] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of a computer software product plus a necessary general-purpose hardware platform, or of course, by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes a number of instructions for enabling a terminal or network-side device to execute the methods described in each embodiment of the present application.
[0115] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms of implementation methods without departing from the purpose of this application and the scope of protection of the claims. These implementation methods are all within the protection of this application.
Claims
1. A video processing method based on a neural network, performed by an electronic device, the method comprising: Obtaining a first offset value and a second offset value, wherein the first offset value refers to an offset value of a parameter value of a luminance quantization parameter relative to a first quantization parameter value, and the second offset value refers to an offset value of a parameter value of a chrominance quantization parameter relative to a second quantization parameter value; Determine a luminance quantization parameter according to the first offset value, and determine a chrominance quantization parameter according to the second offset value; According to the brightness quantization parameter and the chrominance quantization parameter, input data of the neural network is obtained.
2. The method according to claim 1, wherein: The obtaining of the first offset value and the second offset value comprises: Get the brightness quantization parameter index and the chrominance quantization parameter index; According to the correspondence between the quantization parameter index and the offset value, a first offset value corresponding to the brightness quantization parameter index and a second offset value corresponding to the chrominance quantization parameter index are obtained.
3. The method according to claim 2, wherein: The obtaining, according to the correspondence between the quantization parameter index and the offset value, a first offset value corresponding to the luminance quantization parameter index and a second offset value corresponding to the chrominance quantization parameter index comprises: According to the correspondence between the brightness quantization parameter index and the brightness offset value, obtaining a first offset value corresponding to the brightness quantization parameter index; According to the correspondence between the chromaticity quantization parameter index and the chromaticity offset value, a second offset value corresponding to the chromaticity quantization parameter index is obtained.
4. The method according to any one of claims 1 to 3, wherein: The first quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a chrominance quantization parameter; Alternatively, the second quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a brightness quantization parameter.
5. The method according to any one of claims 1 to 4, wherein: The brightness quantization parameter is a sequence-level quantization parameter, or a frame-level quantization parameter, or a block-level quantization parameter; Alternatively, the chrominance quantization parameter is a sequence-level quantization parameter, or a frame-level quantization parameter, or a block-level quantization parameter.
6. The method according to any one of claims 1 to 5, wherein: The input data also includes: a reconstructed image block; or, the input data also includes a predicted image block, a boundary strength, coding block mode information, at least one of a sequence-level quantization parameter and a frame-level quantization parameter, and a reconstructed image block; the method also includes: Perform filtering processing based on the neural network and the input data to obtain a filtered image block; Alternatively, super-resolution reconstruction is performed based on the neural network and the input data to obtain a reconstructed image block, and the resolution of the reconstructed image block is higher than the resolution of the reconstructed image block.
7. A video processing device based on a neural network, comprising: A first acquisition module, configured to acquire a first offset value and a second offset value, wherein the first offset value refers to an offset value of a parameter value of a luminance quantization parameter relative to a first quantization parameter value, and the second offset value refers to an offset value of a parameter value of a chrominance quantization parameter relative to a second quantization parameter value; a determination module, configured to determine a luminance quantization parameter according to the first offset value, and to determine a chrominance quantization parameter according to the second offset value; The second acquisition module is used to obtain input data of the neural network according to the brightness quantization parameter and the chrominance quantization parameter.
8. The device according to claim 7, wherein: The first acquisition module includes: A first acquisition submodule is used to acquire a brightness quantization parameter index and a chrominance quantization parameter index; The second acquisition submodule is used to acquire a first offset value corresponding to the brightness quantization parameter index and a second offset value corresponding to the chrominance quantization parameter index according to a correspondence between the quantization parameter index and the offset value.
9. The device according to claim 8, wherein: The second acquisition submodule includes: A first acquiring unit, configured to acquire a first offset value corresponding to the brightness quantization parameter index according to a correspondence between the brightness quantization parameter index and the brightness offset value; The second acquisition unit is used to acquire a second offset value corresponding to the chromaticity quantization parameter index according to the correspondence between the chromaticity quantization parameter index and the chromaticity offset value.
10. The device according to any one of claims 7 to 9, wherein: The first quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a chrominance quantization parameter; Alternatively, the second quantization parameter value is a parameter value of a sequence-level quantization parameter of an image block, or a parameter value of a frame-level quantization parameter of an image block, or a parameter value of a brightness quantization parameter.
11. The device according to any one of claims 7 to 10, wherein: The brightness quantization parameter is a sequence-level quantization parameter, or a frame-level quantization parameter, or a block-level quantization parameter; Alternatively, the chrominance quantization parameter is a sequence-level quantization parameter, or a frame-level quantization parameter, or a block-level quantization parameter.
12. The device according to any one of claims 7 to 11, wherein: The input data further includes: a reconstructed image block; or the input data further includes a predicted image block, a boundary strength, coding block mode information, at least one of a sequence-level quantization parameter and a frame-level quantization parameter, and a reconstructed image block; the device further includes: The third acquisition module is used to perform filtering processing based on the neural network and the input data to obtain a filtered image block; or to perform super-resolution reconstruction processing based on the neural network and the input data to obtain a reconstructed image block, and the resolution of the reconstructed image block is higher than the resolution of the reconstructed image block.
13. An electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the neural network-based video processing method as described in any one of claims 1 to 6 are implemented.
14. A readable storage medium storing a program or instruction, wherein the program or instruction, when executed by a processor, implements the steps of the neural network-based video processing method according to any one of claims 1 to 6.
15. A chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps of the neural network-based video processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Video frame filtering method, encoding and decoding method, codec and storage medium
CN114157869A
Image encoding / decoding method and apparatus using adaptive color conversion, and method for transmitting bitstream
CN114902666A
Method Of Manufacturing A Redox Flow Battery Hybrid Electrode Using Ultrasonic Welding
KR1020230111723A
Image filter device, image decoding device, and image coding device
WO2019087905A1
Encoding and decoding methods, code stream, encoder, decoder, and storage medium
WO2022227062A1