Image processing method, processing equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2026-04-17
AI Technical Summary
Existing video encoding and decoding methods fail to fully consider the interdependence between images, resulting in the encoder failing to achieve optimal compression performance.
By encoding or decoding the image blocks according to the weight coefficients of the image blocks in the image to be processed, the encoding bit resource allocation is optimized to improve the compression performance of the image encoder.
It realizes that under the current encoding standard that supports larger basic encoding unit sizes, the compression performance of the image encoder is improved and the allocation of encoded bit resources is optimized.
Smart Images

Figure CN121890073A_ABST
Abstract
Description
Image processing method, processing device and storage medium Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, processing device and storage medium. Background Art
[0002] Throughout the development of image coding and decoding technology, improvements to various codec standards have been made in an effort to improve image coding and decoding efficiency from different perspectives. Optimizing the allocation of coding bit resources to enhance image encoder compression performance is also a hot topic in current research.
[0003] During the process of conceiving and implementing this application, the inventors discovered at least the following problem: In some implementations, conventional video encoding and decoding methods perform independent rate-distortion optimization processes on each image. However, because these methods fail to consider the interdependencies between images, the resulting encoding method determined by this rate-distortion optimization approach prevents the encoder from achieving optimal compression performance. Therefore, it is necessary to propose an image encoding and decoding method that considers inter-image dependencies to improve and enhance compression performance during video encoding and decoding.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Technical Solutions
[0005] In response to the above technical problems, the present application provides an image processing method, a processing device and a storage medium, which can improve the compression performance of an image encoder under some current coding standards that support larger basic coding unit sizes.
[0006] The present application provides an image processing method, which can be applied to a processing device (such as an intelligent terminal or a server), comprising the following steps:
[0007] S1: Encoding or decoding the image block according to the weight coefficient of the image block in the image to be processed.
[0008] Optionally, step S1 includes at least one of the following:
[0009] Determine or generate a prediction result according to a weight coefficient of an image block in the image to be processed, and perform encoding or decoding processing on the image block according to the prediction result;
[0010] At least one first parameter is determined according to the weight coefficient, and the image block is encoded according to the first parameter.
[0011] Optionally, the weight coefficient is determined or generated by at least one of the following methods:
[0012] Determining or generating at least one third parameter of the first image block in the image to be processed according to at least one second parameter of the first image block, and determining or generating a weight coefficient according to the third parameter;
[0013] At least one third parameter of the first image region is determined or generated according to at least one second parameter of the first image region, and a weight coefficient is determined or generated according to the third parameter.
[0014] Optionally, a method for determining or generating the second parameter includes at least one of the following:
[0015] The first method is to pre-encode the image to be processed, and determine or generate the second parameter based on the motion compensation prediction error and reconstruction distortion in the pre-encoding process;
[0016] Second method: determining or generating a second parameter according to the motion compensation prediction error, the reconstruction distortion, and the first image group in which the image to be processed is located;
[0017] A third approach is to determine or generate the second parameter based on the motion compensation prediction error, the reconstruction distortion, and the next group of pictures of the first group of pictures.
[0018] Optionally, the method further comprises at least one of the following:
[0019] The second manner further includes: determining or generating the number of motion compensation prediction errors and / or reconstruction distortions according to the first image number of the first image group and the image sequence number of the image to be processed;
[0020] The third approach further includes: determining or generating the number of motion-compensated prediction errors and / or reconstruction distortions according to the number of pictures in the next group of pictures.
[0021] Optionally, the pre-encoding the image to be processed includes at least one of the following:
[0022] Performing a first predictive coding process on the image to be processed according to at least one first parameter;
[0023] Record, copy or restore encoder status.
[0024] Optionally, at least one of the following is also included:
[0025] Determining or generating a first parameter according to an image sequence number of an image to be processed;
[0026] An initial Lagrange multiplier and an initial quantization parameter are determined or generated according to the image to be processed and the first image group in which the image to be processed is located.
[0027] Optionally, the method further comprises at least one of the following:
[0028] The image to be processed is a first type of image;
[0029] Determining weight coefficients of images in the image blocks to be processed based on the image blocks in the second category of images in the first image group;
[0030] In a case where the image to be processed is not a third type of image, determining or generating a Lagrange multiplier and a quantization parameter according to the weight coefficient;
[0031] According to the weight coefficient and the initial Lagrangian multiplier, a Lagrangian multiplier and a quantization parameter are determined or generated.
[0032] The present application also provides a processing device, comprising: a memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of any of the above-mentioned image processing methods are implemented.
[0033] The present application also provides a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-mentioned image processing methods.
[0034] As described above, the image processing method of the present application can be applied to a processing device, including: encoding or decoding image blocks according to weight coefficients of image blocks in the image to be processed. The technical solution of the present application can improve the compression performance of the image encoder. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without inventive work.
[0036] FIG1 is a schematic diagram of the hardware structure of an intelligent terminal for implementing various embodiments of the present application;
[0037] FIG2 is a diagram of a communication network system architecture provided by an embodiment of the present application;
[0038] FIG3 is a schematic diagram of the development process of video coding standards provided in an embodiment of the present application;
[0039] FIG4 is a flowchart showing the steps of an image processing method according to the first embodiment;
[0040] FIG5 is a schematic diagram of a simplified distortion time-domain propagation relationship in low-delay coding according to an embodiment of the present invention;
[0041] FIG6 is a schematic diagram of a simplified time-domain propagation relationship of distortion in random access coding according to an embodiment of the present invention;
[0042] FIG7 is a schematic diagram showing a default quantization parameter QP setting according to a fourth embodiment;
[0043] FIG8 is a schematic diagram showing another default quantization parameter QP setting according to the fourth embodiment;
[0044] FIG9 is a schematic diagram of a performance comparison experiment under an encoder configuration according to the fourth embodiment;
[0045] FIG10 is a schematic diagram of a performance comparison experiment under another encoder configuration according to the fourth embodiment;
[0046] FIG11 is a schematic diagram showing a comparison of rate-distortion curves under a coding configuration according to the fourth embodiment;
[0047] FIG12 is a schematic diagram showing a comparison of rate-distortion curves under another coding configuration according to the fourth embodiment.
[0048] The purpose of this application, its features, and advantages will be further described in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and the accompanying text are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of this application to those skilled in the art by reference to specific embodiments.
[0049] Implementation Methods of the Application
[0050] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0051] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined based on their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0052] It should be understood that although the terms "first," "second," "third," etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if," as used herein, may be interpreted as "upon," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the recited features, steps, operations, elements, components, items, types, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used herein, may be interpreted as inclusive, meaning any one or any combination. For example, “comprising at least one of the following: A, B, C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”; and for another example, “A, B or C” or “A, B and / or C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”. An exception to this definition will occur only when a combination of elements, functions, steps or operations are inherently mutually exclusive in some manner.
[0053] It should be understood that, although the various steps in the flowcharts in the embodiments of the present application are shown in sequence as indicated by the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0054] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0055] It should be noted that in this article, step codes such as S1 are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the order. When implementing the step, those skilled in the art may execute other steps before executing S1, etc., but these should all be within the scope of protection of this application.
[0056] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0057] In the subsequent description, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate the description of the present application and have no specific meaning. Therefore, "module," "component," or "unit" may be used interchangeably.
[0058] The processing device in this application may be a smart terminal or a server. Optionally, the smart terminal may be implemented in various forms. For example, the smart terminal described in this application may include smart terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0059] The subsequent description will be made by taking a mobile terminal as an example. It will be understood by those skilled in the art that, in addition to components specifically used for mobile purposes, the configuration according to the embodiments of the present application can also be applied to fixed-type terminals.
[0060] Please refer to Figure 1, which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 may include components such as an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111. Those skilled in the art will understand that the mobile terminal structure shown in Figure 1 does not limit the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0061] The following is a detailed introduction to the various components of the mobile terminal in conjunction with Figure 1:
[0062] The RF unit 101 can be used to send and receive information or receive signals during calls. Specifically, it receives downlink information from the base station and transmits it to the processor 110 for processing. It also transmits uplink data to the base station. Typically, the RF unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and more. Furthermore, the RF unit 101 can communicate with the network and other devices via wireless communication. The above-mentioned wireless communications can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G and 6G, etc.
[0063] WiFi is a short-range wireless transmission technology. A mobile terminal, through WiFi module 102, enables users to send and receive emails, browse web pages, and access streaming media, providing wireless broadband Internet access. Although FIG1 illustrates WiFi module 102, it is understood that it is not a required component of the mobile terminal and can be omitted as needed without altering the essence of the invention.
[0064] The audio output unit 103 can convert audio data received by the RF unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a talk mode, a recording mode, a voice recognition mode, a broadcast reception mode, or the like. Furthermore, the audio output unit 103 can also provide audio output related to a specific function performed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, or the like.
[0065] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos captured by an image capture device (e.g., a camera) in video capture mode or image capture mode. The processed image frames may be displayed on the display unit 106. The image frames processed by the GPU 1041 may be stored in the memory 109 (or other storage medium) or transmitted via the RF unit 101 or the WiFi module 102. The microphone 1042 may receive sound (audio data) in operating modes such as a phone call mode, a recording mode, and a voice recognition mode, and may process such sound into audio data. In the phone call mode, the processed audio (voice) data may be converted into a format that can be transmitted to a mobile communication base station via the RF unit 101. The microphone 1042 may implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0066] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.
[0067] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0068] The user input unit 107 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile terminal. Optionally, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or near the touch panel 1071) and drive the corresponding connection device according to a pre-set program. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch direction and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 110. It can also receive commands sent by the processor 110 and execute them. In addition, the touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may further include other input devices 1072. Optionally, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, etc., and the specifics are not limited here.
[0069] Optionally, the touch panel 1071 may overlay the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. The processor 110 then provides a corresponding visual output on the display panel 1061 based on the type of touch event. Although in FIG1 , the touch panel 1071 and the display panel 1061 are shown as two separate components to implement the input and output functions of the mobile terminal, in some embodiments, the touch panel 1071 and the display panel 1061 may be integrated to implement the input and output functions of the mobile terminal, which is not limited to this specific embodiment.
[0070] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the mobile terminal 100 or may be used to transmit data between the mobile terminal 100 and an external device.
[0071] Memory 109 can be used to store software programs and various data. Memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, memory 109 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0072] Processor 110 is the control center of the mobile terminal, connecting all components of the mobile terminal using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 109 and accessing data stored in memory 109, it executes various functions of the mobile terminal and processes data, thereby providing overall monitoring of the mobile terminal. Processor 110 may include one or more processing units; preferably, processor 110 may integrate an application processor and a modem processor. Optionally, the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 110.
[0073] The mobile terminal 100 may also include a power supply 111 (such as a battery) for supplying power to various components. Preferably, the power supply 111 may be logically connected to the processor 110 through a power management system, thereby enabling the power management system to manage functions such as charging, discharging, and power consumption.
[0074] Although not shown in FIG. 1 , the mobile terminal 100 may further include a Bluetooth module, etc., which will not be described in detail here.
[0075] To facilitate understanding of the embodiments of the present application, the communication network system on which the mobile terminal of the present application is based is described below.
[0076] Please refer to Figure 2, which is a communication network system architecture diagram provided in an embodiment of the present application. The communication network system is an LTE system of universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203 and an operator's IP service 204, which are connected in sequence.
[0077] Optionally, UE201 may be the above-mentioned terminal 100, which will not be described in detail here.
[0078] E-UTRAN 202 includes eNodeB 2021 and other eNodeBs 2022 . Optionally, eNodeB 2021 may be connected to other eNodeBs 2022 via a backhaul (eg, an X2 interface). eNodeB 2021 is connected to EPC 203 , and eNodeB 2021 may provide access from UE 201 to EPC 203 .
[0079] EPC 203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gate Way) 2034, a PGW (PDN Gate Way) 2035, and a PCRF (Policy and Charging Rules Function) 2036. Optionally, MME 2031 is a control node that processes signaling between UE 201 and EPC 203, providing bearer and connection management. HSS 2032 provides registers for managing functions such as the Home Location Register (not shown) and stores user-specific information such as service features and data rates. All user data can be sent through SGW2034, PGW2035 can provide IP address allocation and other functions for UE 201, PCRF2036 is the policy and charging control policy decision point for service data flow and IP bearer resources, and it selects and provides available policy and charging control decisions for the policy and charging execution function unit (not shown in the figure).
[0080] The IP service 204 may include the Internet, an intranet, an IMS (IP Multimedia Subsystem), or other IP services.
[0081] Although the above introduction takes the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but can also be applied to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., which are not limited here.
[0082] Based on the above-mentioned mobile terminal hardware structure and communication network system, various embodiments of the present application are proposed.
[0083] This application proposes an image processing method that can encode or decode image blocks in the image to be processed based on their weight coefficients. This can improve the compression performance of image encoders under some current coding standards that support larger basic coding unit sizes.
[0084] To facilitate understanding, the following is an explanation of the professional terms that may be involved in this application.
[0085] (1) Video Coding Standards
[0086] With the development of digital video coding and decoding technology, digital video applications have expanded to various fields, including television broadcasting, digital movies, distance education, telemedicine, video surveillance, video conversations, and streaming media transmission. To ensure interoperability between codec products from different manufacturers, a series of video coding standards have been developed.
[0087] The International Telecommunication Union-Telecommunication standardization sector (ITU-T) and the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) were the first two international organizations dedicated to the development of video coding standards. ITU-T initially developed H.261 and H.263, while ISO / IEC developed international video coding standards such as MPEG-1 and MPEG-4 Visual. Subsequently, the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), H.265 / MPEG-H Part 2 High Efficiency Video Coding (HEVC), and H.266 / MPEG-I Part 3 Versatile Video Coding (VVC) standards.
[0088] Major video coding standardization organizations include the Audio Video Coding Standard (AVS) Working Group, established in June 2002 by the Science and Technology Department of the former Ministry of Information Industry of China. This group has developed three generations of AVS standards, each with independent intellectual property rights. Furthermore, Google's open-source codecs VP8, VP9, and VP10, as well as AV1, developed by the Alliance for Open Media (AOM), a joint venture between Google, Intel, Microsoft, Cisco, and other major technology companies, have created more options for multimedia experiences. Figure 3 illustrates the development of video coding standards developed by various organizations, with VVC, AVS3, and AV1 being the most recently released.
[0089] Video coding is a form of lossy source compression, the ultimate goal of which is to minimize distortion in the encoded video within a limited bitrate constraint. Video signals typically contain significant temporal, spatial, and visual redundancy. To achieve high compression ratios, the video coding standard H.261 developed by the International Telecommunication Union (ITU-T) and the more recent standards, such as VVC, AVS3, and AV1, all employ a hybrid video coding framework that incorporates multiple compression tools, including prediction, transform, quantization, and entropy coding. However, the introduction of new technologies has typically resulted in a 200% difference in compression efficiency between successive generations of coding standards. During the encoding process, the input video signal is divided into many small basic coding units (BCUs), such as macroblocks (MBs) in AVC and coding tree units (CTUs) in HEVC and VVC. The encoder then uses rate-distortion optimization (RDO) to select an optimal set of coding parameters for these BCUs, balancing reconstruction distortion with the number of bits required for encoding. For a specific video coding standard, the output bitstream syntax structure and the encoding tools that can be used in the encoder have been determined, but developers can still flexibly adopt different optimization and encoder control strategies to further improve video compression performance.
[0090] (2) Rate-distortion optimization
[0091] Rate-distortion theory is the theoretical foundation of video coding. It defines the maximum compression limit for a given source while allowing for a certain amount of distortion. Accordingly, rate-distortion optimization is a critical technology in video encoders, operating throughout the entire video coding system. Rate-distortion optimization has made significant contributions to compression performance improvements in each generation of video coding standards. However, the independent rate-distortion optimization methods widely used in video encoders are far from achieving optimal compression performance. Due to the use of tools such as intra-frame prediction, inter-frame prediction, motion vector prediction, and context-based entropy coding, coding mode decisions for coding units (CUs) are interdependent. This dependency ultimately manifests as rate-distortion dependency between CUs and coded frames. Existing research has shown that temporal rate-distortion dependency in video coding is much stronger than spatial rate-distortion dependency. Temporal rate-distortion dependency manifests itself as the coding quality of a reference reconstructed pixel affecting the rate-distortion bounds achievable by subsequent coded frames or CUs that reference it. In recent years, several temporally dependent rate-distortion optimization methods have emerged for AVC and HEVC encoders. For example: adaptive quantization parameter cascading at the frame level, adaptive Lagrange multiplier selection at the coding tree unit level, etc.
[0092] (3) Source distortion time domain propagation model
[0093] In order to estimate the impact of the reconstruction distortion of a coding block on the subsequent coding process, it is first necessary to know which coding blocks in the subsequent frames directly or indirectly refer to the current block in inter-frame prediction coding. However, motion compensation in actual video coding uses backward motion search to find the best matching block from the reconstructed reference frame, which may cause the coding distortion of a block to spread to multiple coding blocks in the next frame and indirectly continue to spread to subsequent frames. In addition, HEVC adopts a hierarchical coding structure and a multi-reference frame selection strategy, resulting in a more complex distortion propagation relationship. Based on this, a simplified coding block time domain propagation chain can be established using integer pixel forward motion search based on the H.264 coding reference structure, and a source distortion time domain propagation model is proposed to estimate the distortion propagation factor of the coding block. Subsequently, the source distortion time domain propagation model is extended to the HEVC low-latency coding structure.
[0094] First embodiment
[0095] 4 is a flow chart of an image processing method according to a first embodiment. The image processing method according to the embodiment of the present application can be applied to a processing device, including the following steps:
[0096] S1: Encoding or decoding the image block according to the weight coefficient of the image block in the image to be processed.
[0097] Optionally, in this embodiment, the processing device is taken as a smart terminal for example.
[0098] Optionally, the intelligent terminal can determine or obtain the Lagrange multiplier and quantization parameter of the image block in the image to be processed based on the weight coefficient of the image block; then, the intelligent terminal can encode or decode the image to be processed based on the Lagrange multiplier and quantization parameter. In this way, the values of the Lagrange multiplier and quantization parameter can be flexibly adjusted by the weight coefficient. Since the value of the weight coefficient is determined based on the dependency relationship between the image blocks of different images. Therefore, rate-distortion optimization processing based on the temporal correlation of video images can be achieved. In one embodiment, the weight coefficient is obtained by formula (1):
[0099] ψ j is the weight coefficient of the jth coding tree unit in the frame to be coded, L j is the number of pixel blocks of a predetermined size contained in the jth coding tree unit. Optionally, the predetermined size may be 32×32, 16×16, etc.
[0100] In another embodiment, the weight coefficient can be obtained by formula (2):
[0101] ψ' j is the weight coefficient of the j'th coding unit in the frame to be coded, ω i is the weight factor of the i-th pixel block of predetermined size, L j is the number of pixel blocks of a predetermined size included in the jth coding unit. Optionally, the predetermined size may be 16x16, 8x8, etc.
[0102] Optionally, the image to be processed may be an image frame to be processed that is being encoded or decoded by the smart terminal, such as a frame to be encoded or a frame to be decoded. Optionally, when the image to be processed is an image frame to be processed that is being encoded by the smart terminal, the image blocks in the image to be processed may be coding tree units or coding units in the image frame to be processed. Coding units are obtained by dividing coding tree units.
[0103] Optionally, step S1 may include at least one of the following:
[0104] Determine or generate a prediction result according to a weight coefficient of an image block in the image to be processed, and perform encoding or decoding processing on the image block according to the prediction result;
[0105] At least one first parameter is determined according to the weight coefficient, and the image block is encoded according to the first parameter.
[0106] Optionally, when encoding or decoding an image block in a to-be-processed image based on a weight coefficient of the image block, the smart terminal may first perform prediction processing on the image block based on the weight coefficient of the image block in the to-be-processed image to determine or generate a prediction result. The smart terminal then encodes or decodes the image block based on the prediction result.
[0107] In one embodiment, the weight coefficient is determined, and the Lagrange multiplier and quantization parameter are determined to perform rate-distortion optimization processing. According to the optimal prediction mode determined in the rate-distortion optimization process, predictive coding is performed to determine or generate a prediction result for the image block. The optimal prediction mode can be intra-frame prediction processing and / or inter-frame prediction processing. Afterwards, the residual corresponding to the image block of the prediction result is determined or obtained based on the prediction result, and the residual is encoded or decoded. That is, in the process of encoding or decoding the image block according to the weight coefficient of the image block in the image to be processed, the smart terminal can also first determine or obtain the Lagrange multiplier and quantization parameter of the image block according to the weight coefficient of the image block in the image to be processed; then, the smart terminal determines the Lagrange multiplier and / or quantization parameter as the first parameter, so as to encode or decode the image block according to the first parameter.
[0108] Optionally, the weight coefficient is determined or generated by at least one of the following methods:
[0109] Determining or generating at least one third parameter of the first image block in the image to be processed according to at least one second parameter of the first image block, and determining or generating a weight coefficient according to the third parameter;
[0110] At least one third parameter of the first image region is determined or generated according to at least one second parameter of the first image region, and a weight coefficient is determined or generated according to the third parameter.
[0111] Optionally, the intelligent terminal can determine or generate at least one third parameter of the first image block in the above-mentioned image to be processed based on at least one second parameter of the first image block, and then determine or generate the weight coefficient of the image block in the above-mentioned image to be processed based on the third parameter.
[0112] Optionally, the smart terminal can also determine or generate at least one third parameter of the first image block based on at least one second parameter of the first image area, and then determine or generate the weight coefficient of the image block in the above-mentioned image to be processed based on the third parameter.
[0113] Optionally, the above-mentioned first image block can be an image block that is being encoded or decoded in the image to be processed, or an image block that is similar / adjacent. Optionally, the first image block can also be an image block in the previous or next frame of the image to be processed. Optionally, the above-mentioned first image area can be an image block that is being encoded or decoded in the image to be processed, or an image area that is similar / adjacent. Optionally, the first image area can also be an image area in the previous or next frame of the image to be processed. Optionally, the first image area can also be an image area in a predicted image block obtained by the intelligent terminal through predictive processing of the image to be processed.
[0114] Optionally, the second parameter may be a time domain distortion impact factor or a distortion propagation factor of an image block in the image to be processed during encoding or decoding of the image by the smart terminal. Optionally, the third parameter may be a weight factor of an image block in the image to be processed during encoding or decoding of the image by the smart terminal.
[0115] The above-mentioned time domain distortion impact factor can be the time domain distortion impact factor or distortion propagation factor of a non-key frame, the time domain distortion impact factor or distortion propagation factor of a key frame, or the time domain distortion impact factor or distortion propagation factor of a frame corresponding to other division methods.
[0116] The non-key frames mentioned in the above embodiments refer to other frames in the Group of Picture (GOP) except the key frames. In one embodiment, for the Low Delay (LD) configuration, the key frame is the last frame of the picture group. For the Random Access (RA) configuration, the key frame is a frame whose picture order number (POC) is equal to an integer multiple of the number of pictures in the picture group. For example, if the GOP size is 16, the picture order number of the key frame is 16, 32, etc. The key frame can be an I frame, but not only an I frame. That is, the key frame is the frame of the lowest time domain layer in the time domain hierarchy. For example, a frame with a time domain layer of 0 and a time domain ID temporalID of 0.
[0117] Please refer to Figure 5, which is a schematic diagram of the simplified distortion time domain propagation relationship in low-delay coding according to an embodiment of the invention. In Figure 5, the GOP size is 4, the key frames are images with image sequence numbers 0, 4, 8, etc., and the remaining images are non-key frames. Under the low-delay configuration, the last frame of the previous image group (GOP) is a key frame because the last frame is directly referenced by all frames of the next image group. Please refer to Figure 6, which is a schematic diagram of the simplified distortion time domain propagation relationship in random access coding according to an embodiment of the invention. In Figure 6, the GOP size is 8, the key frames are images with image sequence numbers 0, 8, etc., and the remaining images are non-key frames.
[0118] In implementation mode 1, the weight factors of the image blocks in the image to be processed are shown in equations (3) to (4):
[0119] φ i is the time domain distortion influencing factor or distortion propagation factor, M is the total number of pixel blocks of predetermined size in a frame, is M 1 / φ i The arithmetic mean of i is the weight factor of the i-th pixel block of predetermined size.
[0120] In the second embodiment, the weight factors of the image blocks in the image to be processed are shown in equations (5) to (7):
[0121] φ i is the time domain distortion impact factor or distortion propagation factor, M block is the total number of pixel blocks of predetermined size in a frame, ω average It's M block Initial weight factor The arithmetic mean of i is the weight factor of the i-th pixel block of predetermined size.
[0122] In implementation mode 3, the weight factor of the image block in the image to be processed is shown in formula (8):
[0123] φ i is the time domain distortion impact factor or distortion propagation factor of the j-th coding tree unit.
[0124] Optionally, implementation mode 1 is a method for determining the weight factor in a low-latency configuration, and implementation modes 2 to 3 are methods for determining the weight factor in a random access coding configuration. It should be noted that the pixel block size of the predetermined pixel block size used to determine the weight factor in implementation mode 1 is larger than the pixel block size of the predetermined pixel block size used to determine the weight factor in implementation modes 2 to 3.
[0125] In this embodiment, the image processing method determines the weight factor of the image block according to the time domain distortion influence factor of the image block in the image to be processed during the encoding or decoding process of the image to be processed by the intelligent terminal, and then determines the weight coefficient of the image block based on the weight factor. In this way, the intelligent terminal can determine or obtain the Lagrange multiplier and quantization parameter of the image block according to the weight coefficient of the image block, so as to encode or decode the image block according to the first parameter. That is, in this embodiment, the image processing method adaptively adjusts the Lagrange multiplier and quantization parameter of the coding tree unit by estimating the time domain distortion influence factor of the image block in the video encoding process, thereby optimizing the allocation of coding bit resources and ultimately improving the compression performance of encoding or decoding of the video image.
[0126] Second embodiment
[0127] In this embodiment, the image processing method is still described with the intelligent terminal as the execution subject. Based on the above first embodiment, the determination or generation method of the above second parameter includes at least one of the following:
[0128] The first method is to pre-encode the image to be processed, and determine or generate the second parameter based on the motion compensation prediction error and reconstruction distortion in the pre-encoding process;
[0129] Optionally, the above pre-encoding process is an integer pixel search, and a sub-pixel search is not used in the pre-encoding process. Therefore, the entire pre-encoding process algorithm is relatively simple.
[0130] Optionally, the intelligent terminal may perform pre-coding processing on the image to be processed, thereby determining or generating the second parameter according to the motion compensation prediction error and reconstruction distortion in the pre-coding process on the image to be processed.
[0131] Optionally, the time domain distortion influencing factor of the non-key frame is as shown in formula (9):
[0132] φ i is the distortion factor of the pixel block of the i-th predetermined size, D i and E iare respectively the reconstruction distortion D and the motion compensation prediction MCP error E of the i-th pixel block of predetermined size. Both the reconstruction distortion D and the motion compensation prediction MCP error are measured using the mean square error (MSE).
[0133] Optionally, the time domain distortion influencing factor of the key frame is shown in formula (10):
[0134] Is the number of frames contained in the next GOP. If the next GOP is not the last GOP of the video sequence, then If the next GOP is the last GOP in the video sequence, then Equal to the number of frames remaining.
[0135] Next, the summation terms in equations (9) and (10) are explained. In equation (9), when n=1, D i / E i The 0th order term is 1, and in formula (10), when n = 0, D i / E i The 0th order term of is also 1. This 0th order term is called the distortion impact factor of the current frame.
[0136] In formula (9), when n takes a value other than 1, the corresponding n-1th order term is the propagation factor. That is, in formula (9), n = 2, 3, 4, ..., N GOP -rPOC corresponding to the n-1 term. Similarly, in formula (10), when n takes values other than 0, the corresponding n-1 term is the propagation factor. Among them, in formula (9), when n = 2, the corresponding 1-term is the direct propagation factor. Similarly, in formula (10), when n = 1, the corresponding 1-term is the direct propagation factor. In formula (9), when n takes values other than 1 and 2, the corresponding n-1-term is the indirect propagation factor. That is, in formula (9), n = 3, 4, ...., N GOP -rPOC corresponding to the n-1th order term. Similarly, in formula (10), when n takes values other than 0 and 1, the corresponding nth order term is the indirect propagation factor. That is, in formula (10), n = 2, 3, ...., The corresponding n-order term. That is, in these embodiments, the time domain distortion impact factor includes the distortion impact factor, direct propagation factor, and indirect propagation factor of the current frame. Compared to the time domain distortion impact factor, the time domain distortion propagation factor includes the direct propagation factor and the indirect propagation factor. In one embodiment, equation (10) represents the distortion impact of a key frame on the next GOP.
[0137] Optionally, under the random access coding configuration, the distortion propagation factor is as shown in formula (11):
[0138] φ i is the distortion propagation factor of the i-th pixel block of predetermined size.
[0139] Now let’s explain formula (11). In formula (11), only D i / E i Therefore, Equation (11) does not include the direct propagation factor and only includes D i / E i The indirect propagation factor of the first-order term of . This is determined based on the characteristics of distortion propagation under the random access coding configuration.
[0140] It's important to note that the distortion impact factor adjusts the Lagrange multiplier λ of the coding tree unit. A large impact factor requires a smaller Lagrange multiplier λ; a small impact factor requires a larger one. The distortion impact factor characterizes the impact of distortion on the overall coding process, while the distortion propagation factor refers to the distortion caused by the time domain propagation chain.
[0141] In one embodiment, when the size of the predetermined-size pixel block is 16×16, the reconstruction distortion and motion-compensated prediction (MCP) error of the predetermined-size pixel block are as shown in equations (12) and (13):
[0142] Among them, f c (x,y), f(x,y) and are the coded reconstructed pixels, original pixels, and reference pixels in the current frame, respectively. x and y are the coordinates of a pixel block of a predetermined size.
[0143] Optionally, the intelligent terminal, as an encoder, uses the same formula or algorithm to determine or generate the second parameter for all even-numbered frames to be processed under the random access coding configuration. That is, for the above-mentioned even-numbered frames, regardless of whether they are key frames or non-key frames, the distortion propagation factor is determined or obtained using the method shown in formula (11). In another embodiment, the distortion impact factor under the random access coding configuration can also be determined, and the weight factor of the coding tree unit can be determined based on the distortion impact factor.
[0144] Second method: determining or generating a second parameter according to a motion compensation prediction error, reconstruction distortion, and a frame type of the image to be processed;
[0145] Optionally, the second parameter is determined or generated according to the motion compensation prediction error, the reconstruction distortion, and the corresponding frame type in the first image group where the image to be processed is located. The frame type includes two frame types: key frame and non-key frame.
[0146] Optionally, in the process of determining or generating the second parameter, the intelligent terminal further determines a method for obtaining the second parameter based on the corresponding frame type in the first image group in which the image to be processed is located. In one embodiment, it is determined whether the image to be processed is a key frame in the first image group in which the image to be processed is located. When the image to be processed is a key frame, the second parameter is determined using the first parameter determination method. In another embodiment, it is determined whether the image to be processed is a non-key frame in the first image group in which the image to be processed is located. When the image to be processed is a key frame, the second parameter is determined using the second parameter determination method. The second parameter is a time domain distortion influence factor.
[0147] Optionally, when the image to be processed is a non-key frame, the first parameter determination method may be as shown in formula (10). For example, the number of motion compensation prediction errors and / or reconstruction distortions may be determined or generated based on the number of the first image in the first image group and the image sequence number of the image to be processed.
[0148] Optionally, when the intelligent terminal uses a calculation formula based on the motion compensation prediction error, reconstruction distortion and the number of first images in the first image group to determine or generate the second parameter, it also determines or generates the number of motion compensation prediction errors and / or reconstruction distortions currently needed to calculate the second parameter based on the number of first images in the first image group and the image sequence number of the image to be processed.
[0149] Optionally, when the image to be processed is a key frame, the second parameter determination method may be the method shown in formula (9). For example, the second parameter is determined or generated based on the motion compensation prediction error, the reconstruction distortion, and the next image group of the first image group.
[0150] Optionally, in the process of determining or generating the second parameter, when the image to be processed is a key frame in the first image group in which the image to be processed is located, the smart terminal uses a calculation formula determined based on the motion compensation prediction error, reconstruction distortion and the number of images in the next image group of the first image group to determine or generate the second parameter.
[0151] Optionally, the above-mentioned second parameter determination method further includes:
[0152] Depending on the number of pictures of the next group of pictures, the number of motion compensated prediction errors and / or reconstruction distortions is determined or generated.
[0153] Optionally, when the intelligent terminal uses a calculation formula based on the motion compensation prediction error, reconstruction distortion and the number of images in the next image group of the first image group to determine or generate the second parameter, it also determines or generates the number of motion compensation prediction errors and / or reconstruction distortions currently used to calculate the second parameter based on the number of images in the next image group.
[0154] Optionally, assuming that the above-mentioned image to be processed is a frame to be encoded in the process of encoding a video by the smart terminal, the smart terminal, as a video encoder, can perform the encoding process for the frame to be encoded according to steps 1 to 9 shown below under the low-latency encoding configuration. That is:
[0155] Step 1: Read the frame to be encoded, and according to the default settings of the video encoder in the low-latency layered coding configuration, determine the initial first parameters: the initial quantization parameter QP and the initial Lagrange multiplier λ from the rPOC of the frame to be encoded. Optionally, the rPOC of the frame to be encoded is the position of the frame to be encoded in the current group of pictures GOP, recorded as the relative picture order count (rPOC), and the value of rPOC ranges from 1 to N GOP , N GOP The number of frames contained in a group of pictures GOP, rPOC = N GOP The corresponding frame to be encoded is the key frame.
[0156] Step 2: Based on the initial first parameters: the initial quantization parameter QP and the Lagrange multiplier λ, perform integer-pixel inter-frame prediction encoding on the frame to be encoded, and record the pixel-level motion compensation prediction (MCP) error E and reconstruction distortion D during the encoding process. Optionally, the image block in the frame to be encoded, namely the coding tree unit (CTU), includes at least one 32×32 pixel block.
[0157] Step 3: Calculate the distortion impact factor of the 32×32 pixel block in the frame to be encoded based on the motion compensation prediction MCP error E and the reconstruction distortion D.
[0158] Optionally, when the frame to be encoded is a non-key frame, the following formula (14) is used to calculate the distortion influence factor φ: i :
[0159] Optionally, φ i is the distortion factor of the i-th 32×32 pixel block, D i and E iare the reconstruction distortion D and motion compensation prediction MCP error E of the i-th 32×32 pixel block, respectively. Both the reconstruction distortion D and the motion compensation prediction MCP error are measured using the mean square error (MSE).
[0160] Optionally, when the frame to be encoded is a key frame, the following formula (15) is used to calculate the second parameter, the distortion influence factor φ i :
[0161] Optionally, Is the number of frames contained in the next GOP. If the next GOP is not the last GOP of the video sequence, then If the next GOP is the last GOP in the video sequence, then Equal to the number of frames remaining.
[0162] Step 4: According to the distortion influence factor φ i , the third parameter of the 32×32 pixel block, the weight factor, is calculated by equations (16) and (17).
[0163] Optionally, M is the total number of 32×32 pixel blocks in a frame, is M 1 / φ i The arithmetic mean of i is the weight factor for the i-th 32×32 pixel block.
[0164] Step 5: Based on the third parameter of the 32×32 pixel block - weight factor ω i , the weight coefficient of each coding tree unit CTU in the frame to be coded is calculated by formula (18):
[0165] Optionally, ψ j is the weight coefficient of the jth coding tree unit CTU in the frame to be coded, L j is the number of 32×32 pixel blocks contained in the j-th coding tree unit CTU.
[0166] Step 6: According to the weight coefficient ψ of the coding tree unit CTU j and the initial Lagrange multiplier λ determined in step 1 above, the first parameter of each coding tree unit CTU in the frame to be coded is calculated by equations (19) and (20): Lagrange multiplier and quantization parameter: λ j =ψ j ·λ(19) QP j =F(λ j )(20)
[0167] Optionally, λ j is the Lagrange multiplier of the jth CTU in the frame to be coded, QP j is the quantization parameter of the jth CTU in the frame to be encoded, and F(g) represents a function operator. Optionally, the expression and parameter selection of the function F(g) may be slightly different for different video encoders.
[0168] Step 7: Use the first parameter of the coding tree unit CTU: Lagrange multiplier λ j and quantization parameter QP j Optionally, if the currently encoded frame is not the last frame of the video sequence, return to step 1 to continue processing the next frame; and if the currently encoded frame is the last frame of the video sequence, end encoding.
[0169] Optionally, assuming that the above-mentioned image to be processed is a frame to be encoded in the process of encoding a video by the smart terminal, the smart terminal, as a video encoder, can perform the encoding process for the frame to be encoded according to steps 1 to 9 shown below under the random access encoding configuration. That is:
[0170] Step 1: Read the frame to be encoded. According to the default settings of the video encoder under the random access coding configuration, the initial first parameters are determined by the group of pictures (GOP) size and the time domain level of the frame to be encoded: the initial quantization parameter (QP) and the Lagrange multiplier λ.
[0171] Step 2: If the picture sequence number (POC) of the frame to be encoded is an odd number or the frame to be encoded is an I frame, execute step 9; otherwise, execute step 3.
[0172] Step 3: Using the initial first parameters: the initial quantization parameter QP and the Lagrange multiplier λ, perform integer-pixel inter-frame prediction coding (pre-coding) on the 16×16 pixel block of the frame to be coded, and record the pixel-level motion compensation prediction (MCP) error E and reconstruction distortion D during the coding process. Optionally, the image block (coding tree unit) CTU in the frame to be coded includes at least one 16×16 pixel block.
[0173] Step 4: Based on the motion compensation prediction MCP error E and the reconstruction distortion D, the second parameter of the 16×16 pixel block in the frame to be encoded, the distortion influence factor φ, is calculated by the following formula (21): i Alternatively, regardless of whether the frame to be encoded is a key frame or not, the same formula is used to calculate the distortion influence factor φ i .
[0174] Optionally, φ i is the distortion factor of the i-th 16×16 pixel block, D i and Ei are the reconstruction distortion D and motion compensation prediction MCP error E of the i-th 16×16 pixel block, and both the reconstruction distortion D and the motion compensation prediction MCP error E are measured using MSE.
[0175] Step 5: According to the second parameter - distortion influence factor φ i , the third parameter of the 16×16 pixel block, the weight factor ω, is calculated by equations (22) and (23). i :
[0176] Optionally, M is the total number of 16×16 pixel blocks in a frame to be encoded, is M 1 / φ i The arithmetic mean of i is the weight factor for the i-th 16×16 pixel block.
[0177] Step 6: Based on the third parameter of the 16×16 pixel block - weight factor ω i , the weight coefficient of each coding tree unit CTU in the frame to be coded is calculated by formula (24):
[0178] Optionally, ψ j is the weight coefficient of the jth CTU in the frame to be encoded, L j is the number of 16×16 pixel blocks contained in the j-th CTU.
[0179] Step 7: According to the weight coefficient ψ of the coding tree unit CTU j And the initial Lagrange multiplier λ determined in step 1 above, the first parameter of each coding tree unit CTU in the frame to be coded is calculated by equations (25) and (26): Lagrange multiplier and quantization parameter: λ j =ψ j ·λ (25) QP j =F(λ j ) (26)
[0180] Optionally, λ j is the Lagrange multiplier of the jth coding tree unit CTU in the frame to be coded, QP j is the quantization parameter of the j-th coding tree unit CTU in the frame to be coded, and F(g) represents a function operator. Optionally, for different video encoders, the expression and parameter selection of the function F(g) are slightly different.
[0181] Step 8: Use the first parameter of the coding tree unit CTU: Lagrange multiplier λ j and quantization parameter QP j, encode the frame to be encoded. Then, if the encoded frame is not the last frame to be encoded, return to step 1 to continue processing the next frame, otherwise end the encoding.
[0182] Step 9: Using the initial first parameters obtained in Step 1 above: the initial quantization parameter QP and the initial Lagrange multiplier λ, directly encode the frame to be encoded. If the currently encoded frame is not the last frame to be encoded, return to Step 1 to continue processing the next frame; if the currently encoded frame is the last frame to be encoded, then terminate encoding.
[0183] In this embodiment, the image processing method can estimate the time domain distortion influence factor or distortion propagation factor of different pixel blocks in the encoding process in different ways during the encoding or decoding process of the image to be processed through the intelligent terminal, thereby adaptively adjusting the Lagrange multiplier and quantization parameter of the coding tree unit to optimize the coding bit resource allocation, and ultimately improve the compression performance of encoding or decoding of the video image.
[0184] Third embodiment
[0185] In this embodiment, the image processing method is still described with the intelligent terminal as the execution subject. Based on any of the above embodiments, the above pre-encoding of the image to be processed includes at least one of the following:
[0186] Performing a first predictive coding process on the image to be processed according to at least one first parameter;
[0187] Record, copy or restore encoder status.
[0188] Optionally, during encoding of an image to be processed, in order to adaptively adjust the Lagrange multipliers and quantization parameters of image blocks in the image to be processed, the intelligent terminal may determine or generate an initial first parameter, then perform a first predictive encoding process on the image to be processed based on the initial first parameter, and obtain the motion compensation prediction error and reconstruction distortion used to calculate and determine the above-mentioned second parameter by recording, copying, or restoring the encoder state. In one embodiment, the encoder state includes at least the values of syntax elements in the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), and Adaptation Parameter Set (APS), the value of the quantization parameter, etc.
[0189] Optionally, the image processing method further includes at least one of the following:
[0190] An initial value of the first parameter is determined or generated according to the image sequence number of the image to be processed.
[0191] Optionally, the intelligent terminal may determine or generate the above-mentioned initial first parameter according to the image sequence number of the image to be processed in the first image group, and then perform the first predictive coding process on the image to be processed based on the initial first parameter.
[0192] Optionally, the processing device may further determine an initial quantization parameter and an initial Lagrangian multiplier based on GOP attributes of the first GOP in which the image to be processed is located (e.g., GOP size and the temporal level of the image to be processed). The determined initial quantization parameter and initial Lagrangian multiplier are used as initial first parameters (i.e., initial values of the first parameters) to perform first predictive coding processing on the image to be processed based on the initial first parameters when it is determined that the image to be processed is not an I-frame and the image sequence number of the image to be processed is an even number.
[0193] Optionally, the smart terminal, acting as a video encoder, can determine initial first parameters (initial quantization parameter QP and initial Lagrange multiplier λ) from the rPOC of the frame to be encoded, based on the default settings of the low-latency layered coding configuration. The smart terminal can then perform integer-pixel inter-frame predictive encoding of the 32×32 pixel block of the processed image based on the initial first parameters (initial quantization parameter QP and initial Lagrange multiplier λ), and record pixel-level motion compensation prediction error E and reconstruction distortion D during the encoding process.
[0194] Optionally, the intelligent terminal, acting as a video encoder, may determine, in accordance with the default settings of the random access coding configuration, the initial first parameters: initial quantization parameter QP and Lagrange multiplier λ, based on the size of the first group of pictures (GOP) in which the image to be processed resides and the temporal level of the frame to be encoded. Subsequently, when the image sequence number (POC) of the current image to be processed in the first group of pictures is an even number and the image to be processed is not an I-frame in the first group of pictures, the intelligent terminal uses the initial first parameters: initial quantization parameter QP and Lagrange multiplier λ to perform integer-pixel inter-frame prediction encoding of a 16×16 pixel block on the image to be processed, and records the pixel-level motion compensation prediction (MCP) error E and reconstruction distortion D during the encoding process.
[0195] Optionally, the image processing method further includes:
[0196] The image to be processed is a first type of image.
[0197] Optionally, the first type of image is an image in a first image group in which the image to be processed is located, and has an even image sequence number. Under the random access coding configuration, the intelligent terminal uses the initial first parameter to perform the first predictive coding process on the image to be processed only when the image sequence number of the image to be processed in the first image group is an even number and the image to be processed is not an I-frame in the first image group.
[0198] Optionally, the image processing method further includes:
[0199] The weight coefficients of the images in the image blocks to be processed are determined according to the image blocks in the second category of images in the first image group.
[0200] Optionally, the second type of image can be a key frame image or a non-key frame image in the first image group in which the image to be processed is located. Optionally, when the second type of image is a key frame image, the intelligent terminal can use a calculation formula determined based on the motion compensation prediction error, reconstruction distortion and the number of first images in the first image group to determine or generate the second parameter, and further calculate based on the second parameter to determine the third parameter - the weight factor, and then determine or derive the weight coefficient of each image block in the image to be processed based on the weight factor. Optionally, when the second type of image is a non-key frame image, the intelligent terminal uses a calculation formula determined based on the motion compensation prediction error, reconstruction distortion and the number of images in the next image group of the first image group to determine or generate the second parameter, and further calculate based on the second parameter to determine the third parameter - the weight factor, and then determine or derive the weight coefficient of each image block in the image to be processed based on the weight factor.
[0201] Optionally, the image processing method further includes:
[0202] When the image to be processed is not an image of the third category, the Lagrange multiplier and the quantization parameter are determined or generated according to the weight coefficient.
[0203] Optionally, the third type of image is an I-frame in the first image group in which the image to be processed is located. Under the random access coding configuration, the intelligent terminal uses the initial first parameters to perform the first predictive coding process on the image to be processed only when the image sequence number of the image to be processed in the first image group is an even number and the image to be processed is not an I-frame in the first image group.
[0204] Optionally, the image processing method further includes:
[0205] According to the weight coefficient and the initial Lagrangian multiplier, a Lagrangian multiplier and a quantization parameter are determined or generated.
[0206] Optionally, after determining or generating the weight coefficient of the image block in the above-mentioned image to be processed, the smart terminal can perform calculations based on the weight coefficient and the above-mentioned initial first parameter to determine the Lagrange multiplier and quantization parameter currently used for encoding the image to be processed.
[0207] Optionally, the intelligent terminal determines the weight coefficient ψ of each image block in the image to be processed. j After that, we can further calculate the weight coefficient ψ j and the above initial Lagrange multiplier λ, the first parameter of each image block in the image to be processed is calculated by the following formulas (27) and (28): Lagrange multiplier and quantization parameter: λ j =ψ j ·λ(27) QP j =F(λ j )(28)
[0208] Optionally, λ j is the Lagrange multiplier of the jth CTU in the frame to be coded, QP j is the quantization parameter of the jth CTU in the frame to be encoded, and F(g) represents a function operator. Optionally, the expression and parameter selection of the function F(g) may be slightly different for different video encoders.
[0209] Optionally, the smart terminal may use the first parameter of the image block: Lagrange multiplier λ j and quantization parameter QP j The image to be processed is encoded. Optionally, if the current image to be processed is not the last frame of the video sequence, the intelligent terminal may return to the step of determining the initial first parameter to continue processing the next frame of the image to be processed. If the current image to be processed is the last frame of the video sequence, the intelligent terminal may end encoding the entire video image after completing encoding for the image to be processed.
[0210] In this embodiment, the image processing method flexibly estimates the time domain distortion influence factor or distortion propagation factor of different pixel blocks in the encoding process based on the different encoding configurations of the smart terminal as an encoder during the encoding or decoding process of the image to be processed by the smart terminal, thereby adaptively adjusting the Lagrange multiplier and quantization parameter of the coding tree unit to optimize the allocation of coding bit resources, and ultimately improve the compression performance of encoding or decoding of video images.
[0211] Fourth embodiment
[0212] In this embodiment, the image processing method is still described with the smart terminal as the execution subject. Based on any of the above embodiments, the image processing method is applied to the VTM21.2 low-delay encoding configuration. The smart terminal can use the computer development environment Visual Studio 2019 and implement it based on the reference software VTM21.2 of the general video coding standard. The smart terminal uses two encoding configurations as an encoder: low-delay B-frame (LDB) and low-delay P-frame (LDP). Optionally, in the low-delay encoding LDP configuration of the VTM encoder, a GOP size is 8, the last frame in the GOP is a key frame, and the encoding frame uses a hierarchical QP setting. Optionally, when the encoder input QP is 32, the default QP setting of the 8 coded frames in the GOP under the LDB configuration is shown in Figure 7, and the default QP setting of the 8 coded frames in the image group GOP under the low-delay encoding LDP configuration is shown in Figure 8.
[0213] Optionally, the process of encoding the image to be processed, that is, the frame to be encoded, by the intelligent terminal may be as follows:
[0214] Step 1: Read the frame to be encoded and calculate the relative picture sequence number (rPOC) of the current frame to be encoded based on the picture sequence number (POC) of the frame to be encoded. Optionally, when rPOC = 8, the corresponding frame to be encoded is a key frame, while other rPOC values correspond to non-key frames. Optionally, the intelligent terminal, acting as an encoder, can use rPOC to determine the initial first parameters of the frame to be encoded: the initial quantization parameter (QP) and the Lagrange multiplier (λ), according to the default settings of VTM21.2.
[0215] Step 2: The intelligent terminal performs integer pixel inter-frame prediction encoding of the 32×32 pixel block on the frame to be encoded based on the initial first parameters obtained in step 1: the initial quantization parameter QP and the Lagrange multiplier λ, and records the pixel-level motion compensation prediction MCP error E and reconstruction distortion D during the encoding process.
[0216] Step 3: The intelligent terminal calculates the distortion impact factor φ of the 32×32 pixel block in the frame to be encoded based on the motion compensation prediction MCP error E and reconstruction distortion D obtained in step 2 i .
[0217] Optionally, when the frame to be encoded is a non-key frame, formula (29) is used:
[0218] Among them, φ i is the distortion influencing factor of the i-th 32×32 pixel block, Di and Ei are the reconstruction distortion and MCP error of the i-th 32×32 pixel block, respectively. The reconstruction distortion and MCP error are both measured using MSE.
[0219] Optionally, when the frame to be encoded is a key frame, formula (30) is used:
[0220] in, Is the number of frames contained in the next GOP. If the next GOP is not the last GOP of the video sequence, then If the next GOP is the last GOP of the video sequence, then Equal to the number of frames remaining.
[0221] Step 4: The intelligent terminal calculates the distortion impact factor φ obtained in step 3 i , the weight factor of the 32×32 pixel block is calculated by equations (31) and (32),
[0222] Optionally, M is the total number of 32×32 pixel blocks in a frame, is M 1 / φ i The arithmetic mean of i is the weight factor for the i-th 32×32 pixel block.
[0223] Step 5: The smart terminal obtains the 32×32 pixel block weight factor ω according to step 4 i , the weight coefficient of each image block in the frame to be coded, namely the coding tree unit CTU, is calculated by formula (33).
[0224] Optionally, ψ j is the weight coefficient of the j-th CTU in the frame to be encoded, and Lj is the number of 32×32 pixel blocks contained in the j-th CTU.
[0225] Step 6: The smart terminal calculates the CTU weight coefficient ψ obtained in step 5 j And the initial Lagrangian multiplier λ determined in step 1 above, the Lagrangian multiplier and quantization parameter of each CTU in the frame to be encoded are calculated by equations (34) and (35): λ j =ψ j ·λ (34) QP j =4.3281×log(λ j )+2.1829 (35)
[0226] Optionally, λ j is the Lagrange multiplier of the jth CTU in the frame to be coded, QP j is the quantization parameter of the jth CTU in the frame to be encoded.
[0227] Step 7: The intelligent terminal uses the Lagrange multiplier λ of the coding tree unit CTU obtained in step 6 j and quantization parameter QP j Encode the frame to be encoded. If the currently encoded frame is not the last frame in the video sequence, return to step 1 to continue processing the next frame; if the currently encoded frame is the last frame in the video sequence, end encoding.
[0228] In this embodiment, the image processing method is integrated into VTM21.2. The experimental test uses all 16 standard dynamic range (SDR) videos in Class B, Class C, Class D, and Class E recommended by the VVC Common Test Conditions (CTC). Each video is input at four bitrate points with QP of 22, 27, 32, and 37 according to the CTC test. The image processing method in this embodiment saves bitrate by ( Delta bit-rate (BD-rate) metric, BD-rate represents the percentage of bit rate savings achieved by the test method relative to the benchmark encoder at the same objective quality. A positive value indicates a loss in compression performance, while a negative value indicates an improvement in compression performance.
[0229] Alternatively, as shown in the performance comparison experiment under the LDB encoder configuration in FIG9 and the performance comparison experiment under the LDP encoder configuration in FIG10, it can be seen that the image processing method in this embodiment achieves bit rate savings relative to the VVC benchmark encoder. That is, as shown in the experimental data in FIG9 and FIG10, compared to the VVC benchmark encoder, the image processing method in this embodiment achieves an average bit rate savings of 4.27% under the LDB encoder configuration and an average bit rate savings of 3.64% under the LDP encoder configuration. In addition, the experimental data also shows that the encoding time of the image processing method in this embodiment does not increase significantly compared to the VVC benchmark encoder.
[0230] Optionally, as shown in Figures 11 and 12, Figure 11 is a schematic diagram comparing the rate-distortion curves of the image processing method and the VVC benchmark encoder for the test video Basketball Drill under the LDB encoder configuration in this embodiment, and Figure 12 is a schematic diagram comparing the rate-distortion curves of the image processing method and the VVC benchmark encoder for the test video Arena Of Valor under the LDP encoder configuration in this embodiment. As shown in Figures 11 and 12, the horizontal axis bitrate is the output bit rate, in kbps; the vertical axis Y-PSNR is the peak signal-to-noise ratio of the video luminance component, in dB. Optionally, as shown in Figures 11 and 12, at the same output bit rate, the video encoding quality of the image processing method in this embodiment is significantly better than the video encoding quality of the VVC benchmark encoder. The four rate-distortion points on the curve show that at the same input QP, the output bit rate of the present technical solution is slightly lower than the output bit rate of the VVC benchmark encoder.
[0231] Fifth embodiment
[0232] In this embodiment, the image processing method is still described with the smart terminal as the execution subject. Based on any of the above embodiments, the image processing method is applied to the VTM21.2 random access coding configuration. The smart terminal can use the computer development environment Visual Studio 2019 and implement it based on the reference software VTM21.2 of the general video coding standard. The smart terminal uses the random access (Radom Access, RA) coding configuration as an encoder. Under the RA coding configuration, the GOP size of the VTM encoder is set to 32 or 16 by default. In this embodiment, the GOP size is selected as 16 and the I frame insertion interval is 32. In other embodiments, the interval for inserting I frames is an integer multiple of the number of images in the image group.
[0233] Optionally, the process of encoding the image to be processed (ie, the frame to be encoded) by the intelligent terminal may be as follows:
[0234] Step 1: Initialize the encoder. Get the input quantization parameter QPinput and GOP size NGOP from the encoder configuration file. Set the POC of the starting frame of the video sequence to 0. Then, encode the starting frame of the video sequence using the encoder's default encoding method. Optionally, the input quantization parameter QPinput can be pre-set before the encoding process.
[0235] Step 2: Read the frames to be encoded in a group of pictures GOP, and execute step 3 to process each frame according to the encoding order under the RA encoding configuration.
[0236] Step 3: If the POC of the current frame is an odd number (i.e., POC%2==1) or the current frame is an I frame, the current frame is encoded using the encoder's default method; otherwise, step 4 is executed to pre-encode the current frame. Due to the characteristics of the reference structure / distortion propagation between frames of the random access coding configuration, odd frames will not become reference frames. Since different coding tree units of odd frames are optimized using different Lagrange multipliers, there will be no benefit to subsequent encoding and decoding. Therefore, there is no need to optimize odd frames that are not reference frames. Based on this, only odd frames are encoded using the encoder's default method, and there is no need to perform pre-encoding processing on odd frames. In this embodiment, the propagation factor of the I frame cannot be obtained, so there is no need to perform pre-encoding operations on the I frame. Since this embodiment only performs simplified pre-encoding processing on a portion of the frames, the complexity of the entire encoding process is basically not affected.
[0237] Step 4: Using the encoder's default method, calculate the frame-level quantization parameter QP for the current frame, and then calculate the Lagrange multiplier λ from the QP. Use the above QP and λ to perform integer pixel inter-frame prediction encoding on the 16×16 pixel block of the current frame, and record the pixel-level MCP error E and reconstruction distortion D during the encoding process. Optionally, the calculation formulas for QP and λ are: QP = QP input +ΔQP L (36) λ=W L 2 (QP-12) / 3 (37)
[0238] Optionally, ΔQP L Determined by the input quantization parameter and the temporal level of the current frame, W L is the weight coefficient preset by the encoder.
[0239] The encoding process for integer-pixel inter-frame prediction coding of a 16x16 pixel block in the current frame using the frame-level quantization parameter QP and Lagrange multiplier λ is as follows: For a 16x16 pixel block, a corresponding Lagrange multiplier λ is calculated using a given frame-level quantization parameter QP. The coding distortion and bit rate of the 16x16 pixel block are then calculated for various prediction modes. The prediction mode with the minimum rate-distortion cost is selected for prediction processing and subsequent encoding steps.
[0240] Step 5: Based on the MCP error E and reconstruction distortion D obtained in step 4, the distortion propagation factor of the 16×16 pixel block in the current frame is calculated by equation (38):
[0241] In Equation (38), the same optimization strategy is used for even-numbered frames at different time-domain layers. That is, only direct propagation factors are considered for even-numbered frames at different time-domain layers, without considering indirect propagation factors. The advantage of this method is that, since the distortion effect length is not considered, the calculation is relatively simple and good rate-distortion optimization performance can be achieved.
[0242] It should be noted that in the embodiments of the technical solution of the present application, the block size of the predetermined-size pixel block used in the precoding process under the random access coding configuration is larger than the block size of the predetermined-size pixel block used in the prediction process under the low-latency coding configuration. For example, the block size of the predetermined-size pixel block used in the precoding process under the random access coding configuration is 16x16, while the block size of the predetermined-size pixel block used in the prediction process under the low-latency coding configuration is 32x32. The reason for using different pixel block sizes for different coding configurations is that under the low-latency coding configuration, the temporal distance (the interval between the POC of the current coded frame and the POC of the reference frame) is closer, making it easier to find a matching block with a larger predetermined block size. However, under the random access coding configuration, the temporal distance is farther, increasing the probability of not finding the best matching block. Therefore, a pixel block with a smaller prediction block size is selected to increase the probability of finding a matching block. On the other hand, the propagation factor estimated using a smaller size is more accurate.
[0243] Optionally, under a random access coding configuration, the block size of the pixel block used in the preprocessing process remains unchanged, without further partitioning the pixel block used in the preprocessing process. That is, under a specific coding configuration (e.g., a random access coding configuration or a low-latency coding configuration), the block size of the pixel block is fixed, without further quadtree, ternary, or binary tree partitioning of the pixel block.
[0244] Therefore, the entire pre-encoding process is relatively simple and has little impact on the encoding time of the entire video encoding.
[0245] Optionally, φ i is the distortion propagation factor of the i-th 16×16 pixel block, Di and Ei are the reconstruction distortion and MCP error of the i-th 16×16 pixel block, respectively, and the calculation formula is:
[0246] Among them, f c (x,y), f(x,y) and are the coded reconstructed pixels, original pixels, and reference pixels in the current frame, respectively, and x and y are the coordinates of the 16×16 pixel block.
[0247] Optionally, the above steps 2-5 may be replaced by the following steps 2'-5'.
[0248] Step 2': Read the frames to be encoded from a GOP (Group of Pictures), copy the current encoder state, and record it as encodeS. Then, perform steps 3' and 4' to pre-encode the current GOP. In this implementation, due to the use of pre-encoding, the encoder state changes, potentially affecting subsequent normal encoding. To restore the encoder to its original state for the actual encoding process, the encoder state must be copied.
[0249] Step 3': If the POC of the coded frame is an even number (ie, POC%2==0), encoding is performed using the encoder's default method; otherwise, step 4' is executed.
[0250] Step 4': Using the encoder's default method, calculate the frame-level quantization parameter QP of the encoded frame, and then calculate the Lagrange multiplier λ from the QP. Use the above QP and λ to perform integer pixel inter-frame prediction encoding of 16×16 pixel blocks for frames with an odd POC (i.e., POC%2==1), and record the pixel-level MCP error E and reconstruction distortion D of the encoded frame during the encoding process. Optionally, the calculation formula for QP and λ can be: QP=QP input +ΔQP L (36') λ=W L 2 (QP-12) / 3 (37')
[0251] Compared with step 3 and step 4 in the previous embodiment, step 3' and step 4' pre-code all images in the current group of images. In this way, the time domain spreading factor can be estimated more accurately in the random access coding configuration.
[0252] Optionally, ΔQP L Determined by the input quantization parameter and the temporal level of the coded frame, W L is the weight coefficient pre-set by the encoder. Optionally, during the encoding process of the intelligent terminal, the distortion propagation factor of the 16×16 pixel block in the POC+1 frame (POC+1 is an even number) is calculated based on the MCP error E and reconstruction distortion D recorded above.
[0253] Optionally, φ i is the distortion propagation factor of the i-th 16×16 pixel block in the POC+1 frame, and are the reconstruction distortion and MCP error of the i-th 16×16 pixel block in the POC frame, respectively. The calculation formula is:
[0254] Optionally, f c(x,y), f(x,y) and are the coded reconstructed pixels, original pixels and reference pixels in the POC frame, respectively, and x and y are the coordinates of the 16×16 pixel block.
[0255] In step 4', the propagation factor of the even-numbered frames is estimated using the coding information of the odd-numbered frames (e.g., the pixel-level MCP error E and reconstruction distortion D of the odd-numbered frames). This is done because the temporal distance between the coded frames and the reference frames in the even-numbered image sequence is too far, while the temporal distance between the coded frames and the reference frames in the odd-numbered image sequence is closer. Therefore, using the coding information of the odd-numbered frames instead of the coding information of the even-numbered frames to analyze the effect of distortion propagation is relatively accurate.
[0256] Step 5': After completing the pre-encoding of the last frame in the current GOP through steps 3' and 4', restore the encoder state to encodeS. It should be noted that by pre-encoding the GOP, the distortion propagation factors of the 16×16 pixel blocks in all frames in the current GOP whose POC is an even number have been obtained. Then, the current GOP is encoded for the second time. Optionally, if the POC of the encoded frame is an odd number (i.e., POC%2==1), the encoder default method is used for encoding; otherwise, step 6 is executed. In this embodiment, since the pre-encoding process is utilized, the state of the encoder will change, which may affect the subsequent normal encoding of the encoder. Therefore, after the pre-encoding process is completed, it is necessary to ensure that the parameters in the encoder during formal encoding are not affected by the previous pre-encoding process, and the state of the encoder needs to be restored to the state before the pre-encoding process.
[0257] Step 6: Distortion propagation factor φ obtained from precoding i , calculate the weight factor of the 16×16 pixel block using equations (41), (42) and (43):
[0258] Optionally, Mblock is the total number of 16×16 pixel blocks in a frame, and ωaverage is the number of Mblock The arithmetic mean of i is the weight factor for the i-th 16×16 pixel block.
[0259] Step 7: The 16×16 pixel block weight factor ω obtained in step 6 i , the weight coefficient of each coding tree unit CTU in the frame to be coded is calculated by formula (44),
[0260] Optionally, ψ jis the weight coefficient of the j-th coding tree unit CTU in the frame to be coded, and Lj is the number of 16×16 pixel blocks contained in the j-th coding tree unit CTU.
[0261] Optionally, the above steps 6-7 may be replaced by the following steps 6'-7'.
[0262] Step 6': Distortion propagation factor φ obtained from precoding i , calculate the weight factor of the 16×16 pixel block using formula (41'):
[0263] Step 7': The 16×16 pixel block weight factor ω obtained from step 6' i , the weight coefficient of each coding tree unit CTU in the frame to be coded is calculated by equations (42'), (43') and (44'):
[0264] Optionally, L j is the number of 16×16 pixel blocks contained in the jth coding tree unit (CTU), MCTU is the number of CTUs contained in a frame, ψ j is the weight coefficient of the j-th CTU in the frame to be encoded.
[0265] Step 8: CTU weight coefficient ψ obtained according to the previous steps j and the Lagrange multiplier λ determined in step 4 (or step 4'), the Lagrange multiplier and quantization parameter of each coding tree unit CTU in the frame to be coded are calculated by equations (45) and (46): λ j =ψ j ·λ (45) QP j =4.3281×log(λ j )+2.1829 (46)
[0266] Among them, λ j is the Lagrange multiplier of the jth coding tree unit CTU in the frame to be coded, QP j is the quantization parameter of the j-th coding tree unit CTU in the frame to be encoded.
[0267] Step 9: Use the CTU Lagrange multiplier λ obtained in step 8 j The frame to be encoded is encoded with the quantization parameter QPj.
[0268] Step 10: If the currently processed GOP is not the last GOP in the video sequence, return to step 2 to continue processing the next GOP; otherwise, end the encoding.
[0269] In this embodiment, the image processing method is integrated into VTM21.2. Experimental testing uses all 19 standard SDR videos in ClassA1, ClassA2, ClassB, ClassC, and ClassD recommended by CTC. Each video is input at four bitrate points with QP of 22, 27, 32, and 37 according to the CTC test. The bitrate savings of the image processing method in this embodiment compared to the VVC baseline encoder are measured using BD-rate.
[0270] An embodiment of the present application also provides a processing device (such as the above-mentioned smart terminal or server), including a memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of the image processing method in any of the above-mentioned embodiments are implemented.
[0271] An embodiment of the present application further provides a storage medium on which an image processing program is stored. When the image processing program is executed by a processor, the steps of the image processing method in any of the above embodiments are implemented.
[0272] In the embodiments of the smart terminal and storage medium provided in this application, all technical features of any of the above-mentioned image processing method embodiments may be included. The expanded and explained contents of the specification are basically the same as those of the embodiments of the above-mentioned methods and will not be repeated here.
[0273] An embodiment of the present application further provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the methods in the various possible implementation modes described above.
[0274] An embodiment of the present application also provides a chip, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that a device equipped with the chip executes the methods in the various possible implementation modes as described above.
[0275] It is understood that the above scenarios are merely examples and do not limit the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, those skilled in the art will appreciate that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.
[0276] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0277] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0278] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0279] In this application, the same or similar terminology, technical solutions and / or application scenario descriptions are generally only described in detail the first time they appear. When they appear again later, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, for the same or similar terminology, technical solutions and / or application scenario descriptions that are not described in detail later, you can refer to the previous relevant detailed descriptions.
[0280] In this application, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0281] The various technical features of the technical solution of this application can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0282] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as mentioned above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the method of each embodiment of the present application.
[0283] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a storage disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state storage disk Solid State Disk (SSD)).
[0284] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An image processing method, wherein: Includes steps: S1: Encoding or decoding the image block according to the weight coefficient of the image block in the image to be processed.
2. The method of claim 1, wherein: The step S1 includes at least one of the following: Determine or generate a prediction result according to a weight coefficient of an image block in the image to be processed, and perform encoding or decoding processing on the image block according to the prediction result; At least one first parameter is determined according to the weight coefficient, and the image block is encoded according to the first parameter.
3. The method of claim 2, wherein: The weight coefficient is determined or generated by at least one of the following methods: Determine or generate at least one third parameter of the first image block in the image to be processed according to at least one second parameter of the first image block, and determine or generate a weight coefficient according to the third parameter; At least one third parameter of the first image region is determined or generated according to at least one second parameter of the first image region, and a weight coefficient is determined or generated according to the third parameter.
4. The method of claim 3, wherein: The method for determining or generating the second parameter includes at least one of the following: The first method is to pre-encode the image to be processed, and determine or generate the second parameter according to the motion compensation prediction error and reconstruction distortion in the pre-encoding process; The second method is to determine or generate the second parameter according to the motion compensation prediction error, the reconstruction distortion, and the frame type of the image to be processed.
5. The method of claim 4, wherein: Also includes at least one of the following: The second manner further includes: determining or generating the number of motion compensation prediction errors and / or reconstruction distortions according to the first image number of the first image group and the image sequence number of the image to be processed; The third method further includes: determining or generating the number of motion compensation prediction errors and / or reconstruction distortions according to the number of pictures in the next picture group.
6. The method of claim 4, wherein: The pre-encoding of the image to be processed includes at least one of the following: According to at least one first parameter, performing a first predictive coding process on the image to be processed; Record, copy or restore encoder status.
7. The method of claim 6, wherein: Also includes at least one of the following: An initial value of the first parameter is determined or generated according to the image sequence number of the image to be processed.
8. The method according to any one of claims 1 to 7, wherein: Also includes at least one of the following: The image to be processed is a first type of image; Determining weight coefficients of image blocks in the image to be processed according to image blocks in the second type of image in the first image group; In the case where the image to be processed is not an image of the third category, determining or generating a Lagrange multiplier and a quantization parameter according to the weight coefficient; According to the weight coefficient and the initial Lagrangian multiplier, the Lagrangian multiplier and the quantization parameter are determined or generated.
9. A processing device, wherein: include: A memory and a processor, wherein an image processing program is stored in the memory, and when the image processing program is executed by the processor, the steps of the image processing method according to any one of claims 1 to 8 are implemented.
10. A storage medium, wherein: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 8 are implemented.