Encoding method, decoding method, and system, electronic device and storage medium

By extracting the frequency domain features of the source image from the AI ​​encoding and decoding system and encoding according to the quality level and matching the bit rate, the problems of versatility and flexibility in multi-bit rate scenarios are solved, achieving efficient encoding and saving storage resources.

WO2026129830A1PCT designated stage Publication Date: 2026-06-25CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
Filing Date
2025-10-16
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing AI encoding and decoding systems have poor versatility and flexibility when facing encoding and decoding scenarios with multiple bitrates, and deploying multiple model weights will increase storage resource consumption and computing power consumption.

Method used

By extracting image frequency features from at least two frequency domains of the source image, selecting features whose importance matches the quality level of the source image for encoding, and encoding according to the bit rate matched by the frequency features, multiple model weights are avoided.

Benefits of technology

It achieves adaptive selection of features to be encoded on different frequency components, reducing storage resource consumption and system complexity, while improving encoding efficiency and visual quality, and has strong versatility and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025128156_25062026_PF_FP_ABST
    Figure CN2025128156_25062026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are an encoding method, a decoding method, and a system, an electronic device and a storage medium. The encoding method comprises: extracting from a source image at least two image frequency features in a frequency domain; selecting as a feature to be encoded a feature, the degree of importance of which matches a quality level of the source image, from among the image frequency features, wherein the feature to be encoded refers to an essential feature representation required during encoding on the basis of the quality level; and at a bitrate matching the image frequency feature, encoding the feature to be encoded, so as to obtain a bitstream corresponding to the image frequency feature, wherein the bitrate is directly proportional to a frequency component corresponding to the image frequency feature. In the technical solutions of the embodiments of the present disclosure, features to be encoded can be adaptively selected for different frequency components of a source image on the basis of the quality of the source image, and for the different frequency components, different bitrates are used for encoding, thereby achieving both strong universality and good flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding methods, decoding methods, systems, electronic devices and storage media

[0001] This disclosure claims priority to Chinese Patent Application No. 202411884031.8, filed on December 19, 2024, entitled "Encoding Method, Decoding Method, System, Electronic Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of image processing technology, and in particular to an encoding method, a decoding method, a system, an electronic device, and a storage medium. Background Technology

[0003] Artificial intelligence (AI) encoding and decoding models encode and decode images through neural networks, and AI encoding and decoding models can be obtained through pre-training.

[0004] Generally, an AI codec model only supports encoding and decoding for a fixed bitrate. If it is applied to encoding and decoding scenarios with multiple bitrates, a common approach is to deploy an AI codec model and multiple model weights as an AI codec system. The AI ​​codec system changes the encoding and decoding bitrate of the AI ​​codec model by adjusting the weights used in the AI ​​codec model.

[0005] It is evident that conventional AI encoding and decoding systems have poor versatility and flexibility. Furthermore, deploying multiple model weights requires more storage resources and increases the complexity and computational cost of the AI ​​encoding and decoding system. Summary of the Invention

[0006] To overcome the problems existing in related technologies, this disclosure provides an encoding method, a decoding method, a system, an electronic device, and a storage medium.

[0007] According to a first aspect of the present disclosure, an encoding method is provided for application to an AI encoding model, the method comprising:

[0008] Extract image frequency features from at least two frequency domains of the source image;

[0009] The features selected from the image frequency features whose importance matches the quality level of the source image are used as the features to be encoded. The features to be encoded refer to the necessary feature representations required for encoding based on the quality level.

[0010] The feature to be encoded is encoded according to a bitrate that matches the image frequency feature to obtain a bitstream corresponding to the image frequency feature, wherein the bitrate is proportional to the frequency component corresponding to the image frequency feature.

[0011] According to a second aspect of the present disclosure, a decoding method is provided, applied to an AI decoding model, the method comprising:

[0012] At least two bitstreams are acquired, and the at least two bitstreams correspond one-to-one with at least two image frequency features in the frequency domain; each bitstream is obtained by encoding a feature whose importance in the corresponding image frequency feature matches the quality level of the source image, and the feature to be encoded refers to the necessary feature representation required for encoding based on the quality level;

[0013] For any bitstream, the bitstream is decoded according to a bitrate that matches the bitstream to obtain decoding features. The bitrate that matches the bitstream refers to the bitrate that matches the image frequency features corresponding to the bitstream. The bitrate is proportional to the frequency components corresponding to the image frequency features.

[0014] After obtaining at least two decoding features, the at least two decoding features are fused to obtain the reconstructed image of the source image, wherein the at least two decoding features correspond one-to-one with the at least two bitstreams.

[0015] According to a third aspect of the present disclosure, an encoding / decoding system is provided, the system comprising: an artificial intelligence (AI) encoding model and an AI decoding model, wherein...

[0016] The AI ​​encoding model is used to extract image frequency features in at least two frequency domains from the source image; select features among the image frequency features whose importance matches the quality level of the source image as features to be encoded, wherein the features to be encoded refer to the necessary feature representations required when encoding based on the quality level; encode the features to be encoded according to a bitrate that matches the image frequency features to obtain a bitstream corresponding to the image frequency features, wherein the bitrate is proportional to the frequency components corresponding to the image frequency features;

[0017] The AI ​​decoding model is used to decode any bitstream according to a bitrate that matches the bitstream to obtain decoding features. The bitrate that matches the bitstream refers to the bitrate that matches the image frequency features corresponding to the bitstream. The at least two decoding features obtained by the at least two decoding modules are fused to obtain the reconstructed image of the source image. The at least two decoding features correspond one-to-one with the at least two bitstreams.

[0018] According to a fourth aspect of the present disclosure, an encoding apparatus is provided, the apparatus comprising:

[0019] Extraction unit, used to extract image frequency features in at least two frequency domains of the source image;

[0020] The feature selection unit is used to select features among the image frequency features whose importance matches the quality level of the source image as features to be encoded. The features to be encoded refer to the necessary feature representations required for encoding based on the quality level.

[0021] An encoding unit is used to encode the feature to be encoded according to a bit rate that matches the image frequency feature, so as to obtain a bit stream corresponding to the image frequency feature, wherein the bit rate is proportional to the frequency component corresponding to the image frequency feature.

[0022] According to a fifth aspect of the present disclosure, a decoding apparatus is provided, the apparatus comprising:

[0023] An acquisition unit is used to acquire at least two bitstreams, wherein the at least two bitstreams correspond one-to-one with at least two frequency domain image frequency features; each bitstream is obtained by encoding a feature whose importance in the corresponding image frequency feature matches the quality level of the source image, and the feature to be encoded refers to the necessary feature representation required for encoding based on the quality level;

[0024] A decoding unit is used to decode any given bitstream according to a bitrate that matches the bitstream to obtain decoding features. The bitrate that matches the bitstream refers to the bitrate that matches the image frequency features corresponding to the bitstream. The bitrate is proportional to the frequency components corresponding to the image frequency features.

[0025] A fusion unit is used to fuse at least two decoded features after obtaining them to obtain a reconstructed image of the source image, wherein the at least two decoded features correspond one-to-one with the at least two bitstreams.

[0026] According to a sixth aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being executed by the processor to cause the electronic device to perform the method as described in the first or second aspect.

[0027] According to a seventh aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, the program being executed by a processor to implement the method as described in the first or second aspect.

[0028] According to an eighth aspect of the present disclosure, a computer program product is provided, including instructions that, when executed on a computer, cause the computer to perform the method as described in the first or second aspect.

[0029] The technical solutions provided in this disclosure can include the following beneficial effects:

[0030] In the encoding stage of the source image, at least two frequency domain features of the source image are first extracted. This allows for differentiated processing of different frequency components of the source image, thereby effectively adjusting the bitrate. Furthermore, based on the image frequency features, features whose importance matches the quality level of the source image are selected from each image frequency feature as the features to be encoded. Specifically, for any image frequency feature, the feature whose importance matches the quality level of the source image can refer to the necessary feature representation required for encoding based on the quality level. This reduces the amount of data encoded while maintaining high visual quality. Additionally, in this embodiment, the bitrate is proportional to the frequency domain corresponding to the image frequency feature. Thus, for any image frequency feature to be encoded, it can be encoded at a bitrate matching the image frequency feature to obtain the corresponding bitstream. This facilitates variable bitrate encoding based on different frequency domain components of the source image, improving encoding efficiency while preserving the details of the source image. As can be seen, the technical solution of this disclosure embodiment does not require the deployment of multiple model weights, thereby saving storage resources and reducing the complexity and computing power of the AI ​​encoding and decoding system. Moreover, the technical solution of this disclosure embodiment can adaptively select the features to be encoded for different frequency components of the source image according to the quality of the source image, and use different bit rates for encoding for different frequency components, which is not only highly versatile and flexible.

[0031] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the embodiments of this disclosure will be briefly described below. It should be understood that those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0033] Figure 1 is a schematic diagram of a typical data flow for image encoding and decoding provided in an embodiment of this disclosure;

[0034] Figure 2 is a schematic diagram of an exemplary system architecture of an AI encoding and decoding system provided in an embodiment of this disclosure;

[0035] Figure 3 is an exemplary method flowchart of an encoding method provided in an embodiment of this disclosure;

[0036] Figure 4 is a schematic diagram of an exemplary scenario of image frequency division provided in the embodiments of this disclosure;

[0037] Figure 5 is an exemplary scenario diagram illustrating the relationship between channel parameters and image quality provided in the embodiments of this disclosure;

[0038] Figure 6 is an exemplary method flowchart of a decoding method provided in an embodiment of this disclosure;

[0039] Figure 7 is a schematic diagram of the encoding and decoding data flow in any frequency domain provided in the embodiments of this disclosure;

[0040] Figure 8A is a schematic diagram of an exemplary composition of the encoding device provided in an embodiment of this disclosure;

[0041] Figure 8B is a schematic diagram of an exemplary component of the decoding device provided in an embodiment of this disclosure;

[0042] Figure 9 is an exemplary structural diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0043] The technical solutions of the embodiments of this disclosure will now be described with reference to the accompanying drawings.

[0044] The terminology used in the following embodiments of this disclosure is for the purpose of describing particular embodiments and is not intended to be a limitation of the technical solutions of this disclosure. As used in the specification and appended claims of this disclosure, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise.

[0045] It should also be understood that although the terms "first," "second," etc., may be used in the following embodiments to describe a class of objects, the objects are not limited to these terms. These terms are used to distinguish specific implementations of that class of objects.

[0046] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0047] The following describes the technical scenarios related to the embodiments of this disclosure.

[0048] This disclosure relates to AI-based image encoding and decoding technology. As shown in Figure 1, which illustrates a typical data flow diagram for image encoding and decoding, after acquiring a source image, the source image can be encoded to obtain a compressed bitstream. Then, the bitstream is decoded to obtain a reconstructed image of the source image. The encoding process may include quantization and entropy encoding, and may also include entropy decoding, inverse quantization, and reconstruction.

[0049] Those skilled in the art will understand that the image encoding and decoding process shown in Figure 1 is merely illustrative and does not limit the actual operation of image encoding and decoding. In practical implementations, the image encoding and decoding process may include more procedures than those shown in Figure 1, which are not limited here.

[0050] Quantization refers to reducing the complexity of image data by converting continuous values ​​into discrete values. The quantization process can include dimensionality transformation and determining quantization parameters (such as quantization step size and quantization coefficients). Dimensionality transformation involves converting the image from the spatial domain to the frequency domain, obtaining its frequency domain representation y. This can be achieved using methods such as Discrete Cosine Transform (DCT) or Discrete Fourier Transform (DFT). Furthermore, a quantization matrix can be selected or generated, defining the quantization step size P (quantization parameter) for different frequency coefficients. Then, by dividing the dimensionally transformed coefficients (such as DCT coefficients) by the corresponding quantization matrix value, a quantization factor is obtained. Normalizing the quantization factor approximates the original continuous coefficients as discrete values. A quantization algorithm can satisfy, for example, AdaQ() = Round( / (P*)).

[0051] Entropy coding refers to encoding discrete values ​​obtained from quantization based on probability distribution to reduce redundancy in image data. The entropy coding process may include determining the frequency of each pixel value in the image data, and employing entropy coding algorithms such as Huffman coding, arithmetic coding, or run-length encoding. Based on the probability distribution of pixel values ​​in the image data, shorter codes are used for frequently occurring pixel values, while longer codes are used for less frequently occurring pixel values, resulting in an image bitstream (also called a bit stream) for further compression of the image data.

[0052] It's important to note that in AI-based image encoding and decoding systems, before entropy encoding, hyper-prior encoding and hyper-prior decoding can be used to encode the distribution parameters of the image data as auxiliary information for entropy encoding. Hyper-prior encoding extracts features from the image data and maps these features to a low-dimensional latent space to capture key semantic information of the image and the dependencies between different representations. Hyper-prior decoding reconstructs the latent space representation back into the image space. Then, based on the quality feedback of the reconstructed image, the encoding strategy of entropy encoding is adjusted to optimize the quality of the entropy-encoded image.

[0053] Entropy decoding is the inverse operation of entropy coding, used to reconstruct image data from a bitstream. It can include parsing entropy-coded symbols or codewords from the bitstream, using an entropy decoding algorithm (i.e., the inverse algorithm of entropy coding) to map the codewords back to the original symbols (such as pixel values), and reconstructing the probability distribution of the original symbols based on the context information from the entropy coding process.

[0054] Dequantization is the inverse operation of quantization. It can involve determining the quantization matrix after obtaining the quantization coefficients of the probability distribution from entropy decoding, and multiplying the quantization coefficients by the corresponding values ​​in the quantization matrix to map the quantization coefficients from the discrete levels of the quantization table back to the original continuous value range. Then, an inverse transform (such as inverse DCT) is used to restore the continuous value range to the spatial domain and reconstruct approximate pixel values ​​of the image. Finally, the inversely transformed pixel values ​​are recombine to form complete image data to obtain the reconstructed image.

[0055] Although Figure 1 describes the encoding and decoding processes as a single system, in practical implementations, encoding functionality is typically integrated into the encoder, and decoding functionality is typically integrated into the decoder. The presence and (precise) division of encoder and decoder functionality may vary depending on the specific device and application. Optionally, the encoder and decoder can be located in the same electronic device; alternatively, they can be located in different electronic devices. The electronic devices described herein can include any category of handheld or stationary devices, such as laptops or notebook computers, mobile phones, smartphones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, cameras, in-vehicle devices, display devices, digital media players, video game consoles, video streaming devices (e.g., content service servers or content distribution servers), broadcast receiver devices, broadcast transmitter devices, etc.

[0056] If the image encoding and decoding functions illustrated in Figure 1 are implemented using an AI encoding and decoding model, the aforementioned functions such as quantization, entropy encoding, entropy decoding, dequantization, and reconstruction can all be executed by algorithm modules. Before deployment, the model typically needs to be trained to enable these functions. A common implementation approach is to train an AI encoding and decoding model for a fixed bitrate image encoding and decoding, and this AI encoding and decoding model only supports encoding and decoding at that fixed bitrate. If applied to encoding and decoding scenarios with multiple bitrates, one approach is to deploy multiple model weights, loading the corresponding weights onto the AI ​​encoding and decoding model for different bitrate encoding and decoding tasks. It is evident that conventional AI encoding and decoding systems have poor versatility and flexibility. Deploying multiple model weights requires more storage resources, while loading model weights increases the complexity and computational cost of the AI ​​encoding and decoding system.

[0057] In view of this, the embodiments of this disclosure extract image frequency features from at least two frequency domains of the source image before encoding the source image, and determine the bit rate for different frequency components of the source image, such that the bit rate is proportional to the frequency component. During the encoding process, features whose importance matches the quality level of the source image are selected as features to be encoded, and encoding is performed according to the bit rate matching the image frequency features. In this way, there is no need to deploy multiple model weights, thereby saving storage resources and reducing the complexity and computing power of the AI ​​encoding and decoding system. Furthermore, the technical solution of the embodiments of this disclosure can adaptively select features to be encoded according to different frequency components of the source image, and use different bit rates for encoding different frequency components, which is not only highly versatile but also flexible.

[0058] The system architecture involved in the embodiments of this disclosure is described below.

[0059] Referring to Figure 2, which illustrates an AI encoding / decoding system provided in an embodiment of this disclosure, the system may include an AI encoding model and an AI decoding model. The AI ​​encoding model includes a frequency division module 1000, a first encoding module 2000-1, and an x-th encoding module 2000-2. The AI ​​decoding model includes a first decoding module 3000-1, an x-th decoding module 3000-2, a fusion module 4000, and a reconstruction module 5000. In this embodiment, x can be an integer greater than or equal to 2. In specific implementations, the frequency division module 1000, the first encoding module 2000-1, the x-th encoding module 2000-2, the first decoding module 3000-1, the x-th decoding module 3000-2, the fusion module 4000, and the reconstruction module 5000 can be implemented as hardware components, software components, or a combination of hardware and software. For example, the frequency division module 1000, the first encoding module 2000-1, the xth encoding module 2000-2, the first decoding module 3000-1, the xth decoding module 3000-2, the fusion module 4000, and the reconstruction module 5000 can be deep learning-based algorithm modules.

[0060] It should be understood that the AI ​​encoding and decoding system shown in Figure 2 is merely illustrative and does not limit the AI ​​encoding and decoding system of this disclosure. In actual implementation scenarios, the AI ​​encoding and decoding system may include more modules than those shown in Figure 2, for example, it may also include a first super-prior encoding module and a first super-prior decoding module, etc. No limitation is imposed here.

[0061] The frequency division module 1000 can be used to divide the input source image into x different frequency components, and the x different frequency components can respectively characterize the features of x different frequency domains of the source image. The source image here can include any real-world image or video, or an image or video of a real object. In some embodiments, the frequency division module 1000 can be implemented as x convolutional modules, and the weights, number of channels, and convolutional kernels of these x convolutional modules can be different. Each convolutional module can be used to extract features from one frequency domain of the source image.

[0062] In this way, by performing frequency decomposition on the source image, image information can be analyzed and compressed from a frequency perspective, which can effectively reduce redundant information, improve coding efficiency, and help retain more image details and improve quality.

[0063] Any one of the first encoding modules 2000-1 to the xth encoding module 2000-2 can be used to encode the image frequency features in the corresponding frequency domain to obtain the bitstream of the image in the corresponding frequency domain. The encoding process of the corresponding image frequency features by any encoding module is similar to that shown in the embodiment of FIG1, and may include quantization and entropy encoding. The encoding module may include a quantization module and an entropy encoding module. Optionally, in the embodiments of this disclosure, any encoding module may further include a priori encoding module, a priori decoding module, and a feature selection module. Before performing entropy encoding, the image frequency features can be processed by the priori encoding module and the priori decoding module to obtain mask features for screening important features. Then, the feature selection module combines the mask features to select features for encoding.

[0064] In some embodiments, for a source image, the encoding modules used for encoding the image frequency features of high-frequency components in the first encoding module 2000-1 to the xth encoding module 2000-2 have a higher encoding code rate, while the encoding modules used for encoding the image frequency features of low-frequency components have a lower encoding code rate. This enables variable code rate encoding and helps to preserve more details of the source image.

[0065] In other embodiments, for different source images, the AI ​​encoding and decoding system of this disclosure can adaptively calculate the encoding rate of the first encoding module 2000-1 to the xth encoding module 2000-2 based on the quality level of the source image, which is not only highly versatile but also flexible.

[0066] Any encoding module may include quantization, entropy coding, super-prior coding, super-prior decoding, and feature selection modules, which can be algorithm modules, deep learning-based image processing networks, or models. These algorithm modules, deep learning-based image processing networks, or models can be combined in different ways to achieve the functionality of the encoding module.

[0067] Any one of the first decoding modules 3000-1 to the xth decoding module 3000-2 can be used to decode the bitstream in the corresponding frequency domain to obtain the image frequency features in the corresponding frequency domain. Each decoding module may include an entropy decoding module and an inverse quantization module, used to perform entropy decoding and inverse quantization on the corresponding bitstream to obtain the image frequency features in the corresponding frequency domain. The decoding process of any decoding module is similar to the embodiment illustrated in Figure 1, and will not be described in detail here.

[0068] It should be noted that for any given bitstream, if the decoding bitrate matches the bitrate used to encode the bitstream, then the decoding bitrate used by the decoding module can be the same as the encoded bitstream of the corresponding encoding module. For example, the encoding bitrate of the first encoding module 2000-1 is the same as the decoding bitrate of the first decoding module 3000-1.

[0069] The first decoding module 3000-1 to the xth decoding module 3000-2 respectively decode the image decoding features in the x frequency domains. In order to reconstruct the image, the fusion module 4000 can fuse the image decoding data in the x frequency domains to obtain complete image decoding data. Then, the reconstruction module 5000 can reconstruct an image that matches the source image based on the complete image decoding data.

[0070] In the AI ​​encoding / decoding system environment illustrated in Figure 2, this disclosure provides an encoding method as shown in Figure 3, which includes the following steps:

[0071] In step 101, image frequency features in at least two frequency domains of the source image are extracted.

[0072] In step 102, features whose importance in the image frequency features matches the quality level of the source image are selected as features to be encoded.

[0073] In step 103, the feature to be encoded is encoded according to a bitrate that matches the image frequency feature to obtain a bitstream corresponding to the image frequency feature, wherein the bitrate is proportional to the frequency component corresponding to the image frequency feature.

[0074] As can be seen, by employing the embodiments of this disclosure, at least two frequency domain image frequency features of the source image are first extracted, thereby enabling differentiated processing of different frequency components of the source image and effectively adjusting the bitrate. Furthermore, taking the image frequency features as the main body, features whose importance matches the quality level of the source image are selected from each image frequency feature as the features to be encoded for the corresponding image frequency feature. Specifically, for any image frequency feature, the feature whose importance matches the quality level of the source image can refer to the necessary feature representation required for encoding based on the quality level. This reduces the amount of data encoded and helps maintain high visual quality. In addition, in this embodiment, the bitrate is proportional to the frequency domain corresponding to the image frequency feature. Thus, for any image frequency feature to be encoded, it can be encoded according to the bitrate matching the image frequency feature to obtain the bitstream corresponding to that image frequency feature. This facilitates variable bitrate encoding based on different frequency domain components of the source image, improving encoding efficiency while preserving the details of the source image. As can be seen, the technical solution of this disclosure embodiment does not require the deployment of multiple model weights, thereby saving storage resources and reducing the complexity and computing power of the AI ​​encoding and decoding system. Moreover, the technical solution of this disclosure embodiment can adaptively select the features to be encoded for different frequency components of the source image according to the quality of the source image, and use different bit rates for encoding for different frequency components, which is not only highly versatile and flexible.

[0075] In some embodiments, the AI ​​encoding / decoding system can invoke a frequency division module to extract image frequency features in at least two frequency domains from the source image. During implementation, the frequency division module can extract image features from the source image, convert these image features into frequency domain features, and then extract at least two frequency domain image frequency features from the frequency domain features.

[0076] For example, the frequency division module can use algorithms such as DCT, DFT or wavelet transform to convert the image features into frequency domain features. Then, at least two convolutional modules are used to process the frequency domain features respectively, so that each convolutional module outputs an image frequency feature with one frequency component.

[0077] It should be understood that the embodiments of this disclosure extract image frequency features from at least two frequency domains of the source image and perform encoding compression on a unit basis for each image frequency feature, aiming to balance image quality and compression effect. In some implementation scenarios, if image frequency features from two frequency domains are extracted, although the encoding compression effect is good, it will lead to poor image quality based on the encoded image. If image frequency features from four or more frequency domains are extracted, the image quality based on the encoded image can be maintained at a better level, but due to the large amount of data, the encoding compression effect is not good. In view of this, optionally, the image frequency features from at least two frequency domains in the embodiments of this disclosure can be implemented as three image frequency features: high-frequency component image frequency features, mid-frequency component image frequency features, and low-frequency component image frequency features. The high-frequency component image frequency features represent the texture features of the source image, the mid-frequency component image frequency features represent the contour features of the source image, and the low-frequency component image frequency features represent the color distribution features of the source image.

[0078] For example, referring to the image frequency division scenario illustrated in Figure 4, the frequency division module may include convolution module 1, convolution module 2, and convolution module 3. At least one of the following parameters—the number of convolutional layers, weights, number of channels, and number of convolutional kernels—can be different for each of the convolutional modules 1, 2, and 3. The frequency division module can extract image frequency features of high-frequency components from the source image using convolution module 1, extract image frequency features of mid-frequency components using convolution module 2, and extract image frequency features of low-frequency components using convolution module 3. As shown in Figure 4, the higher the frequency domain, the richer the color information contained in the image frequency features of the corresponding frequency domain components.

[0079] As can be seen, the technical solution disclosed herein, by performing frequency decomposition on the source image, supports the analysis and compression of image information from a frequency perspective, which can effectively reduce redundant information, not only improve coding efficiency, but also help retain more image details and improve quality.

[0080] Furthermore, the AI ​​encoding / decoding system can call at least two encoding modules to encode the aforementioned at least two image frequency features, wherein the at least two encoding modules correspond one-to-one with the at least two image frequency features. The algorithm modules included in each encoding module, as well as the encoding process for the corresponding image frequency features, are similar. The following uses the encoding process of one encoding module for the corresponding image frequency feature as an example to illustrate the image encoding process of this embodiment.

[0081] Before introducing the encoding process, it should be noted that in the source image acquisition stage of this embodiment, the AI ​​encoding / decoding system can obtain the quality value (Q value) of the source image. Furthermore, the AI ​​encoding / decoding system can determine the encoding bitrate of the source image based on its quality value. The Q value refers to the image quality level after reconstruction, which affects the compression quality and detail of the source image. A higher Q value indicates better source image quality, larger image data size, and a lower compression ratio; conversely, a lower Q value indicates worse source image quality, smaller image data size, and a higher compression ratio. Therefore, the AI ​​encoding / decoding system in this embodiment is not only highly versatile but also flexible.

[0082] Furthermore, for any image frequency feature, the encoding module corresponding to the image frequency feature can generate a mask feature of the image frequency feature based on the image frequency feature, the source image, and the quality level. Then, the feature corresponding to the non-mask feature value in the quantized features of the image frequency feature is determined as the feature to be encoded.

[0083] Masking features are used to filter out image frequency features whose importance matches the quality level of the source image. Features whose importance matches the quality level of the source image can refer to the necessary feature representations in the image frequency features that enable the source image to achieve the corresponding quality level. Masking features can include masked feature values ​​and non-masked feature values. Masked feature values ​​are used to identify image frequency features that do not need to be encoded, while non-masked feature values ​​are used to identify image frequency features that need to be encoded.

[0084] The encoding module corresponding to the image frequency features can extract the semantic representation of the corresponding frequency domain from the image frequency features. This semantic representation can be a compressed, low-dimensional spatial feature, which can contain the key information required to reconstruct the image. Further, based on the semantic representation and the image features of the source image, an initial importance feature is generated. Each feature value in the initial importance feature represents the degree of influence of the corresponding pixel on the semantics of the source image. Then, the feature values ​​in the initial importance feature can be adjusted according to the quality level to obtain the mask feature.

[0085] In some embodiments, the encoding module may include a pre-trained importance feature learning model, which can evaluate each feature in the image frequency features based on the semantic representation and the image features of the source image, thereby determining the degree of influence of each feature element in the image frequency features on the quality of the source image, and generating initial importance features based on the degree of influence of each feature element on the quality of the source image.

[0086] For example, semantic representations derived from image frequency features can represent the key information needed to reconstruct the image corresponding to the image frequency features. The importance feature learning model can extract image features from the source image, which may include structural, texture, and content features. Based on these image features, the importance feature learning model can determine the key image features needed to reconstruct the source image. Furthermore, it can assess the impact of each feature element in the image frequency features on the quality of the source image based on the similarity between the semantic representation and the key image features, and generate initial importance features.

[0087] The initial importance feature can be a three-dimensional (3D) tensor. The three dimensions can include position (x, y) and channel (ch). x refers to the coordinate along the width of the image, y refers to the coordinate along the height of the image, and channel refers to the color channel in the image. In a color image, each pixel can contain multiple color channels, and the values ​​of these channels collectively define the color of that point. For example, color channels can be red, green, and blue (RGB) channels, with each channel having a value between 0 and 255. Any feature value includes the position parameter and channel parameter of the corresponding pixel.

[0088] Understandably, the generation of initial importance features depends on the content of the source image. Different source images usually have different important content and structure, and the initial importance features corresponding to the same frequency domain image features of different source images can be different.

[0089] Based on the description of the initial importance features, it can be seen that the specific parameters characterizing image quality can be the channel parameters of each element. For example, please refer to Figure 5, which shows an exemplary scenario diagram of the relationship between channel parameters and image quality. The more feature elements a picture contains, and the richer the channel parameters of these feature elements, the larger the Q value of the image. In this case, the visual quality of the image is higher, and the content details displayed are richer. Conversely, the fewer feature elements a picture contains, and the fewer the channel parameters of these feature elements, the smaller the Q value of the image. In this case, the visual quality of the image is lower, and the content details displayed are less.

[0090] As can be seen, the implementation method of this disclosure supports the capture of important features in the source image based on the content of the source image. On the one hand, it can provide more detailed features for compression and encoding, which is beneficial to improving the image quality. Even for low bitrate images, the reconstructed image can maintain high quality. On the other hand, it can filter out features that do not need to be encoded, thereby improving the compression efficiency.

[0091] The initial importance feature only identifies important and unimportant regions in image frequency features from the perspective of the importance of feature elements, but it cannot characterize the importance of each feature value in important regions. The importance of each feature value can be reflected by the weights of the channel parameters in the feature value.

[0092] Accordingly, adjusting each feature value in the initial importance feature according to the quality level to obtain the mask feature can be implemented as follows: the importance feature learning model calculates the weights of the channel parameters corresponding to each feature value according to the quality level, adjusts each feature value according to the corresponding weight to obtain the importance feature, binarizes each feature value in the importance feature, and uses the binarized feature as the mask feature.

[0093] For example, the mask feature can be a 3D binary mask feature, where the feature value "1" can be a mask feature value, and the feature value "0" can be a non-mask feature value.

[0094] The same feature element has different importance in images of different quality levels. For example, complex texture elements are more important in images with higher quality levels, but less important, or even unnecessary, in images with lower quality levels. Based on this, this implementation allows for adjustments to the weights of each feature value in the initial importance feature set to adapt to different image quality requirements. It can accommodate both the detail required for matching quality and compression efficiency, giving the embodiments of this disclosure good adaptability and scalability.

[0095] Based on the foregoing description of the image encoding process, it can be seen that the image encoding module includes quantization and entropy encoding. Accordingly, in this embodiment of the present disclosure, the encoding module selects the feature corresponding to the non-masked feature value in the image frequency features as the feature to be encoded. This can be achieved by selecting the feature identified by the non-masked feature value from the quantized features of the image frequency features as the feature to be encoded.

[0096] For example, the encoding module can calculate a quantization vector based on the quality level, and quantize the image frequency features based on the quantization vector to obtain the quantized features of the image frequency features. For example, the quantized feature n' can satisfy: n' = AdaQ(n) = Round(n / (Q*Vq)), where n refers to the image frequency feature, Q refers to the quality level Q value, Vq refers to the quantization factor, and * is the quantization vector. Further, the feature identified by the non-masked feature value is selected from the quantized features and determined as the feature to be encoded. satisfy: Where ( ) is the element selection operator, (,) represents the mask feature, refers to the mask feature value in the mask feature, and refers to the non-mask feature value in the mask feature.

[0097] In conventional quantization, the quantization factor is calculated based on the QP value, which is a pre-set fixed value; that is, the degree of quantization remains constant in conventional quantization. In contrast, the technical solution of this disclosure uses the Q value corresponding to the source image to calculate the quantization factor. This allows for more precise control of the quantization process to adapt to different image content and compression requirements, offering good flexibility and scalability. Furthermore, by adaptively controlling the quantization degree based on the Q value, it can maintain key visual information while performing more aggressive compression on less important areas, thereby achieving better visual quality at the same bitrate.

[0098] Furthermore, the encoding method of this disclosure also generates 3D mask features for image frequency features, and after quantizing the image frequency features, further selects features to be encoded based on the 3D mask features. Since the 3D mask features can characterize the channel importance of different regions of the image, the encoding of this disclosure can select features to be encoded from a finer-grained perspective, which is beneficial for achieving variable rate compression and controlling image quality.

[0099] In some embodiments, after obtaining the quality level (i.e., Q-value) of the source image, the AI ​​encoding / decoding system can determine the rated bitrate matching the quality level of the source image. Then, according to a preset bitrate ratio corresponding to the frequency domain and the rated bitrate, it can calculate the bitrate matching each of the at least two image frequency features, where the bitrate ratio corresponding to the higher frequency domain is greater than that corresponding to the lower frequency domain. Afterwards, after the encoding module selects the feature to be encoded from the image frequency features, it can encode the feature to be encoded according to the bitrate matching that image frequency feature to obtain the bitstream corresponding to that image frequency feature. Here, encoding refers to entropy coding.

[0100] For example, the rated bitrate can be the total bitrate budget corresponding to the Q value; the larger the Q value, the larger the rated bitrate.

[0101] In some embodiments, different frequency components can be pre-assigned a fixed percentage of the bit rate. For example, the bit rate of high frequency components is 45%, the bit rate of mid frequency components is 30%, and the bit rate of low frequency components is 25%.

[0102] In other embodiments, the AI ​​codec system can use a rate allocation algorithm to dynamically allocate bitrates to different frequency components based on the Q value. The rate allocation algorithm may include, but is not limited to, Rate-Controlled Distortion Optimization (RCDO) algorithms.

[0103] In conjunction with the foregoing embodiments, the AI ​​encoding / decoding system has divided the source image into at least two frequency domains to obtain image frequency features. Using these frequency features as the primary component, entropy encoding is performed at different coding rates, thereby obtaining bitstreams from at least two frequency domains of the source image. This allows for the analysis and extraction of image components and details at different frequency domain levels, which is beneficial for improving image quality. Furthermore, encoding at a rate matching the image frequency features allows for variable bitrate encoding without increasing model complexity, resulting in strong versatility and high flexibility.

[0104] Corresponding to the foregoing encoding embodiments, this disclosure also provides a decoding method. As shown in FIG6, the decoding method of this disclosure embodiment may include steps 201 to 203.

[0105] Step 201: Obtain at least two bitstreams, wherein the at least two bitstreams correspond one-to-one with the image frequency features in at least two frequency domains.

[0106] Step 202: For any bitstream, decode the bitstream according to the bitrate that matches the bitstream to obtain decoding features.

[0107] The bitrate for matching the bitstream refers to the bitrate that matches the image frequency features corresponding to the bitstream, and the bitrate is proportional to the frequency components corresponding to the image frequency features.

[0108] Step 203: After obtaining at least two decoding features, the at least two decoding features are fused to obtain the reconstructed image of the source image.

[0109] Among them, at least two bitstreams are obtained by encoding at least two image frequency features, and the at least two image frequency features correspond to different frequency domains of the source image. Each bitstream is obtained by encoding a feature whose importance in the corresponding image frequency feature matches the quality level of the source image. Each image frequency feature is encoded by the encoding module corresponding to that image frequency feature.

[0110] The process of encoding any image frequency feature by the corresponding encoding module is detailed in the description of the above embodiments and will not be repeated here.

[0111] As shown in Figure 2, the AI ​​encoding and decoding system in the implementation scenario shown in Figure 6 means that any one of the at least two bitstreams is received and decoded by a decoding module, that is, at least two bitstreams correspond one-to-one with at least two decoding modules.

[0112] Based on the foregoing description of the image decoding process, it can be seen that the decoding module's decoding process for the bitstream includes entropy decoding and inverse quantization. The decoding bitrate used in the entropy decoding process can be the same as the entropy encoding bitrate corresponding to that bitstream. For any bitstream, the bitrate used for entropy decoding of that bitstream can be the bitrate obtained by the corresponding encoding module, that is, the bitrate of the image frequency feature matching corresponding to the bitstream.

[0113] Furthermore, dequantization is the inverse operation of quantization. In this embodiment of the disclosure, the quantization factor is calculated based on the Q value corresponding to the source image. Correspondingly, the quantization factor used in the dequantization process can also be calculated based on the Q value corresponding to the source image. This embodiment of the disclosure will not elaborate further on this.

[0114] It should be understood that the at least two decoding features correspond one-to-one with the at least two bitstreams, and the at least two bitstreams are obtained by feature encoding of different frequency domains of the source image. Correspondingly, the at least two decoding features correspond to different frequency domains of the source image. That is, any decoding feature represents a part of the features of the source image. The decoding feature after fusing the at least two decoding features can represent the complete features of the source image. In this way, the AI ​​encoding and decoding system can reconstruct the source image based on the fused decoding features to obtain the reconstructed image of the source image.

[0115] As can be seen, before encoding the source image, this embodiment extracts image frequency features from at least two frequency domains of the source image and determines the bit rate for different frequency components of the source image, making the bit rate proportional to the frequency domain. During the encoding process, features whose importance matches the quality level of the source image are selected as features to be encoded, and encoding is performed according to the bit rate matching the image frequency features. In this way, there is no need to deploy multiple model weights, thereby saving storage resources and reducing the complexity and computing power of the AI ​​encoding and decoding system. Furthermore, the technical solution of this embodiment can adaptively select features to be encoded according to different frequency components of the source image, and use different bit rates for encoding different frequency components, which is not only highly versatile but also flexible.

[0116] The above embodiments are described from the perspectives of encoding and decoding processes. The embodiments of this disclosure will now be described in conjunction with the composition of an AI encoding / decoding system and exemplary encoding / decoding processes.

[0117] Referring to Figure 7, which shows a schematic diagram of the data flow of encoding and decoding in any frequency domain provided in the embodiments of this disclosure, the AI ​​encoding and decoding model shown in Figure 7 can be any branch model of the AI ​​encoding and decoding system in the embodiments of this disclosure. The branch model is used to encode and decode the image frequency features in the frequency domain.

[0118] It should be noted that the branch model used to process features in other frequency domains in the AI ​​codec system can be similar to that shown in Figure 7. This disclosure will not repeat the description of other branch models in the AI ​​codec system.

[0119] The AI ​​encoding / decoding model shown in Figure 7 may include an encoding module, a super-prior encoding module, a first entropy encoding module, a super-prior decoding module, a quantization module, a feature selection module, a second entropy encoding module, an entropy decoding module, and an inverse quantization module. Each module in this AI encoding / decoding model can be a deep learning model, such as a neural network. After receiving the image corresponding to this AI encoding / decoding model (e.g., an image with high-frequency components), the encoding module encodes the image into a feature representation n (i.e., image frequency features), and transmits the feature representation n to the quantization module and the super-prior encoding module, respectively.

[0120] The quantization module can quantize the image frequency features using the algorithm n' = AdaQ(n) = Round(n / (Q*Vq)) to obtain the quantized feature n' of the image frequency features. Then, the quantized feature n' is transmitted to the feature selection module. The super-prior coding module, the first entropy coding module, and the super-prior decoding module can generate the 3D mask features corresponding to the image frequency features after processing the image frequency features.

[0121] For example, the super-prior encoding module is used to extract super-prior features from image frequency features. Furthermore, it can establish a distribution estimate for each super-prior feature, ensuring independence among the features. This distribution estimate characterizes the semantics of the image frequency features. This distribution estimate is then transmitted to the first entropy encoding module, which uses it to perform arithmetic encoding on the super-prior features, obtaining a binary code stream of super-prior features. This binary code stream is then transmitted to the super-prior decoding module. The super-prior decoding module performs arithmetic decoding on the binary code stream to obtain the recovered super-prior features, and applies a super-decoding neural network to the recovered super-prior features to obtain super-prior information. The advanced prior decoding module also receives the source image and Q-values, extracts image features from the source image, and these features characterize the semantics of the source image, such as structure, texture, and content. Furthermore, the advanced prior decoding module determines the key image features needed to reconstruct the source image based on these features, and evaluates the influence of each feature element in the image frequency features on the quality of the source image based on the similarity between the advanced prior information and the key image features, and generates a 3D importance feature map (i.e., the aforementioned initial importance feature). Further, the advanced prior decoding module can calculate the weights of each feature value in the 3D importance feature map based on the Q-values, so as to dynamically adjust the importance of the channel parameters of each feature value according to the Q-values.

[0122] For example, the weights of each feature value can be implemented as an importance curve, which can be a non-linear curve. The super-prior decoding module can be a convolutional neural network, in which adjustment parameters can be set to adjust the importance curve of the 3D important feature map based on the Q-value. These adjustment parameters can be obtained through pre-training.

[0123] Furthermore, the value range of each feature in the 3D important feature map after weight adjustment is between 0 and 1. The super-prior decoding module can binarize the value of each feature, and the binarized feature is used as the 3D mask feature, and the 3D mask feature is transmitted to the feature selection module.

[0124] Feature selection modules can be implemented, for example, through algorithms: Based on the 3D mask features, select the features to be encoded from the quantized features n'. and the feature to be encoded It is transmitted to the second entropy encoding module.

[0125] It should be noted that the AI ​​encoding / decoding system can determine the bitrate of image frequency features for different frequency components based on the Q-value of the source image, and configure the bitrate of the second entropy encoding module and entropy decoding module in the corresponding branch model. For example, in the AI ​​encoding / decoding model illustrated in Figure 7, which processes image frequency features of high-frequency components, the bitrate corresponding to the high-frequency components calculated by the AI ​​encoding / decoding system based on the Q-value of the source image can be configured for the second entropy encoding module and entropy decoding module in Figure 7. In this way, the second entropy encoding module can process the features to be encoded. Entropy encoding is performed according to the configured bitrate to obtain a bitstream of image frequency features of high-frequency components. This bitstream is then transmitted to the entropy decoding module, which performs entropy decoding to obtain data to be dequantized. This dequantized data is then transmitted to the dequantization module. After processing the dequantized data, the dequantization module outputs image features in the image frequency feature spatial domain. These features are then decoded into image data, which is high-frequency component image data. This image data can be used to fuse with image data from other frequency components and to reconstruct the reconstructed image corresponding to the source image.

[0126] It should be understood that the AI ​​codec model shown in Figure 7 is merely illustrative and does not limit the branch model in the AI ​​codec system of this disclosure. In actual implementation scenarios, the AI ​​codec model may include more or fewer modules than those shown in Figure 7, and multiple modules in Figure 7 may be merged into one module, or some modules in Figure 7 may be split into multiple modules, etc. No limitations are imposed here.

[0127] By adopting the implementation method of the embodiments of this disclosure, the quantization module adaptively captures important features in the source image according to the Q value, and by pre-deploying 3D mask features in the super-prior decoding module, it supports improving the compression efficiency based on the content of the source image while ensuring image quality.

[0128] Since the AI ​​encoding / decoding system is an algorithmic system, it should be trained to execute the above embodiments before being put into use. In some embodiments, the training method may include: inputting a sample image into a network to be trained to obtain at least two predicted frequency features output by the network, wherein the at least two predicted frequency features correspond to different frequency domains of the sample image. Further, a loss function is calculated based on the at least two predicted frequency features, the loss function including at least two loss values, each of which corresponds one-to-one with the at least two predicted frequency features, and any loss value characterizing the loss of the corresponding predicted frequency feature relative to the frequency features in the corresponding frequency domain of the sample image; if the loss function reaches a preset convergence condition, the network to be trained is determined as an encoding model, and the encoding model is used to encode the source image.

[0129] It should be noted that the semantics represented by image frequency features of different frequency components are different and each has its own emphasis. Therefore, the algorithm for calculating the loss value of each predicted frequency feature can be matched with the image semantics corresponding to the predicted frequency feature.

[0130] For example, when the network to be trained includes a high-frequency component feature extraction network, a mid-frequency component feature extraction network, and a low-frequency component feature extraction network, the at least two predicted frequency features include high-frequency component predicted frequency features, mid-frequency component predicted frequency features, and low-frequency component predicted frequency features. The high-frequency component represents edge details and texture information in the image, the mid-frequency component represents the shape and contour of objects in the image, and the low-frequency component represents the overall brightness and color distribution of the image. The perceptual loss value of the high-frequency component predicted frequency features relative to the high-frequency component frequency features of the sample image can be calculated to obtain L. high Calculate the mean squared error (MSE) loss value L of the predicted frequency features of the intermediate frequency components relative to the intermediate frequency component frequency features of the sample image. mid Calculate the color loss value of the predicted low-frequency component frequency features relative to the low-frequency component frequency features of the sample image, and obtain L. low Then, the perceptual loss value, the mean squared error loss value, and the color loss value are weighted and summed, and the result of the weighted sum is used as the loss function. The loss function L, for example, satisfies: L = W high L high +W mid L mid +W low L low Among them, W high W mid and W low These are the weighting coefficients.

[0131] As can be seen, in the encoding stage of the source image in this embodiment, at least two frequency domain image frequency features of the source image are first extracted, thereby enabling differentiated processing of different frequency components of the source image and effectively adjusting the bitrate. Furthermore, based on the image frequency features, features whose importance matches the quality level of the source image are selected from each image frequency feature as the features to be encoded for the corresponding image frequency feature. Specifically, for any image frequency feature, the feature whose importance matches the quality level of the source image can refer to the necessary feature representation required for encoding based on the quality level. This reduces the amount of data encoded and helps maintain high visual quality. In addition, in this embodiment, the bitrate is proportional to the frequency domain corresponding to the image frequency feature. Thus, for any image frequency feature to be encoded, it can be encoded according to the bitrate matching the image frequency feature to obtain the bitstream corresponding to that image frequency feature. This facilitates variable bitrate encoding based on different frequency domain components of the source image, improving encoding efficiency while preserving the details of the source image. As can be seen, the technical solution of this disclosure embodiment does not require the deployment of multiple model weights, thereby saving storage resources and reducing the complexity and computing power of the AI ​​encoding and decoding system. Moreover, the technical solution of this disclosure embodiment can adaptively select the features to be encoded for different frequency components of the source image according to the quality of the source image, and use different bit rates for encoding for different frequency components, which is not only highly versatile and flexible.

[0132] It should be understood that the components illustrated in Figure 2 can be implemented as hardware or a combination of hardware and computer software. Whether the processing steps of any related component are executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can also implement the functions described in the above embodiments using different methods for specific applications, but such implementations should not be considered beyond the scope of this disclosure.

[0133] For example, if the above implementation steps can be implemented by software modules, corresponding to the above encoding method, this disclosure embodiment can also provide an encoding device.

[0134] As shown in Figure 8A, an encoding device is provided, which may include an extraction unit 71, a feature selection unit 72, and an encoding unit 73. This encoding device can be used to perform some or all of the operations of the encoding modules in Figures 2 to 5.

[0135] For example: extraction unit 71 is used to extract image frequency features in at least two frequency domains of the source image; feature selection unit 72 is used to select features among the image frequency features whose importance matches the quality level of the source image as features to be encoded, wherein the features to be encoded refer to the necessary feature representations required when encoding based on the quality level; encoding unit 73 is used to encode the features to be encoded according to a bitrate that matches the image frequency features to obtain a bitstream corresponding to the image frequency features, wherein the bitrate is proportional to the frequency components corresponding to the image frequency features.

[0136] Optionally, the feature selection unit 72 is further configured to generate a mask feature of the image frequency feature based on the image frequency feature, the source image, and the quality level, wherein the mask feature includes a mask feature value and a non-mask feature value; and to determine the feature corresponding to the non-mask feature value in the quantized features of the image frequency feature as the feature to be encoded.

[0137] Optionally, the feature selection unit 72 is further configured to extract the semantic representation of the frequency domain corresponding to the image frequency features from the image frequency features; generate an initial importance feature based on the semantic representation and the image features of the source image, wherein any feature value in the initial importance feature represents the degree of influence of the corresponding pixel on the semantics of the source image; and adjust each feature value in the initial importance feature according to the quality level to obtain the mask feature.

[0138] Optionally, any feature value in the initial importance feature includes the position parameter and channel parameter of the corresponding pixel. The channel parameter represents the parameter of each color channel of the pixel. The feature selection unit 72 is further used to calculate the weight of the channel parameter corresponding to each feature value according to the quality level; adjust each feature value according to the corresponding weight to obtain the importance feature; binarize each feature value in the importance feature, and use the binarized feature as the mask feature.

[0139] Optionally, the feature selection unit 72 is further configured to calculate a quantization vector based on the quality level; quantize the image frequency features based on the quantization vector to obtain quantized features of the image frequency features; and select the features identified by the non-masked feature value from the quantized features to determine the features to be encoded.

[0140] Optionally, the extraction unit 71 is further configured to extract image features of the source image; convert the image features into frequency domain features; and extract image frequency features of high-frequency components, mid-frequency components, and low-frequency components from the frequency domain features, respectively; the image frequency features of high-frequency components characterize the texture features of the source image, the image frequency features of mid-frequency components characterize the contour features of the source image, and the image frequency features of low-frequency components characterize the color distribution features of the source image.

[0141] Optionally, the encoding device further includes a determining unit and a calculating unit. The determining unit is used to determine the rated bit rate matching the quality level. The calculating unit is used to calculate the bit rate matching each of the at least two image frequency features according to the preset bit rate ratio in the frequency domain and the rated bit rate, wherein the bit rate ratio corresponding to the higher frequency domain is greater than the bit rate ratio corresponding to the lower frequency domain.

[0142] Optionally, the encoding device further includes an input unit. The input unit is configured to input a sample image into the network to be trained to obtain at least two predicted frequency features output by the network, wherein the at least two predicted frequency features correspond to different frequency domains of the sample image. The calculation unit is further configured to calculate a loss function based on the at least two predicted frequency features, wherein the loss function includes at least two loss values, each corresponding one-to-one with the at least two predicted frequency features, and any loss value characterizes the loss of the corresponding predicted frequency feature relative to the frequency features in the corresponding frequency domain of the sample image. The determination unit is further configured to, when the loss function reaches a preset convergence condition, determine the network to be trained as an encoding model, wherein the encoding model is used to encode the source image.

[0143] Optionally, when the network to be trained includes a high-frequency component feature extraction network, a mid-frequency component feature extraction network, and a low-frequency component feature extraction network, the at least two predicted frequency features include high-frequency component predicted frequency features, mid-frequency component predicted frequency features, and low-frequency component predicted frequency features. The computing unit is further configured to calculate the perceptual loss value of the high-frequency component predicted frequency features relative to the high-frequency component frequency features of the sample image; calculate the mean square error loss value of the mid-frequency component predicted frequency features relative to the mid-frequency component frequency features of the sample image; calculate the color loss value of the low-frequency component predicted frequency features relative to the low-frequency component frequency features of the sample image; and perform a weighted summation of the perceptual loss value, the mean square error loss value, and the color loss value, using the weighted summation result as the loss function.

[0144] Correspondingly, as shown in Figure 8B, a decoding device is provided, which may include an acquisition unit 81, a decoding unit 82, and a fusion unit 83. This encoding device can be used to perform some or all of the operations of the decoding module in Figure 6.

[0145] For example: Acquisition unit 81 is used to acquire at least two bitstreams, the at least two bitstreams corresponding one-to-one with at least two image frequency features in the frequency domain; each bitstream is obtained by encoding a feature whose importance in the corresponding image frequency feature matches the quality level of the source image, the feature to be encoded refers to the necessary feature representation required when encoding based on the quality level; Decoding unit 82 is used to decode any bitstream according to a bitrate matching the bitstream to obtain decoding features, the bitrate matching the bitstream refers to the bitrate matching the image frequency feature corresponding to the bitstream, the bitrate is proportional to the frequency component corresponding to the image frequency feature; Fusion unit 83 is used to fuse the at least two decoding features after obtaining them to obtain a reconstructed image of the source image, the at least two decoding features corresponding one-to-one with the at least two bitstreams.

[0146] It is understood that the division of units in Figures 8A and 8B is merely a logical functional division. In actual implementation, the functions of these units can be integrated into the hardware entity of the electronic device. Referring to Figure 9, Figure 9 provides an electronic device including a processor 811, a transceiver 812, and a memory 813. These components are connected and communicate via a communication bus 814. The processor 811 can integrate the functions of the modules shown in Figure 2, and the transceiver 812 can be used to acquire source images. The memory 813 includes data and program instructions for the various algorithm modules illustrated in Figure 2. When these program instructions are invoked, the processor 811 executes some or all of the operations shown in Figures 3 to 6.

[0147] For details on the implementation process, please refer to the descriptions related to the electronic devices in Figures 3 to 6, which will not be repeated here.

[0148] This disclosure also provides a computer-readable storage medium storing instructions for reading and writing data, which, when executed on a computer, cause the computer to perform some or all of the steps in the methods described in the foregoing embodiments.

[0149] This disclosure also provides a computer program product including instructions for reading and writing data, which, when run on a computer, causes the computer to perform some or all of the steps in the methods described in the foregoing embodiments.

[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0151] In the embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0152] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0153] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0154] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, smartphone, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0155] Although alternative embodiments of this disclosure have been described, those skilled in the art, upon learning the basic inventive concept, can make further changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this disclosure.

[0156] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this disclosure. It should be understood that the above description is only a specific embodiment of this disclosure and is not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An encoding method characterized by comprising: The method, applied to artificial intelligence (AI) coding models, includes: Extract image frequency features from at least two frequency domains of the source image; The features selected from the image frequency features whose importance matches the quality level of the source image are used as the features to be encoded. The features to be encoded refer to the necessary feature representations required for encoding based on the quality level. The feature to be encoded is encoded according to a bitrate that matches the image frequency feature to obtain a bitstream corresponding to the image frequency feature, wherein the bitrate is proportional to the frequency component corresponding to the image frequency feature.

2. The method of claim 1, wherein, The step of selecting features from the image frequency features whose importance matches the quality level of the source image as the features to be encoded includes: Based on the image frequency features, the source image, and the quality level, a mask feature for the image frequency features is generated, wherein the mask feature includes mask feature values ​​and non-mask feature values. The feature corresponding to the non-masked feature value in the quantized image frequency features is determined as the feature to be encoded.

3. The method of claim 2, wherein, The step of generating a mask feature for the image frequency features based on the image frequency features, the source image, and the quality level includes: Extract the semantic representation of the frequency domain corresponding to the image frequency features from the image frequency features; An initial importance feature is generated based on the semantic representation and the image features of the source image, wherein any feature value in the initial importance feature represents the degree of influence of the corresponding pixel on the semantics of the source image; The mask features are obtained by adjusting the individual feature values ​​in the initial importance feature according to the quality level.

4. The method of claim 3, wherein, Each feature value in the initial importance feature includes the position parameter and channel parameter of the corresponding pixel, whereby the channel parameter characterizes the parameters of each color channel of the pixel. Adjusting each feature value in the initial importance feature according to the quality level includes: Calculate the weights of the channel parameters corresponding to each feature value based on the quality level; Each feature value is adjusted according to its corresponding weight to obtain the importance feature; Each feature value in the importance feature is binarized, and the binarized feature is used as the mask feature.

5. The method of claim 2, wherein, The step of determining the feature corresponding to the non-masked feature value in the image frequency features as the feature to be encoded includes: Calculate the quantization vector based on the quality level; The image frequency features are quantized based on the quantization vector to obtain the quantized features of the image frequency features; The feature identified by the non-masked feature value from the quantized features is determined as the feature to be encoded.

6. The method of claim 1, wherein, The extraction of at least two image frequency features from the source image includes: Extract image features from the source image; The image features are converted into frequency domain features; Image frequency features of high-frequency components, mid-frequency components, and low-frequency components are extracted from the frequency domain features, respectively. The image frequency features of high-frequency components represent the texture features of the source image, the image frequency features of mid-frequency components represent the contour features of the source image, and the image frequency features of low-frequency components represent the color distribution features of the source image.

7. The method of claim 1, wherein, Before encoding the features to be encoded according to a bitrate matching the image frequency features, the method further includes: Determine the rated bitrate to match the quality level; According to the preset frequency domain bit rate ratio and the rated bit rate, calculate the bit rate matching each image frequency feature in the at least two image frequency features, where the bit rate ratio corresponding to the higher frequency domain is greater than the bit rate ratio corresponding to the lower frequency domain. 8.A method for training an AI encoding model, the method comprising: The method includes: The sample image is input into the network to be trained to obtain at least two predicted frequency features output by the network to be trained, and the at least two predicted frequency features correspond to different frequency domains of the sample image. A loss function is calculated based on the at least two predicted frequency features. The loss function includes at least two loss values, which correspond one-to-one with the at least two predicted frequency features. Each loss value represents the loss between the corresponding predicted frequency feature and the frequency feature in the corresponding frequency domain of the sample image. When the loss function reaches the preset convergence condition, the network to be trained is determined as an encoding model, and the encoding model is used to execute the encoding method of any one of claims 1-7.

9. The method of claim 8, wherein, When the network to be trained includes a high-frequency component feature extraction network, a mid-frequency component feature extraction network, and a low-frequency component feature extraction network, the at least two predicted frequency features include high-frequency component predicted frequency features, mid-frequency component predicted frequency features, and low-frequency component predicted frequency features. The calculation of the loss function based on the at least two predicted frequency features includes: Calculate the perceptual loss value of the predicted frequency features of the high-frequency components relative to the frequency features of the high-frequency components of the sample image; Calculate the mean square error loss value of the predicted frequency features of the intermediate frequency components relative to the intermediate frequency features of the sample image; Calculate the color loss value of the predicted frequency features of the low-frequency components relative to the frequency features of the low-frequency components of the sample image; The perceptual loss value, the mean square error loss value, and the color loss value are weighted and summed, and the result of the weighted summation is used as the loss function.

10. A decoding method, comprising: The method, applied to an artificial intelligence (AI) decoding model, includes: At least two bitstreams are acquired, and the at least two bitstreams correspond one-to-one with at least two image frequency features in the frequency domain; each bitstream is obtained by encoding a feature whose importance in the corresponding image frequency feature matches the quality level of the source image, and the feature to be encoded refers to the necessary feature representation required for encoding based on the quality level; For any bitstream, the bitstream is decoded according to a bitrate that matches the bitstream to obtain decoding features. The bitrate that matches the bitstream refers to the bitrate that matches the image frequency features corresponding to the bitstream. The bitrate is proportional to the frequency components corresponding to the image frequency features. After obtaining at least two decoding features, the at least two decoding features are fused to obtain the reconstructed image of the source image, wherein the at least two decoding features correspond one-to-one with the at least two bitstreams.

11. A coding system characterized by The system includes: an artificial intelligence (AI) encoding model and an AI decoding model, wherein, The AI ​​encoding model is used to extract image frequency features in at least two frequency domains from the source image; select features among the image frequency features whose importance matches the quality level of the source image as features to be encoded, wherein the features to be encoded refer to the necessary feature representations required when encoding based on the quality level; encode the features to be encoded according to a bitrate that matches the image frequency features to obtain a bitstream corresponding to the image frequency features, wherein the bitrate is proportional to the frequency components corresponding to the image frequency features; The AI ​​decoding model is used to decode any bitstream according to a bitrate that matches the bitstream to obtain decoding features. The bitrate that matches the bitstream refers to the bitrate that matches the image frequency features corresponding to the bitstream. The at least two decoding features obtained by the at least two decoding modules are fused to obtain the reconstructed image of the source image. The at least two decoding features correspond one-to-one with the at least two bitstreams.

12. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to cause the electronic device to perform the method as described in any one of claims 1-10.

13. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-10.

14. A computer program product, characterised in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-10.