Image processing methods and related device

Through the method of combining feature extraction and inverse transformation networks with neural networks, flexible control of compression code rate in image encoding method and adaptive adjustment of image quality are achieved, and the problem of inability to adjust compression code rate in the prior art is solved, and encoding efficiency and image reconstruction quality are improved.

WO2025152884A1PCT designated stage expired Publication Date: 2025-07-24HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/071985
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-13
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The existing deep learning-based image encoding methods cannot flexibly adjust the compression code rate according to actual needs, resulting in the inability to meet the compression needs of different application scenarios.

Method used

Through feature extraction networks and feature inverse transformation networks, neural networks are used to scale and encode feature maps according to quality factors to achieve flexible control of compression code rates. Combined with progressive decoding technology, it supports step-by-step reconstruction of images under low bandwidth conditions.

Benefits of technology

Adaptive adjustment of image quality at different code rates is realized, encoding efficiency is improved, storage space occupied by the model is reduced, and high image reconstruction quality and user experience are supported at low code rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071985_24072025_PF_FP_ABST
    Figure CN2025071985_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Image processing methods, an image processing method comprising: acquiring an image and a quality factor, the quality factor being related to the compression bitrate of the image; on the basis of the quality factor, obtaining first scaling information by means of a first neural network; performing feature extraction on the image by means of a feature extraction network, so as to obtain a target feature map, wherein the feature extraction network comprises a first feature extraction layer, and the first feature extraction layer is used for performing, by means of the first scaling information, scaling on feature values comprised in a first feature map obtained by means of feature extraction, so as to obtain a second feature map; and encoding the target feature map, so as to obtain a bitstream. In the present application, the distribution of M feature values comprised in the first feature map will change due to the scaling operation performed thereon, so that the compression bitrate of the bitstream obtained by encoding the subsequently obtained feature map is adapted to a required compression bitrate, thus achieving compression bitrate control.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method and related device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 19, 2024, with application number 202410084795.2 and application name “An image processing method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to an image processing method and related equipment. Background Art

[0003] Multimedia data now accounts for the vast majority of Internet traffic. Image data compression plays a crucial role in its storage and efficient transmission. Therefore, image coding is a technology with significant practical value.

[0004] Image coding research has a long history. Researchers have proposed numerous methods and established numerous international standards, such as JPEG, JPEG2000, WebP, and BPG. While these methods are widely used, they are currently experiencing limitations in response to the ever-increasing amount of image data and the emergence of new media types.

[0005] In recent years, researchers have begun to investigate image coding methods based on deep learning. Some have achieved impressive results. For example, Ballé et al. proposed an end-to-end optimized image coding method that surpasses the current state-of-the-art image coding performance, even surpassing the current state-of-the-art traditional coding standard, BPG. However, most current image coding methods based on deep convolutional networks suffer from a flaw: a trained model can only output a single encoding result for a given input image, and cannot achieve the desired compression rate required for practical applications. Summary of the Invention

[0006] In a first aspect, the present application provides an image processing method, the method comprising: obtaining an image and a quality factor, wherein the quality factor is related to the compression bit rate of the image; obtaining first scaling information through a first neural network based on the quality factor; performing feature extraction on the image through a feature extraction network to obtain a target feature map; wherein the feature extraction network includes a first feature extraction layer, the first feature extraction layer being used to scale the feature values ​​included in the first feature map obtained by feature extraction by the first scaling information to obtain a second feature map; and encoding the target feature map to obtain a code stream. In an embodiment of the present application, the distribution of the M feature values ​​included in the first feature map will change due to the scaling operation performed therein, and the compression bit rate of the code stream obtained by encoding the subsequently obtained feature map is adapted to the required compression bit rate, thereby achieving compression bit rate control.

[0007] In one possible implementation, the first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature extraction layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0008] In one possible implementation, the method further includes: obtaining second scaling information through a second neural network based on the quality factor; the first scaling information and the second scaling information are different; the feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is used to scale the feature values ​​included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

[0009] In the embodiment of the present application, for feature maps obtained by different feature extraction layers (or at least two feature extraction layers) in the feature extraction network, the quality factor can be mapped to different scaling information through different neural networks. The resulting scaling information can be adaptively made to have the same size as the corresponding feature map.

[0010] In one possible implementation, the scaling is achieved by a product operation.

[0011] In a possible implementation, the method further includes: obtaining a compressed file of the image according to the code stream and the quality factor.

[0012] In a possible implementation, the quality factor is selected from a plurality of candidate quality factors according to the compression code rate.

[0013] In a second aspect, the present application provides an image processing method, the method comprising: obtaining a compressed file, the compressed file comprising a code stream and a quality factor; obtaining first scaling information through a first neural network based on the quality factor; decoding the code stream to obtain a second feature map; reconstructing the second feature map through a feature inverse transformation network to obtain an image; wherein the feature inverse transformation network comprises a first feature inverse transformation layer, the first feature inverse transformation layer being used to scale the feature values ​​included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

[0014] In one possible implementation, the first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature inverse transformation layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0015] In one possible implementation, the method further includes: obtaining second scaling information through a second neural network based on the quality factor; the first scaling information and the second scaling information are different; the feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is used to scale the eigenvalues ​​included in the third feature map obtained by the feature inverse transformation through the second scaling information to obtain a fourth feature map.

[0016] In one possible implementation, the compressed file further includes prior information, and the code stream includes multiple sub-code streams; decoding the code stream to obtain a second feature map includes: obtaining statistical information of the code stream based on the prior information, the statistical information including a mean; and based on the absence of a target segment code stream in the multiple segment code streams, using the mean of the target segment code stream as the feature map of the target segment code stream in the second feature map.

[0017] In this way, when some code stream segments are missing in the compressed file, the mean value in the statistical information can still be reused as the feature map, thereby ensuring the normal progress of the decoding process.

[0018] In a third aspect, the present application provides an image processing device, comprising:

[0019] An acquisition module, configured to acquire an image and a quality factor, wherein the quality factor is related to a compression bit rate of the image;

[0020] A processing module is used to obtain first scaling information through a first neural network based on the quality factor; perform feature extraction on the image through a feature extraction network to obtain a target feature map; wherein the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the feature values ​​included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; and encode the target feature map to obtain a code stream.

[0021] In one possible implementation, the first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature extraction layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0022] In a possible implementation, the processing module is further configured to:

[0023] obtaining, through a second neural network, second scaling information based on the quality factor, wherein the first scaling information and the second scaling information are different;

[0024] The feature extraction network also includes a second feature extraction layer, which is used to scale the feature values ​​included in the third feature map obtained by feature extraction using the second scaling information to obtain a fourth feature map.

[0025] In one possible implementation, the scaling is achieved by a product operation.

[0026] In a possible implementation, the processing module is further configured to:

[0027] A compressed file of the image is obtained according to the code stream and the quality factor.

[0028] In a possible implementation, the quality factor is selected from a plurality of candidate quality factors according to the compression code rate.

[0029] In a fourth aspect, the present application provides an image processing device, comprising:

[0030] An acquisition module, configured to acquire a compressed file, wherein the compressed file includes a code stream and a quality factor;

[0031] A processing module is used to obtain first scaling information through a first neural network based on the quality factor; decode the code stream to obtain a second feature map; and reconstruct the second feature map through a feature inverse transformation network to obtain an image; wherein the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is used to scale the eigenvalues ​​included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

[0032] In one possible implementation, the first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature inverse transformation layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0033] In a possible implementation, the processing module is further configured to:

[0034] obtaining, through a second neural network, second scaling information based on the quality factor, wherein the first scaling information and the second scaling information are different;

[0035] The feature inverse transformation network also includes a second feature inverse transformation layer, which is used to scale the feature values ​​included in the third feature map obtained by feature inverse transformation using the second scaling information to obtain a fourth feature map.

[0036] In a possible implementation, the compressed file further includes prior information, and the processing module is specifically configured to:

[0037] Obtaining statistical information of the code stream according to the prior information, wherein the statistical information includes a mean;

[0038] Based on the lack of a target segment code stream in the multiple segment code streams, an average value of the target segment code stream is used as a feature map of the target segment code stream in the second feature map.

[0039] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the image processing method described in any one of the first to second aspects above.

[0040] In a sixth aspect, an embodiment of the present application provides a computer program, which, when executed on a computer, enables the computer to execute the image processing method described in any one of the first to second aspects above.

[0041] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting an execution device or a training device to implement the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the execution device or training device. The chip system can be composed of a chip or can include a chip and other discrete devices.

[0042] In an eighth aspect of the present application, a data processing device is provided. The device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is configured to store instructions, the processor is configured to execute the instructions, and the communication interface is configured to communicate with other network elements under the control of the processor. When executed by the processor, the instructions cause the processor to perform the image processing method described in any one of the first and second aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic diagram of the structure of the artificial intelligence main framework;

[0044] FIG2a is a schematic diagram of an application scenario of an embodiment of the present application;

[0045] FIG2 b is a schematic diagram of an application scenario of an embodiment of the present application;

[0046] FIG3 is a schematic diagram of an embodiment of an image processing method provided in an embodiment of the present application;

[0047] FIG4 is a schematic diagram of an embodiment of an image processing method provided in an embodiment of the present application;

[0048] FIG5 is a schematic diagram of a process for scaling a feature map provided by an embodiment of the present application;

[0049] FIG6 is a diagram illustrating a beneficial effect of an embodiment of the present application;

[0050] FIG7 is a diagram illustrating a beneficial effect of an embodiment of the present application;

[0051] FIG8 is a schematic diagram of the progressive image reconstruction result;

[0052] FIG9 is a flowchart illustrating a decoding process;

[0053] FIG10 is a flowchart illustrating progressive decoding;

[0054] FIG11A is a schematic diagram of an image processing process according to an embodiment of the present application;

[0055] FIG11B is a schematic diagram of a training process according to an embodiment of the present application;

[0056] FIG11C is a schematic diagram of an image processing process according to an embodiment of the present application;

[0057] FIG12 is a schematic structural diagram of an image processing device provided in an embodiment of the present application;

[0058] FIG13 is a schematic structural diagram of an image processing device provided in an embodiment of the present application;

[0059] FIG14 is a schematic diagram of a structure of an execution device provided in an embodiment of the present application;

[0060] FIG15 is a schematic structural diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.

[0062] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0063] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0064] First, let's describe the overall workflow of an AI system. See Figure 1, which shows a schematic diagram of the AI ​​framework. This framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0065] (1) Infrastructure

[0066] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.

[0067] (2) Data

[0068] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0069] (3) Data processing

[0070] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0071] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0072] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0073] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0074] (4) General ability

[0075] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0076] (5) Smart products and industry applications

[0077] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.

[0078] This application can be applied to the field of image processing in the field of artificial intelligence. The following will introduce multiple application scenarios of multiple products.

[0079] 1. Image Compression Process Applied to Terminal Devices

[0080] The image compression method provided in the embodiment of the present application can be applied to the image compression process in the terminal device, specifically, it can be applied to the photo album, video surveillance, etc. on the terminal device. Specifically, reference can be made to FIG2a, which is a schematic diagram of the application scenario of the embodiment of the present application. As shown in FIG2a, the terminal device can obtain a picture to be compressed, wherein the picture to be compressed can be a photo taken by a camera or a frame captured from a video. The terminal device can extract features of the obtained picture to be compressed through an artificial intelligence (AI) coding unit in an embedded neural network (neural-network processing unit, NPU), transform the image data into output features with lower redundancy, and generate a probability estimate of each point in the output feature. The central processing unit (CPU) performs arithmetic coding on the extracted output features based on the probability estimate of each point in the output feature, reduces the coding redundancy of the output feature, further reduces the amount of data transmitted during the image compression process, and saves the encoded data in the form of a data file in a corresponding storage location. When the user needs to obtain the file saved in the above storage location, the CPU can obtain and load the above saved file in the corresponding storage location, and obtain the decoded feature map based on arithmetic decoding, and reconstruct the feature map through the AI ​​decoding unit in the NPU to obtain a reconstructed image.

[0081] 2. Image Compression Process on the Cloud Side

[0082] The image compression method provided in the embodiment of the present application can be applied to the image compression process on the cloud side, specifically, it can be applied to functions such as cloud albums on the cloud side server. Specifically, reference can be made to FIG2b, which is a schematic diagram of the application scenario of the embodiment of the present application. As shown in FIG2b, the terminal device can obtain a picture to be compressed, wherein the picture to be compressed can be a photo taken by a camera or a frame captured from a video. The terminal device can use the CPU to perform lossless coding compression on the picture to be compressed to obtain coded data, for example but not limited to any lossless compression method based on the prior art. The terminal device can transmit the coded data to the server on the cloud side. The server can perform corresponding lossless decoding on the received coded data to obtain the image to be compressed. The server can use the AI ​​coding unit in the graphics processing unit (GPU) to extract features of the obtained picture to be compressed, transform the image data into output features with lower redundancy, and generate a probability estimate of each point in the output feature. The CPU performs arithmetic coding on the extracted output features based on the probability estimate of each point in the output feature, reduces the coding redundancy of the output feature, further reduces the amount of data transmitted during the image compression process, and saves the coded data obtained in the form of a data file in the corresponding storage location. When the user needs to obtain the file saved in the above storage location, the CPU can obtain and load the above saved file in the corresponding storage location, and obtain the decoded feature map based on arithmetic decoding, and reconstruct the feature map through the AI ​​decoding unit in the NPU to obtain a reconstructed image. The server can use the CPU to perform lossless encoding on the compressed image to obtain encoded data, such as but not limited to any lossless compression method based on the existing technology. The server can transmit the encoded data to the terminal device, and the terminal device can perform corresponding lossless decoding on the received encoded data to obtain the decoded image.

[0083] In an embodiment of the present application, a step of gaining the eigenvalues ​​in the feature map can be added between the AI ​​encoding unit and the quantization unit, and a step of inversely gaining the eigenvalues ​​in the feature map can be added between the arithmetic decoding and the AI ​​decoding unit. The image processing method in the embodiment of the present application will be described in detail below.

[0084] Since the embodiments of the present application involve the application of a large number of neural networks, for ease of understanding, the relevant terms and concepts of the neural networks that may be involved in the embodiments of the present application are first introduced below.

[0085] (1) Neural Network

[0086] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be:

[0087] Where s = 1, 2, ..., n, where n is a natural number greater than 1, Ws is the weight of Xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0088] (2) Deep Neural Networks

[0089] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. Based on the location of the different layers, the neural network within a DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. Each layer is fully connected, meaning that any neuron in layer i is connected to any neuron in layer i+1.

[0090] Although DNN looks complicated, the work of each layer is actually not complicated. In simple terms, it can be expressed as the following linear relationship: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since DNN has many layers, the coefficient W and the offset vector The number of these parameters is also relatively large. The definitions of these parameters in DNN are as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the third layer index 2 of the output and the second layer index 4 of the input.

[0091] In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as

[0092] It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).

[0093] (3) Convolutional Neural Networks

[0094] A convolutional neural network (CNN) is a deep neural network with a convolutional architecture. It consists of a feature extractor consisting of convolutional layers and subsampling layers, which can be viewed as a filter. A convolutional layer is a layer of neurons that performs convolution processing on the input signal. In a convolutional layer of a CNN, a neuron can only connect to a subset of neurons in adjacent layers. A convolutional layer typically contains several feature planes, each of which can be composed of a rectangular arrangement of neurons. Neurons in the same feature plane share weights, which are referred to as convolution kernels. Shared weights can be understood as extracting image information in a position-independent manner. Convolution kernels can be initialized as matrices of random size, and during CNN training, they can learn to acquire reasonable weights. Furthermore, shared weights have the direct benefit of reducing the number of connections between layers of the CNN, thereby reducing the risk of overfitting.

[0095] (4) Loss function

[0096] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.

[0097] (5) Backpropagation algorithm

[0098] Neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial neural network model during training, reducing the reconstruction error loss of the neural network model. Specifically, the forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial neural network model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0099] Next, the data processing method in the embodiment of the present application will be described from the encoding side. Referring to FIG. 3 , FIG. 3 is a schematic diagram of an embodiment of an image processing method provided by the embodiment of the present application. As shown in FIG. 3 , the image processing method provided by the embodiment of the present application includes:

[0100] 301. Acquire an image and a quality factor, where the quality factor is related to a compression bit rate of the image;

[0101] In the embodiments of the present application, the image is an image to be compressed, wherein the image may be an image captured by the camera of the terminal device, or the image may be an image obtained from within the terminal device (for example, an image stored in the photo album of the terminal device, or a picture obtained by the terminal device from the cloud). It should be understood that the above-mentioned image may be an image requiring image compression, and the present application does not impose any limitation on the source of the image to be processed.

[0102] In an embodiment of the present application, the compression bit rate of the image can be obtained, where the compression bit rate can be specified by the user or determined by the terminal device based on the image, which is not limited here.

[0103] In one possible implementation, a quality factor can be determined based on the compression bit rate. The quality factor can be mapped to scaling information for scaling the feature map of the image. The value of the quality factor can affect the subsequent compression bit rate. Therefore, it is necessary to determine a quality factor that can make the subsequent compression bit rate equal to the required compression bit rate.

[0104] In a possible implementation, the quality factor may be selected within a certain range, for example, from a plurality of candidate quality factors. For example, the candidate quality factors may be the following set: {0.5, 0.925, 1.35, . . . , 9}.

[0105] When selecting, you can first use the maximum and minimum quality factors to pre-compress the image to obtain the bit rate, compare the bit rate with the required bit rate, and then select a quality factor between the maximum and minimum quality factors that can make the bit rate of the pre-compression result closer to the required bit rate. Repeat the above steps until an optimal quality factor is selected, or repeat until the preset number of iterations is reached.

[0106] In a possible implementation, after the image is acquired, the image may be converted into a bitmap.

[0107] 302. Obtain first scaling information through a first neural network according to the quality factor.

[0108] 303. Perform feature extraction on the image through a feature extraction network to obtain a target feature map; wherein the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the feature values ​​included in the first feature map obtained by feature extraction using the first scaling information to obtain a second feature map;

[0109] In one possible implementation, the quality factor can be mapped to scaling information for scaling the feature map of the image through a neural network (such as the first neural network in the embodiment of the present application).

[0110] In an embodiment of the present application, optionally, feature extraction can be performed on the image based on a feature extraction network to obtain a target feature map. In the following, the feature map may also be referred to as a channel feature map image.

[0111] For example, the feature extraction network can be a CNN, which can multiply the upper left 3×3 pixels of the input data (image) by weights and map them to the neurons at the upper left end of the feature map. The weights to be multiplied will also be 3×3. Thereafter, in the same process, the CNN scans the input data (image) one by one from left to right and from top to bottom, and multiplies the weights to map the neurons of the feature map. Here, the 3×3 weights used are called filters or filter kernels. That is, the process of applying a filter in a CNN is the process of performing a convolution operation using a filter kernel, and the extracted result is called a "feature map", wherein the feature map can also be called a multi-channel feature map image, and the term "multi-channel feature map image" can refer to a set of feature map images corresponding to multiple channels. According to an embodiment, a multi-channel feature map image can be generated by a CNN, which is also referred to as a "feature extraction layer" or "convolution layer" of a CNN. The layer of a CNN can define a mapping from output to input. The mapping defined by the layer is executed as one or more filter kernels (convolution kernels) to be applied to the input data to generate a feature map image to be output to the next layer. The input data can be an image or a feature map image of a specific layer.

[0112] During the forward execution, the CNN receives the image and generates a multi-channel feature map image as output (that is, the target feature map in the embodiment of the present application). In addition, during the forward execution, the next layer receives the multi-channel feature map image as input and generates a multi-channel feature map image as output. Then, each subsequent layer will receive the multi-channel feature map image generated in the previous layer and generate the next multi-channel feature map image as output. Finally, by receiving the multi-channel feature map image generated in the (N)th layer.

[0113] At the same time, in addition to applying the convolution kernel operation that maps the input feature map image to the output feature map image, other processing operations may also be performed. Examples of other processing operations may include, but are not limited to, application of activation functions, pooling, resampling, etc.

[0114] It should be noted that the above is only one implementation method of feature extraction for the image, and in practical applications, the specific implementation method of feature extraction is not limited.

[0115] Among them, the feature extraction network may include multiple feature extraction layers (for example, feature extraction layers connected in series), the feature extraction layer can obtain the feature map output by the adjacent previous layer, and perform feature extraction operations to obtain the feature map and input it into the adjacent next feature extraction layer. For example, the multiple feature extraction layers may include a first feature extraction layer and a second feature extraction layer. The first feature extraction layer and the second feature extraction layer are connected, and the feature map obtained by the first feature extraction layer can be input into the second feature extraction layer.

[0116] The feature map may include feature maps of multiple channels, and the feature map of each channel may include multiple feature values.

[0117] In one possible implementation, the first scaling information includes M scaling factors; the first feature map obtained by the first feature extraction network includes M eigenvalues, each scaling factor corresponding to an eigenvalue; and the first feature extraction layer, in addition to obtaining the feature map, may further scale each of the M eigenvalues ​​by the corresponding scaling factor. For example, the scaling operation may be implemented by a multiplication operation.

[0118] The distribution of the M eigenvalues ​​included in the first feature map will change due to the scaling operation performed therein, and the compression code rate of the code stream obtained by encoding the subsequent feature map will be adapted to the required compression code rate, thereby achieving compression code rate control.

[0119] It should be understood that the feature values ​​within the same channel of the first feature map may share the same scaling factor.

[0120] In the embodiment of the present application, for feature maps obtained by different feature extraction layers (or at least two feature extraction layers) in the feature extraction network, the quality factor can be mapped to different scaling information through different neural networks. The resulting scaling information can be adaptively made to have the same size as the corresponding feature map.

[0121] In one possible implementation, second scaling information can be obtained through a second neural network based on the quality factor; the first scaling information and the second scaling information are different; the feature extraction network also includes a second feature extraction layer, and the second feature extraction layer is used to scale the feature values ​​included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

[0122] For example, referring to Figure 5, Figure 5 is a schematic diagram of a process for scaling a feature map provided by an embodiment of the present application, wherein the feature extraction layer 1 can be the feature extraction layer closest to the input image, and the feature extraction layer 1 can perform feature extraction on the image to obtain a first-layer feature map y1. Based on the quality factor, an adaptive channel scaling table can be obtained (the size of the scaling table is consistent with the first-layer feature map y1), and the scaled features can be obtained by element-by-element multiplication. The scaled features can be input into the feature extraction layer 2 to obtain a second-layer feature map y2. Based on the quality factor, an adaptive channel scaling table can be obtained (the size of the scaling table is consistent with the second-layer feature map y2), and the scaled features can be obtained by element-by-element multiplication.

[0123] The above-mentioned scaling information obtained based on the quality factor and the scaling process based on the scaling information can be applied to one or more of the multiple feature extraction layers, which is not limited in the embodiments of the present application.

[0124] 304. Encode the target feature map to obtain a code stream.

[0125] For example, the target feature map may be quantized and entropy encoded to obtain a bitstream.

[0126] In a possible implementation, a compressed file of the image may be obtained according to the code stream and the quality factor.

[0127] For example, the target feature map can be segmented into four non-uniform feature sub-maps y1, y2, y3, and y4, with sizes C1*H*W, C2*H*W, C3*H*W, and C4*H*W, respectively, and C1+C2+C3+C4=C. For example, C=64, C1=8, C2=8, C3=16, and C4=32.

[0128] For example, when encoding the target feature map, the following steps can be performed:

[0129] 1. Context modeling: The four feature subgraphs are passed into the context prediction module to obtain the statistical information required for encoding the four feature subgraphs.

[0130] 2. Feature quantization: quantize the four feature subgraphs into integers.

[0131] 3. Entropy Coding and Data Packaging: Based on the four characteristic subgraphs and their corresponding statistical information, entropy coding is performed in the order y1, y2, y3, and y4 to produce four compressed bitstreams s1, s2, s3, and s4. These s1, s2, s3, and s4 are then concatenated into a single bitstream. The lengths of the four bitstreams and the quality factor entered by the user in step 2 are then written into the header, and the header information is saved. Finally, the bitstreams and header information are packaged into a compressed file.

[0132] In extremely low bitrate scenarios, the reconstructed image quality of traditional codecs is poor. Deep learning-based methods repeatedly train multiple independent models to control the bitrate, which is relatively inefficient in terms of time and resources. This embodiment solves the first problem by using a deep learning architecture to ensure good image reconstruction quality at different bitrates. As for the second problem, this example introduces a quality factor adjustment module in the feature extraction process, scales the feature map layer by layer, and controls the amount of information in the feature map, thereby achieving the effect of single-model bitrate control and solving the second problem.

[0133] In addition, the embodiment of the present application can allow users to adjust the required bit rate according to their needs. For the transmission requirement of a single image compression of less than 800 bytes, the embodiment of the present application can achieve an average of 2.72 adjustments to achieve a specific bit rate, and the encoding time on the CPU is <500ms. The ROM required for the deployment of the embodiment of the present application in the user terminal is low, ROM <40MB. Under low bit rate conditions, the image reconstruction effect is better than similar deep learning algorithms. As shown in Figure 6, based on the test data set Kodak 24 test pictures, LPIPS is used as the reconstructed image quality indicator (the lower the LPIPS, the better the reconstruction quality). At the same bit rate, the LPIPS of this solution is better than the existing solution.

[0134] Next, the data processing method in the embodiment of the present application will be described from the decoding side. Referring to FIG. 4 , FIG. 4 is a schematic diagram of an embodiment of an image processing method provided by the embodiment of the present application. As shown in FIG. 4 , the image processing method provided by the embodiment of the present application includes:

[0135] 401. Obtain a compressed file, where the compressed file includes a bitstream and a quality factor.

[0136] When a compressed file is obtained, the compressed file can be read. The file may include two parts: one part is the compressed code stream, and the other part is header information, which includes prior information, quality factor, and other information (such as the length information of each segmented code stream), etc., which is not limited in the embodiments of this application.

[0137] 402. Obtain first scaling information through a first neural network according to the quality factor.

[0138] 403. Decode the code stream to obtain a second feature map;

[0139] In one possible implementation, the first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature inverse transformation layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0140] In one possible implementation, the compressed file further includes prior information, and the code stream includes multiple sub-code streams; decoding the code stream to obtain a second feature map may include: obtaining statistical information of the code stream based on the prior information, the statistical information including a mean; and based on the absence of a target segment code stream in the multiple segment code streams, using the mean of the target segment code stream as the feature map of the target segment code stream in the second feature map.

[0141] When the bitstream is segmented, during the decompression phase, embodiments of the present application can perform progressive decoding to optimize bandwidth utilization and user experience. Specifically, the encoded image data allows the user to reconstruct the image in stages at the receiving end, thereby achieving progressive display from low quality to high quality. Referring to FIG10 , FIG10 illustrates a process flow for progressive decoding.

[0142] Specifically, the received code stream can be split according to the length of each segment parsed from the header information, and progressive decoding can be performed. The following are the steps for progressive decoding of each segment of the code stream:

[0143] Referring to Figure 11A, the model prior information is input into context prediction module 1, which predicts the statistical information of the first feature subgraph y1: mean m1 and scale parameter scale1. Using statistical information scale1, residual r1 is decoded from bitstream s1 (e.g., through entropy decoding), and y1 is obtained according to r1 + m1 = y1. Subsequently, y1 is passed to context prediction module 2 to obtain m2 and scale2. If the second bitstream s2 has been received, residual r2 is decoded from bitstream s2 by scale2, and r2 + m2 = y2 to obtain y2. If the second bitstream has not been received, r2 cannot be decompressed, and m2 = y2.

[0144] In this way, when some code stream segments are missing in the compressed file, the mean value in the statistical information can still be reused as the feature map, thereby ensuring the normal progress of the decoding process.

[0145] 404. Reconstruct the second feature map through a feature inverse transformation network to obtain an image; wherein the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is used to scale the eigenvalues ​​included in the second feature map obtained by the feature inverse transformation through the first scaling information to obtain a first feature map.

[0146] In one possible implementation, second scaling information can also be obtained through a second neural network based on the quality factor; the first scaling information and the second scaling information are different; the feature inverse transformation network also includes a second feature inverse transformation layer, and the second feature inverse transformation layer is used to scale the eigenvalues ​​included in the third feature map obtained by the feature inverse transformation through the second scaling information to obtain a fourth feature map.

[0147] The decoded feature map can be inverse-transformed layer by layer to obtain the reconstructed image. The quality factor parsed from the header is passed to the quality factor module to generate a scaling table with the same size as the feature map. The scaling table is then multiplied element-by-element with the feature map to produce the scaled feature map. The scaled feature map is then passed to the feature inverse-transformation module of the next layer, where it is combined with the scaling table output by the quality factor module of the next layer to produce the scaled feature map of the next layer. Refer to Figure 9, which illustrates a flowchart of the decoding process.

[0148] Compared to traditional image codecs, this solution utilizes a deep learning architecture that adaptively learns the complex features and structure of images, maintaining high image quality at low bitrates. Compared to similar deep learning algorithms that train multiple models to control bitrate, this solution implements flexible bitrate control by designing an embedded quality factor module within a single model, reducing the model's ROM usage and improving encoding efficiency. Compared to similar deep learning algorithms, this solution utilizes feature map channel segmentation and a progressive decoding process to support progressive image reconstruction.

[0149] The quality factor adjustment module and multi-bitrate adjustment loss function enable precise control of multiple bitrates within a single model, significantly improving encoding efficiency. This allows users to adjust the desired bitrate based on their needs, achieving a specific bitrate with an average of 2.72 adjustments per second, while encoding time on the CPU is less than 500ms. Table 1 shows the statistical results of the compression times.

[0150] Table 1

[0151] The embodiment of this application requires minimal ROM space for deployment. Since this solution only requires a single model, the ROM space can be reduced to <40MB. Furthermore, the reconstruction quality is excellent at low bit rates: especially in extremely low bit rate scenarios, image quality superior to that of traditional encoders can be achieved.

[0152] Refer to Figure 7, which shows a comparison of visualization results of the embodiment of the present application, JPEG, and existing solutions.

[0153] Furthermore, the embodiments of the present application utilize multi-channel segmentation technology to support progressive image reconstruction, allowing images to be displayed in increasing quality from low to high. This feature is particularly beneficial in scenarios where network bandwidth is limited or images are loaded gradually, improving the user experience.

[0154] Refer to FIG8 , which is a schematic diagram of a progressive image reconstruction result.

[0155] Next, we will introduce an example of the training process of the embodiment of the present application, wherein a training data set may be obtained during training, for example, the data set may contain N pictures. Set a quality factor set: Based on experimental experience, for example, a quality factor set may be set. Contains 20 floating point numbers ranging from 0.5 to 9 with an average interval of 0.425. Assign quality factor: Randomly assign a quality factor θ_i to each data x_i in the training dataset, i = 1, 2, ..., N, θ_i is from the quality factor set Random sampling is obtained. The network inference obtains the reconstructed image: the image is passed to the above compression and decompression module to obtain the estimated bitstream length of the reconstructed image and the compressed file. The loss function for multi-bitrate adjustment is calculated by passing the obtained reconstructed image, the original image in the dataset, and the estimated bitstream length of the compressed file into the multi-bitrate conditional loss function for calculation. For example, the loss function can be shown as follows:

[0156] The model updates the weights by backpropagating gradients based on the loss function.

[0157] For example, reference may be made to FIG11B , which is a flowchart illustrating a training process.

[0158] The embodiment of the present application can be applied to low-bandwidth communication scenarios. Referring to FIG11C , FIG11C is a flow chart of a low-bandwidth communication scenario. The entire process of this scenario mainly includes the following steps:

[0159] Data reading and compression: This step occurs at the sending user terminal. The user selects an image from the gallery, and the terminal device reads the image format into a bitmap. This bitmap is then fed into the image codec and compressed to a specific bitrate. This is where the present invention comes into play. Using a single-model, precise bitrate control technique based on deep learning, the bitmap is compressed to a predetermined low bitrate.

[0160] Data transmission: The compressed code stream is sent through the communication equipment of the user terminal.

[0161] Data transmission: The compressed bitstream is transmitted to the service provider's cloud server via a low-bandwidth network. This may be done wired or wirelessly, such as via telephone lines, satellite connections, or cellular networks.

[0162] Data reception: The cloud server receives the compressed code stream.

[0163] Data decompression: On the cloud server, the received data is decompressed into a bitmap and converted into a usable image format. The present invention supports progressive decoding at this stage, reconstructing images from low to high quality based on the completeness of the received data (one to four packets).

[0164] Data transmission: The reconstructed image is sent via the cloud server communication device.

[0165] Data transmission: The reconstructed image is transmitted to the receiving user terminal via a high-bandwidth network.

[0166] Data reception and use: The receiving user terminal receives the reconstructed image. The image can be used for the intended purpose.

[0167] Based on the embodiments corresponding to Figures 1 to 11C , in order to better implement the above-mentioned solutions of the embodiments of the present application, the following also provides related devices for implementing the above-mentioned solutions. Specifically, refer to Figure 12 , which is a schematic structural diagram of an image processing device 1200 provided in an embodiment of the present application. The image processing device 1200 can be a terminal device or a server, and the image processing device 1200 includes:

[0168] An acquisition module 1201 is configured to acquire an image and a quality factor, where the quality factor is related to a compression bit rate of the image;

[0169] The specific description of the acquisition module 1201 can refer to the introduction of step 301 in the above embodiment, and the similarities are not repeated here.

[0170] Processing module 1202 is used to obtain first scaling information through a first neural network based on the quality factor; perform feature extraction on the image through a feature extraction network to obtain a target feature map; wherein the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the feature values ​​included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; and encode the target feature map to obtain a code stream.

[0171] The specific description of the processing module 1202 can refer to the description of steps 302 and 303 in the above embodiment, and the similarities are not repeated here.

[0172] In one possible implementation, the first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature extraction layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0173] In a possible implementation, the processing module 1202 is further configured to:

[0174] obtaining, through a second neural network, second scaling information based on the quality factor, wherein the first scaling information and the second scaling information are different;

[0175] The feature extraction network also includes a second feature extraction layer, which is used to scale the feature values ​​included in the third feature map obtained by feature extraction using the second scaling information to obtain a fourth feature map.

[0176] In one possible implementation, the scaling is achieved by a product operation.

[0177] In a possible implementation, the processing module 1202 is further configured to:

[0178] A compressed file of the image is obtained according to the code stream and the quality factor.

[0179] In a possible implementation, the quality factor is selected from a plurality of candidate quality factors according to the compression code rate.

[0180] Referring to FIG. 13 , FIG. 13 is a schematic diagram of a structure of an image processing apparatus 1300 provided in an embodiment of the present application. The image processing apparatus 1300 may be a terminal device or a server. The image processing apparatus 1300 includes:

[0181] An acquisition module 1301 is configured to acquire a compressed file, where the compressed file includes a bitstream and a quality factor;

[0182] The specific description of the acquisition module 1301 can refer to the introduction of step 401 in the above embodiment, and the similarities are not repeated here.

[0183] Processing module 1302 is used to obtain first scaling information through a first neural network based on the quality factor; decode the code stream to obtain a second feature map; and reconstruct the second feature map through a feature inverse transformation network to obtain an image; wherein the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is used to scale the feature values ​​included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

[0184] The specific description of the acquisition module 1301 can refer to the introduction of steps 402 and 403 in the above embodiment, and the similarities are not repeated here.

[0185] In one possible implementation, the first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, each scaling coefficient corresponds to an eigenvalue, and the first feature inverse transformation layer is used to scale each of the M eigenvalues ​​by the corresponding scaling coefficient.

[0186] In a possible implementation, the processing module 1302 is further configured to:

[0187] obtaining, through a second neural network, second scaling information based on the quality factor, wherein the first scaling information and the second scaling information are different;

[0188] The feature inverse transformation network also includes a second feature inverse transformation layer, which is used to scale the feature values ​​included in the third feature map obtained by feature inverse transformation using the second scaling information to obtain a fourth feature map.

[0189] In a possible implementation, the compressed file further includes prior information, and the processing module 1302 is specifically configured to:

[0190] Obtaining statistical information of the code stream according to the prior information, wherein the statistical information includes a mean;

[0191] Based on the lack of a target segment code stream in the multiple segment code streams, an average value of the target segment code stream is used as a feature map of the target segment code stream in the second feature map.

[0192] Next, a data processing device provided in an embodiment of the present application is introduced. Please refer to Figure 14, which is a structural diagram of a data processing device provided in an embodiment of the present application. Specifically, the data processing device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403 and a memory 1404 (wherein the number of processors 1403 in the data processing device 1400 can be one or more, and Figure 14 takes one processor as an example), wherein the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403 and the memory 1404 may be connected via a bus or other means.

[0193] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0194] Processor 1403 controls the operation of the radar system (including the antenna, receiver 1401, and transmitter 1402). In specific applications, the various components of the radar system are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0195] The image processing method disclosed in the above-mentioned embodiments of the present application (shown in Figures 3 and 4) can be applied to the processor 1403 or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above-mentioned method can be completed by the hardware integrated logic circuit in the processor 1403 or by software instructions. The above-mentioned processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components. The processor 1403 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in memory 1404, and processor 1403 reads information from memory 1404 and, in conjunction with its hardware, completes the steps of the image processing method provided in the above embodiment.

[0196] Receiver 1401 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the radar system. Transmitter 1402 can be used to output digital or character information through the first interface; transmitter 1402 can also be used to send instructions to the disk pack through the first interface to modify the data in the disk pack.

[0197] An embodiment of the present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the image processing method described in the above embodiment.

[0198] An embodiment of the present application further provides a computer-readable storage medium, which stores a program for performing signal processing. When the computer-readable storage medium is run on a computer, the computer executes the image processing method described in the above embodiment.

[0199] The data processing device provided in the embodiment of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip in the execution device to execute the image enhancement method described in the above embodiment, or to enable the chip in the training device to execute the image enhancement method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0200] Specifically, see Figure 15 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. This chip can be represented as a neural network processor NPU 1500. NPU 1500 is mounted on the host CPU (host CPU) as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1503, which is controlled by controller 1504 to extract matrix data from memory and perform multiplication operations.

[0201] In some implementations, arithmetic circuit 1503 includes multiple processing units (PEs). In some implementations, arithmetic circuit 1503 is a two-dimensional systolic array. Arithmetic circuit 1503 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1503 is a general-purpose matrix processor.

[0202] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1501 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1508.

[0203] Unified memory 1506 is used to store input and output data. Weight data is directly transferred to weight memory 1502 through direct memory access controller (DMAC) 1505. Input data is also transferred to unified memory 1506 through DMAC.

[0204] BIU stands for Bus Interface Unit 1510 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1509 .

[0205] The bus interface unit 1510 (BIU) is used for the instruction fetch memory 1509 to obtain instructions from the external memory, and is also used for the storage unit access controller 1505 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0206] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1506 or transfer weight data to the weight memory 1502 or transfer input data to the input memory 1501.

[0207] The vector calculation unit 1507 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0208] In some implementations, the vector calculation unit 1507 can store the processed output vector to the unified memory 1506. For example, the vector calculation unit 1507 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1503, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1507 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1503, such as for use in subsequent layers in a neural network.

[0209] An instruction fetch buffer 1509 connected to the controller 1504 is used to store instructions used by the controller 1504;

[0210] Unified memory 1506, input memory 1501, weight memory 1502, and instruction fetch memory 1509 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0211] The processor mentioned in any of the above may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of programs related to the steps of the data processing method described in the above embodiments.

[0212] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0213] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the method of each embodiment of the present application.

[0214] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0215] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instruction can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website, a computer, a training device or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

Claims

1. An image processing method, characterized in that, The method includes: Obtaining an image and a quality factor, where the quality factor is related to the compression bitrate of the image; According to the quality factor, obtaining first scaling information through a first neural network; Performing feature extraction on the image through a feature extraction network to obtain a target feature map; wherein, the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the eigenvalues included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; Encoding the target feature map to obtain a bitstream.

2. The method according to claim 1, characterized in that, The first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, and each scaling coefficient corresponds to one eigenvalue, and the first feature extraction layer is used to scale each eigenvalue among the M eigenvalues through the corresponding scaling coefficient.

3. The method according to claim 1 or 2, characterized in that, The method further includes: According to the quality factor, obtaining second scaling information through a second neural network; the first scaling information and the second scaling information are different; The feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is used to scale the eigenvalues included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

4. The method according to any one of claims 1 to 3, characterized in that The scaling is implemented through a multiplication operation.

5. The method according to any one of claims 1 to 4, characterized in that The method further includes: According to the bitstream and the quality factor, obtaining a compressed file of the image.

6. The method according to any one of claims 1 to 5, characterized in that, The quality factor is selected from multiple candidate quality factors according to the compression bitrate.

7. An image processing method, characterized in that, The method includes: Obtaining a compressed file, where the compressed file includes a bitstream and a quality factor; According to the quality factor, obtaining first scaling information through a first neural network; Decoding the bitstream to obtain a second feature map; Performing image reconstruction on the second feature map through a feature inverse transformation network to obtain an image; wherein, the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is used to scale the eigenvalues included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

8. The method according to claim 7, characterized in that The first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, and each scaling coefficient corresponds to one eigenvalue, and the first feature inverse transformation layer is used to scale each eigenvalue among the M eigenvalues through the corresponding scaling coefficient.

9. The method according to claim 7 or 8, characterized in that, The method further includes: According to the quality factor, obtaining second scaling information through a second neural network; the first scaling information and the second scaling information are different; The feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is used to scale the eigenvalues included in the third feature map obtained by feature inverse transformation through the second scaling information to obtain a fourth feature map.

10. The method according to any one of claims 7 to 9, characterized in that, The compressed file further includes prior information, and the bitstream includes multiple sub-bitstreams; the decoding the bitstream to obtain a second feature map includes: According to the prior information, obtaining statistical information of the bitstream, where the statistical information includes a mean value; Based on the absence of the target segment stream in the multi-segment stream, the mean value of the target segment stream is used as the feature map of the target segment stream in the second feature map.

11. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire an image and a quality factor, where the quality factor is related to the compression bit rate of the image; A processing module, configured to obtain first scaling information through a first neural network according to the quality factor; perform feature extraction on the image through a feature extraction network to obtain a target feature map; where the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is configured to scale the feature values included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; encode the target feature map to obtain a stream.

12. The device according to claim 11, characterized in that, The first scaling information includes M scaling coefficients; the first feature map includes M feature values, and each scaling coefficient corresponds to a feature value, and the first feature extraction layer is configured to scale each feature value among the M feature values through the corresponding scaling coefficient.

13. The device according to claim 11 or 12, characterized in that, The processing module is further configured to: Obtain second scaling information through a second neural network according to the quality factor; the first scaling information and the second scaling information are different; The feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is configured to scale the feature values included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

14. The device according to any one of claims 11 to 13, characterized in that, The scaling is implemented through a multiplication operation.

15. The device according to any one of claims 11 to 14, characterized in that, The processing module is further configured to: Obtain a compressed file of the image according to the stream and the quality factor.

16. The device according to any one of claims 11 to 15, characterized in that The quality factor is selected from multiple candidate quality factors according to the compression bit rate.

17. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire a compressed file, where the compressed file includes a stream and a quality factor; A processing module, configured to obtain first scaling information through a first neural network according to the quality factor; decode the stream to obtain a second feature map; perform image reconstruction on the second feature map through a feature inverse transformation network to obtain an image; where the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is configured to scale the feature values included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

18. The device according to claim 17, characterized in that, The first scaling information includes M scaling coefficients; the second feature map includes M feature values, and each scaling coefficient corresponds to a feature value, and the first feature inverse transformation layer is configured to scale each feature value among the M feature values through the corresponding scaling coefficient.

19. The device according to claim 17 or 18, characterized in that, The processing module is further configured to: Obtain second scaling information through a second neural network according to the quality factor; the first scaling information and the second scaling information are different; The feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is configured to scale the feature values included in the third feature map obtained by feature inverse transformation through the second scaling information to obtain a fourth feature map.

20. The device according to any one of claims 17 to 19, characterized in that The compressed file further includes prior information, and the processing module is specifically configured to: According to the prior information, statistical information of the bitstream is obtained, and the statistical information includes a mean value; Based on the absence of the target segment bitstream in the multiple segment bitstreams, the mean value of the target segment bitstream is used as the feature map of the target segment bitstream in the second feature map.

21. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers are caused to perform the operations of the method according to any one of claims 1 to 10.

22. A computer program product, characterized in that, It includes computer-readable instructions, and when the computer-readable instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 10.

23. A system includes at least one processor and at least one memory; the processor and the memory are connected through a communication bus and communicate with each other; The at least one memory is used for storing code; The at least one processor is used for executing the code to execute the method according to any one of claims 1 to 10.

24. A chip, characterized in that, It includes at least one processing unit and an interface circuit, the interface circuit is used for providing program instructions or data for the at least one processing unit, and the at least one processing unit is used for executing the program instructions to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image processing method and related equipment

    CN120358361A

  • Image processing method and device, medium and computing equipment

    CN114022361A

  • Variable code rate video compression method, system and device and storage medium

    CN114501013A

  • Layered coding and decoding method and device

    CN114827622A

  • Coding and decoding method, device, equipment, storage medium and computer program product

    CN116778002A