Image processing method and related equipment

Through feature extraction and inverse transformation networks combined with neural network mapping quality factors, the problem that image encoding cannot flexibly adjust the compressed code rate in the prior art is solved, flexible code rate control and progressive image reconstruction are realized, and encoding efficiency and image quality are improved.

CN120358361APending Publication Date: 2025-07-22HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410084795.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing deep learning-based image encoding methods cannot flexibly adjust the compression code rate according to actual needs, resulting in the inability to meet the compression effect requirements of different users.

Method used

Through feature extraction networks and feature inverse transformation networks, the neural network maps the quality factor as scaling information, controls the scaling of the feature map, realizes the adaptation of the compressed code rate, and optimizes the image reconstruction process through progressive decoding technology on the decoding side.

Benefits of technology

It realizes flexible regulation of image compression code rate in a single model, improves coding efficiency, supports progressive image reconstruction, improves image quality and user experience, and performs excellently in low bandwidth conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358361A_ABST
    Figure CN120358361A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method, and the method comprises the steps: obtaining an image and a quality factor, and enabling the quality factor to be related to the compression code rate of the image; obtaining first scaling information through a first neural network according to the quality factor; performing feature extraction on the image through a feature extraction network to obtain a target feature map; wherein the feature extraction network comprises a first feature extraction layer, and the first feature extraction layer is used for scaling feature values included in a first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; and coding the target feature map to obtain a code stream. The distribution of the M characteristic value included in the first characteristic pattern can be changed due to the scaling operation, and the compression code rate of the code stream obtained by encoding the subsequently obtained characteristic pattern is adaptive to the required compression code rate, so that the control of the compression code rate is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to an image processing method and related devices. Background Art

[0002] Nowadays, multimedia data occupies the vast majority of the traffic on the Internet. The compression of image data plays a crucial role in the storage and efficient transmission of multimedia data. Therefore, image coding is a technology with great practical value.

[0003] The research on image coding has a long history. Researchers have proposed a large number of methods and formulated a variety of international standards, such as image coding standards like JPEG, JPEG2000, WebP, BPG, etc. Although these coding methods are widely used at present, for the increasing amount of image data and emerging new media types, these traditional methods show some limitations.

[0004] In recent years, some researchers have started to conduct research on image coding methods based on deep learning. Some researchers have achieved good results. For example, Ballé et al. proposed an end-to-end optimized image coding method, achieving better image coding performance than the current best, even surpassing the current best traditional coding standard BPG. However, most current image codings based on deep convolutional networks have a defect, that is, a trained model can only output one coding result for one input image, and cannot obtain the coding effect with the required compression rate according to actual needs. Summary of the Invention

[0005] In a first aspect, this application provides an image processing method, the method includes: obtaining an image and a quality factor, where the quality factor is related to the compression rate of the image; according to the quality factor, obtaining first scaling information through a first neural network; performing feature extraction on the image through a feature extraction network to obtain a target feature map; where the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the eigenvalue included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; encoding the target feature map to obtain a bitstream. In the embodiments of this application, the distribution of the M eigenvalues included in the first feature map will change due to the scaling operation performed therein, and enable the compression rate of the bitstream obtained by encoding the subsequent obtained feature map to adapt to the required compression rate, thereby realizing the control of the compression rate.

[0006] In a possible implementation, the first scaling information includes M scaling factors; the first feature map includes M feature values, and each scaling factor corresponds to one feature value. The first feature extraction layer is configured to scale each feature value among the M feature values by the corresponding scaling factor.

[0007] In a possible implementation, the method further includes: obtaining second scaling information through a second neural network according to the quality factor; the first scaling information is different from the second scaling information; the feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is configured to scale the feature values included in the third feature map obtained by feature extraction by the second scaling information to obtain a fourth feature map.

[0008] In the embodiments of the present application, for the feature maps obtained by different feature extraction layers (or at least two feature extraction layers) in the feature extraction network, different neural networks can map the quality factor to different scaling information. And the obtained scaling information can adaptively have the same size as the corresponding feature map.

[0009] In a possible implementation, the scaling is implemented through a multiplication operation.

[0010] In a possible implementation, the method further includes: obtaining a compressed file of the image according to the bitstream and the quality factor.

[0011] In a possible implementation, the quality factor is selected from multiple candidate quality factors according to the compression bit rate.

[0012] In a second aspect, the present application provides an image processing method, the method including: obtaining a compressed file, where the compressed file includes a bitstream and a quality factor; obtaining first scaling information through a first neural network according to the quality factor; decoding the bitstream to obtain a second feature map; performing image reconstruction on the second feature map through a feature inverse transformation network to obtain an image; where the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is configured to scale the feature values included in the second feature map obtained by feature inverse transformation by the first scaling information to obtain a first feature map.

[0013] In a possible implementation, the first scaling information includes M scaling factors; the second feature map includes M feature values, and each scaling factor corresponds to one feature value. The first feature inverse transformation layer is configured to scale each feature value among the M feature values by the corresponding scaling factor.

[0014] In a possible implementation, the method further includes: obtaining second scaling information through a second neural network according to the quality factor; the first scaling information is different from the second scaling information; the feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is configured to scale the eigenvalues included in the third feature map obtained by feature inverse transformation through the second scaling information to obtain a fourth feature map.

[0015] In a possible implementation, the compressed file further includes prior information, and the bitstream includes multiple sub-bitstreams; decoding the bitstream to obtain a second feature map includes: obtaining statistical information of the bitstream according to the prior information, where the statistical information includes an average value; based on the absence of a target sub-bitstream in the multiple sub-bitstreams, using the average value of the target sub-bitstream as the feature map of the target sub-bitstream in the second feature map.

[0016] In the above manner, when some bitstream segments in the compressed file are missing, the average value in the statistical information can still be reused as the feature map, ensuring the normal progress of the decoding process.

[0017] In a third aspect, the present application provides an image processing apparatus, where the apparatus includes:

[0018] An acquisition module, configured to acquire an image and a quality factor, where the quality factor is related to the compression bit rate of the image;

[0019] A processing module, configured to obtain first scaling information through a first neural network according to the quality factor; perform feature extraction on the image through a feature extraction network to obtain a target feature map; where the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is configured to scale the eigenvalues included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; and encode the target feature map to obtain a bitstream.

[0020] In a possible implementation, the first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, and each scaling coefficient corresponds to an eigenvalue, and the first feature extraction layer is configured to scale each eigenvalue in the M eigenvalues through the corresponding scaling coefficient.

[0021] In a possible implementation, the processing module is further configured to:

[0022] Obtain second scaling information through a second neural network according to the quality factor; the first scaling information is different from the second scaling information;

[0023] The feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is configured to scale the feature values included in the third feature map obtained by feature extraction by the second scaling information to obtain a fourth feature map.

[0024] In a possible implementation, the scaling is implemented by a multiplication operation.

[0025] In a possible implementation, the processing module is further configured to:

[0026] Obtain a compressed file of the image according to the bitstream and the quality factor.

[0027] In a possible implementation, the quality factor is selected from a plurality of candidate quality factors according to the compression bit rate.

[0028] In a fourth aspect, the present application provides an image processing apparatus, and the apparatus includes:

[0029] An acquisition module, configured to acquire a compressed file, where the compressed file includes a bitstream and a quality factor;

[0030] A processing module, configured to obtain first scaling information through a first neural network according to the quality factor; decode the bitstream to obtain a second feature map; perform image reconstruction on the second feature map through a feature inverse transformation network to obtain an image; where the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is configured to scale the feature values included in the second feature map obtained by feature inverse transformation by the first scaling information to obtain a first feature map.

[0031] In a possible implementation, the first scaling information includes M scaling coefficients; the second feature map includes M feature values, each scaling coefficient corresponds to a feature value, and the first feature inverse transformation layer is configured to scale each feature value among the M feature values by the corresponding scaling coefficient.

[0032] In a possible implementation, the processing module is further configured to:

[0033] Obtain second scaling information through a second neural network according to the quality factor; the first scaling information is different from the second scaling information;

[0034] The feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is configured to scale the feature values included in the third feature map obtained by feature inverse transformation by the second scaling information to obtain a fourth feature map.

[0035] In a possible implementation, the compressed file further includes prior information, and specifically, the processing module is configured to:

[0036] According to the prior information, statistical information of the bitstream is obtained, and the statistical information includes an average value;

[0037] Based on the absence of the target segment bitstream in the multiple segment bitstreams, the average value of the target segment bitstream is used as the feature map of the target segment bitstream in the second feature map.

[0038] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the computer is caused to execute the image processing method according to any one of the first aspect to the second aspect above.

[0039] In a sixth aspect, an embodiment of the present application provides a computer program. When the computer program runs on a computer, the computer is caused to execute the image processing method according to any one of the first aspect to the second aspect above.

[0040] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting an execution device or a training device to implement the functions involved in the above aspects. For example, the processor is used to send or process data and / or information involved in the above method. In a possible design, the chip system further includes a memory for storing necessary program instructions and data for the execution device or the training device. The chip system may be composed of chips or may include chips and other discrete devices.

[0041] In an eighth aspect of the present application, a data processing device is provided. The device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other network elements under the control of the processor. When the instructions are executed by the processor, the processor is caused to execute the image processing method according to any one of the first aspect to the second aspect above. Description of the Drawings

[0042] Figure 1 It is a schematic structural diagram of an artificial intelligence main framework;

[0043] Figure 2a It is a schematic diagram of an application scenario of an embodiment of the present application;

[0044] Figure 2b It is a schematic diagram of an application scenario of an embodiment of the present application;

[0045] Figure 3 It is a schematic diagram of an embodiment of an image processing method provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic diagram of an embodiment of an image processing method provided by an embodiment of the present application;

[0047] Figure 5 Schematic of a process for scaling a feature map provided by an embodiment of the present application;

[0048] Figure 6 Schematic of a beneficial effect of an embodiment of the present application;

[0049] Figure 7 Schematic of a beneficial effect of an embodiment of the present application;

[0050] Figure 8 Schematic of the result of progressive image reconstruction;

[0051] Figure 9 Schematic of a process flow of the decoding process;

[0052] Figure 10 Schematic of the process flow of progressive decoding;

[0053] Figure 11A Schematic of an image processing process provided by an embodiment of the present application;

[0054] Figure 11B Schematic of a training process provided by an embodiment of the present application;

[0055] Figure 11C Schematic of an image processing process provided by an embodiment of the present application;

[0056] Figure 12 Schematic of a structure of an image processing apparatus provided by an embodiment of the present application;

[0057] Figure 13 Schematic of a structure of an image processing apparatus provided by an embodiment of the present application;

[0058] Figure 14 Schematic of a structure of an execution device provided by an embodiment of the present application;

[0059] Figure 15 Schematic of a structure of a chip provided by an embodiment of the present application. Detailed implementation manners

[0060] The embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. The terms used in the embodiments of the present invention are only for explaining the specific embodiments of the present invention and are not intended to limit the present invention.

[0061] The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0062] The terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0063] First, the overall workflow of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 which shows a schematic structural diagram of the artificial intelligence main framework. The above artificial intelligence theme framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.

[0064] (1) Infrastructure

[0065] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and is supported through the basic platform. Communicate with the outside through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.

[0066] (2) Data

[0067] The data on the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the perception data such as force, displacement, liquid level, temperature, humidity, etc.

[0068] (3) Data processing

[0069] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0070] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0071] Reasoning refers to the process of simulating human intelligent reasoning methods in a computer or intelligent system, based on reasoning control strategies, using formalized information for machine thinking and problem-solving. The typical function is search and matching.

[0072] Decision-making refers to the process of making decisions after intelligent information goes through reasoning, and usually provides functions such as classification, sorting, prediction, etc.

[0073] (4) General capabilities

[0074] After the data undergoes the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing, such as it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0075] (5) Intelligent products and industry applications

[0076] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and realize landing applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.

[0077] This application can be applied to the field of image processing in the field of artificial intelligence. Next, multiple application scenarios implemented in products will be introduced.

[0078] I. Image compression process applied to terminal devices

[0079] The image compression method provided by the embodiments of this application can be applied to the image compression process in terminal devices. Specifically, it can be applied to albums, video surveillance, etc. on terminal devices. Specifically, reference can be made to Figure 2a , Figure 2a which is the schematic diagram of the application scenario of the embodiments of this application. As shown in Figure 2aAs shown in , the terminal device can obtain the picture to be compressed, where the picture to be compressed can be a photo taken by a camera or a frame intercepted from a video. The terminal device can extract features from the obtained picture to be compressed through the artificial intelligence (AI) encoding unit in the neural-network processing unit (NPU), transform the image data into output features with lower redundancy, and generate probability estimates for each point in the output features. The central processing unit (CPU) performs arithmetic coding on the extracted output features based on the probability estimates for each point in the output features, reduces the coding redundancy of the output features, further reduces the data transmission volume during the image compression process, and saves the encoded data in the form of a data file at the corresponding storage location. When the user needs to obtain the file saved in the above storage location, the CPU can obtain and load the saved file at the corresponding storage location, obtain the decoded feature map based on arithmetic decoding, and reconstruct the feature map through the AI decoding unit in the NPU to obtain the reconstructed image.

[0080] II. Image Compression Process Applied to the Cloud Side

[0081] The image compression method provided by the embodiments of this application can be applied to the image compression process on the cloud side. Specifically, it can be applied to functions such as cloud albums on cloud side servers. Specifically, reference can be made to Figure 2b , Figure 2b which is a schematic diagram of the application scenario of the embodiments of this application. As shown in Figure 2bAs shown, the terminal device can obtain the picture to be compressed, where the picture to be compressed can be a photo taken by a camera or a frame intercepted from a video. The terminal device can perform lossless encoding and compression on the picture to be compressed through the CPU to obtain encoded data. For example, but not limited to, based on any one of the existing lossless compression methods, the terminal device can transmit the encoded data to the server on the cloud side. The server can perform corresponding lossless decoding on the received encoded data to obtain the picture to be compressed. The server can extract features from the picture to be compressed obtained through the AI encoding unit in the graphics processing unit (GPU), transform the image data into output features with lower redundancy, and generate probability estimates for each point in the output features. The CPU performs arithmetic encoding on the extracted output features based on the probability estimates of each point in the output features, reduces the encoding redundancy of the output features, further reduces the data transmission volume during the image compression process, and saves the encoded data in the corresponding storage location in the form of a data file. When the user needs to obtain the file saved in the above storage location, the CPU can obtain and load the saved file at the corresponding storage location, and obtain the decoded feature map based on arithmetic decoding. The CPU reconstructs the feature map through the AI decoding unit in the NPU to obtain the reconstructed image. The server can perform lossless encoding and compression on the picture to be compressed through the CPU to obtain encoded data. For example, but not limited to, based on any one of the existing lossless compression methods, the server can transmit the encoded data to the terminal device. The terminal device can perform corresponding lossless decoding on the received encoded data to obtain the decoded image.

[0082] In the embodiments of the present application, a step of increasing the gain of the feature values in the feature map can be added between the AI encoding unit and the quantization unit, and a step of performing anti-gain on the feature values in the feature map can be added between the arithmetic decoding and the AI decoding unit. Next, the image processing method in the embodiments of the present application will be described in detail.

[0083] Since the embodiments of the present application involve a large number of applications of neural networks, for the convenience of understanding, the relevant terms and concepts of the neural networks that may be involved in the embodiments of the present application will be introduced below.

[0084] (1) Neural network

[0085] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs and intercept 1 as inputs. The output of this operation unit can be:

[0086]

[0087] Among them, s = 1, 2, ……, n, where n is a natural number greater than 1, Ws is the weight of Xs, b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.

[0088] (2) Deep neural network

[0089] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer.

[0090] Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply operates on the input vector and obtains the output vector through such a simple operation. Since the DNN has many layers, the number of coefficients W and the offset vector is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer.

[0091] In summary, the coefficient from the k-th neuron in the L - 1-th layer to the j-th neuron in the L-th layer is defined as

[0092] It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically speaking, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).

[0093] (3) Convolutional Neural Network

[0094] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only some adjacent layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can obtain reasonable weights through learning. In addition, the direct benefit brought by sharing weights is to reduce the connections between layers of the convolutional neural network while reducing the risk of overfitting.

[0095] (4) Loss Function

[0096] During the process of training a deep neural network, since it is desired that the output of the deep neural network is as close as possible to the value that is truly wanted to be predicted, the weight vectors of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value, and then according to the difference between the two (of course, there is usually an initialization process before the first update, that is, configuring parameters for each layer in the deep neural network in advance). For example, if the predicted value of the network is high, the weight vector is adjusted to make it predict lower, and continuously adjusted until the deep neural network can predict the truly desired target value or a value very close to the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0097] (5) Backpropagation algorithm

[0098] The neural network can use the backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, forward propagating the input signal until the output will generate an error loss, and updating the parameters in the initial neural network model by backpropagating the error loss information, so as to make the error loss converge. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0099] Next, the data processing method in the embodiments of the present application will be introduced from the encoding side. Refer to Figure 3 , Figure 3 is a schematic diagram of an embodiment of an image processing method provided by the embodiments of the present application. As Figure 3 shown, an image processing method provided by the embodiments of the present application includes:

[0100] 301. Obtain an image and a quality factor, where the quality factor is related to the compression bit rate of the image;

[0101] In the embodiments of the present application, the image is an image to be compressed. Among them, the image can be an image captured by the above terminal device through a camera, or the image can also be an image obtained from inside the terminal device (for example, an image stored in the album of the terminal device, or a picture obtained by the terminal device from the cloud). It should be understood that the above image can be an image with an image compression requirement, and the present application does not make any limitation on the source of the image to be processed.

[0102] In the embodiments of the present application, the compression bit rate of an image can be obtained, where the compression bit rate can be specified by the user or determined by the terminal device based on the image, and this is not limited herein.

[0103] In a possible implementation, a quality factor can be determined according to the compression bit rate, and the quality factor can be mapped to scaling information for scaling the feature map of the image. The value of the quality factor can affect the subsequent compression bit rate. Therefore, a quality factor that can make the subsequent compression bit rate equal to the required compression bit rate needs to be determined.

[0104] In a possible implementation, the quality factor can be selected within a certain range. For example, it can be selected from multiple candidate quality factors. For example, the candidate quality factors can be the following set: {0.5, 0.925, 1.35, …, 9}.

[0105] When selecting, the maximum and minimum quality factors can be used first to perform pre-compression on the image to obtain a bit rate, and the obtained bit rate is compared with the required bit rate. Then, a quality factor that can make the bit rate of the pre-compression result closer to the required bit rate is selected between the maximum and minimum quality factors. The above steps are repeated until an optimal quality factor is selected or until a preset number of iterations is reached.

[0106] In a possible implementation, after the image is obtained, the image can be converted into a bitmap.

[0107] 302. Obtain first scaling information through a first neural network according to the quality factor.

[0108] 303. Extract features of the image through a feature extraction network to obtain a target feature map; wherein, the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the feature values included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map.

[0109] In a possible implementation, through a neural network (such as the first neural network in the embodiments of the present application), the quality factor can be mapped to scaling information for scaling the feature map of the image.

[0110] In the embodiments of the present application, optionally, the image can be subjected to feature extraction based on the feature extraction network to obtain a target feature map. Hereinafter, the feature map can also be referred to as a channel feature map image.

[0111] For example, the feature extraction network can be a CNN. The CNN can multiply the upper left 3×3 pixels of the input data (image) by weights and map them to the neurons at the upper left end of the feature map. The weights to be multiplied will also be 3×3. Thereafter, in the same process, the CNN scans the input data (image) one by one from left to right and from top to bottom, and multiplies by the weights to map the neurons of the feature map. Here, the 3×3 weights used are called filters or filter kernels. That is to say, the process of applying the filter in the CNN is the process of performing a convolution operation using the filter kernel, and the extracted result is called a "feature map". Among them, the feature map can also be called a multi-channel feature map image. The term "multi-channel feature map image" can refer to a set of feature map images corresponding to multiple channels. According to an embodiment, a multi-channel feature map image can be generated by the CNN, and the CNN is also called the "feature extraction layer" or "convolution layer" of the CNN. The layer of the CNN can define the mapping from the output to the input. The mapping defined by the layer is performed as one or more filter kernels (convolution kernels) to be applied to the input data to generate a feature map image to be output to the next layer. The input data can be an image or a feature map image of a specific layer.

[0112] During forward execution, the CNN receives an image and generates a multi-channel feature map image as output (that is, the target feature map in the embodiments of the present application). In addition, during forward execution, the next layer receives the multi-channel feature map image as input and generates a multi-channel feature map image as output. Then, each subsequent layer will receive the multi-channel feature map image generated in the previous layer and generate the next multi-channel feature map image as output. Finally, by receiving the multi-channel feature map image generated in the (N)th layer.

[0113] At the same time, in addition to the operation of applying the convolution kernel that maps the input feature map image to the output feature map image, other processing operations can also be performed. Examples of other processing operations can include but are not limited to the application of activation functions, pooling, resampling, etc.

[0114] It should be noted that the above is only one implementation manner for feature extraction of the image. In practical applications, the specific implementation manner of feature extraction is not limited.

[0115] Among them, the feature extraction network can include multiple feature extraction layers (for example, feature extraction layers connected in series). The feature extraction layer can obtain the feature map output by the adjacent previous layer and perform feature extraction operations to obtain a feature map and input it to the adjacent next feature extraction layer. For example, the multiple feature extraction layers can include a first feature extraction layer and a second feature extraction layer. The first feature extraction layer and the second feature extraction layer are connected, and the feature map obtained by the first feature extraction layer can be input into the second feature extraction layer.

[0116] The feature map may include feature maps of multiple channels, and each channel's feature map may include multiple feature values.

[0117] In a possible implementation, the first scaling information includes M scaling coefficients; the first feature map obtained by the first feature extraction network includes M feature values, and each of the scaling coefficients corresponds to a feature value. In addition to obtaining the feature map, the first feature extraction layer may scale each feature value among the M feature values by the corresponding scaling coefficient. For example, the scaling operation may be implemented through a multiplication operation.

[0118] The distribution of the M feature values included in the first feature map will change due to the scaling operation performed therein, and the compression bitrate of the bitstream obtained by encoding the subsequently obtained feature map will be adapted to the required compression bitrate, thereby achieving the control of the compression bitrate.

[0119] It should be understood that the feature values within the same channel of the first feature map may share the same scaling coefficient.

[0120] In the embodiments of the present application, for the feature maps obtained by different feature extraction layers (or at least two feature extraction layers) in the feature extraction network, different neural networks may map the quality factor to different scaling information. And the obtained scaling information may adaptively have the same size as the corresponding feature map.

[0121] In a possible implementation, the second scaling information may be obtained according to the quality factor through a second neural network; the first scaling information and the second scaling information are different; the feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is used to scale the feature values included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

[0122] For example, referring to Figure 5 , Figure 5 FIG. is a schematic diagram of a process for scaling a feature map provided by an embodiment of the present application. Among them, feature extraction layer 1 may be the feature extraction layer closest to the input image. Feature extraction layer 1 may perform feature extraction on the image to obtain the first-layer feature map y1. Based on the quality factor, an adaptive channel scaling table (the size of this scaling table is the same as that of the first-layer feature map y1) may be obtained, and the scaled feature may be obtained by using element-wise multiplication. This scaled feature may be input to feature extraction layer 2 to obtain the second-layer feature map y2. Based on the quality factor, an adaptive channel scaling table (the size of this scaling table is the same as that of the second-layer feature map y2) may be obtained, and the scaled feature may be obtained by using element-wise multiplication.

[0123] The above-mentioned process of obtaining scaling information based on the quality factor and performing scaling based on the scaling information can be applied to one or more of multiple feature extraction layers, which is not limited in the embodiments of this application.

[0124] 304. Encode the target feature map to obtain a bitstream.

[0125] For example, the target feature map can be quantized and entropy encoded to obtain a bitstream.

[0126] In a possible implementation, a compressed file of the image can be obtained according to the bitstream and the quality factor.

[0127] Exemplarily, the target feature map can be channel-split. For example, for a feature map with a size of C*H*W, where C, H, and W are the number of channels, height, and width respectively, the number of channels is split, and the feature map is split into four non-uniform feature submaps y1, y2, y3, y4 with sizes of C1*H*W, C2*H*W, C3*H*W, C4*H*W respectively, and C1 + C2 + C3 + C4 = C. For example: C = 64, C1 = 8, C2 = 8, C3 = 16, C4 = 32.

[0128] Exemplarily, when encoding the target feature map, the following steps can be performed:

[0129] 1. Context modeling: The four feature submaps are input into the context prediction module to obtain the statistical information required for encoding the four feature submaps respectively.

[0130] 2. Feature quantization: Quantize the four feature submaps into integers.

[0131] 3. Entropy encoding and data packaging: Based on the four feature submaps and their corresponding statistical information, entropy encoding is performed in the order of y1, y2, y3, y4 to obtain four compressed bitstreams s1, s2, s3, s4, and s1, s2, s3, s4 are concatenated into a single bitstream in sequence. Subsequently, the length information of the four bitstreams and the quality factor input by the user in step 2 are written in the header, and the header information is saved. Finally, the bitstream and the header information are packaged into a compressed file.

[0132] In the scenario of extremely low bitrates, the reconstructed image quality of traditional codecs is poor. The method based on deep learning retrains multiple independent models to control the bitrate, which is relatively inefficient in terms of time and resource consumption. In this embodiment, a deep learning architecture is used to ensure good image reconstruction quality at different bitrates, thus solving the first problem. Regarding the second problem, in this example, a quality factor adjustment module is introduced in the feature extraction process to perform layer-by-layer scaling on the feature maps and control the information volume of the feature maps, thereby achieving the effect of single-model bitrate control and solving the second problem.

[0133] In addition, the embodiments of the present application can allow users to adjust the required bitrate according to their needs. For the transmission requirement that the compression of a single image is less than 800 bytes, the embodiments of the present application can achieve an average of 2.72 times of regulation to reach a specific bitrate, and the encoding time on the CPU < 500 ms. The ROM required for the deployment of the embodiments of the present application on the user terminal is low, and the ROM < 40 MB. Under low bitrate conditions, the picture reconstruction effect is better than that of similar deep learning algorithms. As Figure 6 shown, based on 24 test pictures of the Kodak test dataset, with LPIPS as the reconstructed image quality index (the lower the LPIPS, the better the reconstruction quality), at the same bitrate, the LPIPS of this solution is better than that of the existing solutions.

[0134] Next, the data processing method in the embodiments of the present application will be introduced from the decoding side. Refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment of an image processing method provided by the embodiments of the present application. As Figure 4 shown, an image processing method provided by the embodiments of the present application includes:

[0135] 401. Obtain a compressed file, where the compressed file includes a bitstream and a quality factor;

[0136] When the compressed file is obtained, the compressed file can be read. The file can include two parts: one part is the bitstream after data compression, and the other part is the header information, including prior information, quality factor, and other information (such as the length information of each segmented bitstream), etc., which are not limited in the embodiments of the present application.

[0137] 402. According to the quality factor, obtain first scaling information through a first neural network;

[0138] 403. Decode the bitstream to obtain a second feature map;

[0139] In a possible implementation, the first scaling information includes M scaling factors; the second feature map includes M feature values, and each scaling factor corresponds to one feature value. The first inverse feature transformation layer is used to scale each feature value in the M feature values by the corresponding scaling factor.

[0140] In a possible implementation, the compressed file further includes prior information, and the bitstream includes multiple sub-bitstreams; when decoding the bitstream to obtain a second feature map, it may include: obtaining statistical information of the bitstream according to the prior information, where the statistical information includes a mean value; based on the absence of a target sub-bitstream in the multiple sub-bitstreams, using the mean value of the target sub-bitstream as the feature map of the target sub-bitstream in the second feature map.

[0141] In the case where the bitstream is segmented, in the decompression stage, the embodiments of the present application can perform progressive decoding to optimize bandwidth utilization and user experience. Specifically, the encoded image data allows users to reconstruct the image in a phased manner at the receiving end, thereby achieving progressive display from low quality to high quality. Refer to Figure 10 , Figure 10 is a schematic diagram of the process of progressive decoding.

[0142] Specifically, the received bitstream can be split according to the lengths of the sub-bitstreams parsed from the header information and progressive decoding can be performed. The following are the steps for progressive decoding of each sub-bitstream:

[0143] Refer to Figure 11A , the model prior information is input into the context prediction module 1, and the statistical information of the first feature sub-map y1 is predicted: the mean value m1 and the scale parameter scale1. The residual r1 is decoded (e.g., by entropy decoding) from the bitstream s1 using the statistical information scale1, and y1 is obtained according to r1 + m1 = y1. Subsequently, y1 is passed into the context prediction module 2 to obtain m2 and scale2. If the second bitstream s2 has been received, the residual r2 is decoded from the bitstream s2 by scale2, and y2 is obtained according to r2 + m2 = y2. If the second bitstream has not been received, r2 cannot be decompressed, and m2 = y2.

[0144] Through the above method, when some sub-bitstream segments in the compressed file are missing, the mean value in the statistical information can still be reused as the feature map, ensuring the normal progress of the decoding process.

[0145] 404. Through the inverse feature transformation network, perform image reconstruction on the second feature map to obtain an image; where the inverse feature transformation network includes a first inverse feature transformation layer, and the first inverse feature transformation layer is used to scale the feature values included in the second feature map obtained by inverse feature transformation by the first scaling information to obtain a first feature map.

[0146] In a possible implementation, second scaling information may also be obtained according to the quality factor through a second neural network; the first scaling information is different from the second scaling information; the feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is configured to scale the eigenvalues included in the third feature map obtained by feature inverse transformation through the second scaling information to obtain a fourth feature map.

[0147] The decoded feature map can be inversely transformed layer by layer to obtain a reconstructed image. The quality factor parsed from the header is passed into the quality factor module to generate a scaling table with the same size as the feature map. Subsequently, the scaling table and the feature map are multiplied element by element to obtain a scaled feature map. The scaled feature map continues to be passed into the next-layer feature inverse transformation module and acts on the scaling table output by the quality factor module of the next layer to obtain the scaled feature map of the next layer. Refer to Figure 9 , Figure 9 for a schematic diagram of a process of the decoding process.

[0148] Compared with traditional image codecs, the present solution adopts a deep learning architecture and can adaptively learn the complex features and structures of images, so as to maintain a high image quality at a low bit rate. Compared with similar deep learning algorithms that train multiple models to control the bit rate, the present solution realizes the function of flexibly controlling the bit rate of a single model by designing an embedded quality factor module in a single model, reduces the ROM occupied by the model, and improves the encoding efficiency. Compared with similar deep learning algorithms, the present solution uses the channel splitting of the feature map to design a progressive decoding process, thereby supporting the progressive image reconstruction function.

[0149] Through the quality factor adjustment module and the multi-bit rate adjustment loss function, multiple bit rates can be accurately controlled within a single model, significantly improving the encoding efficiency, allowing users to adjust the required bit rate according to their needs, reaching a specific bit rate with an average of 2.72 adjustments, and the encoding time on the CPU < 500 ms. As shown in Table 1, Table 1 shows a schematic diagram of the compression times statistical results.

[0150] Table 1

[0151]

[0152] The ROM occupied by the deployment of the embodiments of the present application is small. Since the present solution only needs to use a single model, the ROM occupancy can be < 40 MB. And the reconstruction quality is good at a low bit rate: especially in an extremely low bit rate scenario, an image quality superior to that of traditional encoders can be achieved.

[0153] Refer to Figure 7 , Figure 7 for a comparison of the visualization results of the embodiments of the present application with JPEG and existing solutions.

[0154] In addition, the embodiments of the present application utilize multi-channel segmentation technology to support progressive image reconstruction, enabling the image to be gradually presented in ascending quality. This feature is particularly beneficial in application scenarios with limited network bandwidth or progressive image loading, enhancing the user experience.

[0155] Refer to Figure 8 , Figure 8 which is a schematic illustration of the progressive image reconstruction result.

[0156] Next, a schematic illustration of the training process of the embodiments of the present application will be introduced. During training, a training dataset can be obtained. For example, the dataset contains N pictures. Set the quality factor set: According to experimental experience, exemplarily, the quality factor set contains 20 floating-point numbers with an equal interval of 0.425 from 0.5 to 9. Allocate quality factors: Randomly assign a quality factor θ_i to each data x_i in the training dataset, where i = 1, 2,..., N, and θ_i is randomly sampled from the quality factor set . Network inference to obtain the reconstructed picture: Input the picture into the above compression and decompression module to obtain the reconstructed picture and the estimated result of the bitstream length of the compressed file. Calculate the loss function for multi-bitrate adjustment: Based on the obtained reconstructed picture and the original picture in the dataset, as well as the estimated bitstream length of the compressed file, input them into the multi-bitrate conditional loss function for calculation. For example, the loss function can be as follows:

[0157]

[0158]

[0159]

[0160] The model updates the weights by backpropagating the gradient according to the loss function.

[0161] For example, it can refer to Figure 11B , Figure 11B which is a flow schematic illustration of the training process.

[0162] The embodiments of the present application can be applied to low-bandwidth communication scenarios. Refer to Figure 11C , Figure 11C which is a flowchart of a low-bandwidth communication scenario. The full process of this scenario mainly includes the following steps:

[0163] Reading and Compressing Data: This step occurs at the sending user terminal. The user selects an image from the gallery, and the terminal device reads the image format into a bitmap, which is then sent to an image codec and compressed to a specific bit rate. The present invention plays a role at this stage. Using a single model bit rate precise control technology based on deep learning, the bitmap is compressed to a predetermined low bit rate.

[0164] Sending Data: The compressed bitstream is sent through the communication device of the user terminal.

[0165] Data Transmission: The compressed bitstream is transmitted through a low-bandwidth network to the cloud server of the service provider. This can be done in a wired or wireless manner, such as through a telephone line, satellite connection, or cellular network.

[0166] Receiving Data: The cloud server receives the compressed bitstream.

[0167] Data Decompression: On the cloud server, the received data is decompressed into a bitmap and converted into a usable image format. The present invention supports progressive decoding at this stage. According to the completeness of the received data (from one packet to four packets), the image is reconstructed from low quality to high quality.

[0168] Sending Data: The reconstructed image is sent through the communication device of the cloud server.

[0169] Data Transmission: The reconstructed image is transmitted through a high-bandwidth network to the receiving user terminal.

[0170] Receiving and Using Data: The receiving user terminal receives the reconstructed image. The image can be used for a predetermined purpose.

[0171] At Figures 1 to 11C Based on the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Specifically, refer to Figure 12 , Figure 12 FIG. 1200 is a schematic structural diagram of an image processing apparatus 1200 provided by an embodiment of the present application. The image processing apparatus 1200 may be a terminal device or a server. The image processing apparatus 1200 includes:

[0172] An acquisition module 1201, configured to acquire an image and a quality factor, where the quality factor is related to the compression bit rate of the image;

[0173] Among them, the specific description of the acquisition module 1201 may refer to the introduction of step 301 in the above embodiments, and the similarities will not be elaborated here.

[0174] A processing module 1202, configured to obtain first scaling information through a first neural network according to the quality factor; extract features of the image through a feature extraction network to obtain a target feature map, where the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is configured to scale eigenvalues included in a first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; and encode the target feature map to obtain a bitstream.

[0175] Specific descriptions of the processing module 1202 can refer to the introductions in steps 302 and 303 of the above embodiments, and similarities will not be elaborated here.

[0176] In a possible implementation, the first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, and each scaling coefficient corresponds to an eigenvalue, and the first feature extraction layer is configured to scale each eigenvalue among the M eigenvalues through the corresponding scaling coefficient.

[0177] In a possible implementation, the processing module 1202 is further configured to:

[0178] Obtain second scaling information through a second neural network according to the quality factor; the first scaling information is different from the second scaling information;

[0179] The feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is configured to scale eigenvalues included in a third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

[0180] In a possible implementation, the scaling is implemented through a multiplication operation.

[0181] In a possible implementation, the processing module 1202 is further configured to:

[0182] Obtain a compressed file of the image according to the bitstream and the quality factor.

[0183] In a possible implementation, the quality factor is selected from multiple candidate quality factors according to the compression bit rate.

[0184] Refer to Figure 13 , Figure 13 , which is a schematic structural diagram of an image processing apparatus 1300 provided in an embodiment of the present application. The image processing apparatus 1300 may be a terminal device or a server. The image processing apparatus 1300 includes:

[0185] An acquisition module 1301, configured to acquire a compressed file, where the compressed file includes a bitstream and a quality factor;

[0186] Among them, the specific description of the acquisition module 1301 can refer to the introduction of step 401 in the above embodiment, and the similarities will not be elaborated here.

[0187] The processing module 1302 is configured to obtain first scaling information through a first neural network according to the quality factor; decode the bitstream to obtain a second feature map; and perform image reconstruction on the second feature map through a feature inverse transformation network to obtain an image. The feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is configured to scale the eigenvalue included in the second feature map obtained by the feature inverse transformation through the first scaling information to obtain a first feature map.

[0188] Among them, the specific description of the acquisition module 1301 can refer to the introductions of steps 402 and 403 in the above embodiment, and the similarities will not be elaborated here.

[0189] In a possible implementation, the first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, and each scaling coefficient corresponds to one eigenvalue. The first feature inverse transformation layer is configured to scale each eigenvalue among the M eigenvalues through the corresponding scaling coefficient.

[0190] In a possible implementation, the processing module 1302 is further configured to:

[0191] Obtain second scaling information through a second neural network according to the quality factor; the first scaling information is different from the second scaling information;

[0192] The feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is configured to scale the eigenvalue included in the third feature map obtained by the feature inverse transformation through the second scaling information to obtain a fourth feature map.

[0193] In a possible implementation, the compressed file further includes prior information. Specifically, the processing module 1302 is configured to:

[0194] Obtain statistical information of the bitstream according to the prior information, where the statistical information includes an average value;

[0195] Based on the absence of a target segment bitstream in the multi-segment bitstream, use the average value of the target segment bitstream as the feature map of the target segment bitstream in the second feature map.

[0196] Next, a data processing device provided by an embodiment of the present application will be introduced. Please refer to Figure 14 , Figure 14A schematic structural diagram of the data processing device provided by an embodiment of the present application. Specifically, the data processing device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403, and a memory 1404 (where the number of processors 1403 in the data processing device 1400 can be one or more, Figure 14 and here one processor is taken as an example), where the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403, and the memory 1404 may be connected through a bus or other means.

[0197] The memory 1404 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1403. A part of the memory 1404 may also include a non-volatile random access memory (NVRAM). The memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.

[0198] The processor 1403 controls the operation of the radar system (including the antenna, the receiver 1401, and the transmitter 1402). In a specific application, the various components of the radar system are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.

[0199] The image processing method disclosed in the above embodiments of the present application ( Figure 3 and Figure 4The (as shown) can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1403 or the instructions in the form of software. The above-mentioned processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1403 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1404, and the processor 1403 reads the information in the memory 1404 and combines its hardware to complete the steps of the image processing method provided in the above embodiments.

[0200] The receiver 1401 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the radar system. The transmitter 1402 can be used to output digital or character information through the first interface; the transmitter 1402 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group.

[0201] An embodiment of the present application also provides a computer program product, which, when running on a computer, causes the computer to execute the image processing method described in the above embodiments.

[0202] An embodiment of the present application also provides a computer-readable storage medium, in which a program for signal processing is stored, and when it runs on a computer, it causes the computer to execute the image processing method described in the above embodiments.

[0203] The data processing device provided in the embodiments of this application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute the computer-executable instructions stored in the storage unit to cause the chip in the execution device to execute the image enhancement method described in the above embodiments, or to cause the chip in the training device to execute the image enhancement method described in the above embodiments. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0204] Specifically, please refer to Figure 15 , Figure 15 which is a schematic structural diagram of the chip provided in the embodiments of this application. The chip may be embodied as a neural network processor NPU1500, and the NPU 1500 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1503, and the arithmetic circuit 1503 is controlled by the controller 1504 to extract matrix data from the memory and perform multiplication operations.

[0205] In some implementations, the arithmetic circuit 1503 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 1503 is a two-dimensional systolic array. The arithmetic circuit 1503 may also be a one-dimensional systolic array or other electronic circuits that can perform mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1503 is a general matrix processor.

[0206] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1501 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are saved in the accumulator 1508.

[0207] The unified memory 1506 is used to store input data and output data. The weight data is directly transported through the direct memory access controller (DMAC) 1505, and the DMAC transports it to the weight memory 1502. The input data is also transported to the unified memory 1506 through the DMAC.

[0208] The BIU is the Bus Interface Unit, i.e., the bus interface unit 1510, which is used for the interaction between the AXI bus, the DMAC, and the instruction fetch buffer (IFB) 1509.

[0209] The bus interface unit 1510 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 1509 to obtain instructions from the external memory, and is also used for the storage unit access controller 1505 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0210] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1506, or transfer the weight data to the weight memory 1502, or transfer the input data to the input memory 1501.

[0211] The vector calculation unit 1507 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization (batch normalization), pixel-level summation, upsampling of the feature plane, etc.

[0212] In some implementations, the vector calculation unit 1507 can store the processed output vector in the unified memory 1506. For example, the vector calculation unit 1507 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 1503, such as linear interpolation of the feature plane extracted by the convolutional layer, or a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 1507 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1503, such as for use in subsequent layers in a neural network.

[0213] The instruction fetch buffer 1509 connected to the controller 1504 is used to store the instructions used by the controller 1504;

[0214] The unified memory 1506, the input memory 1501, the weight memory 1502, and the instruction fetch buffer 1509 are all On-Chip memories. The external memory is private to this NPU hardware architecture.

[0215] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program related to the steps of the data processing method described in the above embodiments.

[0216] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0217] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods of the various embodiments of this application.

[0218] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0219] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are wholly or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. An image processing method, characterized in that, The method includes: Obtaining an image and a quality factor, where the quality factor is related to the compression bit rate of the image; According to the quality factor, obtaining first scaling information through a first neural network; Performing feature extraction on the image through a feature extraction network to obtain a target feature map; wherein, the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is used to scale the feature values included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; Encoding the target feature map to obtain a bitstream.

2. The method according to claim 1, wherein The first scaling information includes M scaling coefficients; the first feature map includes M feature values, and each scaling coefficient corresponds to a feature value, and the first feature extraction layer is used to scale each feature value in the M feature values through the corresponding scaling coefficient.

3. The method according to claim 1 or 2, characterized in that, The method further includes: According to the quality factor, obtaining second scaling information through a second neural network; the first scaling information and the second scaling information are different; The feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is used to scale the feature values included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

4. The method according to any one of claims 1 to 3, characterized in that The scaling is implemented through a multiplication operation.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: According to the bitstream and the quality factor, obtaining a compressed file of the image.

6. The method according to any one of claims 1 to 5, characterized in that The quality factor is selected from multiple candidate quality factors according to the compression bit rate.

7. An image processing method, characterized in that, The method includes: Obtaining a compressed file, where the compressed file includes a bitstream and a quality factor; According to the quality factor, obtaining first scaling information through a first neural network; Decoding the bitstream to obtain a second feature map; Performing image reconstruction on the second feature map through a feature inverse transformation network to obtain an image; wherein, the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is used to scale the feature values included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

8. The method according to claim 7, wherein The first scaling information includes M scaling coefficients; the second feature map includes M feature values, and each scaling coefficient corresponds to a feature value, and the first feature inverse transformation layer is used to scale each feature value in the M feature values through the corresponding scaling coefficient.

9. The method according to claim 7 or 8, characterized in that The method further includes: According to the quality factor, obtaining second scaling information through a second neural network; the first scaling information and the second scaling information are different; The feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is used to scale the feature values included in the third feature map obtained by feature inverse transformation through the second scaling information to obtain a fourth feature map.

10. The method according to any one of claims 7 to 9, characterized in that, The compressed file further includes prior information, and the bitstream includes multiple sub-bitstreams; the decoding the bitstream to obtain a second feature map includes: According to the prior information, obtaining statistical information of the bitstream, where the statistical information includes a mean value; Based on the absence of the target segment stream in the multi-segment stream, use the mean value of the target segment stream as the feature map of the target segment stream in the second feature map.

11. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire an image and a quality factor, where the quality factor is related to the compression bit rate of the image; A processing module, configured to obtain first scaling information through a first neural network according to the quality factor; perform feature extraction on the image through a feature extraction network to obtain a target feature map; where the feature extraction network includes a first feature extraction layer, and the first feature extraction layer is configured to scale the eigenvalues included in the first feature map obtained by feature extraction through the first scaling information to obtain a second feature map; encode the target feature map to obtain a bit stream.

12. The device according to claim 11, characterized in that, The first scaling information includes M scaling coefficients; the first feature map includes M eigenvalues, and each scaling coefficient corresponds to an eigenvalue, and the first feature extraction layer is configured to scale each eigenvalue among the M eigenvalues through the corresponding scaling coefficient.

13. The device according to claim 11 or 12, characterized in that, The processing module is further configured to: Obtain second scaling information through a second neural network according to the quality factor; the first scaling information and the second scaling information are different; The feature extraction network further includes a second feature extraction layer, and the second feature extraction layer is configured to scale the eigenvalues included in the third feature map obtained by feature extraction through the second scaling information to obtain a fourth feature map.

14. The device according to any one of claims 11 to 13, characterized in that, The scaling is implemented through a multiplication operation.

15. The device according to any one of claims 11 to 14, characterized in that, The processing module is further configured to: Obtain a compressed file of the image according to the bit stream and the quality factor.

16. The device according to any one of claims 11 to 15, characterized in that The quality factor is selected from multiple candidate quality factors according to the compression bit rate.

17. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire a compressed file, where the compressed file includes a bit stream and a quality factor; A processing module, configured to obtain first scaling information through a first neural network according to the quality factor; decode the bit stream to obtain a second feature map; perform image reconstruction on the second feature map through a feature inverse transformation network to obtain an image; where the feature inverse transformation network includes a first feature inverse transformation layer, and the first feature inverse transformation layer is configured to scale the eigenvalues included in the second feature map obtained by feature inverse transformation through the first scaling information to obtain a first feature map.

18. The device according to claim 17, characterized in that, The first scaling information includes M scaling coefficients; the second feature map includes M eigenvalues, and each scaling coefficient corresponds to an eigenvalue, and the first feature inverse transformation layer is configured to scale each eigenvalue among the M eigenvalues through the corresponding scaling coefficient.

19. The device according to claim 17 or 18, characterized in that The processing module is further configured to: Obtain second scaling information through a second neural network according to the quality factor; the first scaling information and the second scaling information are different; The feature inverse transformation network further includes a second feature inverse transformation layer, and the second feature inverse transformation layer is configured to scale the eigenvalues included in the third feature map obtained by feature inverse transformation through the second scaling information to obtain a fourth feature map.

20. The device according to any one of claims 17 to 19, characterized in that The compressed file further includes prior information, and the processing module is specifically configured to: Based on the prior information, statistical information of the bitstream is obtained, and the statistical information includes a mean value; Based on the absence of the target segment bitstream in the multi-segment bitstreams, the mean value of the target segment bitstream is used as the feature map of the target segment bitstream in the second feature map.

21. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers are caused to perform the operations of the method according to any one of claims 1 to 10.

22. A computer program product, characterized in that, It includes computer-readable instructions, and when the computer-readable instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 10.

23. A system includes at least one processor and at least one memory; the processor and the memory are connected through a communication bus and communicate with each other; The at least one memory is used for storing code; The at least one processor is used for executing the code to execute the method according to any one of claims 1 to 10.

24. A chip, characterized in that, It includes at least one processing unit and an interface circuit, the interface circuit is used for providing program instructions or data for the at least one processing unit, and the at least one processing unit is used for executing the program instructions to implement the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Image processing methods and related device

    WO2025152884A1