Image processing method and related device

The image processing method addresses the limitation of fixed bit rates in deep convolutional networks by using adjustable gain values for feature extraction and entropy coding, achieving flexible and efficient compression rate control.

JP7767536B2Active Publication Date: 2025-11-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024151474
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-07
Filing Date
2024-09-03
Publication Date
2025-11-11
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

Existing image coding methods based on deep convolutional networks typically output a single coding result for a specific type of input image, making it difficult to achieve the desired compression bit rate based on actual requirements.

Method used

An image processing method that involves feature extraction, quantization, and entropy coding with adjustable target gain values to control compression bit rates, allowing for flexible bit rate management and accurate encoding.

Benefits of technology

Enables precise control over compression bit rates, ensuring the encoded data meets predetermined conditions and maintains a desired level of information entropy, thereby improving the adaptability and efficiency of image coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767536000027
    Figure 0007767536000027
  • Figure 0007767536000028
    Figure 0007767536000028
  • Figure 0007767536000029
    Figure 0007767536000029
Patent Text Reader

Abstract

To provide an image processing method with which compression bit rate control is realized in the same compression model, and an image processing device.SOLUTION: An image processing method includes steps of: acquiring a first image; carrying out feature extraction on the first image to obtain at least one first feature map including N first feature values (N is an integer); acquiring a target compression bit rate, which corresponds to M target gain values, where each target gain value corresponds to one first feature value (M is a positive integer smaller than or equal to N); processing the corresponding first feature values according to the M target gain values to obtain M second feature values; and performing quantization and entropy coding on the first feature map including the M second feature values to obtain coded data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to Chinese Patent Application No. 202010082808.4, entitled "IMAGE PROCESSING METHOD AND RELATED DEVICE," filed with the State Intellectual Property Office of the People's Republic of China on February 7, 2020, which is hereby incorporated by reference in its entirety.

[0002] The present application relates to the field of artificial intelligence, and in particular to image processing methods and related devices. [Background technology]

[0003] Today, multimedia data accounts for a large portion of Internet traffic. Image data compression plays a crucial role in the storage and efficient transmission of multimedia data. Therefore, image coding is a technology of great practical value.

[0004] Image coding has been studied for a long time. Researchers have proposed a large number of methods and developed various international standards such as JPEG, JPEG2000, WebP, and BPG. Although these coding methods are now widely applied, these traditional methods exhibit some limitations in the face of increasing amounts of image data and the constantly emerging new media types.

[0005] In recent years, researchers have begun to study image coding methods based on deep learning. Some researchers have already achieved good results. For example, Balle et al. proposed an end-to-end optimal image coding method that surpasses the current best image coding performance and even surpasses the current best conventional coding standard, BPG. However, most image coding methods based on deep convolutional networks currently have a drawback: a trained model can only output one coding result for one type of input image. As a result, it is difficult to achieve the coding effect with a target compression bit rate based on actual requirements. Summary of the Invention [Means for solving the problem]

[0006] The present application provides an image processing method for implementing compression bit rate control in a uniform compression model.

[0007] According to a first aspect, the present application provides an image processing method, the method comprising: The method includes the steps of: acquiring a first image; performing feature extraction on the first image to obtain at least one first feature map, where the at least one first feature map includes N first feature values, where N is a positive integer; acquiring a target compression bit rate, where the target compression bit rate corresponds to M target gain values, each target gain value corresponding to one first feature value, where M is a positive integer less than or equal to N; processing the corresponding first feature values ​​based on the M target gain values ​​to obtain M second feature values, where the original first feature map may be replaced with the at least one processed first feature map in this embodiment of the present application; and performing quantization and entropy coding on the at least one processed first feature map to obtain encoded data, where the at least one processed first feature map includes the M second feature values. In the above scheme, different target gain values ​​are set for different target compression bit rates to implement compression bit rate control.

[0008] In an optional design of the first aspect, an information entropy of the quantized data obtained by quantizing the at least one processed first feature map satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0009] In an optional design of the first embodiment, a higher target compression bit rate indicates a higher information entropy of the quantized data.

[0010] In an optional design of the first aspect, the difference between the compression bit rate corresponding to the encoded data and the target compression bit rate falls within a preset range.

[0011] In an optional design of the first aspect, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values, respectively.

[0012] In an optional design of the first aspect, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, and P is a positive integer less than or equal to M.

[0013] In an optional design of the first aspect, the method comprises: determining M target gain values ​​corresponding to the target compression bit rate based on a target mapping relationship, wherein the target mapping relationship is used to indicate a correlation between the compression bit rate and the M target gain values; the target mapping relationship includes a plurality of compression bit rates, a plurality of gain vectors, and correlations between the plurality of compression bit rates and the plurality of gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates and the M target gain values ​​are elements of one of the plurality of gain vectors; or If the goal mapping relationship includes a goal function mapping relationship, and the input of the goal function relationship includes a target compression bit rate, then the output of the goal function relationship includes M target gain values.

[0014] In an optional design of the first aspect, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to M first gain values, the second compression bit rate corresponds to M second gain values, and the M target gain values ​​are obtained by performing an interpolation operation on the M first gain values ​​and the M second gain values.

[0015] In an optional design of the first aspect, the M first gain values ​​include a first target gain value, the M second gain values ​​include a second target gain value, and the M target gain values ​​include a third target gain value, wherein the first target gain value, the second target gain value, and the third target gain value correspond to the same feature value among the M first feature values, and the third target gain value is obtained by performing an interpolation operation on the first target gain value and the second target gain value.

[0016] In an optional design of the first aspect, the first image includes a target object, and the M first feature values ​​are feature values ​​in at least one feature map that correspond to the target object.

[0017] In an optional design of the first aspect, each of the M target gain values ​​corresponds to one inverse gain value, and the inverse gain value is used to process the feature value obtained in the decoding process of the encoded data, and the product of each of the M target gain values ​​and the corresponding inverse gain value falls within a preset range.

[0018] In an optional design of the first aspect, the method further includes: performing entropy decoding on the encoded data to obtain at least one second feature map, where the at least one second feature map includes N third feature values, each third feature value corresponding to one first feature value; obtaining M target inverse gain values, where each target inverse gain value corresponds to one third feature value; performing gain processing on the corresponding third feature values ​​based on the M target inverse gain values ​​to obtain M fourth feature values; and performing image reconstruction on the at least one second feature map obtained after the inverse gain processing to obtain a second image, where the at least one second feature map obtained after the inverse gain processing includes the M fourth feature values.

[0019] In an optional design of the first aspect, the M fourth feature values ​​are obtained by individually performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values.

[0020] In an optional design of the first aspect, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, and P is a positive integer less than or equal to M.

[0021] In an optional design of the first aspect, the method further includes determining M target inverse gain values ​​corresponding to the target compression bit rate based on a target mapping relationship, where the target mapping relationship is used to indicate a correlation between the compression bit rate and the inverse gain vector.

[0022] In an optional design of the first aspect, the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and a correlation between the plurality of compression bit rates and the plurality of inverse gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates, and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors.

[0023] In an optional design of the first aspect, the target mapping relationship comprises an objective function mapping relationship, where an input of the objective function relationship comprises the target compression bit rate and an output of the objective function relationship comprises the M target inverse gain values.

[0024] In an optional design of the first aspect, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0025] In an optional design of the first aspect, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range.

[0026] In an optional design of the first aspect, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0027] In an optional design of the first aspect, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, the M target inverse gain values ​​include a third target inverse gain value, the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M third feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0028] According to a second aspect, the present application provides an image processing method, the method comprising: The method includes the steps of: obtaining encoded data; performing entropy decoding on the encoded data to obtain at least one second feature map, where the at least one second feature map includes N third feature values, where N is a positive integer; obtaining M target inverse gain values, where each target inverse gain value corresponds to one third feature value, where M is a positive integer less than or equal to N; processing the corresponding third feature values ​​based on the M target inverse gain values ​​to obtain M fourth feature values; and performing image reconstruction based on the at least one processed second feature map to obtain a second image, where the at least one processed second feature map includes the M fourth feature values.

[0029] In an optional design of the second aspect, the M fourth feature values ​​are obtained by performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values, respectively.

[0030] In an optional design of the second aspect, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, and P is a positive integer less than or equal to M.

[0031] In an optional design of the second aspect, the method comprises: obtaining a target compression bit rate; determining M target inverse gain values ​​corresponding to the target compression bit rate based on a target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bit rate and the inverse gain vector; the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and correlations between the plurality of compression bit rates and the plurality of inverse gain vectors; the target compression bit rate is one of a plurality of compression bit rates and the M target inverse gain values ​​are elements of a plurality of inverse gain vectors; or If the target mapping relationship includes a target function mapping relationship, and the input of the target function relationship includes a target compression bit rate, then the output of the target function relationship includes M target inverse gain values.

[0032] In an optional design of the second aspect, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0033] In an optional design of the second aspect, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0034] In an optional design of the second aspect, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, the M target inverse gain values ​​include a third target inverse gain value, the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0035] According to a third aspect, the present application provides an image processing method, the method comprising: acquiring a first image; performing feature extraction on the first image based on the encoding network to obtain at least one first feature map, wherein the at least one first feature map includes N first feature values, where N is a positive integer; obtaining a target compression bit rate, the target compression bit rate corresponding to M initial gain values ​​and M initial inverse gain values, each initial gain value corresponding to one first feature value, each initial inverse gain value corresponding to one third feature value, and M being a positive integer less than or equal to N; processing the corresponding first feature values ​​based on the M initial gain values ​​to obtain M second feature values, respectively; performing quantization and entropy coding on the at least one processed first feature map based on a quantization network and an entropy coding network to obtain coded data and a bit rate loss, wherein the at least one first feature map obtained after gain processing includes M second feature values; performing entropy decoding on the encoded data based on an entropy decoding network to obtain at least one second feature map, wherein the at least one second feature map includes M third feature values, each third feature value corresponding to one first feature value; processing the corresponding third feature values ​​based on the M initial inverse gain values ​​to obtain M fourth feature values, respectively; performing image reconstruction on at least one processed second feature map based on the decoding network to obtain a second image, wherein the at least one processed feature map includes M fourth feature values; obtaining a distortion loss of the second image relative to the first image; performing joint training on a first encoding / decoding network, M initial gain values, and M initial inverse gain values ​​by using a loss function until an image distortion value between the first image and the second image reaches a first predetermined degree, where the image distortion value is related to a bitrate loss and a distortion loss, and the encoding / decoding network includes a coding network, a quantization network, an entropy coding network, and an entropy decoding network; outputting a second encoding / decoding network, M target gain values, and M target inverse gain values, wherein the second encoding / decoding network is a model obtained after iterative training is performed on the first encoding / decoding network, and the M target gain values ​​and the M target inverse gain values ​​are obtained after iterative training is performed on the M initial gain values ​​and the M initial inverse gain values; Includes.

[0036] In an optional design of the third aspect, information entropy of the quantized data obtained by quantizing at least one first feature map obtained after gain processing satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0037] In an optional design of the third aspect, the preset condition includes at least indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0038] In an optional design of the third aspect, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values, respectively.

[0039] In an optional design of the third aspect, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, and P is a positive integer less than or equal to M.

[0040] In an optional design of the third aspect, the first image includes a target object, and the M first feature values ​​are feature values ​​in at least one feature map that correspond to the target object.

[0041] In an optional design of the third aspect, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range, and the product of each of the M initial gain values ​​and the corresponding initial inverse gain value falls within a preset range.

[0042] According to a fourth aspect, the present application provides an image processing device, comprising: an acquisition module configured to acquire a first image; a feature extraction module configured to perform feature extraction on the first image to obtain at least one first feature map, the at least one first feature map comprising N first feature values, where N is a positive integer; an acquisition module acquires a target compression bit rate, the target compression bit rate corresponds to M target gain values, each target gain value corresponds to one first feature value, M is a positive integer less than or equal to N; a gain module configured to process each of the corresponding first feature values ​​based on the M target gain values ​​to obtain M second feature values; a quantization and entropy coding module configured to perform quantization and entropy coding on the at least one processed first feature map to obtain coded data, wherein the at least one processed first feature map includes M second feature values; Equipped with.

[0043] In an optional design of the fourth aspect, an information entropy of the quantized data obtained by quantizing the at least one processed first feature map satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0044] In an optional design of the fourth aspect, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0045] In an optional design of the fourth aspect, the difference between the compression bit rate corresponding to the encoded data and the target compression bit rate falls within a preset range.

[0046] In an optional design of the fourth aspect, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values, respectively.

[0047] In an optional design of the fourth aspect, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, and P is a positive integer less than or equal to M.

[0048] In an optional design of the fourth aspect, the apparatus comprises: a determining module configured to determine M target gain values ​​corresponding to the target compression bitrate based on a target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bitrate and the M target gain values; the target mapping relationship includes a plurality of compression bit rates, a plurality of gain vectors, and correlations between the plurality of compression bit rates and the plurality of gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates and the M target gain values ​​are elements of one of the plurality of gain vectors; or If the goal mapping relationship includes a goal function mapping relationship, and the input of the goal function relationship includes a target compression bit rate, then the output of the goal function relationship includes M target gain values.

[0049] In an optional design of the fourth aspect, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to M first gain values, the second compression bit rate corresponds to M second gain values, and the M target gain values ​​are obtained by performing an interpolation operation on the M first gain values ​​and the M second gain values.

[0050] In an optional design of the fourth aspect, the M first gain values ​​include a first target gain value, the M second gain values ​​include a second target gain value, and the M target gain values ​​include a third target gain value, wherein the first target gain value, the second target gain value, and the third target gain value correspond to the same feature value among the M first feature values, and the third target gain value is obtained by performing an interpolation operation on the first target gain value and the second target gain value.

[0051] In an optional design of the fourth aspect, the first image includes a target object, and the M first feature values ​​are feature values ​​in at least one feature map that correspond to the target object.

[0052] In an optional design of the fourth aspect, each of the M target gain values ​​corresponds to one inverse gain value, and the inverse gain value is used to process the feature value obtained in the decoding process of the encoded data, and the product of each of the M target gain values ​​and the corresponding inverse gain value falls within a predetermined range.

[0053] In an optional design of the fourth aspect, the apparatus comprises: a decoding module configured to perform entropy decoding on the encoded data to obtain at least one second feature map, the at least one second feature map including N third feature values, each third feature value corresponding to one first feature value; the obtaining module is further configured to obtain M target inverse gain values, each target inverse gain value corresponding to one third feature value; This device is an inverse gain module configured to perform gain processing on corresponding third feature values ​​based on the M target inverse gain values, respectively, to obtain M fourth feature values; a reconstruction module configured to perform image reconstruction on at least one second feature map obtained after the inverse gain processing to obtain a second image, wherein the at least one second feature map obtained after the inverse gain processing includes M fourth feature values; Further provided are:

[0054] In an optional design of the fourth aspect, the M fourth feature values ​​are obtained by performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values, respectively.

[0055] In an optional design of the fourth aspect, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, and P is a positive integer less than or equal to M.

[0056] In the optical design of the fourth aspect, the determination module The method is further configured to determine M target inverse gain values ​​corresponding to the target compression bit rate based on the target mapping relationship, wherein the target mapping relationship is used to indicate a correlation between the compression bit rate and the inverse gain vector.

[0057] In an optional design of the fourth aspect, the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and a correlation between the plurality of compression bit rates and the plurality of inverse gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates, and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors.

[0058] In an optional design of the fourth aspect, the target mapping relationship comprises an objective function mapping relationship, where an input of the objective function relationship comprises the target compression bit rate and an output of the objective function relationship comprises M target inverse gain values.

[0059] In an optional design of the fourth aspect, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0060] In an optional design of the fourth aspect, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range.

[0061] In an optional design of the fourth aspect, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0062] In an optional design of the fourth aspect, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, the M target inverse gain values ​​include a third target inverse gain value, the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0063] According to a fifth aspect, the present application provides an image processing device, comprising: a capture module configured to capture the encoded data; a decoding module configured to perform entropy decoding on the encoded data to obtain at least one second feature map, the at least one second feature map including N third feature values, where N is a positive integer; The acquisition module is further configured to acquire M target inverse gain values, each target inverse gain value corresponding to one third feature value, where M is a positive integer less than or equal to N; an inverse gain module configured to process each corresponding third feature value based on the M target inverse gain values ​​to obtain M fourth feature values; a reconstruction module configured to perform image reconstruction on the at least one processed second feature map to obtain a second image, wherein the at least one processed second feature map includes M fourth feature values; Equipped with.

[0064] In an optional design of the fifth aspect, the M fourth feature values ​​are obtained by individually performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values.

[0065] In an optional design of the fifth aspect, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, and P is a positive integer less than or equal to M.

[0066] In an optional design of the fifth aspect, the acquisition module is further configured to acquire a target compression bit rate; This device is a determining module for determining M target inverse gain values ​​corresponding to a target compression bit rate based on a target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bit rate and the inverse gain vector; the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and correlations between the plurality of compression bit rates and the plurality of inverse gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors; or If the target mapping relationship includes a target function mapping relationship, and the input of the target function relationship includes a target compression bit rate, then the output of the target function relationship includes M target inverse gain values.

[0067] In an optional design of the fifth aspect, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0068] In an optional design of the fifth aspect, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0069] In an optional design of the fifth aspect, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, the M target inverse gain values ​​include a third target inverse gain value, the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0070] According to a sixth aspect, the present application provides an image processing device, comprising: an acquisition module configured to acquire a first image; a feature extraction module configured to perform feature extraction on the first image based on the encoding network to obtain at least one first feature map, the at least one first feature map including N first feature values, where N is a positive integer; the obtaining module is further configured to obtain a target compression bit rate, the target compression bit rate corresponding to M initial gain values ​​and M initial inverse gain values, each initial gain value corresponding to one first feature value, each initial inverse gain value corresponding to one third feature value, and M is a positive integer less than or equal to N; a gain module configured to process each of the corresponding first feature values ​​based on the M initial gain values ​​to obtain M second feature values; a quantization and entropy coding module configured to perform quantization and entropy coding on the at least one processed first feature map based on a quantization network and an entropy coding network to obtain coded data and a bit rate loss, wherein the at least one first feature map obtained after the gain processing includes M second feature values; a decoding module configured to perform entropy decoding on the encoded data based on the entropy decoding network to obtain at least one second feature map, the at least one second feature map including M third feature values, each third feature value corresponding to one first feature value; an inverse gain module configured to process each of the corresponding third feature values ​​based on the M initial inverse gain values ​​to obtain M fourth feature values; a reconstruction module configured to perform image reconstruction on the at least one processed second feature map based on the decoding network to obtain a second image, wherein the at least one processed feature map includes M fourth feature values; the acquisition module is further configured to acquire a distortion loss of the second image relative to the first image; a training module configured to perform joint training on a first encoding / decoding network, M initial gain values, and M initial inverse gain values ​​by using a loss function until an image distortion value between the first image and the second image reaches a first predetermined degree, the image distortion value being related to a bitrate loss and a distortion loss, and the encoding / decoding network including a coding network, a quantization network, an entropy coding network, and an entropy decoding network; an output module configured to output a second encoding / decoding network, M target gain values, and M target inverse gain values, wherein the second encoding / decoding network is a model obtained after iterative training is performed on the first encoding / decoding network, and the M target gain values ​​and the M target inverse gain values ​​are obtained after iterative training is performed on the M initial gain values ​​and the M initial inverse gain values; Equipped with.

[0071] In an optional design of the sixth aspect, the information entropy of the quantized data obtained by quantizing at least one first feature map obtained after gain processing satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0072] In an optional design of the sixth aspect, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0073] In an optional design of the sixth aspect, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values, respectively.

[0074] In an optional design of the sixth aspect, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, and P is a positive integer less than or equal to M.

[0075] In an optional design of the sixth aspect, the first image includes a target object, and the M first feature values ​​are feature values ​​in at least one feature map that correspond to the target object.

[0076] In an optional design of the sixth aspect, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range, and the product of each of the M initial gain values ​​and the corresponding initial inverse gain value falls within a preset range.

[0077] According to a seventh aspect, an embodiment of the present application provides an execution device. The execution device may include a memory, a processor, and a bus system. The memory is configured to store a program, and the processor is configured to execute the program in the memory, the program: acquiring a first image; performing feature extraction on the first image to obtain at least one first feature map, wherein the at least one first feature map includes N first feature values, where N is a positive integer; obtaining a target compression bit rate, the target compression bit rate corresponding to M target gain values, each target gain value corresponding to one first feature value, where M is a positive integer less than or equal to N; processing each of the corresponding first feature values ​​based on the M target gain values ​​to obtain M second feature values; performing quantization and entropy coding on the at least one processed first feature map to obtain coded data, wherein the at least one processed first feature map includes M second feature values; Includes.

[0078] In an optional design of the seventh aspect, the execution device is a virtual reality VR device, a mobile phone, a tablet computer, a notebook computer, a server, or an intelligent wearable device.

[0079] In a seventh aspect of the present application, the processor may be further configured to perform the steps of the first aspect or any possible implementation of the first aspect. For details, please refer to the first aspect. The details will not be described again here.

[0080] According to an eighth aspect, an embodiment of the present application provides an execution device. The execution device may include a memory, a processor, and a bus system. The memory is configured to store a program, and the processor is configured to execute the program in the memory, the program: obtaining encoded data; performing entropy decoding on the encoded data to obtain at least one second feature map, wherein the at least one second feature map includes N third feature values, where N is a positive integer; obtaining M target inverse gain values, each target inverse gain value corresponding to one third feature value, where M is a positive integer less than or equal to N; processing the corresponding third feature values ​​respectively based on the M target inverse gain values ​​to obtain M fourth feature values; performing image reconstruction on the at least one processed second feature map to obtain a second image, wherein the at least one processed second feature map includes M fourth feature values; Includes.

[0081] In an optional design of the eighth aspect, the execution device is a virtual reality VR device, a mobile phone, a tablet computer, a notebook computer, a server, or an intelligent wearable device.

[0082] In the eighth aspect of the present application, the processor may be further configured to perform the steps of the second aspect or any possible implementation of the second aspect. For details, please refer to the second aspect. The details will not be described again here.

[0083] According to a ninth aspect, an embodiment of the present application provides a training device. The training device may include a memory, a processor, and a bus system. The memory is configured to store a program, and the processor is configured to execute the program in the memory, the program: acquiring a first image; performing feature extraction on the first image based on the encoding network to obtain at least one first feature map, wherein the at least one first feature map includes N first feature values, where N is a positive integer; obtaining a target compression bit rate, the target compression bit rate corresponding to M initial gain values ​​and M initial inverse gain values, each initial gain value corresponding to one first feature value, each initial inverse gain value corresponding to one third feature value, and M being a positive integer less than or equal to N; processing the corresponding first feature values ​​based on the M initial gain values ​​to obtain M second feature values, respectively; performing quantization and entropy coding on the at least one processed first feature map based on a quantization network and an entropy coding network to obtain coded data and a bit rate loss, wherein the at least one first feature map obtained after gain processing includes M second feature values; performing entropy decoding on the encoded data based on an entropy decoding network to obtain at least one second feature map, wherein the at least one second feature map includes M third feature values, each third feature value corresponding to one first feature value; processing the corresponding third feature values ​​based on the M initial inverse gain values ​​to obtain M fourth feature values, respectively; performing image reconstruction on at least one processed second feature map based on the decoding network to obtain a second image, wherein the at least one processed feature map includes M fourth feature values; obtaining a distortion loss of the second image relative to the first image; performing joint training on a first encoding / decoding network, M initial gain values, and M initial inverse gain values ​​by using a loss function until an image distortion value between the first image and the second image reaches a first predetermined degree, where the image distortion value is related to a bitrate loss and a distortion loss, and the encoding / decoding network includes a coding network, a quantization network, an entropy coding network, and an entropy decoding network; outputting a second encoding / decoding network, M target gain values, and M target inverse gain values, wherein the second encoding / decoding network is a model obtained after iterative training is performed on the first encoding / decoding network, and the M target gain values ​​and the M target inverse gain values ​​are obtained after iterative training is performed on the M initial gain values ​​and the M initial inverse gain values; Includes.

[0084] In a ninth aspect of the present application, the processor may be further configured to perform the steps of the third aspect or any possible implementation of the third aspect. For details, please refer to the third aspect. The details will not be described again here.

[0085] According to a tenth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to perform the image processing method according to any one of the first to third aspects.

[0086] According to an eleventh aspect, an embodiment of the present application provides a computer program which, when run on a computer, enables the computer to perform the image processing method according to any one of the first to third aspects.

[0087] According to a twelfth aspect, the present application provides a chip system. The chip system includes a processor configured to support an execution device or a training device in performing the functions of the aforementioned aspects, for example, in transmitting or processing data and / or information in the aforementioned methods. In a possible design, the chip system further includes a memory. The memory is configured to store necessary program instructions and data for the execution device or training device. The chip system may include a chip, or may include a chip and other discrete components.

[0088] An embodiment of the present application provides an image processing method, comprising: acquiring a first image; performing feature extraction on the first image to obtain at least one first feature map; the at least one first feature map including N first feature values, where N is a positive integer; obtaining a target compression bit rate; the target compression bit rate corresponding to M target gain values, each target gain value corresponding to one first feature value, where M is a positive integer less than or equal to N; processing the corresponding first feature value based on the M target gain values ​​to obtain M second feature values; and performing quantization and entropy coding on the at least one processed first feature map to obtain encoded data, where the at least one processed first feature map includes the M second feature values. In the above-described method, different target gain values ​​are set for different target compression bit rates to implement compression bit rate control. [Brief explanation of the drawings]

[0089] [Figure 1] 1 is a schematic diagram of the structure of the artificial intelligence main framework. [Figure 2a] FIG. 1 illustrates an application scenario according to an embodiment of the present application. [Figure 2b] FIG. 1 illustrates an application scenario according to an embodiment of the present application. [Figure 3] 1 illustrates an embodiment of an image processing method according to an embodiment of the present application; [Figure 4] FIG. 1 illustrates a CNN-based image processing process. [Figure 5a] FIG. 10 illustrates information entropy distribution of feature maps at compressed bit rates according to an embodiment of the present application. [Figure 5b] FIG. 10 illustrates information entropy distribution of feature maps at compressed bit rates according to an embodiment of the present application. [Figure 6] FIG. 1 illustrates an objective function mapping relationship according to an embodiment of the present application. [Figure 7] FIG. 1 illustrates an embodiment of an image processing method according to an embodiment of the present application. [Figure 8] FIG. 1 illustrates an image compression procedure according to an embodiment of the present application. [Figure 9] FIG. 1 illustrates the compression effect according to an embodiment of the present application. [Figure 10] FIG. 1 illustrates a training process according to an embodiment of the present application. [Figure 11] FIG. 2 illustrates an image processing process according to an embodiment of the present application. [Figure 12] 1 is a diagram illustrating a system architecture of an image processing system according to an embodiment of the present invention. [Figure 13] 1 is a schematic flow chart of an image processing method according to an embodiment of the present application. [Figure 14] 1 is a schematic diagram of the structure of an image processing device according to an embodiment of the present invention; [Figure 15] 1 is a schematic diagram of the structure of an image processing device according to an embodiment of the present invention; [Figure 16] 1 is a schematic diagram of the structure of an image processing device according to an embodiment of the present invention; [Figure 17] 1 is a schematic diagram of the structure of an execution device according to an embodiment of the present application; [Figure 18] 1 is a schematic diagram of the structure of a training device according to an embodiment of the present application; [Figure 19] 1 is a schematic diagram of the structure of a chip according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0090] Hereinafter, embodiments of the present invention will be described with reference to the drawings in the embodiments of the present invention. The terms used in the embodiments of the present invention are merely used to describe specific embodiments of the present invention and are not intended to limit the present invention.

[0091] Hereinafter, embodiments of the present application will be described with reference to the drawings. Those skilled in the art can learn that the technical solutions provided in the embodiments of the present application can also be applied to similar technical problems as technology evolves and new scenarios emerge.

[0092] In the specification, claims, and accompanying drawings of this application, terms such as "first," "second," etc. are intended to distinguish between similar objects, but do not necessarily indicate a particular order or sequence. Terms used in this manner are interchangeable under appropriate circumstances, and it should be understood that this is merely a distinguishing method used to describe objects having the same attributes in the embodiments of this application. In addition, the terms "include," "have," and any other variations thereof are meant to cover non-exclusive inclusions, and thus a process, method, system, product, or device that includes a set of units is not necessarily limited to those units, but may include other units not expressly listed or that are inherent to such process, method, product, or device.

[0093] First, the overall operating procedure of the AI ​​system is described. Figure 1 is a schematic diagram of the AI ​​main framework structure. The following describes the aforementioned AI main framework from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the overall process from data acquisition to data processing. For example, the process may be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a "data-information-knowledge-wisdom" condensation process. The "IT value chain" reflects the value AI brings to the information technology industry, from the infrastructure and information underlying human intelligence (technology provision and processing implementation) to the system's industrial ecological process.

[0094] (1) Infrastructure The infrastructure provides computational support for the artificial intelligence system and communicates with the outside world through the use of a base platform. The infrastructure communicates with the outside world through the use of sensors. Computational power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, or hardware acceleration chips such as FPGAs). The base platform includes the security and support of related platforms such as distributed computing frameworks and networks, which may include cloud storage and computing, interconnection and interaction networks, etc. For example, sensors communicate with the outside world to acquire data, and the data is provided to intelligent chips in the distributed computing system provided by the base platform for computation.

[0095] (2) Data Data from the higher layers of the infrastructure represent data sources in the field of artificial intelligence, including graphs, images, voice, text, and more related to Internet of Things data from traditional devices, service data from existing systems, and sensory data such as force, displacement, liquid level, temperature, humidity, etc.

[0096] (3) Data processing Data processing typically includes methods such as data training, machine learning, deep learning, search, inference, and decision-making.

[0097] Machine learning and deep learning can refer to the use of symbolic and formalized intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0098] Inference is the process of simulating intelligent human reasoning in a computer or intelligent system, using formalized information to solve problems based on inference control policies. Typical functions of inference are searching and matching.

[0099] Decision making is a process in which decisions are made after intelligent information reasoning, usually providing functions such as classification, ranking, and prediction.

[0100] (4) General abilities After the above-mentioned data processing is performed on the data, some general capabilities may be further formed based on the data processing results, such as algorithms or general systems, to perform translation, text analysis, computer vision processing, voice recognition, image recognition, etc.

[0101] (5) Intelligent products and industrial applications Intelligent products and industrial applications refer to the products and applications of artificial intelligence systems in various fields, which encapsulate the entire artificial intelligence solution, commercialize intelligent information decision-making, and realize ground-breaking applications. Its application fields mainly include intelligent terminals, intelligent transportation, intelligent medical care, autonomous driving, safe cities, etc.

[0102] The present application may be applied to the image processing field in the field of artificial intelligence, and the following describes several application scenarios of the product landing.

[0103] I. Application to image compression processing in terminal devices The image compression method provided in the embodiments of the present application may be applied to an image compression process in a terminal device, specifically to an album, video surveillance, etc. in the terminal device. For details, please refer to FIG. 2a. FIG. 2a illustrates an application scenario according to an embodiment of the present application. As shown in FIG. 2a, the terminal device can acquire a photo to be compressed. The photo to be compressed may be a photo taken by a camera or a photo frame extracted from a video. The terminal device can use an artificial intelligence (AI) encoding unit in an embedded neural network processing unit (NPU) to perform feature extraction on the acquired photo to be compressed, convert the image data into output features with less redundancy, and generate probability estimates of points in the output features. A central processing unit (CPU) uses the probability estimates of each point in the output features to perform arithmetic coding on the extracted output features, reducing the coding redundancy of the output features and further reducing the amount of data transmission in the image compression process. The encoded data obtained after encoding is stored in a corresponding storage location in the form of a data file. When a user needs to retrieve a file stored in a storage location, the CPU can retrieve and load the stored file from the corresponding storage location, obtain a decoded feature map based on arithmetic decoding, and perform reconstruction on the feature map by using the AI ​​decoding unit in the NPU to obtain a reconstructed image.

[0104] 2. Application to image compression processing on the cloud The image compression method provided in the embodiments of the present application may be applied to an image compression process on the cloud side, specifically, to a function such as a cloud album in a cloud-side server. For details, see FIG. 2b. FIG. 2b illustrates an application scenario according to an embodiment of the present application. As shown in FIG. 2b, a terminal device may acquire a photo to be compressed. The photo to be compressed may be a photo taken by a camera or a photo frame extracted from a video. The terminal device may perform lossless encoding compression on the photo to be compressed by using a CPU to obtain encoded data. The lossless encoding compression may be performed based on, for example, but is not limited to, any lossless compression method in the prior art. The terminal device may send the encoded data to a cloud-side server. The server may perform corresponding lossless decoding on the received encoded data to obtain an image to be compressed. The server may use an AI encoding unit in a graphics processing unit (GPU) to perform feature extraction on the acquired photo to be compressed, converting image data into output features with lower redundancy and generating probability estimates of points in the output features. The CPU performs arithmetic coding on the extracted output features by using the probability estimates of points in the output features to reduce the coding redundancy of the output features and further reduce the amount of data transmission in the image compression process, and stores the coded data obtained after coding in a corresponding storage location in the form of a data file. When a user needs to retrieve the file stored in the storage location, the CPU retrieves and loads the stored file from the corresponding storage location, obtains a decoded feature map based on arithmetic decoding, and performs reconstruction on the feature map by using an AI decoding unit in the NPU to obtain a reconstructed image.The server can perform lossless encoding compression on the photo to be compressed by using a CPU to obtain encoded data, and the lossless encoding compression can be performed based on, for example, but is not limited to, any lossless compression method in the prior art. The server can send the encoded data to a terminal device, and the terminal device can perform corresponding lossless decoding on the received encoded data to obtain a decoded image.

[0105] In this embodiment of the present application, a step of performing gain processing on the feature values ​​in the feature map may be added between the AI ​​encoding unit and the quantization unit, and a step of performing inverse gain processing on the feature values ​​in the feature map may be added between the arithmetic decoding unit and the AI ​​decoding unit. Next, an image processing method in an embodiment of the present invention will be described in detail.

[0106] Since the embodiments of the present application relate to a large amount of neural network applications, for ease of understanding, the following first describes the relevant terms and concepts of neural networks that may be used in the embodiments of the present application.

[0107] (1) Neural Networks The neural network may include neurons, which may be computational units that use xs and intercept 1 as inputs, and the output of the computational units may be:

number

[0108] (2) Deep Neural Networks A deep neural network (DNN), also known as a multi-layer neural network, may be understood as a neural network with multiple hidden layers. DNNs are divided based on the location of different layers. The neural network within a DNN is classified into three types: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are hidden layers. The layers are fully connected. Specifically, any neuron in the i-th layer is always connected to any neuron in the (i+1)-th layer.

[0109] Although DNNs appear very complex, the function of each layer is not complicated. In other words, DNNs are linear relationships as follows:

number

number

number

number

number

number

number

number

[0110] In conclusion, the coefficient from the kth neuron in the (L-1)th layer to the jth neuron in the Lth layer is

number

[0111] Note that there is no parameter W in the input layer. In a deep neural network, the more hidden layers there are, the more capable the network is at describing complex cases in the real world. Theoretically, a model with more parameters is more complex and has a larger "capacity," which indicates that the model can complete more complex learning tasks. Training a deep neural network is the process of learning weight matrices, and the ultimate goal of training is to obtain the weight matrices of all layers of the trained deep neural network (weight matrices formed by the vector W of multiple layers).

[0112] (3) Convolutional Neural Networks A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A CNN includes a feature extractor, which includes a convolutional layer and a subsampling layer. The feature extractor can be considered a filter. A convolutional layer is a neuron layer in a CNN that performs convolutional processing on an input signal. In a convolutional layer of a CNN, a neuron may be connected to only a portion of neurons in an adjacent layer. A convolutional layer generally includes several feature planes, each of which includes neurons arranged in several rectangles. Neurons in the same feature plane share weights, which are convolution kernels. Weight sharing may be understood as a position-independent method of extracting image information. The convolution kernels may be initialized in the form of a matrix of random size. During the training process of a CNN, appropriate weights for the convolution kernels can be obtained through learning. Additionally, weight sharing is advantageous because it reduces the connections between layers of a convolutional neural network, reducing the risk of overfitting.

[0113] (4) Loss function In the process of training a deep neural network, the current network prediction may be compared with the expected target value, since the output of the deep neural network is expected to be as close as possible to the actual predicted value. Then, the weight vectors in each layer of the neural network are updated based on the difference between the current prediction and the target value. (Usually, there is an initialization process before the first update; in other words, parameters are preset for each layer of the deep neural network.) For example, if the network prediction is large, the weight vector is adjusted to lower the prediction until the deep neural network can predict the actual expected target value or a value close to the actual expected target value. Therefore, how to obtain the difference between the predicted value and the target value through comparison must be predefined. This is the loss function or objective function. Loss functions and objective functions are important formulas used to measure the difference between the predicted value and the target value. The loss function is used as an example. A larger output value (loss) of the loss function indicates a larger difference. Therefore, training a deep neural network is a process of minimizing loss as much as possible.

[0114] (5) Backpropagation Algorithm During the training process, the neural network can correct the parameter values ​​in the initial neural network model by using the back propagation (BP) algorithm, thereby making the reconstruction error loss of the neural network model smaller and smaller. Specifically, the input signal is forward-transferred until an error loss is generated at the output, and the parameters of the initial neural network model are updated based on the back propagation error loss information to make the error loss smaller. The back propagation algorithm is a backpropagation operation that mainly relies on the error loss, and aims to obtain the optimal neural network model parameters, such as the weight matrix.

[0115] The embodiments of the present application first provide an explanation by using an example in which the application scenario is a terminal device.

[0116] For example, the terminal device may be a mobile phone, a tablet computer, a notebook computer, or an intelligent wearable device, and the terminal device may perform compression processing on the captured photo. In another example, the terminal device may be a virtual reality (VR) device. As another example, the embodiments of the present application may also be applied to intelligent monitoring. In intelligent monitoring, a camera may be configured. In this case, in intelligent monitoring, a photo to be compressed may be captured by using the camera. It should be understood that the embodiments of the present application may also be applied to other scenarios in which image compression needs to be performed. Other application scenarios will not be listed one by one here.

[0117] 3 shows an embodiment of an image processing method according to an embodiment of the present application. As shown in FIG. 3, the image processing method provided in this embodiment of the present application includes the following steps:

[0118] 301. Acquire a first image.

[0119] In this embodiment of the present application, the first image is an image to be compressed. The first image may be an image taken by the aforementioned terminal device by using a camera, or the first image may be an image acquired from the terminal device (for example, an image stored in the album of the terminal device, or a photo acquired from the cloud by the terminal device). It should be understood that the first image may be an image having image compression requirements, and the source of the image to be processed is not limited in the present application.

[0120] 302. Perform feature extraction on the first image to obtain at least one first feature map, where the at least one first feature map includes N first feature values, where N is a positive integer.

[0121] In this embodiment of the present application, optionally, the terminal device can perform feature extraction on the first image based on CNN to obtain at least one first feature map. Hereinafter, the first feature map may also be referred to as a channel-wise feature map, and each semantic channel corresponds to one first feature map (channel-wise feature map).

[0122] In this embodiment of the present application, Fig. 4 shows a CNN-based image processing process. Fig. 4 shows a first image 401, a CNN 402, and a plurality of first feature maps 403. The CNN 402 can include a plurality of CNN layers.

[0123] For example, the CNN 402 may multiply the top-left 3x3 pixels of the input data (first image) by a weight and map it to the top-left neuron of the first feature map. The weight is also 3x3. Then, in a similar process, the CNN 402 scans the input data (first image) from left to right and top to bottom, multiplying the input data by a weight and mapping it to the neuron of the feature map. In this specification, the 3x3 weights used are referred to as a filter or filter core. That is, the process of applying a filter to the CNN 402 is a process of performing a convolution operation using a filter core, and the extracted result is referred to as a "first feature map." The first feature map may also be referred to as a multi-channel-wise feature map, and the term "multi-channel-wise feature map" may refer to a set of feature maps corresponding to multiple channels. According to one embodiment, the multi-channel-wise feature map may be generated by the CNN 402, which is also referred to as a "feature extraction layer" or "convolutional layer" of the CNN. A layer of the CNN may define a mapping from output to input. The mapping defined by a layer is implemented as one or more filter cores (convolution cores) applied to the input data, producing a feature map that is output to the next layer. The input data can be an image or a feature-mapped image for a particular layer.

[0124] See Figure 4. During forward execution, a CNN 402 receives a first image 401 and generates a multi-channel-wise feature map 403 as an output. Additionally, during forward execution, a next layer 402 receives a multi-channel-wise feature map 403 as an input and generates a multi-channel-wise feature map 403 as an output. Each subsequent layer then receives the multi-channel-wise feature map generated in the previous layer and generates the next multi-channel-wise feature map as an output. Finally, the multi-channel-wise feature map generated in the (N)th layer is received.

[0125] Additionally, other processing operations may be performed in addition to applying a convolution core to map input feature maps to output feature maps. Examples of other processing operations may include, but are not limited to, applying activation functions, pooling, resampling, etc.

[0126] It should be noted that the above is only one embodiment for performing feature extraction on the first image, and in practical applications, the specific embodiment of feature extraction is not limited.

[0127] In this embodiment of the present application, in the above-mentioned method, an original image (first image) is transformed into another space (at least one first feature map) by using a CNN convolutional neural network. Optionally, there are 192 feature maps, i.e., there are 192 semantic channels, and each semantic channel corresponds to one first feature map. In this embodiment of the present application, the at least one first feature map may be in the form of a three-dimensional tensor, and the size of the tensor may be 192×w×h, where w×h is the width and length of the matrix corresponding to the first feature map of a single channel.

[0128] In this embodiment of the present application, feature extraction may be performed on the first image to obtain a plurality of feature values. At least one first feature map may include some or all of the plurality of feature values. Gain processing may not be performed on feature maps corresponding to some semantic channels that have a relatively small impact on the compression result. In this case, the at least one first feature map may include some of the plurality of feature values.

[0129] In this embodiment of the present application, the at least one first feature map includes N first feature values, where N is a positive integer.

[0130] 303. Obtain a target compression bit rate, where the target compression bit rate corresponds to M target gain values, each target gain value corresponds to one first feature value, and M is a positive integer less than or equal to N.

[0131] In this embodiment of the present application, the terminal device can obtain a target compression bit rate. The target compression bit rate may be specified by a user or may be determined by the terminal device based on the first image. This is not limited here.

[0132] In this embodiment of the present application, the target compression bit rate corresponds to M target gain values, each target gain value corresponds to one first feature value, and M is a positive integer less than or equal to N. That is, there is a certain correlation between the target compression bit rate and the M target gain values, and after obtaining the target compression bit rate, the terminal device can determine the M corresponding target gain values ​​based on the obtained target compression bit rate.

[0133] Optionally, in one embodiment, the terminal device can determine M target gain values ​​corresponding to the target compression bitrate based on the target mapping relationship. The target mapping relationship is used to indicate the correlation between the compression bitrate and the M target gain values. The target mapping relationship may be a pre-stored mapping relationship. After obtaining the target compression bitrate, the terminal device can directly find the target mapping relationship corresponding to the target compression bitrate in a corresponding storage location.

[0134] Optionally, in one embodiment, the target mapping relationship may include a plurality of compression bit rates, a plurality of gain vectors, and a correlation between the plurality of compression bit rates and the plurality of gain vectors, where the target compression bit rate is one of the plurality of compression bit rates, and the M target gain values ​​are elements of one of the plurality of gain vectors.

[0135] In this embodiment of the present application, the target mapping relationship may be a preset table or another form. The target mapping relationship includes a plurality of compression bit rates and a gain vector corresponding to the compression bit rates. The gain vector may include a plurality of elements, where each compression bit rate corresponds to M target gain values, and the M target gain values ​​are elements included in the gain vector corresponding to each compression bit rate.

[0136] Optionally, in one embodiment, the target mapping relationship may comprise a target function mapping relationship, where an input of the target function relationship comprises a target compression bit rate and an output of the target function relationship comprises M target gain values.

[0137] In this embodiment of the present application, the target mapping relationship may be a preset target function mapping relationship or another form. The target function mapping relationship may indicate at least a correspondence relationship between a compression bit rate and a gain value. When the input of the target function relationship includes a target compression bit rate, the output of the target function relationship includes M target gain values.

[0138] It should be noted that in this embodiment of the present application, some or all of the M target gain values ​​may be the same. In this case, a number less than M may be used to indicate a target gain value corresponding to a first feature value among the M target feature values. For example, in one embodiment, the at least one first feature map includes a first target feature map, which includes P first feature values, all of which correspond to the same target gain value, where P is a positive integer less than or equal to M. That is, the P first feature values ​​are feature values ​​of the same semantic channel and correspond to the same target gain value. In this case, the P first feature values ​​may be indicated by using one gain value.

[0139] In another embodiment, if the gain values ​​of the first feature values ​​corresponding to each semantic channel are the same, the M first gain values ​​may be represented by using the same number of target gain values ​​as the number of semantic channels. Specifically, if there are 192 semantic channels (first feature maps), the M first gain values ​​may be represented by using 192 gain values.

[0140] In this embodiment of the present application, the first feature values ​​included in all or a portion of the at least one first feature map may correspond to the same target gain value. In this case, the at least one first feature map includes a first target feature map, which includes P first feature values, all of which correspond to the same target gain value, where P is a positive integer less than or equal to M. That is, the first target feature map is one of the at least one first feature map, which includes P first feature values, all of which correspond to the same target gain value.

[0141] In this embodiment of the present application, the N first feature values ​​may be all feature values ​​included in at least one first feature map. When M is equal to N, this corresponds to all feature values ​​included in the at least one first feature map each having a corresponding target gain value. When M is less than N, this corresponds to a portion of feature values ​​included in the at least one first feature map each having a corresponding target gain value. In one embodiment, when the number of first feature maps is greater than 1, all feature values ​​included in each of the portion of the at least one first feature map each have a corresponding target gain value, and a portion of feature values ​​included in each of the portion of the at least one first feature map each have a corresponding target gain value.

[0142] Optionally, in an embodiment, the first image includes a target object, and the M first feature values ​​are feature values ​​in at least one feature map that correspond to the target object.

[0143] In this embodiment of the present application, in some scenarios, the M first feature values ​​are feature values ​​corresponding to one or more target objects among the N first feature values, for example, for video content captured on a monitor, gain processing may not be performed on areas where the scene is relatively stationary, and gain processing may be performed on content of objects or people passing through the area.

[0144] 304. Process the corresponding first feature values ​​respectively based on the M target gain values ​​to obtain M second feature values.

[0145] In this embodiment of the present application, after the target compression bit rate and M target gain values ​​corresponding to the target compression bit rate are obtained, the corresponding first feature values ​​may be processed based on the M target gain values ​​respectively to obtain M second feature values. In one embodiment, the M second feature values ​​may be obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values ​​respectively, i.e., the corresponding second feature values ​​may be obtained after the first feature values ​​are multiplied by the corresponding target gain values.

[0146] In this embodiment of the present application, to implement the effects of different compression bit rates in the same AI compression model, different target gain values ​​may be obtained for different obtained target compression bit rates. After the corresponding first feature values ​​are processed based on the M target gain values ​​to obtain M second feature values, the distribution of the N first feature values ​​included in the at least one feature map corresponding to the original first image changes due to the M first feature values ​​that are subjected to gain processing.

[0147] In this embodiment of the present application, FIGS. 5a and 5b show the distribution of feature maps for different compression bit rates according to an embodiment of the present application. Different compression bit rates are represented by using different bits per pixel (bpp). bpp represents the number of bits used to store each pixel, and a smaller bpp indicates a smaller compression bit rate. FIG. 5a shows the distribution of N first feature values ​​when bpp is 1. FIG. 5b shows the distribution of N first feature values ​​when bpp is 0.15. The output features (N first feature values) of the encoding network of a model with a higher compression bit rate have a larger variance in the statistical histogram, and therefore, the information entropy obtained after quantization is greater. Therefore, provided that different compression bit rates correspond to different target gain values, gain processing is performed to different degrees on the N first feature values ​​based on different target compression bit rates, thereby enabling the reconstruction effect of multiple bit rates to be implemented in a single AI compression model. Specifically, the selection rules for M target gain values ​​are as follows: The larger the target compression bit rate, the more dispersed the N first feature values ​​obtained after the corresponding first feature values ​​are processed based on the M target gain values, respectively, and therefore the larger the information entropy obtained after quantization.

[0148] In this embodiment of the present application, after feature extraction is performed on the first image, all extracted first feature maps need to be processed to obtain multiple first feature maps. Feature values ​​included in the multiple first feature maps correspond to the same target gain value. In this case, all feature values ​​included in the multiple first feature maps are multiplied by the corresponding target gain value to change the distribution of the N first feature values ​​included in the multiple first feature maps. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0149] In this embodiment of the present application, after feature extraction is performed on the first image, all extracted first feature maps need to be processed to obtain multiple first feature maps. The feature values ​​included in each of the multiple first feature maps correspond to the same target gain value, i.e., each first feature map corresponds to one target gain value. In this case, to change the distribution of the N first feature values ​​included in the multiple first feature maps, the feature values ​​included in each of the multiple first feature maps are multiplied by the corresponding target gain value. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0150] In this embodiment of the present application, after feature extraction is performed on the first image, all extracted first feature maps need to be processed to obtain multiple first feature maps. The feature values ​​included in each of the portions of the first feature maps correspond to the same target gain value, and the feature values ​​included in each of the remaining portions of the first feature maps correspond to different target gain values; that is, each of the portions of the first feature maps corresponds to one target gain value, and each of the remaining portions of the first feature maps corresponds to multiple target gain values ​​(different feature values ​​in the same feature map may correspond to different target gain values). In this case, to change the distribution of the N first feature values ​​included in the multiple first feature maps, the feature values ​​included in each of the portions of the multiple first feature maps are multiplied by the corresponding target gain value, and the feature values ​​included in the remaining portions of the multiple first feature maps are multiplied by the corresponding target gain value. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0151] In this embodiment of the present application, after feature extraction is performed on the first image to obtain multiple first feature maps, some of the extracted first feature maps need to be processed (gain processing may not be performed on first feature maps corresponding to some semantic channels that have a relatively small impact on the compression result). The number of extracted first feature maps that need to be processed is greater than 1. The feature values ​​included in each of the multiple first feature maps correspond to the same target gain value, i.e., each first feature map corresponds to one target gain value. In this case, to change the distribution of the N first feature values ​​included in the multiple first feature maps, the feature values ​​included in each of the multiple first feature maps are multiplied by the corresponding target gain value. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0152] In this embodiment of the present application, after feature extraction is performed on the first image, some of the extracted first feature maps need to be processed to obtain multiple first feature maps (gain processing may not be performed on first feature maps corresponding to some semantic channels that have a relatively small impact on the compression result). The number of extracted first feature maps that need to be processed is greater than 1. Feature values ​​included in each of the part of the first feature maps correspond to the same target gain value, and feature values ​​included in each of the remaining parts of the first feature maps correspond to different target gain values; that is, each of the part of the first feature maps corresponds to one target gain value, and each of the remaining parts of the first feature maps corresponds to multiple target gain values ​​(different feature values ​​in the same feature map may correspond to different target gain values). In this case, to change the distribution of the N first feature values ​​included in the multiple first feature maps, the feature values ​​included in each of the part of the multiple first feature maps are multiplied by the corresponding target gain value, and the feature values ​​included in the remaining parts of the multiple first feature maps are multiplied by the corresponding target gain value. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0153] In this embodiment of the present application, after feature extraction is performed on the first image to obtain multiple first feature maps, some of the extracted first feature maps need to be processed (gain processing may not be performed on first feature maps corresponding to some semantic channels that have a relatively small impact on the compression result). The number of extracted first feature maps that need to be processed is equal to 1, and the feature values ​​included in the first feature maps correspond to the same target gain value, i.e., each first feature map corresponds to one target gain value. In this case, to change the distribution of the N first feature values ​​included in the multiple first feature maps, the feature values ​​included in the first feature maps are multiplied by the corresponding target gain value. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0154] In this embodiment of the present application, after feature extraction is performed on the first image to obtain multiple first feature maps, some of the extracted first feature maps need to be processed (gain processing may not be performed on first feature maps corresponding to some semantic channels that have a relatively small impact on the compression result). The number of extracted first feature maps that need to be processed is equal to 1, and the feature values ​​included in the first feature maps correspond to different target gain values, i.e., the first feature map corresponds to multiple target gain values ​​(different feature values ​​in the same feature map may correspond to different target gain values). In this case, to change the distribution of the N first feature values ​​included in the multiple first feature maps, the feature values ​​included in the first feature map are multiplied by the corresponding target gain value. A larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0155] It should be noted that the gain processing may be performed only on the first feature value included in the first feature map.

[0156] It should be noted that when the same scale of gain processing is performed on the feature values ​​of the semantic channels, i.e., when the first feature values ​​included in the multiple first feature maps corresponding to all the semantic channels correspond to the same target gain value, the information entropy of the N first feature values ​​may be changed, but the compression effect is relatively low. Therefore, the basic gain calculation unit is set at the semantic channel level (the first feature values ​​included in each of the first feature maps corresponding to at least two of all the semantic channels correspond to different target gain values) or the feature value level (at least two of all the first feature values ​​included in the first feature maps corresponding to the semantic channels correspond to different target gain values), so that a relatively good compression effect can be achieved.

[0157] The following describes how to obtain M target gain values ​​that can implement the aforementioned technical effects.

[0158] 1.Manual decision method In this embodiment of the present application, the objective function mapping relationship may be determined manually. If the first feature values ​​included in the first feature maps corresponding to each semantic channel correspond to the same target gain value, the input of the objective function mapping relationship may be the semantic channel and the target compression bit rate, and the output of the objective function mapping relationship is the corresponding target gain value (since the first feature values ​​included in the first feature maps correspond to the same target gain value, all target gain values ​​corresponding to the semantic channels can be represented using one target gain value). For example, the target gain value corresponding to each semantic channel may be determined using a linear function, a quadratic function, a cubic function, or a quartic function. FIG. 6 illustrates an objective function mapping relationship according to one embodiment of the present application. As shown in FIG. 6, the objective function mapping relationship is a linear function, the input of the function is the semantic channel sequence number (e.g., semantic channel sequence numbers 1 to 192), and the output of the function is the target mapping function. Each target compression bit rate corresponds to a different objective function mapping relationship. A larger target compression bit rate corresponds to a smaller slope of the objective function mapping relationship. The approximate distribution law of a quadratic or cubic nonlinear function is similar, and the details will not be described here.

[0159] In this embodiment of the present application, the target gain values ​​corresponding to the M first feature values ​​may be manually determined, and the specific setting manner is not limited in the present application, provided that a larger target compression bit rate indicates a more dispersed distribution of the N first feature values.

[0160] 2. Training method In this embodiment of the present application, obtaining M target gain values ​​corresponding to each target compression bit rate through training needs to be combined with the processing on the decoding side, so obtaining M target gain values ​​corresponding to each target compression bit rate through training will be described in detail in subsequent embodiments, and the details will not be described here.

[0161] 305. Quantize and entropy encode the at least one processed first feature map to obtain encoded data, where the at least one processed first feature map includes M second feature values.

[0162] In this embodiment of the present application, the corresponding first feature values ​​may be processed based on M target gain values ​​respectively to obtain M second feature values, and then quantization and entropy coding may be performed on the at least one processed first feature map to obtain encoded data, where the at least one processed first feature map includes the M second feature values.

[0163] In this embodiment of the present application, the N first feature values ​​are converted into quantization centers according to a specified rule to facilitate subsequent entropy encoding. The quantization operation can convert the N first feature values ​​from floating-point numbers into a bitstream (e.g., a bitstream using specific-bit integers, such as 8-bit integers or 4-bit integers). In some embodiments, the quantization operation may be performed on the N first feature values ​​by, but is not limited to, rounding.

[0164] In this embodiment of the present application, the information entropy of the quantized data obtained by quantizing at least one processed first feature map satisfies a predetermined condition, and the predetermined condition is related to a target compression bit rate. Specifically, a larger target compression bit rate indicates a larger information entropy of the quantized data.

[0165] In this embodiment of the present application, probability estimates of points in the output features may be obtained by using an entropy estimation network, and entropy coding is performed on the output features by using the probability estimates to obtain a binary bitstream. It should be noted that the entropy coding process in this application may use existing entropy coding techniques, and the details will not be described in this application.

[0166] In this embodiment of the present application, the difference between the compression bit rate corresponding to the encoded data and the target compression bit rate falls within a preset range. The preset range may be selected in practical applications. As long as the difference between the compression bit rate corresponding to the encoded data and the target compression bit rate falls within an acceptable range, the specific preset range is not limited in the present application.

[0167] In this embodiment of the present application, after the encoded data is obtained, the encoded data may be sent to a terminal device for decompression. In this case, the image processing device for decompression can decompress the data. Alternatively, the terminal device for compression can store the encoded data in a storage device. When the encoded data is needed, the terminal device can obtain the encoded data from the storage device and decompress the encoded data.

[0168] Optionally, in an embodiment, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to M first gain values, the second compression bit rate corresponds to M second gain values, and the M target gain values ​​are obtained by performing an interpolation operation on the M first gain values ​​and the M second gain values. In this embodiment of the present application, the M first values ​​include a first target gain value, the M second gain values ​​include a second target gain value, and the M target gain values ​​include a third target gain value, the first target gain value, the second target gain value, and the third target gain value correspond to the same feature value among the M first feature values, and the third target gain value is obtained by performing an interpolation operation on the first target gain value and the second target gain value.

[0169] In this embodiment of the present application, compression effects for multiple compression bit rates may be implemented in a single model. Specifically, to implement compression effects for different compression bit rates, different target gain values ​​may be set correspondingly for multiple target compression bit rates. Then, to obtain a new gain value for any compression effect within the compression bit rate range, an interpolation operation may be performed on the target gain values ​​by using an interpolation algorithm. Specifically, the M first gain values ​​include a first target gain value, the M second gain values ​​include a second target gain value, and the M target gain values ​​include a third target gain value, where the first target gain value, the second target gain value, and the third target gain value correspond to the same feature value among the M first feature values, and the third target gain value is obtained by performing an interpolation operation on the first target gain value and the second target gain value. The interpolation operation may be performed based on the following formula: m l =[(m i ) l ·(m j ) 1-l ], where m l represents the third target gain value, and m i represents the first target gain value, and mj represents the second target gain value, and m l , m i , and m j correspond to the same feature value, and l∈(0,1) is an adjustment factor, which may be determined based on the size of the target compressed bitrate.

[0170] In this embodiment of the present application, after M target gain values ​​corresponding to each of a plurality of compression bit rates are obtained, when compression corresponding to the target compression bit rate is performed, two groups of target gain values ​​(each group including M target gain values) corresponding to two compression bit rates adjacent to the target compression bit rate may be determined from the plurality of compression bit rates, and the above-mentioned interpolation process is performed on the two groups of target gain values ​​to obtain the M target gain values ​​corresponding to the target compression bit rate. In this embodiment of the present application, any compression effect of the AI ​​compression model in the compression bit rate interval may be implemented.

[0171] In this embodiment of the present application, each of the M target gain values ​​corresponds to one inverse gain value, and the inverse gain value is used to process the feature value obtained in the decoding process of the encoded data, and the product of each of the M target gain values ​​and the corresponding inverse gain value falls within a preset range. The inverse gain process on the decoding side will be described in a subsequent embodiment, and the details will not be described here.

[0172] This embodiment of the present application provides an image processing method, comprising: acquiring a first image; performing feature extraction on the first image to obtain at least one first feature map; the at least one first feature map including N first feature values, where N is a positive integer; obtaining a target compression bit rate; the target compression bit rate corresponding to M target gain values, each target gain value corresponding to one first feature value, where M is a positive integer less than or equal to N; processing the corresponding first feature value based on the M target gain values ​​to obtain M second feature values; and performing quantization and entropy coding on the at least one processed first feature map to obtain encoded data; the at least one processed first feature map including the M second feature values. In the above-described method, different target gain values ​​are set for different target compression bit rates to implement compression bit rate control.

[0173] 7 shows an embodiment of an image processing method according to an embodiment of the present application. As shown in FIG. 7, the image processing method provided in this embodiment includes the following steps:

[0174] 701. Obtain encoded data.

[0175] In this embodiment of the present application, the encoded data obtained in FIG. 3 and the corresponding embodiment may be obtained.

[0176] In this embodiment of the present application, after the encoded data is obtained, the encoded data may be transmitted to a terminal device for decompression. In this case, the image processing device for decompression can obtain and decompress the encoded data. Alternatively, the terminal device for compression can store the encoded data in a storage device. When the encoded data is needed, the terminal device can obtain the encoded data from the storage device and decompress the encoded data.

[0177] 702. Perform entropy decoding on the encoded data to obtain at least one second feature map, where the at least one second feature map includes N third feature values, where N is a positive integer.

[0178] In this embodiment of the present application, the encoded data may be decoded by using a conventional entropy decoding technique to obtain reconstructed output features (at least one second feature map), where the at least one second feature map includes N third feature values.

[0179] It should be noted that the at least one second feature map in this embodiment of the present application may be the same as the at least one processed first feature map described above.

[0180] 703. Obtain M target inverse gain values, each target inverse gain value corresponding to one third feature value, where M is a positive integer less than or equal to N.

[0181] Optionally, in one embodiment, a target compression bitrate may be obtained, and M target inverse gain values ​​corresponding to the target compression bitrate may be determined based on a target mapping relationship. The target mapping relationship is used to indicate a correlation between the compression bitrate and the inverse gain vector. The target mapping relationship includes a plurality of compression bitrates, a plurality of inverse gain vectors, and a correlation between the plurality of compression bitrates and the plurality of inverse gain vectors, where the target compression bitrate is one of the plurality of compression bitrates and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors, or the target mapping relationship includes a target function mapping relationship, where an input of the target function relationship includes the target compression bitrate and an output of the target function relationship includes the M target inverse gain values.

[0182] In this embodiment of the present application, the target inverse gain value may be obtained in the step of obtaining the target gain value in the embodiment corresponding to Fig. 3. This is not limited here.

[0183] Optionally, in one embodiment, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, and P is a positive integer less than or equal to M.

[0184] Optionally, in an embodiment, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0185] 704. Process the corresponding third feature values ​​respectively based on the M target inverse gain values ​​to obtain M fourth feature values.

[0186] In this embodiment of the present application, the M fourth feature values ​​may be obtained by individually performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values. Specifically, in this embodiment of the present application, to obtain the M fourth feature values, the M third feature values ​​in at least one second feature map are respectively multiplied by the corresponding inverse gain values, so that the at least one second feature map obtained after the inverse gain processing includes the M fourth feature values. The inverse gain processing is combined with the gain processing in the embodiment corresponding to FIG. 3, so that successful image analysis can be ensured.

[0187] 705. Perform image reconstruction on the at least one processed second feature map to obtain a second image, where the at least one processed second feature map includes M fourth feature values.

[0188] In this embodiment of the present application, after the M fourth feature values ​​are obtained, image reconstruction may be performed on the at least one processed second feature map to obtain a second image. The at least one processed second feature map includes the M fourth feature values. The at least one second feature map is analyzed in the above-mentioned manner and reconstructed into a second image.

[0189] Optionally, in one embodiment, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to M first inverse gain values, the second compression bit rate corresponds to M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values. In this embodiment of the present application, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, and the M target inverse gain values ​​include a third target inverse gain value, the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0190] In this embodiment of the present application, each of the M target gain values ​​corresponds to one inverse gain value, and the inverse gain value is used to process the feature value obtained in the decoding process of the encoded data, and the product of each of the M target gain values ​​and the corresponding inverse gain value falls within a preset range, that is, there is a specific value relationship between the target gain value and the inverse gain value corresponding to the same feature value as follows: The product of the two values ​​falls within a preset range. The preset range may be a value range close to the value "1", and is not limited here.

[0191] This embodiment of the present application provides an image processing method. Encoded data is obtained, and entropy decoding is performed on the encoded data to obtain at least one second feature map, where the at least one second feature map includes N third feature values, where N is a positive integer, and M target inverse gain values ​​are obtained, where each target inverse gain value corresponds to one third feature value, where M is a positive integer less than or equal to N. To obtain M fourth feature values, the corresponding third feature values ​​are respectively processed based on the M target inverse gain values. Image reconstruction is performed on the at least one processed second feature map to obtain a second image, where the at least one processed second feature map includes the M fourth feature values. In the above-described method, different target inverse gain values ​​are set for different target compression bit rates to implement compression bit rate control.

[0192] Next, the architecture of a variational autoencoder (VAE) is used as an example to explain the image compression method provided in the embodiments of this application. A variational autoencoder is an autoencoder used for data compression or noise reduction.

[0193] FIG. 8 illustrates an image compression procedure according to an embodiment of the present application.

[0194] This embodiment provides an explanation by using an example in which the target gain values ​​corresponding to the same semantic channel are the same, and the target inverse gain values ​​corresponding to the same semantic channel are the same. There are 192 semantic channels, and training needs to be performed at four specified code points (four compression bit rates) during training. Each compression bit rate corresponds to one target gain vector and one target inverse gain vector. The target gain vector m i is a vector of size 192 × 1, corresponding to the compression bit rate.

number

number

[0195] 801. Obtain output features y after a first image enters the encoding network.

[0196] 802. Output features obtained after gain processing

number

[0197] 803.Features

number

number

[0198] 804. Obtain probability estimates of points in the output features by using an entropy estimation module, and perform entropy coding on the output features by using the probability estimates to obtain a binary bitstream.

[0199] 805. Reconstructed output features

number

[0200] 806. To obtain the output feature y' obtained after the inverse gain processing, the output feature

number

number

[0201] 807. After the output features enter the decoding network, the output features y′ are analyzed and reconstructed into a second image.

[0202] Please refer to Figure 9. The left graph in Figure 9 shows a comparison of the rate-distortion performance of the single model of this embodiment (non-dashed line) with the rate-distortion performance of four compression models separately trained using the VAE method of the prior art (dashed line) under the condition that a multi-scale structural similarity index measure (MS-SSIM) is used as the evaluation metric, where the abscissa is BPP and the ordinate is MS-SSIM. The right graph in Figure 9 shows a comparison of the rate-distortion performance of the single model of this embodiment (non-dashed line) with the rate-distortion performance of four compression models separately trained using the VAE method of the prior art (dashed line) under the condition that peak signal-to-noise ratio (PSNR) is used as the evaluation metric, where the abscissa is BPP and the ordinate is PSNR. In this embodiment, on the premise that the number of model parameters is basically consistent with the number of model parameters of a single model in the VAE method, it can be seen that the compression effect of any bit rate can be implemented based on both evaluation indexes, the compression effect is not worse than the multi-model implementation effect of the VAE method, and the model storage amount can be reduced by N times (N is the number of models required in the VAE method to implement the compression effect of different bit rates in this embodiment of the present invention).

[0203] 10 shows the training process according to one embodiment of the present application. As shown in FIG. 10, the loss function of the model in this embodiment is as follows: loss=l d +β·l r , where l d is the distortion loss of the second image relative to the first image, calculated based on the evaluation index, and l r is the bitrate loss (or called bitrate estimate) obtained by the entropy estimation network through calculation, and β is a Lagrangian coefficient for adjusting the tradeoff between distortion loss and bitrate estimate.

[0204] To obtain the gain matrix and inverse gain matrix {M,M'} corresponding to different compression bit rates, the model training process may be shown in Figure 10. The Lagrangian coefficient β in the loss function is constantly transformed in the model training process, and the corresponding gain and inverse gain vectors

number

[0205] For example, the compression effects of four compression bit rates can be implemented in a single model. The four gain vectors obtained by training are multiplied by the corresponding inverse gain vectors. The multiplication results of the corresponding elements in the target gain vector and the target inverse gain vector corresponding to different compression bit rates are approximately equal, so that the following relationship can be obtained:

number

number

number

[0206] To implement continuous bitrate adjustment in a single model, in this embodiment, the following derivation may be performed by using the above formula:

number

[0207] In this embodiment of the present application, an interpolation operation may be performed on four adjacent gain and inverse gain vector pairs obtained through training to obtain a new gain and inverse gain vector pair.

[0208] To obtain the gain matrix M matching different compression bit rates, the training process is as follows: In this embodiment, the Lagrangian coefficients in the loss function are constantly transformed in the model training process, and the corresponding gain vector m i and the inverse gain vector

number

number

[0209] In this embodiment of the present application, the gain vector m i and the inverse gain vector

number

[0210] In this embodiment, on the premise that the number of model parameters basically matches the number of model parameters of a single VAE method model, the compression effect of any bit rate can be implemented, the compression effect is not worse than the effect of independent training at each bit rate, and the model storage amount can be reduced by N times (N is the number of models required to implement the compression effect of different bit rates of this embodiment of the present invention in the VAE method).

[0211] Please note that only VAE is used above as an architecture for explanation. In practical application, the image compression method may be further applied to another AI compression model architecture (for example, an auto-encoder or another image compression model). This is not limited in this application.

[0212] 12 is a diagram of a system architecture of an image processing system according to one embodiment of the present application. In FIG. 12, an image processing system 200 includes an execution device 210, a training device 220, a database 230, a client device 240, and a data storage system 250. The execution device 210 includes a computation module 211.

[0213] The database 230 stores a set of first images. The training device 220 generates target models / rules 201 used to process the first images, and performs iterative training on the target models / rules 201 by using the first images in the database to obtain mature target models / rules 201. This embodiment of the present application provides an explanation by using an example in which the target models / rules 201 includes a second encoding / decoding network, M target gain values ​​and M target inverse gain values ​​corresponding to each compression bit rate.

[0214] The second encoding / decoding network and the M target gain values ​​and M target inverse gain values ​​corresponding to each compression bit rate obtained by training device 220 may be applied to different systems or devices, such as a mobile phone, a tablet computer, a notebook computer, a VR device, or a surveillance system. Execution device 210 may access or store data, instructions, etc. in data storage system 250. Data storage system 250 may be located within execution device 210, or data storage system 250 may be external memory to execution device 210.

[0215] The computing module 211 performs feature extraction on the first image received by the client device 240 by using a second encoding / decoding network to obtain at least one first feature map, the at least one first feature map including N first feature values, where N is a positive integer; obtain a target compression bit rate; the target compression bit rate corresponds to M target gain values, each target gain value corresponding to one first feature value, where M is a positive integer less than or equal to N; process the corresponding first feature values ​​based on the M target gain values ​​to obtain M second feature values; and perform quantization and entropy coding on the at least one processed first feature map to obtain encoded data, where the at least one processed first feature map includes the M second feature values.

[0216] The calculation module 211 further performs entropy decoding on the encoded data by using a second encoding / decoding network to obtain at least one second feature map, where the at least one second feature map includes N third feature values, where N is a positive integer, obtains M target inverse gain values, where each target inverse gain value corresponds to one third feature value, where M is a positive integer less than or equal to N, processes the corresponding third feature values ​​based on the M target inverse gain values ​​to obtain M fourth feature values, and performs image reconstruction on the at least one processed second feature map to obtain a second image, where the at least one processed second feature map includes the M fourth feature values.

[0217] In some embodiments of the present application, see Figure 12. The execution device 210 and the client device 240 may be independent devices. An I / O interface 212 is configured in the execution device 210 to exchange data with the client device 240. A "user" can input a first image into the I / O interface 212 by using the client device 240, and the execution device 210 returns the second image to the client device 240 by using the I / O interface 212 to provide the second image to the user.

[0218] It should be noted that FIG. 12 is merely a schematic diagram of the architecture of an image processing system according to one embodiment of the present invention, and the location of devices, components, modules, etc. shown in the diagram does not constitute any limitation. For example, in some other embodiments of the present application, the execution device 210 may be configured within the client device 240. For example, if the client device is a mobile phone or tablet computer, the execution device 210 may be a module configured to process array images within a host central processing unit (Host CPU) of the mobile phone or tablet computer, or the execution device 210 may be a graphics processing unit (GPU) or neural network processing unit (NPU) within the mobile phone or tablet computer. The GPU or NPU is installed in the host central processing unit as a coprocessor, and the host central processing unit assigns tasks to the GPU or NPU.

[0219] With reference to the foregoing description, the following begins with describing the specific implementation procedure of the training stage of the image processing method provided in the embodiment of the present application.

[0220] 1. Training Phase For details, please refer to Figure 13. Figure 13 is a schematic flow chart of an image processing method according to an embodiment of the present application. The image processing method provided in this embodiment of the present application may include the following steps:

[0221] 1301. Acquire a first image.

[0222] 1302. Perform feature extraction on the first image based on the encoding network to obtain at least one first feature map, where the at least one first feature map includes N first feature values, where N is a positive integer.

[0223] 1303. Obtain a target compression bit rate, where the target compression bit rate corresponds to M initial gain values ​​and M initial inverse gain values, each initial gain value corresponds to one first feature value, each initial inverse gain value corresponds to one third feature value, and M is a positive integer less than or equal to N.

[0224] 1304. Process the corresponding first feature values ​​based on the M initial gain values ​​respectively to obtain M second feature values.

[0225] 1305. Perform quantization and entropy coding on the at least one processed first feature map based on a quantization network and an entropy coding network to obtain coded data and bit rate loss, where the at least one first feature map obtained after gain processing includes M second feature values.

[0226] 1306. Perform entropy decoding on the encoded data based on the entropy decoding network to obtain at least one second feature map, where the at least one second feature map includes M third feature values, and each third feature value corresponds to one first feature value.

[0227] 1307. Process the corresponding third feature values ​​respectively based on the M initial inverse gain values ​​to obtain M fourth feature values.

[0228] 1308. Perform image reconstruction on the at least one processed second feature map based on the decoding network to obtain a second image, where the at least one processed feature map includes M fourth feature values.

[0229] 1309. Obtain the distortion loss of the second image relative to the first image.

[0230] 1310. Perform joint training on a first encoding / decoding network, M initial gain values, and M initial inverse gain values ​​by using a loss function until an image distortion value between a first image and a second image reaches a first predetermined degree, wherein the image distortion value is related to a bitrate loss and a distortion loss, and the encoding / decoding network includes a coding network, a quantization network, an entropy coding network, and an entropy decoding network.

[0231] 1311. A second encoding / decoding network, outputting M target gain values ​​and M target inverse gain values, wherein the second encoding / decoding network is a model obtained after iterative training is performed on the first encoding / decoding network, and the M target gain values ​​and M target inverse gain values ​​are obtained after iterative training is performed on the M initial gain values ​​and M initial inverse gain values.

[0232] Please refer to the description in the above embodiment for a specific description of steps 1301 to 1311. This is not limited here.

[0233] Optionally, an information entropy of the quantized data obtained by quantizing the at least one processed first feature map satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0234] Optionally, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0235] Optionally, the M second feature values ​​are obtained by performing multiplication operations on the M initial gain values ​​and the corresponding first feature values ​​respectively.

[0236] Optionally, the M fourth feature values ​​are obtained by performing multiplication operations on the M initial inverse gain values ​​and the corresponding third feature values ​​respectively.

[0237] Optionally, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range, and the product of each of the M initial gain values ​​and the corresponding initial inverse gain value falls within a preset range.

[0238] According to the embodiment corresponding to FIGS. 1 to 13, in order to better implement the aforementioned solutions in the embodiments of the present application, the following further provides a related device configured to implement the aforementioned solutions. Please refer to FIG. 14 for details. FIG. 14 is a schematic diagram of the configuration of an image processing device 1400 according to an embodiment of the present application. The image processing device 1400 may be a terminal device or a server, and the image processing device 1400: an acquisition module 1401 configured to acquire a first image; a feature extraction module 1402 configured to perform feature extraction on the first image to obtain at least one first feature map, the at least one first feature map comprising N first feature values, where N is a positive integer; the obtaining module 1401 is further configured to obtain a target compression bit rate, the target compression bit rate corresponding to M target gain values, each target gain value corresponding to one first feature value, M being a positive integer less than or equal to N; a gain module 1403 configured to process the corresponding first feature values, respectively, based on the M target gain values ​​to obtain M second feature values; a quantization and entropy coding module 1404 configured to perform quantization and entropy coding on the at least one processed first feature map to obtain coded data, wherein the at least one processed first feature map includes M second feature values; Equipped with.

[0239] Optionally, an information entropy of the quantized data obtained by quantizing the at least one processed first feature map satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0240] Optionally, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0241] Optionally, the difference between the compression bit rate corresponding to the encoded data and the target compression bit rate falls within a preset range.

[0242] Optionally, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values ​​respectively.

[0243] Optionally, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, where P is a positive integer less than or equal to M.

[0244] Optionally, the apparatus comprises: a determination module configured to determine M target gain values ​​corresponding to a target compression bit rate based on a target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bit rate and the M target gain values; the target mapping relationship includes a plurality of compression bit rates, a plurality of gain vectors, and correlations between the plurality of compression bit rates and the plurality of gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates and the M target gain values ​​are elements of one of the plurality of gain vectors; or If the goal mapping relationship includes a goal function mapping relationship, and the input of the goal function relationship includes a target compression bit rate, then the output of the goal function relationship includes M target gain values.

[0245] Optionally, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first gain values, and the second compression bit rate corresponds to the M second gain values, and the M target gain values ​​are obtained by performing an interpolation operation on the M first gain values ​​and the M second gain values.

[0246] Optionally, the M first gain values ​​include a first target gain value, the M second gain values ​​include a second target gain value, and the M target gain values ​​include a third target gain value, wherein the first target gain value, the second target gain value, and the third target gain value correspond to the same feature value among the M first feature values, and the third target gain value is obtained by performing an interpolation operation on the first target gain value and the second target gain value.

[0247] Optionally, the first image includes a target object, and the M first feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0248] Optionally, each of the M target gain values ​​corresponds to one inverse gain value, and the inverse gain value is used to process the feature value obtained in the decoding process of the encoded data, and the product of each of the M target gain values ​​and the corresponding inverse gain value falls within a preset range.

[0249] Optionally, the apparatus comprises: a decoding module configured to perform entropy decoding on the encoded data to obtain at least one second feature map, the at least one second feature map including N third feature values, each third feature value corresponding to one first feature value; the obtaining module is further configured to obtain M target inverse gain values, each target inverse gain value corresponding to one third feature value; This device is an inverse gain module configured to perform gain processing on corresponding third feature values ​​based on the M target inverse gain values, respectively, to obtain M fourth feature values; a reconstruction module configured to perform image reconstruction on at least one second feature map obtained after the inverse gain processing to obtain a second image, wherein the at least one second feature map obtained after the inverse gain processing includes M fourth feature values; Further provided are:

[0250] Optionally, the M fourth feature values ​​are obtained by performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values ​​respectively.

[0251] Optionally, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, where P is a positive integer less than or equal to M.

[0252] Optionally, the decision module: The method is further configured to determine M target inverse gain values ​​corresponding to the target compression bit rate based on the target mapping relationship, wherein the target mapping relationship is used to indicate a correlation between the compression bit rate and the inverse gain vector.

[0253] Optionally, the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and a correlation relationship between the plurality of compression bit rates and the plurality of inverse gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates, and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors.

[0254] Optionally, the target mapping relationship comprises an objective function mapping relationship, where an input of the objective function relationship comprises a target compression bit rate and an output of the objective function relationship comprises M target inverse gain values.

[0255] Optionally, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0256] Optionally, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range.

[0257] Optionally, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, and the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0258] Optionally, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, and the M target inverse gain values ​​include a third target inverse gain value, wherein the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0259] This embodiment of the present application provides an image processing apparatus 1400. The acquisition module 1401 acquires a first image. The feature extraction module 1402 performs feature extraction on the first image to obtain at least one first feature map, where the at least one first feature map includes N first feature values, where N is a positive integer. The acquisition module 1401 acquires a target compression bit rate, where the target compression bit rate corresponds to M target gain values, where each target gain value corresponds to one first feature value, where M is a positive integer less than or equal to N. The gain module 1403 processes corresponding first feature values ​​based on the M target gain values ​​to obtain M second feature values. The quantization and entropy coding module 1404 performs quantization and entropy coding on the at least one processed first feature map to obtain encoded data, where the at least one processed first feature map includes M second feature values. In the above-mentioned scheme, different target gain values ​​are set for different target compression bit rates to implement compression bit rate control.

[0260] 15 is a schematic diagram of the configuration of an image processing device 1500 according to an embodiment of the present invention. The image processing device 1500 may be a terminal device or a server. an acquisition module 1501 configured to acquire encoded data; a decoding module 1502 configured to perform entropy decoding on the encoded data to obtain at least one second feature map, the at least one second feature map including N third feature values, where N is a positive integer; The obtaining module 1501 is further configured to obtain M target inverse gain values, each target inverse gain value corresponding to one third feature value, where M is a positive integer less than or equal to N; an inverse gain module 1503 configured to process the corresponding third feature values, respectively, based on the M target inverse gain values ​​to obtain M fourth feature values; a reconstruction module 1504 configured to perform image reconstruction on the at least one processed second feature map to obtain a second image, wherein the at least one processed second feature map includes M fourth feature values; Equipped with.

[0261] Optionally, the M fourth feature values ​​are obtained by performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values ​​respectively.

[0262] Optionally, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, where P is a positive integer less than or equal to M.

[0263] Optionally, the acquisition module is further configured to acquire a target compression bitrate; This device is a determining module for determining M target inverse gain values ​​corresponding to a target compression bit rate based on a target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bit rate and the inverse gain vector; the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and correlations between the plurality of compression bit rates and the plurality of inverse gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors; or If the target mapping relationship includes a target function mapping relationship, and the input of the target function relationship includes a target compression bit rate, then the output of the target function relationship includes M target inverse gain values.

[0264] Optionally, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0265] Optionally, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, and the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0266] Optionally, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, and the M target inverse gain values ​​include a third target inverse gain value, wherein the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0267] This embodiment of the present invention provides an image processing apparatus. An acquisition module 1501 acquires encoded data. A decoding module 1502 performs entropy decoding on the encoded data to obtain at least one second feature map, where the at least one second feature map includes N third feature values, where N is a positive integer. The acquisition module 1501 acquires M target inverse gain values, where each target inverse gain value corresponds to one third feature value, where M is a positive integer less than or equal to N. An inverse gain module 1503 processes the corresponding third feature values ​​based on the M target inverse gain values ​​to obtain M fourth feature values. A reconstruction module 1504 performs image reconstruction on the at least one processed second feature map to obtain a second image, where the at least one processed second feature map includes the M fourth feature values. In the above-described scheme, different target gain values ​​are set for different target compression bit rates to implement compression bit rate control.

[0268] 16 is a schematic diagram of the configuration of an image processing device 1600 according to an embodiment of the present application. The image processing device 1600 may be a terminal device or a server. an acquisition module 1601 configured to acquire a first image; a feature extraction module 1602 configured to perform feature extraction on the first image based on the encoding network to obtain at least one first feature map, the at least one first feature map including N first feature values, where N is a positive integer; the obtaining module 1601 is further configured to obtain a target compression bit rate, the target compression bit rate corresponding to M initial gain values ​​and M initial inverse gain values, each initial gain value corresponding to one first feature value, each initial inverse gain value corresponding to one third feature value, and M is a positive integer less than or equal to N; a gain module 1603 configured to process the corresponding first feature values ​​based on the M initial gain values, respectively, to obtain M second feature values; a quantization and entropy coding module 1604 configured to perform quantization and entropy coding on the at least one processed first feature map based on a quantization network and an entropy coding network to obtain coded data and a bit rate loss, wherein the at least one first feature map obtained after the gain processing includes M second feature values; a decoding module 1605 configured to perform entropy decoding on the encoded data based on the entropy decoding network to obtain at least one second feature map, the at least one second feature map including M third feature values, each third feature value corresponding to one first feature value; an inverse gain module 1606 configured to process the corresponding third feature values, respectively, based on the M initial inverse gain values ​​to obtain M fourth feature values; a reconstruction module 1607 configured to perform image reconstruction on the at least one processed second feature map based on the decoding network to obtain a second image, wherein the at least one processed feature map includes M fourth feature values; the acquisition module 1601 is further configured to acquire a distortion loss of the second image relative to the first image; a training module 1608 configured to perform joint training on a first encoding / decoding network, M initial gain values, and M initial inverse gain values ​​by using a loss function until an image distortion value between the first image and the second image reaches a first preset degree, the image distortion value being related to a bitrate loss and a distortion loss, and the encoding / decoding network includes a coding network, a quantization network, an entropy coding network, and an entropy decoding network; an output module 1609 configured to output a second encoding / decoding network, M target gain values, and M target inverse gain values, wherein the second encoding / decoding network is a model obtained after iterative training is performed on the first encoding / decoding network, and the M target gain values ​​and the M target inverse gain values ​​are obtained after iterative training is performed on the M initial gain values ​​and the M initial inverse gain values; Includes.

[0269] Optionally, an information entropy of the quantized data obtained after quantizing the at least one first feature map obtained after the gain processing satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate, and N is a positive integer greater than or equal to M.

[0270] Optionally, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0271] Optionally, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values ​​respectively.

[0272] Optionally, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, where P is a positive integer less than or equal to M.

[0273] Optionally, the first image includes a target object, and the M first feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0274] Optionally, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range, and the product of each of the M initial gain values ​​and the corresponding initial inverse gain value falls within a preset range.

[0275] The following describes an execution device provided in an embodiment of the present application. FIG. 17 is a schematic diagram of the structure of an execution device according to an embodiment of the present application. The execution device 1700 may specifically be represented as a virtual reality (VR) device, a mobile phone, a tablet computer, a notebook computer, an intelligent wearable device, a monitoring data processing device, etc. This is not limited here. The image processing device described in the embodiment corresponding to FIG. 14 and FIG. 15 may be implemented in the execution device 1700 to perform the functions of the image processing device of the embodiment corresponding to FIG. 14 and FIG. 15. Specifically, the execution device 1700 includes a receiver 1701, a transmitter 1702, a processor 1703, and a memory 1704 (there may be one or more processors 1703 in the execution device 1700, and one processor is used as an example in FIG. 17). The processor 1703 may include an application processor 17031 and a communication processor 17032. In some embodiments of the present application, the receiver 1701, the transmitter 1702, the processor 1703, and the memory 1704 may be connected by using a bus or in another manner.

[0276] The memory 1704 may include read-only memory and random-access memory and may provide instructions and data to the processor 1703. A portion of the memory 1704 may further include non-volatile random access memory (NVRAM). The memory 1704 stores processor-executable operating instructions, executable modules, data structures, a subset thereof, or an extended set thereof. The operating instructions may include various operating instructions for performing various operations.

[0277] The processor 1703 controls the operation of the execution device. During a particular application, the components of the execution device are coupled to each other through the use of a bus system. In addition to a data bus, the bus system may further include a power bus, a control bus, a status signal bus, etc. However, for clarity of explanation, the various types of buses in the figures are labeled as bus systems.

[0278] The methods disclosed in the above embodiments of the present application may be applied to or implemented by the processor 1703. The processor 1703 may be an integrated circuit chip and have signal processing capabilities. In the implementation process, the steps in the above methods may be implemented by using hardware integrated logic circuits in the processor 1703 or by using instructions in the form of software. The processor 1703 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, or may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor 1703 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the methods disclosed with reference to the embodiments of the present application may be directly performed and completed by a hardware decoding processor, or may be performed and completed by using a combination of hardware modules and software modules in the decoding processor. The software modules may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is located in the memory 1704, and the processor 1703 reads information in the memory 1704 and completes the steps of the aforementioned method in combination with the processor's hardware.

[0279] The receiver 1701 may be configured to receive input digital or textual information and generate signal inputs related to relevant setting and function control of the execution device. The transmitter 1702 may be configured to output the digital or textual information via the first interface. The transmitter 1702 may further be configured to send instructions to the disk group via the first interface to modify data in the disk group. The transmitter 1702 may further include a display device, such as a display screen.

[0280] In this embodiment of the present application, optionally, the processor 1703 is configured to execute an image processing method executed by an execution device in the embodiments corresponding to Figures 9 to 11. Specifically, the application processor 17031 acquires a first image; performing feature extraction on the first image to obtain at least one first feature map, the first feature map including N first feature values, where N is a positive integer; obtain a target compression bit rate, the target compression bit rate corresponding to M target gain values, each target gain value corresponding to one first feature value, M being a positive integer less than or equal to N; processing the corresponding first feature values ​​respectively based on the M target gain values ​​to obtain M second feature values; performing quantization and entropy coding on the at least one processed first feature map including the M second feature values ​​to obtain coded data; It is structured as follows.

[0281] Optionally, an information entropy of the quantized data obtained by quantizing the at least one processed first feature map satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0282] Optionally, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0283] Optionally, the difference between the compression bit rate corresponding to the encoded data and the target compression bit rate falls within a preset range.

[0284] Optionally, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values ​​respectively.

[0285] Optionally, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, where P is a positive integer less than or equal to M.

[0286] Optionally, the application processor 17031 determining M target gain values ​​corresponding to the target compression bitrate based on the target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bitrate and the M target gain values; the target mapping relationship includes a plurality of compression bit rates, a plurality of gain vectors, and correlations between the plurality of compression bit rates and the plurality of gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates and the M target gain values ​​are elements of one of the plurality of gain vectors; or the target mapping relationship comprises a target function mapping relationship, and the input of the target function relationship comprises a target compression bit rate, and the output of the target function relationship comprises M target gain values; It is further structured as follows.

[0287] Optionally, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first gain values, and the second compression bit rate corresponds to the M second gain values, and the M target gain values ​​are obtained by performing an interpolation operation on the M first gain values ​​and the M second gain values.

[0288] Optionally, the M first gain values ​​include a first target gain value, the M second gain values ​​include a second target gain value, and the M target gain values ​​include a third target gain value, wherein the first target gain value, the second target gain value, and the third target gain value correspond to the same feature value among the M first feature values, and the third target gain value is obtained by performing an interpolation operation on the first target gain value and the second target gain value.

[0289] Optionally, the first image includes a target object, and the M first feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0290] Optionally, each of the M target gain values ​​corresponds to one inverse gain value, and the inverse gain value is used to process the feature value obtained in the decoding process of the encoded data, and the product of each of the M target gain values ​​and the corresponding inverse gain value falls within a preset range.

[0291] Optionally, the application processor 17031 performing entropy decoding on the encoded data to obtain at least one second feature map, the at least one second feature map including N third feature values, each third feature value corresponding to one first feature value; obtaining M target inverse gain values, each target inverse gain value corresponding to one third feature value; performing gain processing on the corresponding third feature values ​​based on the M target inverse gain values ​​to obtain M fourth feature values; performing image reconstruction on the at least one second feature map obtained after the inverse gain processing to obtain a second image, the at least one second feature map obtained after the inverse gain processing including the M fourth feature values; It is further structured as follows.

[0292] Optionally, the M fourth feature values ​​are obtained by performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values ​​respectively.

[0293] Optionally, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, where P is a positive integer less than or equal to M.

[0294] Optionally, the application processor 17031 is further configured to determine M target inverse gain values ​​corresponding to the target compression bit rate based on a target mapping relationship, the target mapping relationship being used to indicate a correlation between the compression bit rate and the inverse gain vector.

[0295] Optionally, the target mapping relationship includes a plurality of compression bit rates, a plurality of inverse gain vectors, and a correlation relationship between the plurality of compression bit rates and the plurality of inverse gain vectors, wherein the target compression bit rate is one of the plurality of compression bit rates, and the M target inverse gain values ​​are elements of one of the plurality of inverse gain vectors.

[0296] Optionally, the target mapping relationship comprises an objective function mapping relationship, where an input of the objective function relationship comprises a target compression bit rate and an output of the objective function relationship comprises M target inverse gain values.

[0297] Optionally, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0298] Optionally, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range.

[0299] Optionally, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, and the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0300] Optionally, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, and the M target inverse gain values ​​include a third target inverse gain value, wherein the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0301] Specifically, the application processor 17031: Obtain the encoded data, performing entropy decoding on the encoded data to obtain at least one second feature map, the at least one second feature map including N third feature values, where N is a positive integer; Obtain M target inverse gain values, each target inverse gain value corresponding to one third feature value, where M is a positive integer less than or equal to N; respectively processing the corresponding third feature values ​​according to the M target inverse gain values ​​to obtain M fourth feature values; performing image reconstruction on the at least one processed second feature map to obtain a second image, wherein the at least one processed second feature map includes M fourth feature values; It is structured as follows.

[0302] Optionally, the M fourth feature values ​​are obtained by performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values ​​respectively.

[0303] Optionally, the at least one second feature map includes a second target feature map, the second target feature map including P third feature values, all of the P third feature values ​​corresponding to the same target inverse gain value, where P is a positive integer less than or equal to M.

[0304] Optionally, the application processor 17031 The method is further configured to: obtain a target compression bitrate; and determine M target inverse gain values ​​corresponding to the target compression bitrate based on a target mapping relationship, wherein the target mapping relationship is used to indicate a correlation between the compression bitrate and the inverse gain vector, the target mapping relationship including a plurality of compression bitrates, a plurality of inverse gain vectors, and correlations between the plurality of compression bitrates and the plurality of inverse gain vectors, the target compression bitrate being one of the plurality of compression bitrates and the M target inverse gain values ​​being one element of the plurality of inverse gain vectors; or the target mapping relationship includes an objective function mapping relationship, wherein an input of the objective function relationship includes the target compression bitrate and an output of the objective function relationship includes the M target inverse gain values.

[0305] Optionally, the second image includes a target object, and the M third feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0306] Optionally, the target compression bit rate is greater than the first compression bit rate and less than the second compression bit rate, the first compression bit rate corresponds to the M first inverse gain values, and the second compression bit rate corresponds to the M second inverse gain values, and the M target inverse gain values ​​are obtained by performing an interpolation operation on the M first inverse gain values ​​and the M second inverse gain values.

[0307] Optionally, the M first inverse gain values ​​include a first target inverse gain value, the M second inverse gain values ​​include a second target inverse gain value, and the M target inverse gain values ​​include a third target inverse gain value, wherein the first target inverse gain value, the second target inverse gain value, and the third target inverse gain value correspond to the same feature value among the M first feature values, and the third target inverse gain value is obtained by performing an interpolation operation on the first target inverse gain value and the second target inverse gain value.

[0308] An embodiment of the present application further provides a training device. FIG. 18 is a schematic diagram of the structure of a training device according to an embodiment of the present application. The image processing device described in the embodiment corresponding to FIG. 16 may be deployed in a training device 1800 to perform the functions of the image processing device in the embodiment corresponding to FIG. 16. Specifically, the training device 1800 is implemented by one or more servers. The training device 1800 may vary considerably due to different configurations or performances and may include one or more central processing units (CPUs) 1822 (e.g., one or more processors), memory 1832, and one or more storage media 1830 (e.g., one or more mass storage devices) that store applications 1842 or data 1844. The memory 1832 and the storage medium 1830 may be temporary or persistent storage devices. The program stored in the storage medium 1830 may include at least one module (not shown), and each module may include a series of instruction operations for the training device. Additionally, the central processing unit 1822 is arranged to communicate with a storage medium 1830 and is capable of executing a series of instructions in the storage medium 1830 in the training device 1800 .

[0309] The training device 1800 may further include one or more power sources 1826, one or more wired or wireless network interfaces 1850, one or more input / output interfaces 1858, and / or one or more operating systems 1841, such as Windows Server™, Mac OS X™, Unix™, Linux™, or FreeBSD™.

[0310] In this embodiment of the present application, the central processing unit 1822 is configured to execute the image processing method performed by the image processing device of the embodiment corresponding to Figure 16. Specifically, the central processing unit 1822: Acquire a first image, performing feature extraction on the first image based on the encoding network to obtain at least one first feature map, wherein the at least one first feature map includes N first feature values, where N is a positive integer; obtain a target compression bit rate, the target compression bit rate corresponding to M initial gain values ​​and M initial inverse gain values, each initial gain value corresponding to one first feature value, each initial inverse gain value corresponding to one third feature value, M being a positive integer less than or equal to N; processing the corresponding first feature values ​​based on the M initial gain values ​​to obtain M second feature values; performing quantization and entropy coding on the at least one processed first feature map based on a quantization network and an entropy coding network to obtain coded data and a bit rate loss, wherein the at least one first feature map obtained after the gain processing includes M second feature values; performing entropy decoding on the encoded data based on an entropy decoding network to obtain at least one second feature map, wherein the at least one second feature map includes M third feature values, each third feature value corresponding to one first feature value; respectively processing the corresponding third feature values ​​according to the M initial inverse gain values ​​to obtain M fourth feature values; performing image reconstruction on the at least one processed second feature map based on the decoding network to obtain a second image, wherein the at least one processed second feature map includes M fourth feature values; Obtain a distortion loss of the second image relative to the first image; performing joint training on a first encoding / decoding network, M initial gain values, and M initial inverse gain values ​​by using a loss function until an image distortion value between the first image and the second image reaches a first predetermined degree, wherein the image distortion value is related to a bitrate loss and a distortion loss; and the encoding / decoding network includes a coding network, a quantization network, an entropy coding network, and an entropy decoding network; outputting a second encoding / decoding network, M target gain values, and M target inverse gain values, wherein the second encoding / decoding network is a model obtained after iterative training is performed on the first encoding / decoding network, and the M target gain values ​​and the M target inverse gain values ​​are obtained after iterative training is performed on the M initial gain values ​​and the M initial inverse gain values; It is structured as follows.

[0311] Optionally, an information entropy of the quantized data obtained by quantizing the at least one first feature map obtained after the gain processing satisfies a predetermined condition, the predetermined condition being related to a target compression bit rate.

[0312] Optionally, the preset conditions include at least: This includes indicating that the higher the target compression bit rate, the greater the information entropy of the quantized data.

[0313] Optionally, the M second feature values ​​are obtained by performing multiplication operations on the M target gain values ​​and the corresponding first feature values ​​respectively.

[0314] Optionally, the at least one first feature map includes a first target feature map, the first target feature map including P first feature values, all of the P first feature values ​​corresponding to the same target gain value, where P is a positive integer less than or equal to M.

[0315] Optionally, the first image includes a target object, and the M first feature values ​​are feature values ​​in the at least one feature map that correspond to the target object.

[0316] Optionally, the product of each of the M target gain values ​​and the corresponding target inverse gain value falls within a preset range, and the product of each of the M initial gain values ​​and the corresponding initial inverse gain value falls within a preset range.

[0317] An embodiment of the present application further provides a computer program product, which, when run on a computer, enables the computer to perform the steps performed by the execution device in the method described in the previous embodiment shown in Figure 17, or enables the computer to perform the steps performed by the training device in the method described in the previous embodiment shown in Figure 18.

[0318] An embodiment of the present application further provides a computer-readable storage medium, which stores a program for performing signal processing. When the program runs on a computer, the computer is enabled to perform the steps performed by the execution device in the method described in the previous embodiment shown in Figure 17, or the computer is enabled to perform the steps performed by the training device in the method described in the previous embodiment shown in Figure 18.

[0319] The execution device, training device, or terminal device provided in the embodiments of the present application may specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor. The communication unit may be, for example, an input / output interface, pins, or circuits. The processing unit may execute computer-executable instructions stored in the storage unit to cause the chip in the execution device to perform the image processing method described in the embodiment shown in FIGS. 3 to 7, or the chip in the training device to perform the image processing method described in the embodiment shown in FIG. 13. Optionally, the storage unit is a storage unit within the chip, such as a register or cache. Alternatively, the storage unit may be a storage unit located outside the chip at the wireless access device end, such as a read-only memory (ROM) or another type of static storage device capable of storing static information and instructions, or a random access memory (RAM).

[0320] For details, please refer to Figure 19. Figure 19 is a schematic diagram of the structure of a chip according to an embodiment of the present application. The chip may be represented as a neural network processing unit NPU 2000. The NPU 2000 is mounted on a host CPU as a coprocessor, and the host CPU assigns tasks to the NPU. The core part of the NPU is an arithmetic circuit 2003, which is controlled by a controller 2004 to extract matrix data from a memory and perform multiplication operations.

[0321] In some embodiments, the arithmetic circuitry 2003 includes multiple process engines (PEs). In some embodiments, the arithmetic circuitry 2003 is a two-dimensional systolic array. The arithmetic circuitry 2003 may be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some embodiments, the arithmetic circuitry 2003 is a general-purpose matrix processor.

[0322] For example, it is assumed that there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves data corresponding to matrix B from the weight memory 2002 and buffers the data in each PE in the arithmetic circuit. The arithmetic circuit obtains data of matrix A from the input memory 2001, performs a matrix operation between the data and matrix B, and stores the partial or final result of the obtained matrix in an accumulator 2008.

[0323] The unified memory 2006 is configured to store input data and output data. Weight data is transferred directly to the weight memory 2002 by using a Direct Memory Access Controller (DMAC) 2005. Input data is also transferred to the unified memory 2006 by using the DMAC.

[0324] The BIU is a Bus Interface Unit, ie, a Bus Interface Unit 2010, configured to interact with a DMAC and an Instruction Fetch Buffer (IFB) 2009 by using an AXI bus.

[0325] The bus interface unit 2010 (BIU for short) is further configured such that the instruction fetch buffer 2009 is configured to retrieve instructions from an external memory, and the direct memory access controller 2005 is configured to retrieve raw data of the input matrix A or the weight matrix B from the external memory.

[0326] The DMAC is mainly configured to transfer input data in the external memory DDR to the unified memory 2006, transfer weight data to the weight memory 2002, or transfer input data to the input memory 2001.

[0327] The vector calculation unit 2007 includes multiple arithmetic processing units. If necessary, further processing such as vector multiplication, vector addition, exponential operation, logarithm operation, value comparison, etc. is performed on the output of the arithmetic circuit. The vector calculation unit 2007 is mainly configured to perform network calculations such as batch normalization, pixel-level summation, and feature surface upsampling for non-convolutional / fully connected layers in a neural network.

[0328] In some implementations, the vector calculation unit 2007 can store the processed output vector in the unified memory 2006. For example, the vector calculation unit 2007 can apply a linear and / or nonlinear function to the output of the calculation circuitry 2003, such as performing linear interpolation on feature surfaces extracted in a convolutional layer. In another example, the linear and / or nonlinear function is applied to the vector of accumulated values ​​to generate activation values. In some implementations, the vector calculation unit 2007 generates a normalized value, a pixel-level sum, or a normalized value and a pixel-level sum. In some implementations, the processed output vector can be used as an activated input to the calculation circuitry 2003, such as in a subsequent layer of a neural network.

[0329] An instruction fetch buffer 2009 coupled to the controller 2004 is configured to store instructions used by the controller 2004 .

[0330] The unified memory 2006, the input memory 2001, the weight memory 2002, and the instruction fetch buffer 2009 are all on-chip memories. The external memory is dedicated to the hardware architecture of the NPU.

[0331] The processor referred to in any of the above may be a general purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits configured to control program execution of the method according to the first aspect.

[0332] In addition, please note that the described device embodiments are merely examples. Units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, and may be located in one location or distributed over multiple network units. Some or all of the modules may be selected according to actual requirements to achieve the objectives of the solutions of the embodiments. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that the modules have communication connections with each other, which may be specifically implemented as one or more communication buses or signal cables.

[0333] Based on the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by software in addition to necessary general-purpose hardware, or by dedicated hardware including, of course, application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, any function that can be performed by a computer program can be easily implemented by using corresponding hardware, and the specific hardware structure used to achieve the same function may be in various forms, such as an analog circuit, a digital circuit, or a dedicated circuit. However, in the present application, a software program implementation is most often a better implementation. Based on this understanding, the technical solution of the present application, essentially or in part contributing to the prior art, may be implemented in the form of a software product. The software product is stored in a readable storage medium such as a computer floppy disk, a USB flash drive, a removable hard disk, a ROM, a RAM, a magnetic disk, or an optical disk, and includes several instructions for instructing a computer device (which may be a personal computer, a training device, or a network device) to perform the method described in the embodiments of the present application.

[0334] All or part of the foregoing embodiments may be implemented using software, hardware, firmware, or any combination thereof. If software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product.

[0335] A computer program product includes one or more computer instructions. When the computer program instructions are loaded into a computer and executed, all or part of the procedures or functions according to the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, optical fiber, or digital subscriber line (DSL)) or wireless (e.g., infrared, radio frequency, or microwave) transmission. The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device, such as a training device or data center, that integrates one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a DVD), a semiconductor medium (e.g., a solid-state disk (SSD)), etc. [Explanation of symbols]

[0336] 200 Image Processing System 201 Target Model / Rules 210 Execution Device 211 Computational Module 212 I / O interfaces 220 Training equipment 230 databases 240 Client Device 250 Data Storage System 401 First Image 402 CNN / layer 403 Multi-Channel-Wise Feature Maps 1400 Image Processing Device 1401 Acquisition Module 1402 Feature Extraction Module 1403 Gain Module 1404 Entropy Coding Module 1500 Image Processing Device 1501 Acquisition Module 1502 Decryption Module 1503 Inverse Gain Module 1504 Reconstruction Module 1600 Image Processing Device 1601 Acquisition Module 1602 Feature Extraction Module 1603 Gain Module 1604 Entropy Coding Module 1605 Decryption Module 1606 Inverse Gain Module 1607 Reconstruction Module 1608 Training Module 1609 Output Module 1700 Execution Device 1701 Receiver 1702 Transmitter 1703 processor 1704 memory 1800 training equipment 1822 Central Processing Unit 1826 power supply 1830 storage medium 1832 memory 1841 Operating Systems 1842 Applications 1844 Data 1850 Wireless Network Interface 1858 Output Interface 2000 Neural Network Processing Unit (NPU) 2001 input memory 2002 Weight Memory 2003 Arithmetic circuit 2004 Controller 2005 Direct Memory Access Controller 2006 Unified Memory 2007 Vector Computation Unit 2008 Accumulator 2009 Instruction Fetch Buffer 2010 Bus Interface Unit 17031 Application Processor 17032 Communications Processor

Claims

1. 1. An image encoding method, comprising: acquiring a first image; performing feature extraction on the first image to obtain a first feature map, the first feature map including N first feature values, where N is a positive integer; obtaining M target gain values ​​corresponding to a target compression bit rate based on the objective function mapping relationship, where each target gain value corresponds to one first feature value, and M is a positive integer less than or equal to N; obtaining a processed first feature map based on the M target gain values ​​and corresponding first feature values, wherein the processed first feature map includes M second feature values; encoding the processed first feature map to obtain a bitstream; A method comprising:

2. The method of claim 1 , wherein a difference between the compressed bit rate corresponding to the bitstream and the target compressed bit rate falls within a preset range.

3. The method of claim 1 or 2, wherein the M second feature values ​​are obtained by individually performing multiplication operations on the M target gain values ​​and the corresponding first feature values.

4. The step of obtaining M target gain values ​​corresponding to a target compression bit rate based on an objective function mapping relationship, 4. The method of claim 1, further comprising obtaining the M target gain values ​​based on the objective function mapping relationship, wherein an input of the objective function mapping relationship comprises the target compression bitrate, and an output of the objective function mapping relationship comprises the M target gain values.

5. 5. The method of claim 1, wherein the target compression bit rate is greater than a first compression bit rate and less than a second compression bit rate, the first compression bit rate corresponds to M first gain values, the second compression bit rate corresponds to M second gain values, and the M target gain values ​​are obtained by performing an interpolation operation on the M first gain values ​​and the M second gain values.

6. 6. The method of claim 1, wherein the first image includes a target object, and first feature values ​​corresponding to the M target gain values ​​are feature values ​​in the first feature map that correspond to the target object.

7. 7. The method according to claim 1, wherein each of the M target gain values ​​corresponds to an inverse gain value, and the inverse gain value is used to process feature values ​​obtained in a decoding process of the bitstream, and a product of each of the M target gain values ​​and the corresponding inverse gain value falls within a preset range.

8. The method of claim 7 , wherein the product of each of the M target gain values ​​and the corresponding inverse gain value is one.

9. An image decoding method, comprising: obtaining a bitstream; decoding the bitstream to obtain a second feature map, the second feature map including N third feature values, where N is a positive integer; obtaining M target inverse gain values ​​corresponding to a target compression bit rate based on the objective function mapping relationship, where each target inverse gain value corresponds to one third feature value, and M is a positive integer less than or equal to N; obtaining a processed second feature map based on the M target inverse gain values ​​and corresponding third feature values, wherein the processed second feature map includes M fourth feature values; performing image reconstruction on the processed second feature map to obtain a second image; A method comprising:

10. The method of claim 9 , wherein the M fourth feature values ​​are obtained by individually performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values.

11. The step of obtaining M target inverse gain values ​​corresponding to a target compression bit rate based on a target function mapping relationship, comprising: obtaining the target compression bit rate; obtaining the M target inverse gain values ​​based on the objective function mapping relationship; and wherein when an input of the objective function mapping relationship comprises the target compression bit rate, an output of the objective function mapping relationship comprises the M target inverse gain values.

12. An image encoding device, an acquisition module configured to acquire a first image; a feature extraction module configured to perform feature extraction on the first image to obtain a first feature map, the first feature map comprising N first feature values, where N is a positive integer; the obtaining module is further configured to obtain M target gain values ​​corresponding to a target compression bit rate based on the objective function mapping relationship, each target gain value corresponding to one first feature value, where M is a positive integer less than or equal to N; The image encoding device, a gain module configured to obtain a processed first feature map based on the M target gain values ​​and corresponding first feature values, wherein the processed first feature map includes M second feature values; an entropy coding module configured to code the processed first feature map to obtain a bitstream; An apparatus comprising:

13. The apparatus of claim 12 , wherein the M second feature values ​​are obtained by individually performing multiplication operations on the M target gain values ​​and the corresponding first feature values.

14. An image decoding device, a capture module configured to capture the bitstream; a decoding module configured to decode the bitstream to obtain a second feature map, the second feature map including N third feature values, where N is a positive integer; the obtaining module is further configured to obtain M target inverse gain values ​​corresponding to a target compression bit rate based on the objective function mapping relationship, each target inverse gain value corresponding to one third feature value, where M is a positive integer less than or equal to N; The image decoding device an inverse gain module configured to obtain a processed second feature map based on the M target inverse gain values ​​and corresponding third feature values, wherein the processed second feature map includes M fourth feature values; a reconstruction module configured to perform image reconstruction on the processed second feature map to obtain a second image; An apparatus comprising:

15. The apparatus of claim 14 , wherein the M fourth feature values ​​are obtained by individually performing multiplication operations on the M target inverse gain values ​​and the corresponding third feature values.

16. 9. An image coding device comprising a memory and a processor coupled to each other, wherein the processor invokes program code stored in the memory to perform the method of any one of claims 1 to 8.

17. 12. An image decoding device comprising a memory and a processor coupled to each other, wherein the processor invokes program code stored in the memory to perform the method of any one of claims 9 to 11.

18. 12. A computer program comprising computer executable instructions for storage on a non-transitory medium, the computer program comprising, when executed by a processor, causing the processor to perform the method of any one of claims 1 to 11.

19. A chip system comprising a processor and a memory, the memory configured to store program instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 11.

20. 12. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1 to 11.