An image encoding method, storage medium, and terminal device

By acquiring image feature maps, saliency feature maps and attention maps to generate mask feature maps, the problem of allocating the same bits of the significant image content and non-striking image content in the prior art is solved, and the effect of reconstructing images is improved.

CN113965756BActive Publication Date: 2025-07-29WUHAN TCL CORP RES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010706744.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-21
Publication Date
2025-07-29
Estimated Expiration
2040-07-21

AI Technical Summary

Technical Problem

The prior art assigns the same bits to the salient image content and the non-strient image content in image encoding, resulting in poor reconstructed image effects.

Method used

By acquiring the image feature map, significance feature map and attention map, a mask feature map is generated, and the information amount of each channel in the coded feature map is determined based on the significance feature and attention feature, and different bits are allocated.

Benefits of technology

The amount of information about the significant image content in the encoded file is improved and the effect of reconstructing the image is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113965756B_ABST
    Figure CN113965756B_ABST
Patent Text Reader

Abstract

The present application discloses an image encoding method, a storage medium, and a terminal device. The method includes obtaining an image feature map, a saliency feature map, and an attention map; determining a mask feature map based on the attention map and the saliency feature map; generating an encoding feature map corresponding to the image to be encoded according to the image feature map and the mask feature map; and obtaining an encoding file corresponding to the image to be encoded according to the encoding feature map. The mask feature map determined by the present application according to the saliency feature map and the attention map enables the encoding feature map to reflect the saliency features and attention features of the image to be encoded. In this way, the information amount of each channel in the encoding feature map is determined according to the saliency features and attention features, and then different bit positions are allocated to channels with different information amounts, thereby improving the image information of the salient image content in the encoding file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and particularly relates to an image encoding method, a storage medium, and a terminal device. Background Art

[0002] Image encoding refers to the technology of representing the original image pixel matrix with fewer bytes. Usually, in order to save storage space when storing images, or to improve the image transmission speed when transmitting images. Currently, when performing entropy encoding on an image, the same number of bits is allocated to significant image content (such as foreground content, etc.) and non-significant image content (such as background content, etc.), so that the amount of image information carrying important image content in the encoded file is the same as that of non-important image content. However, in the human visual system, the human eye is more sensitive to foreground content. Therefore, the encoded file using the same number of bits will affect the reconstructed image and result in a poor effect. Summary of the Invention

[0003] The technical problem to be solved by the present application is to provide an image encoding method, a storage medium, and a terminal device in view of the deficiencies of the prior art.

[0004] To solve the above technical problem, in the first aspect of the embodiments of the present application, an image encoding method is provided. The method includes:

[0005] Obtain an image feature map, a saliency feature map, and an attention map corresponding to the image to be encoded;

[0006] Based on the attention map and the saliency feature map, determine a mask feature map corresponding to the image to be encoded;

[0007] Based on the image feature map and the mask feature map, generate an encoded feature map corresponding to the image to be encoded;

[0008] Based on the encoded feature map, obtain an encoded file corresponding to the image to be encoded.

[0009] In one implementation, each channel in the mask feature map corresponds one-to-one with each channel in the saliency feature map, and there are at least two channels in the mask feature map that contain different amounts of image information.

[0010] In one implementation, the image encoding method applies an image encoding model.

[0011] In one implementation, the image encoding model includes a feature map extraction module, a saliency extraction module, and an attention module; the obtaining of the image feature map, the saliency feature map, and the attention map corresponding to the image to be encoded specifically includes:

[0012] The feature map extraction module determines an image feature map corresponding to the image to be encoded based on the image to be encoded;

[0013] The saliency extraction module determines a saliency feature map of the image to be encoded based on the image to be encoded;

[0014] The attention module determines an attention map corresponding to the image to be encoded based on the image feature map.

[0015] In one implementation, the attention module includes a first attention unit, a second attention unit, and a fusion unit; the attention module determining an attention map corresponding to the image to be encoded based on the image feature map specifically includes:

[0016] The first attention unit determines a first attention map corresponding to the image to be encoded based on the image feature map;

[0017] The second attention unit determines a second attention map corresponding to the image to be encoded based on the image feature map;

[0018] The fusion unit determines an attention map corresponding to the image to be encoded based on the first attention map and the second attention map.

[0019] In one implementation, the image encoding model includes a mask module, and determining a mask feature map corresponding to the image to be encoded based on the attention map and the saliency feature map specifically includes:

[0020] The mask module determines an intermediate feature map based on the attention map and the saliency feature map, where the intermediate feature map is a single-channel feature map;

[0021] The mask module determines a mask feature map corresponding to the image to be encoded based on the intermediate feature map.

[0022] In one implementation, the image scale of the attention image is the same as the image scale of the saliency feature map; the mask module determining an intermediate feature map based on the attention map and the saliency feature map specifically includes:

[0023] For each pixel point in the attention map, a candidate pixel point corresponding to the pixel point is obtained, where the pixel position of the candidate pixel point in the saliency feature map corresponds to the pixel position of the pixel point in the attention map;

[0024] The pixel value of the pixel point is adjusted based on the pixel value of the candidate pixel point, and the adjusted pixel value is used as the pixel value of the pixel point to obtain an adjusted attention map;

[0025] Determine the intermediate feature map based on the adjusted attention map.

[0026] In one implementation, the mask module determines the mask feature map corresponding to the image to be encoded based on the intermediate feature map, which specifically includes:

[0027] The mask module determines a multi-channel feature map, where the image size of the multi-channel feature map is the same as that of the intermediate feature map;

[0028] For each channel in the multi-channel feature map, the mask module adjusts the pixel values of each pixel point in the channel based on the channel number of the channel and the intermediate feature map;

[0029] Use the adjusted multi-channel feature map as the mask feature map.

[0030] In one implementation, the mask module determines the pixel values of each pixel point in the channel based on the channel number of the channel and the intermediate feature map, which specifically includes:

[0031] For each pixel point in the channel, the mask module determines the target pixel value corresponding to the pixel point, where the target pixel value is the pixel value of the target pixel point, and the pixel position of the target pixel point in the intermediate feature map corresponds to the pixel position of the pixel point in the channel;

[0032] The mask module determines the pixel value of the pixel point according to the target pixel value and the channel number of the channel.

[0033] In one implementation, before the mask module determines a multi-channel feature map, the method includes:

[0034] The mask module adjusts the image size of the intermediate feature map and uses the adjusted intermediate feature map as the intermediate feature map, where the image size of the adjusted intermediate feature map is the same as that of the image feature map.

[0035] In one implementation, the image encoding module includes a quantization module; after generating the encoded feature map corresponding to the image to be encoded based on the image feature map and the mask feature map, the method includes:

[0036] The quantization module generates a quantized feature map of the image to be encoded based on the encoded feature map and uses the quantized feature map as the encoded feature map.

[0037] In a second aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any one of the above-mentioned image encoding methods.

[0038] In a third aspect of the embodiments of the present application, a terminal device is provided, which includes: a processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory;

[0039] The communication bus realizes the connection and communication between the processor and the memory;

[0040] When the processor executes the computer-readable program, the steps in any one of the above-mentioned image encoding methods are implemented.

[0041] Beneficial effects: Compared with the prior art, the embodiments of the present application provide an image encoding method, a storage medium, and a terminal device. The image encoding method includes obtaining an image feature map and a saliency feature map corresponding to an image to be encoded, and determining an attention map of the image to be encoded based on the image feature map; determining a mask feature map corresponding to the image to be encoded based on the attention map and the saliency feature map; generating an encoded feature map corresponding to the image to be encoded according to the image feature map and the mask feature map; and obtaining an encoded file corresponding to the image to be encoded according to the encoded feature map. The mask feature map determined according to the saliency feature map and the attention map in the present application enables the encoded feature map to reflect the saliency feature and the attention feature of the image to be encoded. In this way, the amount of information in each channel of the encoded feature map is determined according to the saliency feature and the attention feature, and then different bit positions are allocated to channels with different amounts of information, so as to improve the image information of the salient image content in the encoded file, and further improve the image effect of the reconstructed image reconstructed according to the encoded file. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative labor, other drawings can be obtained based on these drawings.

[0043] Figure 1 It is a flowchart of the image encoding method provided by the present application.

[0044] Figure 2 It is a schematic diagram of the principle of the image encoding method provided by the present application.

[0045] Figure 3Schematic diagram of the principle of the feature extraction module in the image encoding method provided by this application.

[0046] Figure 4 Schematic diagram of the principle of the attention module in the image encoding method provided by this application.

[0047] Figure 5 Schematic diagram of the working principle of the mask module in the image encoding method provided by this application.

[0048] Figure 6 Schematic diagram of the principle of generating a reconstructed image from the encoded file obtained by the image encoding method provided by this application.

[0049] Figure 7 Schematic diagram of the structural principle of the decoding module provided by this application.

[0050] Figure 8 Schematic diagram of the structural principle of the residual block in the decoding module provided by this application.

[0051] Figure 9 Schematic diagram of the structural principle of the upsampling module in the decoding module provided by this application.

[0052] Figure 10 Schematic diagram of the reconstructed image generated by collecting the encoded file obtained by the image encoding method provided by this application.

[0053] Figure 11 Schematic diagram of the reconstructed image generated from the encoded file directly encoded from the image feature map.

[0054] Figure 12 Schematic diagram of the structural principle of the terminal device provided by this application. Detailed implementation manners

[0055] This application provides an image encoding method, a storage medium, and a terminal device. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following further describes this application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain this application and are not used to limit this application.

[0056] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0057] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0058] The inventors have found through research that with the continuous development of deep learning technology, deep convolutional networks have been widely applied to image compression methods. However, in the currently commonly used network model based on autoencoders, after obtaining the encoded feature map through the network model of the autoencoder, the encoded features are usually quantized, and entropy coding is performed on the quantized encoded feature map. And when performing entropy coding, the same number of bits are assigned to the significant image content (such as foreground content, etc.) and non-significant image content (such as background content, etc.) in the quantized encoded feature map, which makes the amount of image information carrying the significant image content in the encoded file obtained by encoding the same as the amount of image information of the non-significant image content. However, in the human visual system, the human eye is more sensitive to the content corresponding to the significant image region. Thus, the encoded file using the same number of bits in this way will affect the reconstruction result, and the reconstructed image has a poor effect.

[0059] To solve the above problems, in the embodiments of the present application, an image feature map and a saliency feature map corresponding to an image to be encoded are obtained through an image encoding model, and an attention map of the image to be encoded is determined based on the image feature map; a mask feature map corresponding to the image to be encoded is determined based on the attention map and the saliency feature map; an encoded feature map corresponding to the image to be encoded is generated according to the image feature map and the mask feature map; and an encoded file corresponding to the image to be encoded is obtained according to the encoded feature map. The mask feature map determined according to the saliency feature map and the attention map in the present application enables the encoded feature map to reflect the saliency features and attention features of the image to be encoded. In this way, the information volume of each channel in the encoded feature map is determined according to the saliency features and attention features, and then different bit positions are allocated to channels with different information volumes, so that the image information of the salient image content in the encoded file can be improved, and further the image effect of the reconstructed image reconstructed according to the encoded file can be improved.

[0060] For example, the embodiments of the present application can be applied to the scenario of an electronic device configured with an image encoding model. In this scenario, first, the electronic device can collect an image to be encoded, and obtain an image feature map and a saliency feature map corresponding to the image to be encoded through the image encoding model, and determine an attention map of the image to be encoded based on the image feature map; determine a mask feature map corresponding to the image to be encoded based on the attention map and the saliency feature map; generate an encoded feature map corresponding to the image to be encoded according to the image feature map and the mask feature map; and obtain an encoded file corresponding to the image to be encoded according to the encoded feature map.

[0061] It should be noted that the above application scenarios are only shown for the convenience of understanding the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0062] Next, with reference to the accompanying drawings, the content of the application will be further described through the description of the embodiments.

[0063] The present embodiment provides an image encoding method, as Figure 1 and Figure 2 shown, the method includes:

[0064] S10. Obtain an image feature map, a saliency feature map, and an attention map corresponding to the image to be encoded.

[0065] Specifically, the image to be encoded may be an image obtained by an electronic device (e.g., a server) running an image encoding method. The image to be encoded may be an original image collected by an image acquisition device. The image acquisition device may be configured in the electronic device running the image encoding method or on other external devices, and the original image obtained is sent to the electronic device running the image encoding method by the external device. In a possible implementation of this embodiment, the image acquisition device is configured in the electronic device running the image encoding method. In this way, after the electronic device acquires an image, the image can be directly encoded. On the one hand, the storage space occupied by the image can be reduced. On the other hand, when the image needs to be sent to other external devices, the encoded image can be sent to improve the image transmission speed.

[0066] The image feature map is the global feature map of the image to be encoded, which contains all the image features of the image to be encoded. Each channel in the image feature map is a local feature map of the image to be encoded, containing local features of the image to be encoded. Thus, the importance levels of the local features carried by each channel in the image to be encoded are different. For example, in the human visual system, the human eye is more sensitive to the content corresponding to the salient region (e.g., the foreground region, etc.). Therefore, the content corresponding to the salient region in the image to be encoded is more important than the content corresponding to the non-salient region in the image to be encoded. Thus, the importance levels of the image content of the image to be encoded can be divided into level 1 and level 2 according to the salient region and the non-salient region. Among them, level 1 corresponds to the salient region, level 2 corresponds to the non-salient region, and the importance corresponding to level 1 is higher than the importance corresponding to level 2. And the amount of image information in the channel carrying the content corresponding to the salient region in the image features is more than the amount of image information in the channel carrying the content corresponding to the non-salient region.

[0067] Furthermore, the saliency feature map is used to reflect the foreground feature information of the image to be encoded. The saliency feature map is a single-channel feature map, and the value range of the pixel points in the saliency feature map can be [0, 1]. The pixel points with a pixel value of 1 are used to represent the first part of the foreground region in the feature map to be encoded, and the pixel points with a pixel value of 0 represent the first part of the background region in the feature map to be encoded; among the pixel points with pixel values between 0 and 1, some pixel points represent the second part of the foreground region, and some pixel points represent the second part of the background region. The first part of the foreground region and the second part of the foreground region constitute the foreground region of the image to be encoded, and the first part of the background region and the second part of the background region constitute the background region of the image to be encoded. Among them, the second part of the foreground region is the boundary region of the foreground region, the second part of the background region is the boundary region of the background region, and the second part of the foreground region intersects with the second part of the background region.

[0068] Further, in an implementation manner of this embodiment, the image encoding method is applied to an image encoding module, an attention mechanism is embedded in the image encoding module, and the attention map is determined based on the image feature map through the attention mechanism, and is used to highlight the foreground information in the image feature map and suppress the background information. Among them, the attention mechanism is an information enhancement mechanism based on pixels for specific targets (for example, objects in the image to be encoded, such as people, etc.) in the feature map spectrum in deep learning by using an attention mechanism similar to the human eye. The attention mechanism is a mechanism that can enhance the target information (for example, foreground information, etc.) in the image feature map. After processing the image feature map based on the attention mechanism, the target information in the image feature map will be enhanced. The image feature map after attention enhancement can realize the enhancement of pixel-level information based on the target.

[0069] Further, in an implementation manner of this embodiment, as Figure 2 shown, the image encoding model includes a feature map extraction module, a saliency extraction module, and an attention module; the obtaining of the image feature map, saliency feature map, and attention map corresponding to the image to be encoded specifically includes:

[0070] S11. The feature map extraction module determines the image feature map corresponding to the image to be encoded based on the image to be encoded;

[0071] S12. The saliency extraction module determines the saliency feature map of the image to be encoded based on the image to be encoded;

[0072] S13. The attention module determines the attention map corresponding to the image to be encoded based on the image feature map.

[0073] Specifically, in step S11, the feature extraction module is used to obtain the feature map of the image to be encoded, the input item of the feature extraction module is the image to be encoded, and the output item is the image feature map. Among them, the feature extraction module may include several cascaded convolution modules, and the image feature map corresponding to the image to be encoded is generated through several cascaded convolution modules. In a specific implementation manner of this embodiment, as Figure 3As shown, the feature extraction module includes six convolutional modules, denoted as the first convolutional module, the second convolutional module, the third convolutional module, the fourth convolutional module, the fifth convolutional module, and the sixth convolutional module. Among them, in the first convolution, the convolution kernel is 7*7, the number of convolution kernels is 60, the stride is 1, and the padding pixels are 3; the normalization operation is InstanceNorm normalization; the non-linear transformation function is the relu function; in the second convolutional module, the convolution kernel is 3*3, the number of convolution kernels is 120, the stride is 1, and the padding pixels are 3; the normalization operation is InstanceNorm normalization; the non-linear transformation function is the relu function; in the third convolutional module, the convolution kernel is 3*3, the number of convolution kernels is 240, the stride is 2, and the padding pixels are 1; the normalization operation is InstanceNorm normalization; the non-linear transformation function is the relu function; in the fourth convolutional module, the convolution kernel is 3*3, the number of convolution kernels is 480, the stride is 2, and the padding pixels are 1; the normalization operation is InstanceNorm normalization; the non-linear transformation function is the relu function; in the fifth convolutional module, the convolution kernel is 3*3, the number of convolution kernels is 960, the stride is 2, and the padding pixels are 1; the normalization operation is InstanceNorm normalization; the non-linear transformation function is the relu function; in the sixth convolutional module, the convolution kernel is 3*3, the number of convolution kernels is C, the stride is 1, and the padding pixels are 1; the normalization operation is InstanceNorm normalization. Among them, C can be determined according to actual needs. In a possible implementation manner of this embodiment, C is set to 16.

[0074] Based on this, the number of channels of the image feature map is 16, that is, the image feature map is a 16-channel feature map. The acquisition process of the image feature map can be as follows: The image to be encoded is input into the first convolutional module, and a 60-channel feature map A is output through the convolutional module; the feature map A is input into the second convolutional module, and a 120-channel feature map B is output through the second convolutional module; the feature map B is input into the third convolutional module, and a 240-channel feature map C is output through the third convolutional module; the feature map C is input into the fourth convolutional module, and a 480-channel feature map D is output through the fourth convolutional module; the feature map D is input into the fifth convolutional module, and a 960-channel feature map E is output through the fifth convolutional module; the feature map E is input into the sixth convolutional module, and a 16-channel image feature map is output through the sixth convolutional module.

[0075] Further, in the step S12, the saliency module is used to extract the saliency features of the image to be encoded and generate a saliency feature map based on the extracted saliency features. The saliency module can be a traditional machine learning model or a deep learning model. In an implementation manner of this embodiment, since the deep learning model is a machine learning model that can simulate the neural structure of the human brain, it has powerful image processing capabilities and forward propagation, making the learning result closest to the human brain result. Therefore, the detection algorithm model adopts a deep learning model.

[0076] Further, the saliency module can adopt a saliency detection network model, and the saliency detection network model can include an encoding-decoding unit and a residual unit; the encoding-decoding unit is used to extract the initial saliency feature map corresponding to the image to be encoded, where the alignment degree between the boundary of the foreground region in the initial saliency feature map and the boundary of the real foreground region of the image to be encoded is less than a preset threshold (for example, 95%, etc.), and the residual unit is used to correct the initial saliency feature map to obtain a saliency feature map, where the alignment degree between the boundary of the foreground region in the initial saliency feature map and the boundary of the real foreground region of the image to be encoded is greater than or equal to the preset threshold. Of course, in practical applications, the saliency module can also adopt existing saliency detection network models, such as the BasNet network, etc.

[0077] Further, in the step S13, the attention module is used to obtain the attention map corresponding to the image to be encoded, where the input item of the attention module is the image feature map, and the output item of the attention module is the attention map. The attention map is a spatial attention map, and the attention model highlights the high-level target semantic features corresponding to the global receptive field by reducing the feature map size of the image features; then it enlarges the image size of the image feature map to amplify the activated foreground salient region in the image feature map, so as to highlight the differential features between the foreground region and the background region, and obtain a spatial attention feature map.

[0078] In an implementation manner of this embodiment, as Figure 4 shown, the attention module includes a first attention unit, a second attention unit, and a fusion unit; the attention module determines the attention map corresponding to the image to be encoded based on the image feature map, which specifically includes:

[0079] The first attention unit determines the first attention map corresponding to the image to be encoded based on the image feature map;

[0080] The second attention unit determines the second attention map corresponding to the image to be encoded based on the image feature map;

[0081] The fusion unit determines the attention map corresponding to the image to be encoded based on the first attention map and the second attention map.

[0082] Specifically, the input item of the first attention unit is the image feature map, and the output item is the second attention map. The first attention unit includes a cascaded first convolutional sub-unit and second convolutional sub-unit. Among them, the first convolutional sub-unit includes operations of convolution operation and normalization operation. The parameters of the convolution operation are that the convolution kernel is 1*K, the number of convolution kernels is C / / 2, the stride is 1, and the padding pixels are (0, K / / 2), where / / is the integer division symbol, C is the number of channels of the image feature map, and the normalization operation is BatchNorm. The parameters of the convolution operation in the second convolutional sub-unit are K*1, the number of convolution kernels is 1, the stride is 1, and the padding pixels are (K / / 2, 0), where / / is the integer division symbol, and the normalization operation is BatchNorm.

[0083] Further, the input item of the second attention unit is the image feature map, and the output item is the second attention map. The second attention unit includes a cascaded third convolutional sub-unit and fourth convolutional sub-unit. Among them, the third convolutional sub-unit includes operations of convolution operation and normalization operation. The parameters of the convolution operation are that the convolution kernel is K*1, the number of convolution kernels is C / / 2, the stride is 1, and the padding pixels are (K / / 2, 0), where / / is the integer division symbol, C is the number of channels of the image feature map, and the normalization operation is BatchNorm. The parameters of the convolution operation in the second convolutional sub-unit are 1*K, the number of convolution kernels is 1, the stride is 1, and the padding pixels are (0, K / / 2), where / / is the integer division symbol, and the normalization operation is BatchNorm.

[0084] Further, after obtaining the first attention map and the second attention map, determine the attention map corresponding to the image to be encoded based on the first attention map and the second attention map, where the image scale of the first attention map is the same as the image scale of the second attention map. It can be understood that the image size of the first attention map is the same as the image size of the second attention map, and the number of channels of the first attention map is the same as the number of channels of the second attention map. In one implementation manner of this embodiment, the image size of the first attention map and the image size of the second attention map are both the same as the image size of the image feature map. For example, the image size of the image feature map is 224*224, and the image sizes of the first attention map and the second attention map are both 224*224. In addition, in a specific implementation manner, the determining the attention map corresponding to the image to be encoded based on the first attention map and the second attention map can specifically be: adding the first attention map and the second attention map to obtain a fused feature map; then performing an activation operation on the fused feature map, and using the activated fused feature map as the attention map.

[0085] Further, the activation operation may be a Sigmoid activation operation. The Sigmoid activation operation is a non-linear activation function. The Sigmoid activation function is used to limit the pixel values of the pixel points in the attention map between 0 and 1, so that the significant information in the significant feature map in the intermediate feature map determined based on the attention map and the significant feature map remains unchanged, and the background information can be suppressed.

[0086] S20. Based on the attention map and the significant feature map, determine the mask feature map corresponding to the image to be encoded.

[0087] Specifically, the mask feature map is used to reflect the significant image information of each channel in the significant feature map. The significant image information refers to the image information that needs to be retained in each channel during encoding (for example, the image information of the foreground area). Correspondingly, the redundant image information refers to the image information of each channel during encoding (for example, the image information of the background area). In addition, each channel in the mask feature map is used to reflect the significant image information of the target channel corresponding to the channel. The target channel is the channel in the significant feature map, and the channel number of the target channel in the significant feature map is the same as the channel number of the channel in the mask feature map. Thus, the number of the first channels of the mask feature map is the same as the number of the second channels of the significant feature map. For example, if the number of the second channels of the significant feature map is 16, then the number of channels of the mask feature map is 16. Channel 10 in the mask feature map is used to reflect the significant image information carried by channel 10 in the significant feature map.

[0088] Further, in an implementation manner of this embodiment, the image encoding model includes a mask module. As Figure 5 shown, determining the mask feature map corresponding to the image to be encoded based on the attention map and the significant feature map specifically includes:

[0089] The mask module determines an intermediate feature map based on the attention map and the significant feature map, where the intermediate feature map is a single-channel feature map;

[0090] The mask module determines the mask feature map corresponding to the image to be encoded based on the intermediate feature map.

[0091] Specifically, the intermediate feature map is a single-channel feature map, which is used to reflect the significant region information in the feature map to be encoded and suppress the background region of the feature map to be encoded. Among them, the image scale of the attention image is the same as that of the significant feature map, and both the image scale of the attention image and the image size of the significant feature map are the same as the image scale of the intermediate feature map. For example, if the image scale of the attention map is 256*256*1, then the image scale of the significant feature map is 256*256*1, and correspondingly, the image scale of the intermediate feature map is 256*256*1.

[0092] In an implementation manner of this embodiment, the mask module determines the intermediate feature map based on the attention map and the significant feature map, specifically including:

[0093] For each pixel point in the attention map, obtain the corresponding candidate pixel point;

[0094] Based on the pixel value of the candidate pixel point, adjust the pixel value of this pixel point, and use the adjusted pixel value as the pixel value of this pixel point to obtain the adjusted attention map;

[0095] Based on the adjusted attention map, determine the intermediate feature map.

[0096] Specifically, the pixel position of the candidate pixel point in the significant feature map corresponds to the pixel position of this pixel point in the attention map. It can be understood that for each pixel point in the attention map, obtain the pixel position of this pixel point in the attention map. After obtaining the pixel position, select the pixel point corresponding to this pixel position in the significant feature based on this pixel position as the candidate pixel point of this pixel point. For example, for the pixel point B in this channel, the pixel position of the pixel point B in the attention map is (50, 50), then the pixel point position of the candidate pixel point in the significant feature map is (50, 50). In addition, after obtaining the candidate pixel point, the adjustment process of adjusting the pixel value of this pixel point based on the pixel value of the candidate pixel point may include: calculating the average value of the pixel value of the candidate pixel point and the pixel value of this pixel point, and adjusting the pixel value of this pixel point based on the average value. For example, use the average value as the pixel value of this pixel point; or, perform weighted processing on the pixel value of the candidate pixel point and the pixel value of this pixel point, and use the pixel value obtained by the weighted processing as the pixel value of this pixel point; or, select the maximum pixel value among the pixel value of the candidate pixel point and the pixel value of this pixel point, and use the maximum pixel value as the pixel value of this pixel point. In addition, from the fact that the value range of the pixel values of each pixel point in the significant feature map and the value range of the pixel values of each pixel point in the attention map are both [0, 1], it can be known that the value range of the pixel values of each pixel point in the intermediate feature map is [0, 1].

[0097] Further, in an implementation manner of this embodiment, the mask module determines the mask feature map corresponding to the image to be encoded based on the intermediate feature map, which specifically includes:

[0098] A10. The mask module determines a multi-channel feature map, where the image size of the multi-channel feature map is the same as the image size of the intermediate feature map;

[0099] A20. For each channel in the multi-channel feature map, the mask module adjusts the pixel values of each pixel point in this channel based on the channel number of this channel and the intermediate feature map;

[0100] A30. The adjusted multi-channel feature map is used as the mask feature map.

[0101] Specifically, in step A10, the image size of the multi-channel feature map is the same as the image size of the intermediate feature map, and the number of channels of the multi-channel feature map is the same as the number of channels of the image feature map. For example, if the image size of the intermediate feature map is 128*128 and the number of channels of the image feature map is 16, then the image size of the multi-channel feature map is 128*18 and the number of channels is 16. In addition, since the mask feature map is used to adjust the amount of image information carried by each channel in the image feature map, the image size of the mask feature map needs to be the same as the image size of the image feature map. However, the saliency feature map is determined based on the image to be encoded and is used to reflect the salient image content area and non-salient image content area in the image to be encoded, so the image size of the saliency feature map is the same as the image size of the image to be encoded.

[0102] Based on this, before determining the multi-channel feature map based on the intermediate feature map, the image size of the intermediate feature map can be adjusted so that the image size of the intermediate feature map is the same as the image size of the image feature map. Correspondingly, in an implementation manner of this embodiment, before the mask module determines a multi-channel feature map, the method includes:

[0103] The mask module adjusts the image size of the intermediate feature map and uses the adjusted intermediate feature map as the intermediate feature map, where the image size of the adjusted intermediate feature map is the same as the image size of the feature map.

[0104] Specifically, before adjusting the image size of the intermediate feature map, the first image size of the intermediate feature map and the second image size of the image feature map can be obtained first, and the adjustment ratio can be determined according to the first image size and the second image size, and the image size of the intermediate feature map can be adjusted based on this adjustment ratio. For example, if the first image size of the intermediate feature map is 256*256 and the second image size of the image feature map is 32*32, then the adjustment ratio is 256 / 32 = 8. Here, the adjustment ratio is the ratio of the width of the first image size to the width of the second image size, or the ratio of the height of the first image size to the height of the second image size, where the ratio of the width to the height of the first image size is the same as the ratio of the width to the height of the second image size. In addition, after determining the adjustment ratio, downsampling is performed on the saliency feature map based on the adjustment ratio to obtain the adjusted intermediate feature map, where the downsampling step size is the adjustment ratio. For example, if the adjustment ratio is 16, then the downsampling step size is 16.

[0105] Furthermore, in the step A20, the pixel values of each pixel point in the multi-channel feature map can all be preset values, such as 1, 0, etc. The channel number is the channel number of each channel in the multi-channel feature map. The channel numbers in the multi-channel feature map are natural numbers starting from 0, and the channel numbers corresponding to two adjacent channels are consecutive. It can be understood that the channel number of the first channel in the multi-channel feature map in the channel direction is 0, the channel number of the second channel is 1, and so on. The channel number of the last channel is C - 1, where C is the number of channels in the multi-channel feature map. In other words, when the number of channels in the multi-channel feature map is C, the channel numbers of the multi-channel feature map are 0, 1,..., C - 1. For example, when the number of channels C in the multi-channel feature map is 4, the channel numbers are 0, 1, 2, and 3.

[0106] Furthermore, in an implementation manner of this embodiment, the mask module determines the pixel value of each pixel point in the channel based on the channel number of the channel and the intermediate feature map, which specifically includes:

[0107] For each pixel point in the channel, the mask module determines the target pixel value corresponding to the pixel point;

[0108] The mask module determines the pixel value of the pixel point according to the target pixel value and the channel number of the channel.

[0109] Specifically, the target pixel value is the pixel value of the target pixel point in the intermediate feature map, and the pixel position of the target pixel point in the intermediate feature map corresponds to the pixel position of this pixel point in this channel. It can be understood that for each pixel point in this channel, the pixel position of this pixel point in this channel is obtained, where the pixel position refers to the coordinate information corresponding to the position of the pixel point in this channel. For example, for the pixel point with the coordinate information (0, 0) corresponding to the position in channel 0, the pixel position of this pixel point in channel 0 is (0, 0); after obtaining the pixel position, based on this pixel position, the pixel point corresponding to this pixel position is selected in the intermediate feature map, and the pixel value of the selected pixel point is used as the target pixel value corresponding to this pixel point. For example, for pixel point A in this channel, the pixel position of pixel point A in this channel is (50, 50), then the pixel point position of the target pixel value corresponding pixel point in the saliency feature map is (50, 50).

[0110] Furthermore, the value range of the pixel values of the pixel points in the intermediate feature map is 0 - 1, the number of channels of the multi-channel feature map is a preset number, and the pixels of each pixel point in each channel of the multi-channel feature map are determined based on the pixel value and the channel number. To improve the correlation between the pixel value and the channel number of the pixel points in the saliency feature map, before adjusting this channel based on the channel number of this channel and the intermediate feature map, it is necessary to adjust the pixel values of each pixel point in the intermediate feature map to a preset interval, where the preset interval is the value range of the pixel values of each pixel point in the adjusted single-channel feature map, the upper limit value of the preset interval is determined based on the number of channels of the multi-channel feature map, and the lower limit value of the preset interval is 0. For example, if the number of channels of the multi-channel feature map is C, then the preset interval is [0, C], and the pixel values of each pixel point in the adjusted intermediate feature map are all within the range of [0, C]. In a specific implementation manner of this embodiment, when the intermediate feature map is a normalized single-channel feature map, adjusting the pixel values of each pixel point in the intermediate feature map to the preset interval can multiply the pixel value of each pixel point in the single-channel feature map by the number of channels of the multi-channel feature map. For example, if the number of channels of the multi-channel feature map is C, then for the pixel value of each pixel point in the saliency feature map, multiply the pixel value of this pixel point by C, and use the obtained product as the pixel value corresponding to the target pixel value. In this way, the pixel values of each pixel point in the intermediate feature map are adjusted to the preset interval, so that the importance degree of each pixel point in each intermediate feature map on each channel of the multi-channel feature map can be determined, thereby determining the saliency region information in each channel.

[0111] Further, after obtaining the target pixel value corresponding to the pixel point, according to the target pixel value and the channel number corresponding to the channel map where the pixel point is located, calculate the pixel value corresponding to the pixel point, and use the calculated pixel value as the pixel value corresponding to the pixel point. Thus, the pixel values of each pixel point in each channel of the multi-channel feature map can be adjusted, and the adjusted multi-channel feature map is used as the mask feature map. In a specific implementation manner of this embodiment, the calculation formula for the pixel value of each pixel point in each channel of the multi-channel feature map can be:

[0112]

[0113] where k is the channel number of the channel in the multi-channel feature map, k = 0, 1, 2..., C - 1, and C is the number of channels of the multi-channel feature map; i, j represent the position of the pixel point in the channel; m i,j,k is the pixel value of the pixel point; y i,j is the pixel value of the target pixel value at the i, j position in the intermediate feature map.

[0114] In this embodiment, by mapping the value range of the pixel value of each pixel point in the intermediate feature map from [0, 1] to [0, C], and then corresponding each channel of the intermediate feature map to each channel of the multi-channel feature map according to the above formula, and using the pixel value y i,j of each pixel point and the channel number k of each channel to determine the pixel value of each pixel point in each channel k, it can be ensured that each channel k carries different significant image information, so that each channel in the encoded encoded feature map carries different amounts of significant image information. During encoding, adaptive encoding can be performed according to the amount of significant image information carried, so that channels carrying more useful information can be allocated more bit positions, and thus more significant image information is retained in the encoded file.

[0115] S30. Generate an encoded feature map corresponding to the image to be encoded based on the image feature map and the mask feature map.

[0116] Specifically, the encoded feature map is a feature map for encoding. After obtaining the encoded feature map, the encoded feature map can be encoded to obtain an encoded file corresponding to the image to be encoded. The mask feature map is used to reflect the significant image information of each channel in the image feature map. Thus, when determining the encoded feature map according to the image feature map and the mask feature map, information filtering can be performed on each channel in the image feature map through the mask feature map to remove the non-significant image information carried by each channel in the image feature map.

[0117] In addition, filtering the information of each channel in the image feature map through the mask feature map means that for each target channel in the image feature map, determining the reference channel corresponding to the target channel, performing an element-wise multiplication operation between each channel in the image feature map and its corresponding reference channel, and using the image obtained from the element-wise multiplication operation as the encoded feature map. The channel number of the reference channel in the mask feature map is the same as the channel number of the target channel in the image feature map, and for each target channel in the image feature map, there is a reference channel in the mask feature map corresponding to the target channel. This is because the number of channels in the mask feature map is the same as the number of channels in the image feature map, and the channel numbering method for each channel in the mask feature map is the same as the channel numbering method for each channel in the image feature map. For example, if the number of channels in the image feature map is C and the channel encoding method is 0, 1,..., C - 1, then the number of channels in the mask feature map is C and the channel encoding method is 0, 1,..., C - 1; based on this, the target channel with channel number 5 in the image feature map corresponds to the reference channel with channel number 5 in the mask feature map. In this embodiment, using the mask feature map to screen the image feature map can improve the information of the significant regions in the encoded feature map, filter the information of the non-significant regions, and increase the amount of information in the significant regions of the encoded feature map.

[0118] Further, after obtaining the mask feature map, the mask feature map can be adjusted so that while the mask feature map carries the information of the salient region, it retains some information of the non-salient region. In this way, while including the information amount of the salient region information in the encoded feature map, the encoded feature map can also include some non-salient region information, so that a reconstructed image is obtained based on the encoded file corresponding to the encoded file. While ensuring the improvement of the image details of the salient region, the reconstructed image can carry the image content of the non-salient region. Based on this, in an implementation manner of this embodiment, the adjustment process of the mask feature map may specifically include: for each channel in the mask feature map, add the pixel value of each pixel point in the channel to a first preset value to obtain the added pixel value; then divide the added pixel value corresponding to each pixel point in the channel by a second preset value, and use the quotient obtained by the division as the pixel value corresponding to the pixel point to obtain the adjusted mask feature map. Any one of the first pixel values of the pixel points representing the salient region in the obtained mask feature map is greater than all the second pixel values. The second pixel value is used to represent the non-salient region, and the second pixel value is not zero. In this way, in the encoded image determined based on the mask image and the image feature map, while highlighting the content of the salient region (retaining more information, that is, allocating more bit positions), the content of the non-salient region can also be retained (retaining less information, that is, allocating fewer bit positions), so that the decoded image can be complete and enhance the details and textures of the salient region. In an implementation manner of this embodiment, the first preset value may be 1, and the second preset value may be 2.

[0119] Further, in order to reduce the data volume of the encoded feature map, after the image encoding model according to the image feature map and the mask feature map, the encoded feature map can be quantized, and the encoded feature map is obtained according to the quantized feature map. Correspondingly, in an implementation manner of this embodiment, the image encoding module includes a quantization module. After generating the encoded feature map corresponding to the image to be encoded based on the image feature map and the mask feature map, the method includes:

[0120] The quantization module generates a quantization feature map of the image to be encoded based on the encoded feature map, and uses the quantization feature map as the encoded feature map.

[0121] Specifically, the quantization refers to dividing the value range of each pixel point of the encoded feature map into several intervals, and setting the values of all pixel points in each interval to the same value. The quantization of the encoded feature map can adopt existing quantization methods that can achieve image quantization. In one implementation of this embodiment, the quantization method for the encoded feature map can be to quantize the encoded feature map using a clustering quantization method. The process of quantizing the encoded feature map using the clustering quantization method can be as follows: Given the clustering quantization center points, calculate the distance between each pixel point in the encoded feature map and the quantization center points, and take the minimum distance among all the obtained distances as the quantization value. Among them, the calculation formula for the distance between each pixel point and the quantization center points can be:

[0122] Q(input_x i ):=argmin j (input_x i -c j ),

[0123] where input_x i represents the i-th data of the input encoded feature map, and c j represents the j-th component of the clustering quantization center points C={c1,c2,...,c L}, j ∈ [1, L], and L is a positive integer.

[0124] Furthermore, in order to ensure error backpropagation, it is necessary to perform soft quantization processing on the distance between each pixel point and the quantization center points first, and then perform hard quantization processing. The processing method of the soft quantization processing is:

[0125]

[0126] The processing process of the hard quantization processing is:

[0127] stop_gradient(Q(input_x i )-soft_Q(input_x i ))+soft_Q(input_x i )

[0128] where stop_gradient(·) stops gradient calculation.

[0129] In addition, after quantization processing, round the quantized distances, determine the quantization value according to the rounded distances, and finally quantize the encoded feature map according to the quantization value to obtain the quantized encoded feature map.

[0130] S40. Based on the encoded feature map, obtain the encoded file corresponding to the image to be encoded.

[0131] Specifically, the encoded file is obtained by encoding the encoded feature map, and when encoding the encoded feature map, entropy encoding can be used for encoding. The encoded file can be losslessly compressed from the encoded feature map by entropy encoding. Among them, the entropy encoding can adopt various existing encoding methods, for example, Huffman encoding or arithmetic encoding, etc. Of course, it is worth noting that when encoding the encoded feature map by entropy encoding, an adaptive encoding method based on the amount of information carried by the channels with significant image information is adopted. It can be understood that when encoding the encoded feature map, the number of bits corresponding to each channel is determined based on the amount of information carried by each channel in the encoded feature map, and the corresponding number of bits is respectively assigned to each channel during encoding. In addition, the number of bits corresponding to the channel is positively correlated with the amount of information carried by the channel with significant image information, that is, the more the amount of information carried by the channel with significant image information, the larger the number of bits corresponding to the channel; conversely, the less the amount of information carried by the channel with significant image information, the smaller the number of bits corresponding to the channel.

[0132] Based on the above image encoding method, this embodiment can also provide an image decoding method, as Figure 6 shown, the image decoding method is applied to an image decoding model, and the image decoding method includes:

[0133] The image decoding module determines the reconstructed image corresponding to the encoded file based on the encoded file, where the encoded file is encoded based on the above encoding method.

[0134] Specifically, as Figure 7 shown, the image decoding model includes a first convolution module A, a second convolution module A, and a reconstruction module. The first convolution module A includes a first convolution unit, a second convolution unit, and a third convolution unit. Among them, in the first convolution unit, the convolution kernel is 3*3, the number of convolution kernels is 480, the stride is 1, and the padding pixel is 1; the normalization operation is the InstanceNorm operation; the non-linear transformation function is the relu function; in the second convolution unit, the convolution kernel is 3*3, the number of convolution kernels is 960, the stride is 1, and the padding pixel is 1; the normalization operation is the InstanceNorm operation; the non-linear transformation function is the relu function; in the third convolution unit, the convolution kernel is 3*3, the number of convolution kernels is 960, the stride is 1, and the padding pixel is 1; the normalization operation is the InstanceNorm operation; the non-linear transformation function is the relu function.

[0135] The second convolution module A includes 9 residual blocks, as Figure 8As shown, each residual block includes a fourth convolutional unit and a fifth convolutional unit. In the fourth convolutional unit, the convolutional kernel is 3*3, the number of convolutional kernels is 960, the stride is 1, and the padding pixels are 1; the normalization operation is the InstanceNorm operation; the non-linear transformation function is the relu function. In the fifth convolutional unit, the convolutional kernel is 3*3, the number of convolutional kernels is 960, the stride is 1, and the padding pixels are 1; the normalization operation is the InstanceNorm operation. After the feature map is output through the fifth convolutional unit, the input term of the fourth convolutional unit is added to the output term of the fifth convolutional unit through a shortcut operation to obtain the output term corresponding to each residual block.

[0136] The reconstruction module includes four cascaded upsampling modules and a sixth convolutional unit; as Figure 9 shown, the upsampling module includes an upsampling unit and a seventh convolutional unit. In the upsampling unit, 2-fold bilinear interpolation upsampling is performed, where the number of convolutional kernels corresponding to the upsampling is 480; in the seventh convolutional unit, the convolutional kernel is 3*3, the number of convolutional kernels is 480, the stride is 1, and the normalization operation is the InstanceNorm operation; the non-linear transformation function is the relu function. In the sixth convolutional unit, the convolutional kernel is 7*7, the number of convolutional kernels is 3, the stride is 1, and the padding pixels are 3.

[0137] In summary, this embodiment provides an image coding method. The image coding method obtains an image feature map and a saliency feature map corresponding to an image to be coded through an image coding model; determines a mask feature map corresponding to the image to be coded based on the saliency feature map; generates a coding feature map corresponding to the image to be coded according to the image feature map and the mask feature map; and finally obtains a coding file corresponding to the image to be coded according to the coding feature map. In this application, the mask feature map determined according to the saliency feature map is used to determine the image information included in each channel of the image feature map, so that the saliency image information carried by each channel in the coding feature map is different. In this way, different bit positions can be allocated to different channels according to the saliency image information during image coding, so that the channels including more saliency image information occupy more bit positions, improving the image information of the saliency image content in the coding file, and further improving the image effect of the reconstructed image reconstructed according to the coding file. For example, as Figure 10 shown, the first reconstructed image generated from the coding file obtained by using the coding method provided in this embodiment, and as Figure 11 shown, the second reconstructed image generated from the coding file directly using the image feature map for equal-width coding. It can be seen that the image details of the parrot area as the saliency area in the first reconstructed image are better than the image details of the parrot area as the saliency area in the second reconstructed image.

[0138] Based on the above image encoding method, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the image encoding method as described in the above embodiment.

[0139] Based on the above image encoding method, the present application also provides a terminal device, as Figure 12 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.

[0140] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0141] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as program instructions or modules corresponding to the method in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the method in the above embodiment.

[0142] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes can also be transient storage media.

[0143] In addition, the specific processes of loading and executing multiple instructions by the above-mentioned storage medium and the instruction processor in the terminal device have been described in detail in the above method and will not be repeated here.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. An image encoding method, characterized in that, The method includes: Obtaining an image feature map, a saliency feature map, and an attention map corresponding to the image to be encoded; Determining a mask feature map corresponding to the image to be encoded based on the attention map and the saliency feature map; Generating an encoded feature map corresponding to the image to be encoded based on the image feature map and the mask feature map; Obtaining an encoded file corresponding to the image to be encoded based on the encoded feature map; The image feature map is a global feature map of the image to be encoded, and each channel in the image feature map is a local feature map of the image to be encoded; The saliency feature map is used to reflect the foreground feature information of the image to be encoded, where the saliency feature map is a single-channel feature map; The image encoding method applies an image encoding model; The image encoding model includes a mask module, and the determining the mask feature map corresponding to the image to be encoded based on the attention map and the saliency feature map specifically includes: The mask module determines an intermediate feature map based on the attention map and the saliency feature map, where the intermediate feature map is a single-channel feature map used to reflect the saliency region information in the image to be encoded and suppress the background region of the image to be encoded; The mask module determines a multi-channel feature map, where the image size of the multi-channel feature map is the same as that of the intermediate feature map; For each channel in the multi-channel feature map, the mask module adjusts the pixel values of each pixel point in the channel based on the channel number of the channel and the intermediate feature map, so that each channel carries different saliency image information, and the channels carrying more useful information are allocated more bit positions, and more saliency image information is retained in the encoded file obtained by encoding; Taking the adjusted multi-channel feature map as the mask feature map.

2. The image encoding method according to claim 1, wherein Each channel in the mask feature map corresponds one-to-one with each channel in the saliency feature map, and there are at least two channels in the mask feature map that contain different amounts of image information.

3. The image encoding method according to claim 1, wherein The image encoding model includes a feature map extraction module, a saliency extraction module, and an attention module; The obtaining the image feature map, the saliency feature map, and the attention map corresponding to the image to be encoded specifically includes: The feature map extraction module determines an image feature map corresponding to the image to be encoded based on the image to be encoded; The saliency extraction module determines a saliency feature map corresponding to the image to be encoded based on the image to be encoded; The attention module determines an attention map corresponding to the image to be encoded based on the image feature map.

4. The image encoding method according to claim 3, wherein The attention module includes a first attention unit, a second attention unit, and a fusion unit; the attention module determining the attention map corresponding to the image to be encoded based on the image feature map specifically includes: The first attention unit determines a first attention map corresponding to the image to be encoded based on the image feature map; The second attention unit determines a second attention map corresponding to the image to be encoded based on the image feature map; The fusion unit determines an attention map corresponding to the image to be encoded based on the first attention map and the second attention map.

5. The image encoding method according to claim 1, wherein The mask module determines a mask feature map corresponding to the image to be encoded based on the intermediate feature map.

6. The image encoding method according to claim 5, characterized in that, The image scale of the attention image is the same as the image scale of the saliency feature map; The mask module determining the intermediate feature map based on the attention map and the saliency feature map specifically includes: For each pixel point in the attention map, obtain the corresponding candidate pixel point, where the pixel position of the candidate pixel point in the saliency feature map corresponds to the pixel position of this pixel point in the attention map; Adjust the pixel value of this pixel point based on the pixel value of the candidate pixel point, and use the adjusted pixel value as the pixel value of this pixel point to obtain an adjusted attention map; Determine the intermediate feature map based on the adjusted attention map.

7. The image encoding method according to claim 1, wherein The mask module determining the pixel value of each pixel point in this channel based on the channel number of this channel and the intermediate feature map specifically includes: For each pixel point in this channel, the mask module determines the corresponding target pixel value, where the target pixel value is the pixel value of the target pixel point, and the pixel position of the target pixel point in the intermediate feature map corresponds to the pixel position of this pixel point in this channel; The mask module determines the pixel value of this pixel point according to the target pixel value and the channel number of this channel.

8. The image encoding method according to claim 1, characterized in that, Before the mask module determines a multi-channel feature map, the method includes: The mask module adjusts the image size of the intermediate feature map, and uses the adjusted intermediate feature map as the intermediate feature map, where the image size of the adjusted intermediate feature map is the same as the image size of the image feature map.

9. The image encoding method according to claim 1, wherein The image encoding module includes a quantization module; After generating an encoded feature map corresponding to the image to be encoded based on the image feature map and the mask feature map, the method includes: The quantization module generates a quantized feature map of the image to be encoded based on the encoded feature map, and uses the quantized feature map as the encoded feature map.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the image encoding method according to any one of claims 1-9.

11. A terminal device, characterized in that, Including: A processor, a memory and a communication bus; a computer-readable program executable by the processor is stored on the memory; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, it implements the steps in the image encoding method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Region-of-interest image coding and decoding system and method based on deep learning

    CN109889839A

  • Method for detecting image salient target

    CN110956185A

  • Panoramic image segmentation method and device and electronic equipment

    CN111292334A

  • Pedestrian re-identification method and device based on mask alignment and attention mechanism

    CN111353385A