Image Encoding Method, Decoding Method, Device, Readable Medium, and Electronic Device
The image encoding method enhances compression efficiency by selectively compressing key regions using regional gradient calculation and a visual transformation model, addressing the inefficiencies of uniform encoding in existing technologies and enabling machine vision tasks.
Patent Information
- Application Number
- JP2025501514
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-15
- Filing Date
- 2023-07-14
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-07-14
AI Technical Summary
Current image/video encoding technologies based on convolutional neural networks uniformly encode all regions of an image, failing to distinguish key and non-key regions, leading to inefficient compression and inability to perform machine vision tasks effectively.
An image encoding method that utilizes regional gradient calculation and a visual transformation model to selectively compress key regions and discard non-key regions, allowing flexible control of compression ratios.
Improves image compression efficiency by retaining important image information and enabling effective machine vision tasks through selective and controllable compression.
Smart Images

Figure 2025522069000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This disclosure is filed based on a Chinese patent application with application number 202210837739.2, filing date July 15, 2022, and invention title "Image Encoding Method, Decoding Method, Apparatus, Readable Medium, and Electronic Device", and the priority of the Chinese patent application is claimed, and all the contents of the Chinese patent application are incorporated herein by reference.
[0002] This disclosure belongs to the field of artificial intelligence technology, and specifically relates to an image encoding method, a decoding method, an apparatus, a readable medium, and an electronic device.
Background Art
[0003] Conventional image / video encoding is aimed at human visual tasks and is mostly used for entertainment purposes, emphasizing the fidelity, high frame rate, sharpness, etc. of video data signals. With the rapid development of 5G, big data, and artificial intelligence, under the background of image / video big data applications, media content, such as images and videos, is widely applied in fields such as intelligent visual tasks such as target detection, target tracking, image classification, image segmentation, pedestrian re - identification, etc. These intelligent visual tasks are also called intelligent tasks for machine vision.
[0004] It should be noted that the information disclosed in the above background art part is only used to enhance the understanding of the background of this disclosure, and thus may include information that does not constitute prior art known to those skilled in the art.
Summary of the Invention
[0005] Other features and advantages of this disclosure will become apparent from the following detailed description or will be partially acquired through the practice of this disclosure.
[0006] According to one aspect of the embodiments of the present disclosure, an image encoding method is provided. The image encoding method includes: acquiring an original image, performing a blocking process to obtain a plurality of image blocks; calculating a gradient value of pixels in each image block, and selecting important region blocks from the plurality of image blocks based on the gradient value of the pixels; and inputting the important region blocks and position information of the important region blocks in the original image into a visual transformation model for encoding to generate a bitstream.
[0007] According to one aspect of the embodiments of the present disclosure, an image encoding apparatus is provided. The image encoding apparatus includes an acquisition module, a calculation module, and an encoding module. The acquisition module is configured to acquire an original image, perform a blocking process, and obtain a plurality of image blocks. The calculation module is configured to calculate a gradient value of pixels in each image block and select important region blocks from the plurality of image blocks based on the gradient value of the pixels. The encoding module is configured to input the important region blocks and position information of the important region blocks in the original image into a visual transformation model for encoding to generate a bitstream.
[0008] In some embodiments of the present disclosure, the calculation module is further configured to calculate a gradient value of pixels in each image block, calculate an average gradient value of each image block based on the gradient value of the pixels, sort the plurality of image blocks based on the average gradient value, and determine, as the important region blocks, the image blocks among the plurality of image blocks whose average gradient value is greater than or equal to a predetermined value.
[0009] In some embodiments of the present disclosure, the encoding module is further configured to input the important region blocks into the visual transformation model to output an encoded visible patch and a mask token, generate an image token based on the encoded visible patch, the mask token, and position information of the important region blocks in the original image, and generate the bitstream based on the image token.
[0010] In some embodiments of the present disclosure, the acquisition module further acquires an original image of n×n, where n is a positive integer, evenly blocks the n×n original image into m×m according to non-overlapping regions, and sets the size of each image block to
Number
[0011] In some embodiments of the present disclosure, the calculation module further discards an image block among the plurality of image blocks whose gradient average value is smaller than a predetermined value, and by setting the predetermined value, the number of discarded image blocks and the preset compression ratio α of the image are
Number
[0012] According to one aspect of the embodiments of the present disclosure, an image decoding method is provided, which decodes the encoding by the above image encoding method. The image decoding method includes receiving a bitstream generated by encoding, and decoding the bitstream, and obtaining a reconstructed image through normalization, a multi-head self-attention mechanism, and multi-layer perceptron processing of the decoding result.
[0013] According to one aspect of the embodiments of the present disclosure, an image decoding apparatus is provided. The image decoding apparatus includes a receiving module and a decoding module. The receiving module is configured to receive a bitstream generated by encoding, and the decoding module is configured to decode the bitstream and obtain a reconstructed image through normalization, a multi-head self-attention mechanism, and multi-layer perceptron processing of the decoding result.
[0014] According to one aspect of the embodiments of the present disclosure, there is provided a computer-readable medium storing a computer program that, when executed by a processor, implements the image encoding method or the image decoding method in the above aspect.
[0015] According to one aspect of the embodiments of the present disclosure, there is provided an electronic device including a processor and a memory for storing executable instructions of the processor, wherein the processor is configured to execute the image encoding method or the image decoding method in the above aspect by executing the executable instructions.
[0016] According to one aspect of the embodiments of the present disclosure, there is provided a computer program product or a computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, and the computer device executes an image encoding method or an image decoding method as in the above technical solution.
[0017] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and are not intended to limit the present disclosure.
Brief Description of the Drawings
[0018] Here, the drawings are incorporated into the specification, showing embodiments that conform to the present disclosure, and are used to explain the principles of the present disclosure together with the specification. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those skilled in the art can also obtain other drawings based on these drawings without creative labor.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0019] Next, with reference to the drawings, exemplary embodiments will be described in more detail. However, it should be understood that the exemplary embodiments can be implemented in various forms and are not limited to the examples described herein. In contrast, these embodiments are provided to make the present disclosure more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0020] Furthermore, the described features, structures, or characteristics can be incorporated into one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will recognize that they can implement the technical approach of the present disclosure without one or more of the specific details, or can adopt other methods, components, devices, steps, etc. In other cases, well-known methods, devices, implementations, or operations cannot be illustrated or described in detail so as not to obscure the aspects of the present disclosure.
[0021] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, may be implemented in one or more hardware modules or integrated circuits, or may be implemented in different networks and / or processor devices and / or microcontroller devices.
[0022] The flowcharts shown in the drawings are only illustrative explanations and do not have to include all contents and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be merged or partially merged, so the actual execution order may change according to the actual situation.
[0023] With the popularization of machine vision intelligent tasks such as image classification, video object detection, object tracking, image segmentation, and pedestrian re-identification, in the current related technical solutions, when adopting the image / video encoding and decoding technology based on convolutional neural networks, this method is disadvantageous for image encoding / decoding because it uniformly encodes all regions of the entire image.
[0024] In contrast, the present disclosure provides an image encoding method, a decoding method, an apparatus, a readable medium, and an electronic device. In the technical solution provided by the embodiments of the present disclosure, first, the original image is blocked and combined using block region gradient calculation and a visual conversion model. In this way, this method selectively and controllably compresses different information regions of the compressed image, retains as much as possible for the key regions where the information in the image / video is concentrated, compresses as little as possible, and for the non-key regions where the information in the image is sparse, compresses as much as possible, improves the image compression efficiency, realizes flexible coding rate control under a unified solution, and thereby realizes improving the image compression efficiency to a certain extent.
[0025] Referring to FIG. 1, FIG. 1 schematically shows an exemplary image encoding / decoding system architecture block diagram.
[0026] The system includes a data acquisition module 101, an encoder module 102, and a decoder module 103. Here, the data acquisition module 101 is used to acquire the image / video 1001 and transmit it to the encoder module 102. In the encoder module 102, a convolutional neural network is utilized to encode the image / video 1001 into a bitstream 1002, and the bitstream 1002 may be used to transmit to the decoder module 103 at the other end. Similarly, in the decoder module 103, a convolutional neural network can be used to reconstruct the bitstream into an image / video 1003, input the reconstructed image / video 1003 as a human vision task 104, and finally, a result 1004 is obtained after the calculation by the human vision task.
[0027] Adopting this method has the following technical problems. The encoding of the image / video by the encoder and decoder uniformly encodes all regions of the entire image and cannot distinguish the key regions and non-key regions of the image itself. After uniformly encoding all regions of the entire image, the key regions of the image are greatly compressed, losing important information of the image. In the current method, each region of the image cannot be selectively compressed. That is, in the current image encoding / decoding method, image blocks in non-important regions of the image cannot be discarded during image encoding. Therefore, these encoders designed based on the deep convolutional neural network structure are not flexible in controlling the compression ratio. Also, since the current encoding / decoding system method is for human vision tasks, it cannot successfully complete the machine vision intelligent analysis task when facing machine vision tasks.
[0028] To solve the above problems, the present disclosure redesigns the encoder and decoder modules in an encoding system for machine vision intelligence analysis tasks. In the design of the encoder, a Transformer image encoder based on regional gradient information is proposed. In the decoder design, a decoder based on the Transformer module (Block) is proposed. Referring to FIG. 2, FIG. 2 schematically shows an exemplary system architecture block diagram to which the technical solution of the present disclosure is applied. This system architecture includes a data collection module, which is used to collect images / videos (S201) to obtain the original image 2001. Then, the original image is input into the encoder module 2002, and the original image is sequentially subjected to blocking processing (S202), gradient calculation (S203), important region calculation (S204), and vision Transformer encoding (S205) to output a bitstream. The bitstream output by the encoder module 2002 is reconstructed (S206) into an image / video through the decoder module 2003 using a converter module. The reconstructed image / video is used as the input for the machine vision task. Finally, the result 2004 is obtained through machine vision task calculation (S207).
[0029] To achieve selective compression of different regions of an image / video, when the present disclosure encodes image / video data, the designed encoder combines regional gradient calculation and a converter module, and the design of the image decoder is only the converter module. When encoding an image, the image is subjected to blocking calculation, the gradient value is calculated for the image pixels within each block region, the average value of the gradient calculation values for each region is calculated, all the image blocks are sorted based on the average value of the gradient calculation values, and the sorted lower blocks are discarded. The sorted upper image blocks are input into the subsequent converter module, and the images in other image block regions are discarded as they are. By controlling the ratio of the discarded images, the compression ratio can be flexibly controlled.
[0030] Hereinafter, an image encoding method, a decoding method, an apparatus, a readable medium, and an electronic device provided by the present disclosure will be described in detail based on specific embodiments.
[0031] Referring to FIG. 3, FIG. 3 schematically shows a step flow of an image encoding method according to an embodiment of the present disclosure. The image encoding method can be executed by a controller and mainly can include the following steps S301 to S303.
[0032] In step S301, an original image is acquired, block processing is performed, and a plurality of image blocks are acquired.
[0033] In some embodiments, an image / video is acquired by a data acquisition module to obtain an original image, and the original image is subjected to block processing. For example, the size of the original image is n×n, and the n×n image is evenly blocked into m×m image blocks according to non-overlapping regions, and the size of each image block is
Number
[0034] In step S302, the gradient value of the pixels in each image block is calculated, and important region blocks are selected from the plurality of image blocks based on the gradient value of the pixels.
[0035] In some embodiments, it is advantageous to calculate the gradient value of the pixels in each block and select important region blocks based on the gradient value of the pixels. Thereby, selective and controllable compression can be performed on different information regions of the compressed image, retaining as much as possible for the key regions where the information in the image / video is concentrated and performing as little compression as possible. On the other hand, for the non-key regions where the information in the image is sparse, as much compression as possible is performed to improve the image compression efficiency and achieve flexible coding rate control under a unified solution.
[0036] In step S303, the important region block and the position information in the original image of the important region block are input into the visual conversion model for encoding to generate a bitstream.
[0037] In the technical solution provided by the embodiments of the present disclosure, first, the original image is blocked and combined with block region gradient calculation and a visual conversion model. In this way, this method selectively and controllably compresses different information regions of the compressed image, retains as much as possible for the key regions where the information in the image / video is concentrated and compresses as little as possible. On the other hand, for the non-key regions where the information in the image is sparse, as much compression as possible is performed to improve the image compression efficiency and achieve flexible coding rate control under a unified solution.
[0038] In some embodiments of the present disclosure, calculating the gradient value of the pixels in each image block and selecting important region blocks from a plurality of image blocks based on the gradient value of the pixels includes calculating the gradient value of the pixels in each image block, calculating the gradient average value of each image block based on the gradient value of the pixels, sorting a plurality of image blocks based on the gradient average value, and determining the image blocks whose gradient average value is greater than or equal to a predetermined value among the plurality of image blocks as important region blocks.
[0039] In this way, the important regions of information focus can be sorted and selected based on the gradient calculation average value, the non-important region blocks in the image can be discarded, and image compression can be realized.
[0040] In some embodiments, when selecting important region blocks, for each pixel (x, y) in each image block, the gradients in the x - direction and y - direction are calculated respectively. The gradient calculation in the x - direction is as shown in the following formula (1).
[0041]
Number
[0042] The gradient calculation in the y - direction is as shown in the following formula (2).
[0043]
Number
[0044] The gradient value g x in the x - direction of the pixel (x, y) and the gradient value g y in the y - direction are calculated as follows in the following formula (3):
[0045]
Number
[0046] Here, g(x, y) is the gradient calculation value of (x, y). Next, the gradient calculation values of all pixels within the block region are calculated as follows in the following formula (4).
[0047]
Number
[0048] Here, d(i, j) is the average gradient value of all pixels within each block. The ranges of the values of i and j are both 0 to m - 1.
[0049] Referring to FIG. 5, FIG. 5 schematically shows a schematic diagram of the gradient average value of each image block to which the technical solution of the present disclosure is applied. The original image 5002 with a size of n×n is evenly blocked into m×m 504 according to non-overlapping regions (S502), and the size of each image block is [Number] . Next, all image blocks are sorted from largest to smallest according to the value of d(i,j), that is [Number] . The p image blocks with the smallest sorted d(i,j) values are discarded, and the number of the remaining image blocks is n×n - p. These remaining image blocks are used as important region blocks.
[0050] In this way, this method selectively and controllably compresses different information regions of the compressed image, retains as much as possible for the key regions where the information in the image / video is concentrated, and compresses as little as possible. On the other hand, for the non-key regions where the information in the image is sparse, it compresses as much as possible to improve the image compression efficiency and achieve flexible coding rate control under a unified solution.
[0051] In one embodiment of the present disclosure, inputting the important region block and the position information in the original image of the important region block into the visual conversion model for encoding to generate a bitstream includes inputting the important region block into the visual conversion model and outputting an encoded visible patch and a mask token, and generating an image token based on the encoded visible patch, the mask token and the position information in the original image of the important region block, and generating a bitstream based on the image token.
[0052] Referring to FIG. 6, FIG. 6 schematically shows a schematic diagram of an encoder module to which the technical solution of the present disclosure is applied. After obtaining the non-overlapping region segmentation result 6002 of the original image, gradient calculation (S602) is performed, and then important region calculation (S604) is performed to determine the important region blocks, and the important region blocks that have not been discarded and their position information in the original image are input together into the Vision Transformer model 6004. Here, the patch embedding (Patch Embeddings) and position embedding (Positional Embeddings) 60042 information of the important region blocks are input into the encoder 60044 module of the Vision Transformer model 6004.
[0053] Calculating using a plurality of encoder modules of the Vision Transformer model 6004, p patch information and its position information of the same number as the input are obtained, and at this time, re-sorting is performed based on the position information to obtain d×d image blocks of the same number as the original image size. Among these d×d image blocks, the patch information of the important region blocks that have not been discarded so far is calculated by the Vision Transformer model 6004 and is called Encoded Visible Patches. The rest are sorted based on the position information and are called Mask Tokens.
[0054] In this way, the video encoding system method of the related technical solution is aimed at human vision tasks, and when it is aimed at machine vision tasks, it cannot successfully complete the machine vision intelligent analysis task. The technical proposal of this embodiment is aimed at machine vision tasks and can better perform the machine vision intelligent analysis task.
[0055] In some embodiments of the present disclosure, a predetermined value can be set so that the number of discarded image blocks and the preset compression ratio α of the image satisfy Equation (5).
[0056]
Equation
[0057] Here, p is the number of discarded image blocks.
[0058] In this way, by controlling the ratio of the image blocks to be discarded after blocking the image, the compression rate can be flexibly controlled.
[0059] According to one aspect of the embodiments of the present disclosure, an image decoding method is provided, which is an image decoding method for decoding the encoding by the above image encoding method, including receiving a bitstream generated by encoding, decoding the bitstream, and outputting a reconstructed image after normalizing the decoding result, passing through a multi-head self-attention mechanism and a multi-layer perceptron process.
[0060] Referring to FIG. 7, FIG. 7 schematically shows a schematic diagram of a decoder module to which the technical solution of the present disclosure is applied. Referring to FIG. 6, after obtaining the output encoded visible patch 70022 and the mask token 70024 of the visual transformation model based on the region gradient information, these two parts are combined with the position information of the original image for position embedding (S702) and added, and the added result is input to a decoder constructed by a converter module 7004 for decoding (S704). In the decoder, the converter module 7004 is composed of a normalization (Normalize) layer 70042 and 70046, a multi-head self-attention (Multi-head Self Attention) layer 70044, and a multi-layer neural network (Multi-Layer Perceptron, MLP, also called a multi-layer perceptron) module 70048.
[0061] After inputting the information of the image block of the vector t output by the normalization layer 70042 into the multi-head self-attention layer 70044, the weight matrix W in the multi-head self-attention layer 70044 t , and the attention weight matrix of each head
Number
[0062] Next, multiply the image block vector t by the attention weight matrix of each header
Number
Number
[0063]
Number
[0064] Each image block vector t is calculated as shown in the following formulas (7) and (8) corresponding to the attention of each header:
[0065]
Number
[0066]
Number
[0067] In the above formula, h t is the header corresponding to the image block vector t for each attention,
Number
Number
Number
[0068] In Equation (8),
Number
Number
[0069] To facilitate the understanding of the technical solution of the present disclosure, referring to FIG. 8, FIG. 8 schematically shows a schematic diagram of a codec flow to which the technical solution of the present disclosure is applied.
[0070] On the encoding side, in step S801, the original image (or video) 8002 is subjected to block calculation, and an n×n image is evenly blocked into m×m according to non-overlapping regions.
[0071] In step S802, for each pixel (x, y) in the image block, using Equations (1) and (2), the gradient in its x direction and the gradient in its y direction are calculated.
[0072] In step S803, using Equation (3), the gradient calculation value of the pixel (x, y) is calculated.
[0073] In step S804, using Equation (4), the gradient average value of all pixels in each image block is calculated.
[0074] In step S805, all image blocks are sorted according to the value of d(i, j). Discard p blocks with small sorted d(i, j) values, and the calculation of the compression ratio α satisfies Equation (5).
[0075] In step S806, an image token is generated based on the result of the deletion operation. For example, the encoded visible patch includes a mask token.
[0076] In step S807, a bitstream 8004 is generated based on the image token.
[0077] On the decoding side, in step S808, the encoded visible patch, mask token, and position embedding information are obtained based on the bitstream 8006.
[0078] In step S809, the data with the encoded visible patch and mask token position-embedded is normalized. In step S810, multi-head self-attention is calculated using (6) to (8).
[0079] In step S811, the multi-head self-attention calculation result is normalized.
[0080] In step S812, the normalized result is calculated by a multi-layer perceptron.
[0081] In step S813, the reconstructed image / video 8008 is output.
[0082] The present disclosure designs an image codec using a method based on region gradient calculation in image encoding and decoding, that is, selectively compresses image content information, calculates the gradient, gradient calculation value, and average gradient calculation value for a block of images, selects important information blocks based on the average gradient calculation value, proposes an idea of calculating key regions based on the average gradient calculation value information, sorts and selects important regions of information focus, further discards unimportant regions in the image, realizes selective compression of the image, and can flexibly control the compression ratio by controlling the ratio of image blockification after the image is blockified. Also, all the video encoding system methods of related technical solutions are for human visual tasks, and in the case of machine vision tasks, they cannot successfully perform machine vision intelligent analysis tasks. The system proposed in the present disclosure is for machine vision tasks and can better achieve machine vision intelligent analysis tasks.
[0083] Note that in the accompanying drawings, each step of the method in the present disclosure is described in a specific order, but this does not require or imply that these steps must be executed in that specific order, or that the desired result cannot be achieved without executing all the illustrated steps. Additionally or alternatively, some steps can be omitted, multiple steps can be merged into one step execution, or one step can be decomposed into multiple step executions.
[0084] The following introduces embodiments of the apparatus of the present disclosure, which can be used to execute the image encoding method or the image decoding method in the above-described embodiments of the present disclosure. FIG. 9 is a schematic diagram showing a configuration block diagram of an image encoding apparatus provided according to an embodiment of the present disclosure. As shown in FIG. 9, the image encoding apparatus 900 can include an acquisition module 901, a calculation module 902, and an encoding module 903.
[0085] The acquisition module 901 is configured to acquire the original image, perform blockification processing, and acquire a plurality of image blocks. The calculation module 902 is configured to calculate the gradient value of the pixels in each image block and select important region blocks from a plurality of image blocks based on the gradient values of the pixels. The encoding module 903 is configured to input the important region block and the position information in the original image of the important region block into the visual transformation model for encoding to generate a bitstream.
[0086] In some embodiments of the present disclosure, the calculation module 902 is further configured to calculate the gradient value of the pixels in each image block, calculate the average gradient value of each image block based on the gradient values of the pixels, sort a plurality of image blocks based on the average gradient value, and determine, as important region blocks, the image blocks among the plurality of image blocks whose average gradient value is greater than or equal to a predetermined value.
[0087] In some embodiments of the present disclosure, the encoding module 903 is further configured to input the important region block into the visual transformation model, output an encoded visible patch and a mask token, generate an image token based on the encoded visible patch, the mask token, and the position information in the original image of the important region block, and generate a bitstream based on the image token.
[0088] In some embodiments of the present disclosure, according to the above technical solution, the acquisition module 901 is further configured to acquire an original image of n×n, where n is a positive integer, block the original image of n×n into m×m equally according to non-overlapping regions, and the size of each image block is
Number
[0089] In some embodiments of the present disclosure, the calculation module is further configured to discard the image blocks among the plurality of image blocks whose average gradient value is smaller than a predetermined value. Here, by setting the predetermined value, the number of image blocks to be discarded and the preset compression ratio α of the image are
[0090]
Number
[0091] According to one aspect of the embodiments of the present disclosure, an image decoding apparatus is provided. The image decoding apparatus includes a receiving module and a decoding module. The receiving module is configured to receive a bitstream generated by being encoded, and the decoding module is configured to decode the bitstream and output a reconstructed image through normalizing the decoding result, a multi-head self-attention mechanism, and a multi-layer perceptron process.
[0092] Specific details of the image encoding apparatus or the image decoding apparatus provided in each embodiment of the present disclosure are described in detail in the corresponding method embodiments that will not be further described herein.
[0093] FIG. 10 schematically shows a computer system configuration block diagram of an electronic device for realizing the embodiments of the present disclosure.
[0094] It should be noted that the computer system 1000 of the electronic device shown in FIG. 10 is an example and does not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0095] As shown in FIG. 10, the computer system 1000 includes a central processing unit 1001 (Central Processing Unit, CPU) that can execute various appropriate operations and processes according to a program stored in a read-only memory 1002 (Read-Only Memory, ROM) or a program loaded from a storage unit 1008 into a random access memory 1003 (Random Access Memory, RAM). The random access memory 1003 also stores various programs and data necessary for the operation of the system. The central processing unit 1001, the read-only memory 1002, and the random access memory 1003 are connected to each other via a bus 1004. An input / output interface 1005 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1004.
[0096] The input / output interface 1005 includes an input unit 1006 including a keyboard, a mouse, etc., an output unit 1007 such as a cathode ray tube (Cathode Ray Tube, CRT), a liquid crystal display (Liquid Crystal Display, LCD), a speaker, etc., a storage unit 1008 including a hard disk, etc., and a communication unit 1009 including a network interface card such as a local network card, a modem, etc. The communication unit 1009 executes communication processing via a network such as the Internet. A driver 1010 is also connected to the input / output interface 1005 as necessary. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in the drive 1010 as necessary, and a computer program that is easy to read from it is installed in the storage unit 1008 as necessary.
[0097] In particular, according to the embodiments of the present disclosure, the processes described in various method flowcharts can be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product including a computer program carried on a computer-readable medium including program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via the communication unit 1009 and / or installed from the removable media 1011. When this computer program is executed by the central processor 1001, various functions limited to the system of the present disclosure are executed.
[0098] Note that the computer-readable medium shown in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any arbitrary combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of any of the above, but is not limited thereto. More specific examples of the computer-readable storage medium include an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above, but is not limited thereto. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program that is instructed to be executed by a system, apparatus, or device, or used in combination therewith. On the other hand, in the present disclosure, the computer-readable signal medium can include a data signal that propagates in a baseband that carries a computer-readable program code or as part of a carrier wave. The data signal propagated in this way can take various forms including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transmit a program used by or in combination with a system, apparatus, or device. The program code included in the computer-readable medium can be transmitted through any suitable medium including, but not limited to, wireless, wired, or any suitable combination of the above.
[0099] Flowcharts and block diagrams in the drawings showing the realizable architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, program segment, or portion of code that includes one or more executable instructions for realizing a given logical function. It should also be noted that in alternative implementations, the functions shown in the blocks may occur in an order different from the order shown in the drawings. For example, two blocks represented consecutively may actually be executed substantially in parallel, or may be executed in the reverse order depending on the related functions. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be realized in a dedicated hardware-based system that executes a given function or operation, or can be realized in a combination of dedicated hardware and computer instructions.
[0100] In addition, in the above detailed description, some modules or units of the device for executing operations were mentioned, but such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied in a plurality of modules or units.
[0101] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described in this specification can be realized by software, and can also be realized by combining the necessary hardware with the software. Therefore, the technical aspects according to the embodiments of the present disclosure may be embodied in the form of a software product stored in a non-volatile storage medium (which may be a CD-ROM, USB disk, mobile hard disk, etc.) that includes some instructions for executing the method according to the embodiments of the present disclosure.
[0102] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, including known common general knowledge or conventional technical means in the art not disclosed in the present disclosure, in accordance with the general principles of the present disclosure.
[0103] It should be understood that the present disclosure is not limited to the exact structures shown in the above-described drawings, and various modifications and changes are possible without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image encoding method, comprising: obtaining an original image, performing a blocking process, and obtaining a plurality of image blocks; calculating a gradient value of pixels in each image block, and selecting important region blocks from the plurality of image blocks based on the gradient value of the pixels; inputting the important region blocks and position information of the important region blocks in the original image into a visual transformation model for encoding to generate a bitstream. An image encoding method characterized by the above.
2. The calculating a gradient value of pixels in each image block and selecting important region blocks from the plurality of image blocks based on the gradient value of the pixels includes: calculating a gradient value of pixels in each image block, and calculating an average gradient value of each image block based on the gradient value of the pixels; sorting the plurality of image blocks based on the average gradient value, and determining, as the important region blocks, image blocks among the plurality of image blocks whose average gradient value is greater than or equal to a predetermined value. The image encoding method according to Claim 1, characterized by the above.
3. The inputting the important region blocks and position information of the important region blocks in the original image into a visual transformation model for encoding to generate a bitstream includes: inputting the important region blocks into the visual transformation model to output encoded visible patches and mask tokens; generating image tokens based on the encoded visible patches, the mask tokens, and the position information of the important region blocks in the original image, and generating the bitstream based on the image tokens. The image encoding method according to Claim 1 or 2, characterized by the above.
4. The obtaining an original image, performing a blocking process, and obtaining a plurality of image blocks includes: obtaining an n×n original image, where n is a positive integer; uniformly blocking the n×n original image into m×m according to non-overlapping regions, and setting the size of each image block to 【Number 1】 , where m is a positive integer and n>m. The image encoding method according to Claim 2 is characterized by the above.
5. The calculating a gradient value of pixels in each image block and selecting important region blocks from the plurality of image blocks based on the gradient value of the pixels further includes: discarding image blocks among the plurality of image blocks whose average gradient value is smaller than a predetermined value. By setting the predetermined value, the number of discarded image blocks and the preset compression ratio α of the image satisfy 【Number 2】 where p is the number of discarded image blocks The image encoding method according to claim 4, characterized in that
6. An image decoding method for decoding the encoding by the image encoding method according to any one of claims 1 to 5, comprising: receiving the bitstream generated by encoding; decoding the bitstream and obtaining a reconstructed image through normalization, a multi-head self-attention mechanism, and a multi-layer perceptron process The image decoding method characterized by that.
7. An image encoding apparatus including an acquisition module, a calculation module, and an encoding module, wherein the acquisition module is configured to acquire an original image, perform a blocking process, and acquire a plurality of image blocks; the calculation module is configured to calculate the gradient value of the pixels in each image block and select important region blocks from the plurality of image blocks based on the gradient value of the pixels; the encoding module is configured to input the important region blocks and the position information of the important region blocks in the original image into a visual conversion model for encoding and generate a bitstream The image encoding apparatus characterized by that.
8. An image decoding apparatus including a reception module and a decoding module, wherein the reception module is configured to receive the bitstream generated by encoding; the decoding module is configured to decode the bitstream and obtain a reconstructed image through normalization, a multi-head self-attention mechanism, and a multi-layer perceptron process The image decoding apparatus characterized by that.
9. A computer-readable medium storing a computer program that, when executed by a processor, implements the image encoding method according to any one of claims 1 to 5 or the image decoding method according to claim 6 The computer-readable medium characterized by that.
10. An electronic device including a processor and a memory storing executable instructions of the processor, wherein the processor is configured to execute the image encoding method according to any one of claims 1 to 5 or the image decoding method according to claim 6 by executing the executable instructions The electronic device characterized by that.
Citation Information
Patent Citations
Data compression circuit
JP1994350992A
Methods for encoding and decoding picture and corresponding devices
JP2015211466A
Encoder, decoder, learning device and program
JP2020022145A
Block-Based Fast Image Compression
US20070201751A1
Methods for encoding and decoding a picture and corresponding devices
US20150312590A1