Image lossy compression method based on depth image prior
Through a deep image prior method, the network parameters of the convolutional neural network are used to represent images, and the problems of low image quality and low versatility under extremely high compression ratio are solved, and the effect of restoring high-quality images at high compression rates is achieved.
Patent Information
- Application Number
- CN202411858236.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The prior art loses more image information under extremely high compression ratios, resulting in lower image quality, and low versatility of image compression algorithms based on deep learning, making it difficult to process image types not included in the training data.
A lossy image compression method based on the depth image prior is proposed. By constructing a convolutional neural network, the network structure is used as the prior information of the image, the network parameters are used as the low-dimensional representation of the image, and iteratively calculates through the gradient descent optimizer to obtain the optimal set of network parameters to realize the lossy compression of the image.
While maintaining universality, high-quality images can be restored at high compression rates, solving the problems of image information loss and low universality.
Smart Images

Figure CN119941878A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an image lossy compression method based on depth image prior. Background Art
[0002] Among image lossy compression technologies, the most widely used are the JPEG series of algorithms. The JPEG algorithm mainly uses a joint coding method of predictive coding (DPCM), discrete cosine transform (DCT) and entropy coding to remove redundant image and color data, and can compress images into a very small storage space, using less disk space to obtain better image quality. However, under extremely high compression ratios (extremely low bpp), the image information is lost more, and the image may have image quality problems such as aliasing, color blocks, and small squares visible to the naked eye caused by the blocking process, resulting in low image quality after decompression.
[0003] With the development of deep learning technology, image compression algorithms based on deep learning have begun to emerge in academia. This type of algorithm generally designs a convolutional neural network with an "encoder-decoder" structure. The first half of the network is the encoder, which can output a set of low-dimensional encoded data after the image is input to achieve image compression, while the second half of the network is the decoder, which can restore the image after the encoded data is input, thereby achieving image decompression. The designed network needs to be trained with a large amount of image data in advance. After training, the network can still maintain high image quality at a very high compression rate when compressing images. However, it is difficult for the training data to cover all types of images, and the compression effect may not be good for image types that are not included; in addition, the compression end and the decompression end need to hold the same trained network parameters at the same time to restore the compressed data normally, making it difficult to publicly distribute the compressed data on a large scale. Therefore, the image compression algorithm based on deep learning has low universality.
[0004] The JPEG series of algorithms are relatively common, but in the case of extremely high compression ratios (extremely low bpp), the image information is lost more, and the image quality restored after decompression is ultimately low.
[0005] Deep learning-based algorithms can maintain high image quality at extremely high compression rates, but the compression effect may not be good for image types not included in the training data. In addition, the need to distribute network parameters in advance makes compressed data difficult to disseminate publicly. These two points make this type of algorithm less versatile. Summary of the invention
[0006] The main purpose of the embodiment of the present invention is to propose a lossy image compression method based on depth image prior, which can maintain a high image quality of the restored image at a high compression rate while maintaining versatility.
[0007] To achieve the above objective, an embodiment of the present invention provides a method for image lossy compression based on depth image prior, comprising the following steps:
[0008] Constructing a convolutional neural network, using the network structure of the convolutional neural network as prior information of the image, and using the network parameters of the convolutional neural network as a low-dimensional representation of the image;
[0009] Through iterative calculations by the gradient descent optimizer, the image fitting process is completed to obtain the optimal set of network parameters;
[0010] According to the optimal network parameter set, the compressed image data is used as the image data, and then after inputting the tensor into the convolutional neural network, the corresponding image is output, thereby realizing lossy compression of the image.
[0011] In some embodiments, after inputting a tensor into the convolutional neural network, the expression for outputting the corresponding image is:
[0012]
[0013] Among them, the tensor input to the convolutional neural network is The output image is H0×W0 and H×W represent the input and output space sizes, respectively, k0 and k out Represent the number of input and output channels respectively, and λ represents all the parameters of the network; if the tensor Z0 is fixed, then f is regarded as a function of λ, with the network structure itself as the prior information of the image, and the network parameter λ as the parameter of the image, so that different images It is represented by the network parameter θ; when the parameter amount of θ is smaller than the image When the spatial domain parameter quantity is , the lossy compression of the image is achieved through this representation.
[0014] In some embodiments, the network structure of the convolutional neural network is composed of a set of progressive upsampling structures, and finally the output layer outputs the reconstructed image;
[0015] Among them, the upsampling layer consists of 1×1 convolution, upsampling operator, linear rectification activation function and channel normalization;
[0016] The upsampling operator is implemented using bilinear interpolation or bicubic interpolation. The output tensor of the i-th upsampling layer is expressed as:
[0017] Zi =cn(relu(U i Z i-1 θ i ))
[0018] Among them, cn(·) represents channel normalization, U i represents the upsampling operator, represents the weight of the 1×1 convolution kernel of the i-th layer, Z i-1 Represents the input tensor of the i-th layer;
[0019] The output layer consists of a 1×1 convolution kernel and a Sigmoid activation function, and its output is expressed as:
[0020]
[0021] Among them, d means that the network uses a d-layer upsampling structure. Represents the weight of the 1×1 convolution kernel of the output layer;
[0022] When the number of channels k of the output tensor of each upsampling layer i =k, the total number of network parameters is determined by the following formula:
[0023] N=dk 2 +2dk+kk out
[0024] Among them, 2dk is the sum of the parameters of the two parameters in all channel normalization operations, k out is the number of channels of the output image. When the image is a grayscale image, k out =1; when the image is a color image, k out =3.
[0025] In some embodiments, for small-sized images, the compression and restoration process of the convolutional neural network includes the following steps:
[0026] In the initial state, the convolutional neural network does not need to be trained, and the network parameters of the convolutional neural network are randomly initialized;
[0027] When implementing image compression, the optimal network parameters that can approximate the original image I are obtained by solving the optimization model Among them, the expression of the optimization model is: in,‖·‖ p is the L1 norm or L2 norm; f(Z0;θ) represents the expression of convolutional neural network;
[0028] The network parameters are iteratively updated by using the gradient descent optimizer in neural network training until the optimal network parameters are obtained.
[0029] Save the optimal network parameters Achieve lossy compression of small-sized images;
[0030] When restoring the image, read the optimal network parameters saved during compression And perform a forward propagation calculation of the network Get the restored image.
[0031] In some embodiments, for large-size images, the compression and restoration process of the convolutional neural network includes the following steps:
[0032] Divide the original large-size image into blocks, and then process each block separately;
[0033] When compressing an image, the block size is defined as a square M×M;
[0034] Record the size of the original image H×W; fill the image to the size H′×W′ by mirroring the right and bottom sides of the original image, so that H′ and W′ are both integer multiples of M;
[0035] Divide the image into several blocks to form a block set B;
[0036] Execute the small-size image compression process for all blocks respectively, and obtain a set Θ consisting of the optimal network parameters obtained for each block;
[0037] Save the set Θ to achieve lossy compression of large-size images;
[0038] When restoring the image, each set of parameters in the network parameter set Θ is used for forward propagation to obtain the image block set B;
[0039] The blocks of the image block set B are arranged in sequence into an image of size H′×W′; and the image is cropped according to the original image size H×W to obtain a restored image.
[0040] In some embodiments, the method further comprises:
[0041] The image compression and restoration phase includes the following steps:
[0042] After completing the image compression calculation process, the network parameters are converted from the original single-precision floating-point type to the half-precision floating-point type and then stored.
[0043] Another aspect of the embodiment of the present invention further provides an image lossy compression system based on depth image prior, comprising:
[0044] The first module is used to construct a convolutional neural network, use the network structure of the convolutional neural network as the prior information of the image, and use the network parameters of the convolutional neural network as the low-dimensional representation of the image;
[0045] The second module is used to perform iterative calculations through a gradient descent optimizer to complete the image fitting process and obtain the optimal network parameter set;
[0046] The third module is used to use the optimal set of network parameters as the compressed data of the image, and then realize lossy compression of the image by inputting a tensor into the convolutional neural network and outputting a corresponding image.
[0047] To achieve the above objective, another aspect of an embodiment of the present invention provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.
[0048] To achieve the above objective, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0049] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.
[0050] The embodiments of the present invention include at least the following beneficial effects: the present invention provides a method for image lossy compression based on deep image priors, which constructs a convolutional neural network, uses the network structure of the convolutional neural network as the prior information of the image, and uses the network parameters of the convolutional neural network as the low-dimensional representation of the image; performs iterative calculations through a gradient descent optimizer to complete the image fitting process and obtain the optimal network parameter set; uses the optimal network parameter set as the compressed image data, and then realizes the lossy compression of the image by inputting a tensor into the convolutional neural network and outputting the corresponding image. The embodiments of the present invention can maintain a high image quality when the image is restored at a high compression rate while maintaining universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;
[0052] Figure 2is a flow chart of the overall steps provided by an embodiment of the present invention;
[0053] Figure 3 Schematic diagram of the network structure of a convolutional neural network provided by an embodiment of the present invention;
[0054] Figure 4 is a small-size image compression flow chart provided by an embodiment of the present invention;
[0055] Figure 5 is a small-size image restoration flow chart provided by an embodiment of the present invention;
[0056] Figure 6 This is a large-size image compression process provided by an embodiment of the present invention;
[0057] Figure 7 is a large-size image restoration flow chart provided by an embodiment of the present invention;
[0058] Figure 8 is an example diagram of a compression result of a small-size image provided by an embodiment of the present invention;
[0059] Fig. 9 is an example diagram of a large-size image segmentation process provided by an embodiment of the present invention;
[0060] Fig.10 is an example diagram of the compression result of a large-size image provided by an embodiment of the present invention;
[0061] Fig.11 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.
[0063] It is understood that the terms "first", "second", etc. used in the present invention may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0064] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0066] Before describing the embodiments of the present invention in detail, some related technologies involved in the embodiments of the present invention are first described as follows:
[0067] Grayscale image: An image with only one sampled color per pixel, usually displayed as a grayscale from the darkest black to the brightest white. The most common grayscale image uses an 8-bit unsigned integer to represent the grayscale value of each pixel, and the grayscale value of each pixel ranges from [0,255].
[0068] Color image: Each color image consists of three color channels: red (R), green (G), and blue (B). Similar to grayscale images, each color channel uses grayscale values to represent the brightness of each color. Common color images use 8-bit unsigned integers to represent grayscale values in each channel, with a value range of [0,255]. Each pixel occupies a total of 24 bits of binary data length.
[0069] Floating point numbers: Computers generally use floating point format to store and process decimals, expressing decimals in scientific notation. In the IEEE754-2019 standard definition, floating point numbers contain three parts: sign bit, exponent bit, and mantissa bit. Depending on the binary length used, common floating point numbers include double-precision floating point numbers (FP64), single-precision floating point numbers (FP32), and half-precision floating point numbers (FP16), which occupy 64 bits, 32 bits, and 16 bits of binary length respectively. The number of binary bits occupied by the three parts are 1:11:52, 1:8:23, and 1:5:10 respectively.
[0070] Bits Per Pixel: Bits Per Pixel (bpp) is a unit that describes the number of binary bits used to represent color information for each pixel in a digital image. For uncompressed images, a 24-bit RGB color image occupies 24bpp, and an 8-bit grayscale image occupies 8bpp. For compressed images, the bpp value is the average number of bits per pixel actually required after compression.
[0071] Lossless compression: Compression uses the statistical redundancy of data to completely restore the original data without causing any distortion.
[0072] Lossy compression: It takes advantage of the fact that humans are insensitive to certain frequency components in images and allows a certain amount of information to be lost during the compression process. Although the original data cannot be completely restored, the lost part has little impact on the understanding of the original image.
[0073] The image lossy compression method based on deep image prior provided by the embodiment of the present invention relates to the field of image processing technology. The image lossy compression method based on deep image prior provided by the embodiment of the present invention can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements an image lossy compression method based on deep image prior, etc., but is not limited to the above forms.
[0074] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0075] like Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected to a network wirelessly or wired to complete data transmission and exchange.
[0076] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0077] In addition, the server 101 can also be a node server in the blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.
[0078] The terminal 102 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. The terminal 102 may also be a vehicle-mounted terminal of various device types described above, but is not limited thereto. The terminal 102 and the server 101 may be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the present invention.
[0079] Based on the example Figure 1In the implementation environment shown, an embodiment of the present invention provides an image lossy compression method based on depth image prior. The following is explained using the image lossy compression method based on depth image prior applied to the server 101 as an example. It can be understood that the method can also be applied to the terminal 102.
[0080] Reference Figure 2 , Figure 2 The flowchart of the image lossy compression method based on depth image prior applied to the server provided by the embodiment of the present invention, the execution subject of the method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 , the method may include the following steps:
[0081] Constructing a convolutional neural network, using the network structure of the convolutional neural network as prior information of the image, and using the network parameters of the convolutional neural network as a low-dimensional representation of the image;
[0082] Through iterative calculations by the gradient descent optimizer, the image fitting process is completed to obtain the optimal set of network parameters;
[0083] According to the optimal network parameter set, the compressed image data is used as the image data, and then after inputting the tensor into the convolutional neural network, the corresponding image is output, thereby realizing lossy compression of the image.
[0084] In some embodiments, after inputting a tensor into the convolutional neural network, the expression for outputting the corresponding image is:
[0085]
[0086] Among them, the tensor input to the convolutional neural network is The output image is H0×W0 and H×W represent the input and output space sizes, respectively, k0 and k out Represent the number of input and output channels respectively, and θ represents all the parameters of the network; if the tensor Z0 is fixed, then f is regarded as a function of θ, with the network structure itself as the prior information of the image, and the network parameter θ as the parameter of the image, so that different images It is represented by the network parameter θ; when the parameter amount of θ is smaller than the image When the spatial domain parameter quantity is , the lossy compression of the image is achieved through this representation.
[0087] In some embodiments, the network structure of the convolutional neural network is composed of a set of progressive upsampling structures, and finally the output layer outputs the reconstructed image;
[0088] Among them, the upsampling layer consists of 1×1 convolution, upsampling operator, linear rectification activation function and channel normalization;
[0089] The upsampling operator is implemented using bilinear interpolation or bicubic interpolation. The output tensor of the i-th upsampling layer is expressed as:
[0090] Z i =cn(relu(U i Z i-1 θ i ))
[0091] Among them, cn(·) represents channel normalization, U i represents the upsampling operator, represents the weight of the 1×1 convolution kernel of the i-th layer, Z i-1 Represents the input tensor of the i-th layer;
[0092] The output layer consists of a 1×1 convolution kernel and a Sigmoid activation function, and its output is expressed as:
[0093]
[0094] Among them, d means that the network uses a d-layer upsampling structure. Represents the weight of the 1×1 convolution kernel of the output layer;
[0095] When the number of channels k of the output tensor of each upsampling layer i =k, the total number of network parameters is determined by the following formula:
[0096] N=dk 2 +2dk+kk out
[0097] Among them, 2dk is the sum of the parameters of the two parameters in all channel normalization operations, k out is the number of channels of the output image. When the image is a grayscale image, k out =1; when the image is a color image, k out =3.
[0098] In some embodiments, for small-sized images, the compression and restoration process of the convolutional neural network includes the following steps:
[0099] In the initial state, the convolutional neural network does not need to be trained, and the network parameters of the convolutional neural network are randomly initialized;
[0100] When implementing image compression, the optimal network parameters that can approximate the original image I are obtained by solving the optimization model Among them, the expression of the optimization model is: in,‖·‖p is the L1 norm or L2 norm; f(Z0;θ) represents the expression of convolutional neural network;
[0101] The network parameters are iteratively updated by using the gradient descent optimizer in neural network training until the optimal network parameters are obtained.
[0102] Save the optimal network parameters Achieve lossy compression of small-sized images;
[0103] When restoring the image, read the optimal network parameters saved during compression And perform a forward propagation calculation of the network Get the restored image.
[0104] In some embodiments, for large-size images, the compression and restoration process of the convolutional neural network includes the following steps:
[0105] Divide the original large-size image into blocks, and then process each block separately;
[0106] When compressing an image, the block size is defined as a square M×M;
[0107] Record the size of the original image H×W; fill the image to the size H′×W′ by mirroring the right and bottom sides of the original image, so that H′ and W′ are both integer multiples of M;
[0108] Divide the image into several blocks to form a block set B;
[0109] Execute the small-size image compression process for all blocks respectively, and obtain a set Θ consisting of the optimal network parameters obtained for each block;
[0110] Save the set Θ to achieve lossy compression of large-size images;
[0111] When restoring the image, each set of parameters in the network parameter set Θ is used for forward propagation to obtain the image block set B;
[0112] The blocks of the image block set B are arranged in sequence into an image of size H′×E′; and the image is cropped according to the original image size H×W to obtain a restored image.
[0113] In some embodiments, the method further comprises:
[0114] The image compression and restoration phase includes the following steps:
[0115] After completing the image compression calculation process, the network parameters are converted from the original single-precision floating-point type to the half-precision floating-point type and then stored.
[0116] The specific implementation process of the embodiment of the present invention is described in detail below with reference to the accompanying drawings of the specification, taking a specific application scenario as an example:
[0117] 1. Low-dimensional representation of images:
[0118] Given a convolutional neural network f, input a tensor to the network And output an image This process can be expressed as
[0119]
[0120] Among them, H0×W0 and H×W represent the input and output space sizes, k0 and k out Represent the number of input and output channels respectively, and θ represents all the parameters of the network. If the tensor Z0 is fixed, f can be regarded as a function of θ, with the network structure itself as the prior information of the image and the network parameter θ as the parameter of the image, so that different images It can be represented by the network parameter θ. When the parameter amount of θ is smaller than the image When the spatial domain parameter quantity is , this representation method can be used to achieve lossy compression of the image.
[0121] 2. Network structure of convolutional neural network:
[0122] The present invention uses a convolutional network structure with insufficient parameters, which is composed of a set of progressive upsampling structures, and finally the reconstructed image is output by the output layer. The network structure is as follows: Figure 3 As shown. The upsampling layer consists of 1×1 convolution, upsampling operator, linear rectified unit (ReLU) and channel normalization (CN). The upsampling operator can use bilinear interpolation or bicubic interpolation. The output tensor of the i-th upsampling layer can be expressed as
[0123] Z i =cn(relu(U i Z i-1 θ i ))#(2)
[0124] Among them, cn(·) represents channel normalization, U i represents the upsampling operator, represents the weight of the 1×1 convolution kernel of the i-th layer, Z i-1represents the input tensor of the i-th layer. The output layer consists of a 1×1 convolution kernel and a Sigmoid activation function, and its output can be expressed as
[0125]
[0126] Among them, d means that the network uses a d-layer upsampling structure. Represents the weight of the 1×1 convolution kernel of the output layer.
[0127] When the number of channels k of the output tensor of each upsampling layer i =k, the total number of network parameters can be determined by the following formula:
[0128] N=dk 2 +2dk+kk out #(4)
[0129] Among them, 2dk is the sum of the parameters of the two parameters in all channel normalization operations, k out is the number of channels of the output image. When the image is a grayscale image, k out =1, for color images, k out =3.
[0130] 3. Compression and recovery of small size images:
[0131] In the initial state, the network does not need to be trained in advance, and the network parameters θ are randomly initialized. When the present invention realizes image compression, the goal is to make the image output by the network The original image I is approximated as much as possible, that is, the original image I is fitted through the network f(Z0;θ). Therefore, it is necessary to solve the optimization model to obtain the optimal network parameters that can approximate the original image I. The optimization model can be expressed as
[0132]
[0133] in,‖·‖ p is the L1 norm (p=1) or L2 norm (p=2). The network parameters are iteratively updated using the gradient descent optimizer commonly used in neural network training until the optimal network parameters are obtained. because The number of parameters of is less than that of the original image, so After saving, the image is lossy compressed. When restoring the image, you only need to read the optimal network parameters saved during compression. And perform a forward propagation calculation of the network The image can be output. The complete compression process is as follows Figure 4 As shown, the recovery process is as follows Figure 5 shown.
[0134] 4. Compression and recovery of large-size images:
[0135] In order to compress large-size images and unify the calculation process and reuse parameter settings when compressing images of different sizes, the original image needs to be divided into blocks first, and then each block is processed separately.
[0136] When compressing an image, the block size is defined as a square M×M. First, record the size of the original image H×W; fill the image to size H′×W′ by mirroring the right and bottom edges of the original image, so that H′ and W′ are both integer multiples of M; divide the image into several blocks to form a block set B; perform the compression process of small-size images on all blocks, and obtain a set Θ composed of the optimal network parameters obtained for each block; finally save Θ. The compression process of large-size images is as follows: Figure 6 shown.
[0137] When restoring the image, use each set of parameters in the network parameter set Θ for forward propagation to obtain an image block set B; arrange the blocks of the image block set B into an image of size H′×W′; crop the original image according to the size H×W to obtain the restored image. The restoration process of large-size images is as follows: Figure 7 shown.
[0138] 5. Storage of compressed data:
[0139] In the image compression and restoration stage, in order to maintain the accuracy of forward propagation and gradient calculation, all parameters are calculated using single-precision floating-point type (fp32, each value occupies 32 bits). However, the original image is generally stored using an 8-bit unsigned integer type (uint8). If the fp32 type is used to store the parameters after compression, it is likely that the space occupied will be larger than the original image, and the compression will lose its meaning. In order to take into account both accuracy and space occupied, after completing the aforementioned image compression calculation process, the network parameters need to be converted from the original single-precision floating-point type to a half-precision floating-point type (fp16, each value occupies 16 bits) before storage. For example, for a 256×256 pixel color image block, when k=32 and d=5 are selected, the space occupied by storing the network parameters is (5×32 2 +2×5×32+3×32)×16=88576 bits, corresponding to a bpp of 85576÷(256×256)≈1.352. It should be noted that since large-size images need to be padded with data at the edges to an integer multiple of the image block size, the padded part is invalid data, so the actual valid data bpp of large-size images may be slightly larger.
[0140] The following uses experimental data as an example to illustrate the technical effect of the method of the embodiment of the present invention:
[0141] In the embodiments of the present invention, all results use the following parameter settings: the upsampling operator uses bicubic interpolation, the input tensor Z0 is randomly generated and fixed, and obeys the uniform distribution U[-0.2, 0.2], the optimization model uses the smooth L1 norm, the gradient optimizer uses the Adam optimizer, the initial value of the learning rate is 1e-2 and gradually decreases to 1e-4, and the number of iterations is fixed to 20,000.
[0142] 1. Small size images:
[0143] This example will show the compression result of a 256×256 pixel color image. The algorithm is tested under the conditions of k=24, k=32 and k=64, and the corresponding bpp are 0.779, 1.352 and 5.203 respectively. The original image and the compressed image results are shown in Figure 2. Figure 8 shown.
[0144] 2. Large size images:
[0145] This example will show the compression result of a 1020×678 pixel image. The block size is set to 256×256. The algorithm is tested under k=24, k=32, and k=64 conditions. The corresponding actual effective bpp are 0.886, 1.537, and 5.917 respectively. Fig. 9 As shown, the original image and the compressed image results are as follows Fig.10 shown.
[0146] In summary, the embodiment of the present invention proposes an image compression algorithm based on deep image priors, which utilizes a part of deep learning technology, constructs an under-parameterized neural network, and uses the network structure as a prior information of an image without prior training of the neural network. The network parameters are used as a low-dimensional representation of the image, and after the image is fitted by iterative calculation through a gradient descent optimizer, the obtained optimal network parameters are used as the compressed image data. Under the premise of maintaining versatility, the present invention can still maintain a high image quality when the image is restored at a high compression rate.
[0147] Another aspect of the embodiment of the present invention further provides an image lossy compression system based on depth image prior, comprising:
[0148] The first module is used to construct a convolutional neural network, use the network structure of the convolutional neural network as the prior information of the image, and use the network parameters of the convolutional neural network as the low-dimensional representation of the image;
[0149] The second module is used to perform iterative calculations through a gradient descent optimizer to complete the image fitting process and obtain the optimal network parameter set;
[0150] The third module is used to use the optimal set of network parameters as the compressed data of the image, and then realize lossy compression of the image by inputting a tensor into the convolutional neural network and outputting a corresponding image.
[0151] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0152] The embodiment of the present invention further provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned image lossy compression method based on depth image prior when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0153] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0154] See also Fig.11 , Fig.11 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0155] The processor 1101 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention;
[0156] The memory 1102 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1102 can store an operating system and other applications. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1102, and the processor 1101 calls and executes the image lossy compression method based on depth image prior according to the embodiment of the present invention;
[0157] Input / output interface 1103, used to implement information input and output;
[0158] The communication interface 1104 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0159] A bus 1105 that transmits information between various components of the device (e.g., the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104);
[0160] The processor 1101 , the memory 1102 , the input / output interface 1103 and the communication interface 1104 are connected to each other in communication within the device via the bus 1105 .
[0161] An embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the image lossy compression method based on depth image prior is implemented.
[0162] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0163] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0164] It should be noted that in various specific embodiments of the present invention, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present invention needs to obtain the user's sensitive personal information, it will obtain the user's separate permission or consent through a pop-up window or jump to a confirmation page, and after clearly obtaining the user's separate permission or consent, it will obtain the necessary user-related data for the normal operation of the embodiment of the present invention.
[0165] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.
[0166] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0167] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0169] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0170] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0171] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0172] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0173] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including multiple instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store programs.
[0175] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present invention is not limited thereby. Any modification, equivalent substitution and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. A method for image lossy compression based on depth image prior, characterized in that: The following steps are involved: Constructing a convolutional neural network, using the network structure of the convolutional neural network as prior information of the image, and using the network parameters of the convolutional neural network as a low-dimensional representation of the image; Through iterative calculations by the gradient descent optimizer, the image fitting process is completed to obtain the optimal set of network parameters; According to the optimal network parameter set, the compressed image data is used as the image data, and then after inputting the tensor into the convolutional neural network, the corresponding image is output, thereby realizing lossy compression of the image.
2. The image lossy compression method based on depth image prior according to claim 1, characterized in that: After inputting a tensor into the convolutional neural network, the expression for outputting the corresponding image is: Among them, the tensor input to the convolutional neural network is The output image is H0×W0 and H×W represent the input and output space sizes, respectively, k0 and k out Represent the number of input and output channels respectively, and θ represents all the parameters of the network; if the tensor Z0 is fixed, then f is regarded as a function of θ, with the network structure itself as the prior information of the image, and the network parameter θ as the parameter of the image, so that different images It is represented by the network parameter θ; when the parameter of θ is smaller than the image When the spatial domain parameter quantity is , the lossy compression of the image is achieved through this representation.
3. The image lossy compression method based on depth image prior according to claim 1, characterized in that: The network structure of the convolutional neural network consists of a set of progressive upsampling structures, and finally the output layer outputs the reconstructed image; Among them, the upsampling layer consists of 1×1 convolution, upsampling operator, linear rectification activation function and channel normalization; The upsampling operator is implemented using bilinear interpolation or bicubic interpolation. The output tensor of the i-th upsampling layer is expressed as: WITH i =cn(relu(U i WITH i-1 θ i )) Among them, cn(·) represents channel normalization, U i represents the upsampling operator, represents the weight of the 1×1 convolution kernel of the i-th layer, Z i-1 Represents the input tensor of the i-th layer; The output layer consists of a 1×1 convolution kernel and a Sigmoid activation function, and its output is expressed as: Among them, d means that the network uses a d-layer upsampling structure. Represents the weight of the 1×1 convolution kernel of the output layer: When the number of channels k of the output tensor of each upsampling layer i =k, the total number of network parameters is determined by the following formula: N=dk 2 +2dk+kk out Among them, 2dk is the sum of the parameters of the two parameters in all channel normalization operations, k out is the number of channels of the output image. When the image is a grayscale image, k out =1; when the image is a color image, k out =3.
4. The image lossy compression method based on depth image prior according to claim 1, characterized in that: For small-sized images, the compression and restoration process of the convolutional neural network includes the following steps: In the initial state, the convolutional neural network does not need to be trained, and the network parameters of the convolutional neural network are randomly initialized; When implementing image compression, the optimal network parameters that can approximate the original image I are obtained by solving the optimization model Among them, the expression of the optimization model is: Among them, ||·|| p is the L1 norm or L2 norm; f(Z0;θ) represents the expression of convolutional neural network; The network parameters are iteratively updated by using the gradient descent optimizer in neural network training until the optimal network parameters are obtained. Save the optimal network parameters Achieve lossy compression of small-sized images; When restoring the image, read the optimal network parameters saved during compression And perform a forward propagation calculation of the network Get the restored image.
5. The image lossy compression method based on depth image prior according to claim 1, characterized in that: For large-size images, the compression and recovery process of the convolutional neural network includes the following steps: Divide the original large-size image into blocks, and then process each block separately; When compressing an image, the block size is defined as a square M×M; Record the size of the original image H×W; fill the image to the size H'×W' by mirroring the right and bottom sides of the original image, so that H' and W' are both integer multiples of M; Divide the image into several blocks to form a block set B; Execute the small-size image compression process for all blocks respectively, and obtain a set Θ consisting of the optimal network parameters obtained for each block; Save the set Θ to achieve lossy compression of large-size images; When restoring the image, each set of parameters in the network parameter set Θ is used for forward propagation to obtain the image block set B; The blocks of the image block set B are arranged in sequence into an image of size H′×W′; and the image is cropped according to the original image size H×W to obtain a restored image.
6. The image lossy compression method based on depth image prior according to claim 1, characterized in that: The method further comprises: The image compression and restoration phase includes the following steps: After completing the image compression calculation process, the network parameters are converted from the original single-precision floating-point type to the half-precision floating-point type and then stored.
7. A lossy image compression system based on deep image prior, characterized in that: include: The first module is used to construct a convolutional neural network, use the network structure of the convolutional neural network as the prior information of the image, and use the network parameters of the convolutional neural network as the low-dimensional representation of the image; The second module is used to perform iterative calculations through a gradient descent optimizer to complete the image fitting process and obtain the optimal network parameter set; The third module is used to use the optimal set of network parameters as the compressed data of the image, and then realize lossy compression of the image by inputting a tensor into the convolutional neural network and outputting a corresponding image.
8. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
SAR image compression method based on convolutional neural network
CN111681293A
Image lossless / near lossless compression method based on deep learning
CN114359422A
Spectral calculation imaging method and device combining depth prior and learnable imaging model
CN115795221A
Self-supervised multi-focus image fusion method and model construction method thereof
CN118537238A