Anti-interference image compression encoding method, decoding method and device
Through deep learning wavelet transformation and entropy coding technology, the problem of image compression encoding in satellite communications is solved, and the anti-interference progressive compression and efficient decoding are realized, which improves the stability and decoding quality of image transmission.
Patent Information
- Application Number
- CN202510469563.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-11
AI Technical Summary
The existing image compression encoding algorithms are susceptible to data loss in satellite communication, resulting in decoding failure and unable to effectively deal with signal instability in satellite communication.
The image is processed by deep learning wavelet transformation, decomposed into multiple subbands, and entropy encoding is performed in the encoding order, combining context modeling and interleaved entropy encoding, and target code stream is generated, and decoding is achieved through deep learning wavelet inverse transformation and image enhancement processing.
It realizes progressive compression encoding with anti-interference capability in satellite communication, reduces algorithm complexity and encoding time, and improves the stability and decoding quality of image transmission.
Smart Images

Figure CN120302060A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to an anti-interference image compression and encoding method, a decoding method, and a device. Background Art
[0002] In the field of satellite communication, the communication bandwidth between communication devices such as mobile phones and satellites is extremely limited. Under such conditions, when transmitting images, the images can only be compressed and divided into several data packets for transmission.
[0003] In the prior art, deep learning image compression algorithms, such as DPICT, CTC, ProgDTD, etc., can be used to compress and encode images. However, the above algorithms use progressive encoding methods and are not applicable in the field of satellite communication. The JPEG2000 compression algorithm can also perform progressive encoding on images. However, if packet loss occurs during transmission, the decoded image will be severely interfered. Summary of the Invention
[0004] The main purpose of this application is to provide an anti-interference image compression and encoding method, a decoding method, and a device, aiming to solve the technical problem that the existing image compression and decoding algorithms cannot be decoded due to data loss during satellite data transmission.
[0005] To achieve the above purpose, this application provides an anti-interference image compression and encoding method, including:
[0006] Input an image;
[0007] Perform deep learning wavelet transform processing on the image to obtain multiple subbands;
[0008] Determine the encoding order corresponding to each coding segment included in each subband;
[0009] Perform entropy encoding on each coding segment according to the encoding order to obtain a target bitstream.
[0010] Optionally, the performing deep learning wavelet transform processing on the image to obtain multiple subbands includes:
[0011] Split the image in the row direction to obtain a high-frequency subband and a low-frequency subband;
[0012] Split the low-frequency subband in the column direction to obtain a first subband and a second subband; the first subband represents the low-frequency part in the low-frequency subband, and the second subband represents the high-frequency part in the low-frequency subband;
[0013] Split the high-frequency subband in the column direction to obtain a third subband and a fourth subband; the third subband represents the low-frequency part in the high-frequency subband, and the fourth subband represents the high-frequency part in the high-frequency subband.
[0014] Optionally, determining the coding order corresponding to each coded segment included in each sub-band includes:
[0015] Performing quantization processing on each of the sub-bands;
[0016] Obtaining the decomposition level, sub-band type, and the bit plane to which each sub-band after quantization processing belongs;
[0017] Determining the coding order corresponding to each coded segment included in each sub-band according to the decomposition level, the sub-band type, and the bit plane to which it belongs.
[0018] Optionally, performing entropy coding on each of the coded segments to obtain a target bitstream includes:
[0019] Performing context modeling on each of the coded segments to determine the context information corresponding to each pixel in each of the coded segments;
[0020] Performing probability estimation based on the context information corresponding to each pixel to obtain an estimated value corresponding to each pixel;
[0021] Performing interleaved entropy coding on each of the coded segments according to the estimated value corresponding to each pixel to obtain a target bitstream.
[0022] Optionally, performing context modeling on each of the coded segments to determine the context information corresponding to each pixel in each of the coded segments includes:
[0023] For each pixel, determining the pixel category corresponding to the pixel based on the coding situation of adjacent pixels;
[0024] Determining the context information corresponding to the pixel according to the pixel category corresponding to the pixel and the number of significant pixels among the adjacent pixels.
[0025] Optionally, performing interleaved entropy coding on each of the coded segments according to the estimated value corresponding to each pixel includes:
[0026] For each pixel, determining a probability threshold value associated with the estimated value corresponding to the pixel;
[0027] Determining the coding method corresponding to the pixel based on the probability threshold value;
[0028] Performing interleaved entropy coding on the pixel by means of the coding method.
[0029] In addition, to achieve the above object, the present application provides an anti-interference image compression and decoding method, including:
[0030] Inputting a target bitstream;
[0031] Entropy decode the target bitstream according to the preset decoding order to obtain multiple subbands;
[0032] Perform an inverse deep learning wavelet transform on the multiple subbands to obtain an image;
[0033] Perform image enhancement processing on the image to obtain a target image.
[0034] Optionally, the performing image enhancement processing on the image to obtain a target image includes:
[0035] Extract features from the image through a primary feature extraction layer to obtain initial features;
[0036] Obtain multi-scale features through a deep residual nested module based on the initial features;
[0037] Scale the feature map through an upsampling module, where the feature map is generated based on the multi-scale features;
[0038] Obtain a target image through a reconstruction module based on the feature map.
[0039] In addition, to achieve the above object, the present application also provides an anti-interference image compression and encoding device, and the anti-interference image compression and encoding device includes:
[0040] A first input module for inputting an image;
[0041] A transformation module for performing a deep learning wavelet transform on the image to obtain multiple subbands;
[0042] A determination module for determining the encoding order corresponding to each encoding segment included in each subband;
[0043] An encoding module for performing entropy encoding on each encoding segment according to the encoding order to obtain a target bitstream.
[0044] In addition, to achieve the above object, the present application also provides an anti-interference image compression and decoding device, and the anti-interference image compression and decoding device includes:
[0045] A second input module for inputting a target bitstream;
[0046] A decoding module for performing entropy decoding on the target bitstream according to the preset decoding order to obtain multiple subbands;
[0047] An inverse transformation module for performing an inverse deep learning wavelet transform on the multiple subbands to obtain an image;
[0048] An enhancement module for performing image enhancement processing on the image to obtain a target image.
[0049] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solutions:
[0050] The computer device includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps of any one of the anti-interference image compression encoding method and the anti-interference image compression decoding method proposed in the embodiments of the present application are implemented.
[0051] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solutions:
[0052] A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of any one of the anti-interference image compression encoding method and the anti-interference image compression decoding method proposed in the embodiments of the present application are implemented.
[0053] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0054] The embodiments of the present application provide an anti-interference image compression encoding method, a decoding method and a device. The anti-interference image compression encoding method includes: inputting an image; performing deep learning wavelet transform processing on the image to obtain a plurality of subbands; determining the encoding order corresponding to each encoding segment included in each subband; and performing entropy encoding on each encoding segment according to the encoding order to obtain a target bitstream. The encoding method provided by the embodiments of the present application adopts an anti-interference entropy encoding algorithm, making the algorithm have a progressive compression effect and anti-interference ability, and better coping with the problem of signal instability in satellite communication. In addition, in the embodiments of the present application, only by performing deep learning wavelet transform processing on the image and performing entropy encoding according to the encoding order, the compression encoding of the image can be realized, reducing the algorithm complexity and thus reducing the encoding duration. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 is a flowchart of the anti-interference image compression encoding method provided by the embodiments of the present application;
[0057] Figure 2 is a schematic flowchart of the deep learning wavelet transform processing provided by the embodiments of the present application;
[0058] Figure 3 It is a schematic diagram of pixel saliency classification provided by an embodiment of the present application;
[0059] Figure 4 It is a schematic flowchart of interleaved entropy coding provided by an embodiment of the present application;
[0060] Figure 5 It is a flowchart of an anti-interference image compression and decoding method provided by an embodiment of the present application;
[0061] Figure 6 It is a schematic flowchart of interleaved entropy decoding provided by an embodiment of the present application;
[0062] Figure 7 It is a schematic flowchart of deep learning inverse wavelet transform processing provided by an embodiment of the present application;
[0063] Figure 8 It is a schematic structural diagram of a deep residual nesting module provided by an embodiment of the present application;
[0064] Figure 9 It is a schematic structural diagram of an RCAB module provided by an embodiment of the present application;
[0065] Figure 10 It is a schematic structural diagram of an anti-interference image compression and encoding device provided by an embodiment of the present application;
[0066] Figure 11 It is a schematic structural diagram of an anti-interference image compression and decoding device provided by an embodiment of the present application.
[0067] Figure 12 It is a basic structural block diagram of a computer device according to an embodiment of the present application. Detailed implementation manners
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0069] References to "embodiments" in this specification mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0070] To enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0071] Please refer to Figure 1 , Figure 1 which is a flowchart of an anti-interference image compression and encoding method provided by an embodiment of the present application. It should be noted that the anti-interference image compression and encoding method provided by the embodiment of the present application can be applied to the application scenario of mobile terminal and satellite communication, and the mobile terminal realizes image compression encoding and decoding.
[0072] As Figure 1 shown, the anti-interference image compression and encoding method provided by the embodiment of the present application includes:
[0073] S110, input an image.
[0074] S120, perform deep learning wavelet transform processing on the image to obtain multiple subbands.
[0075] In this step, the encoding end receives the image and performs deep learning wavelet transform processing on the image to obtain multiple subbands.
[0076] S130, determine the encoding order corresponding to each coding segment included in each subband.
[0077] In this step, after obtaining multiple subbands, the pixel values of each subband are decomposed into multiple bit planes, that is, bit plane encoding is performed on each coding segment included in each subband to determine the encoding order corresponding to each coding segment. For specific implementation manners, please refer to the subsequent embodiments.
[0078] S140, perform entropy encoding on each coding segment according to the encoding order to obtain a target bitstream.
[0079] In this step, after determining the encoding order corresponding to each coding segment in each subband, entropy encoding is performed on each coding segment according to the encoding order to obtain a target bitstream.
[0080] Among them, the above-mentioned entropy encoding of each coding segment includes context modeling and interleaved entropy encoding of the coding segment. For specific implementation manners, please refer to the subsequent embodiments.
[0081] The encoding method provided by the embodiments of the present application adopts an anti-interference entropy encoding algorithm, enabling the algorithm to have a progressive compression effect and anti-interference ability, and better coping with the problem of signal instability in satellite communication. In addition, in the embodiments of the present application, only by performing deep learning wavelet transform processing on the image and performing entropy encoding in the encoding order, the compression encoding of the image can be achieved, reducing the algorithm complexity and thus reducing the encoding duration.
[0082] Optionally, the performing deep learning wavelet transform processing on the image to obtain multiple subbands includes:
[0083] Splitting the image in the row direction to obtain a high-frequency subband and a low-frequency subband;
[0084] Splitting the low-frequency subband in the column direction to obtain a first subband and a second subband; the first subband represents the low-frequency part in the low-frequency subband, and the second subband represents the high-frequency part in the low-frequency subband;
[0085] Splitting the high-frequency subband in the column direction to obtain a third subband and a fourth subband; the third subband represents the low-frequency part in the high-frequency subband, and the fourth subband represents the high-frequency part in the high-frequency subband.
[0086] In this embodiment, a predictor and an updater can be used to perform deep learning wavelet transform processing on the image to obtain multiple subbands.
[0087] In the convolutional layer of the predictor, a conventional convolutional layer is only applied to a part of the input channels for spatial feature extraction, and the remaining channels are kept unchanged. Keeping the remaining channels unchanged instead of deleting them from the feature map. Without loss of generality, it is considered that the input and output feature maps have the same number of channels.
[0088] Specifically, please refer to Figure 2 , as Figure 2 shown, splitting the input image in the row direction to obtain an even part and an odd part, where the above even part can be understood as the low-frequency subband, and the above odd part can be understood as the high-frequency subband. Performing N-order lifting on the low-frequency subband, and then splitting it in the column direction to obtain the split even part and odd part. Performing N-order lifting on the split even part to obtain the first subband ( Figure 2 "LL" shown in Figure 2 ), and performing N-order lifting on the split odd part to obtain the second subband ( "HL" shown in
[0089] ), that is, the first subband represents the low-frequency part in the low-frequency subband, and the second subband represents the high-frequency part in the low-frequency subband.Perform an N - th order lifting on the high - frequency sub - band, then split it in the column direction to obtain the split even part and odd part. Perform an N - th order lifting on the split even part to obtain the third sub - band ( Figure 2 "LH" shown in Figure 2 ). Perform an N - th order lifting on the split odd part to obtain the fourth sub - band (
[0090] "HH" shown in ). That is, the third sub - band represents the low - frequency part in the high - frequency sub - band, and the fourth sub - band represents the high - frequency part in the high - frequency sub - band.
[0091] In other embodiments, the image can also be divided multiple times to obtain multiple sub - bands, not limited to obtaining 4 sub - bands in this embodiment.
[0092] Optionally, determining the coding order corresponding to each coding segment included in each sub - band includes:
[0093] Perform quantization processing on each sub - band;
[0094] Obtain the decomposition level, sub - band type, and the bit - plane to which each sub - band corresponds after quantization processing;
[0095] Determine the coding order corresponding to each coding segment included in each sub - band according to the decomposition level, the sub - band type, and the bit - plane to which it belongs.
[0096] In this embodiment, a quantizer can be used to adapt to the statistical characteristics of wavelet coefficients to perform quantization processing on sub - bands.
[0097] Furthermore, obtain the decomposition level, sub - band type, and the bit - plane to which each sub - band corresponds. The above - mentioned decomposition level refers to the number of stages of wavelet transform performed on the image. The above - mentioned sub - band types include LL, HL, LH, and HH. The above - mentioned bit - planes include the high - bit plane and the low - bit plane.
[0098] Based on the above - mentioned decomposition level, sub - band type, and bit - plane, determine the coding order corresponding to each coding segment included in each sub - band. An optional implementation method is that the coding order of coding segments corresponding to a higher decomposition level precedes that of coding segments corresponding to a lower decomposition level. The coding order of coding segments corresponding to a higher bit - plane precedes that of coding segments corresponding to a lower bit - plane. The coding order of coding segments corresponding to the sub - band type LL precedes that of coding segments corresponding to the sub - band type HL; the coding order of coding segments corresponding to the sub - band type HL precedes that of coding segments corresponding to the sub - band type LH; the coding order of coding segments corresponding to the sub - band type LH precedes that of coding segments corresponding to the sub - band type HH.
[0099] For coded segments belonging to the same subband, the coding order corresponding to the coded segment with a higher bit plane is limited to that with a lower bit plane. If two coded segments belong to the same bit plane, the coding order is determined based on the decomposition levels corresponding to the two coded segments. If the decomposition levels corresponding to the two coded segments are the same, it is determined based on the subband types corresponding to the two coded segments.
[0100] In this embodiment, the coding process has the characteristic of gradual refinement. For each additional bit plane, the quantization step size is halved, and the reconstruction quality is improved accordingly. This design supports progressive compression, enabling the image to be coded in the order of importance, thus supporting progressive transmission and decoding. In terms of distortion control, the quality of the reconstructed image is controlled by using subband weights, while considering the influence of the non - orthogonality of the wavelet transform.
[0101] Optionally, the entropy coding of each coded segment to obtain the target bitstream includes:
[0102] Performing context modeling on each coded segment to determine the context information corresponding to each pixel in each coded segment;
[0103] Performing probability estimation based on the context information corresponding to each pixel to obtain the estimated value corresponding to each pixel;
[0104] According to the estimated value corresponding to each pixel, performing interleaved entropy coding on each coded segment to obtain the target bitstream.
[0105] The concept of context refers to considering the already - coded information around a bit when coding that bit, in order to better predict the value of the current bit. By using this context information, the coding accuracy and compression efficiency can be improved. In bit - plane coding, the context information usually comes from the already - coded bits of the current pixel and its neighboring pixels. The context of a pixel is determined by the already - coded bits of the pixel and its eight nearest neighbors in the same subband segment. By assigning a category to each pixel to summarize the information of the already - coded bits of the pixel, the context information can be effectively captured.
[0106] In this embodiment, performing context modeling on each coded segment to determine the context information corresponding to each pixel in each coded segment. For the specific implementation method, please refer to the subsequent embodiments.
[0107] After obtaining the context information corresponding to each pixel, an estimator p i , p i is generated by the context modeler, where p represents the probability that the pixel bit is equal to zero.
[0108] Optionally, each context information maintains two counts: the number of zero bits and the total number of bits. The ratio of these counts represents the probability that the pixel bit to which the context information belongs is equal to zero. Each context is set with initial values. The initial value of the number of zero bits is 2, and the initial value of the total number of bits is 4, corresponding to a zero probability of 1 / 2. For each bit encountered, the total count of the context increases. If the bit is 0, the zero-bit count also increases. When the total count reaches a specified value, both counts are rescaled by dividing by 2. Optionally, a threshold can be set to 500. In addition, during the encoding process, the context counter is updated in real time, and the effective dynamic range of the counter is maintained through the scaling mechanism to ensure that the probability estimation has both memory and can adapt to changes in local statistical characteristics, achieving a balance between compression efficiency and computational complexity and providing accurate probability input for subsequent arithmetic coding.
[0109] After obtaining the estimated value corresponding to each pixel, each coding segment is subjected to interleaved entropy coding to obtain the target bitstream. Among them, interleaved entropy coding dynamically resolves the input data stream into a combination of multiple variable-length codewords and uses the characteristics of each component code to achieve efficient compression. Its core features include: both the input codeword and the output codeword are of variable length; different component codes are optimized for data segments with specific statistical characteristics; and prefix freedom and exhaustiveness are ensured to make the encoding and decoding unambiguous.
[0110] Optionally, context modeling is performed on each of the coding segments to determine the context information corresponding to each pixel in each coding segment, including:
[0111] For each pixel, based on the coding situation of adjacent pixels, determine the pixel category corresponding to the pixel;
[0112] According to the pixel category corresponding to the pixel and the number of significant pixels among the adjacent pixels, determine the context information corresponding to the pixel.
[0113] In this embodiment, based on the coding situation of adjacent pixels, the pixel category corresponding to each pixel can be determined.
[0114] Please refer to Figure 3 , as Figure 3 shown, the categories of pixels are defined as follows based on the significance of pixels:
[0115] Category 0 (not significant): All the encoded numerical bits of the pixel are 0.
[0116] Category 1 (just significant): The first "1" bit of the pixel has been encoded.
[0117] Category 2 (partially significant): Two numerical bits of the pixel have been encoded.
[0118] Category 3 (fully significant): Three or more numerical bits of the pixel have been encoded.
[0119] That is to say, if all the encoded value bits of a certain pixel are 0, for example Figure 3 the first to the third columns shown, then it is determined that the pixel is not significant, and the corresponding pixel category is 0. If the first "1" bit of a certain pixel has been encoded, for example Figure 3 the fourth column shown, then it is determined that the pixel is just significant, and the corresponding pixel category is 1. If two value bits of a certain pixel have been encoded, for example Figure 3 the fifth column shown, then it is determined that the pixel is partially significant, and the corresponding pixel category is 2. If three or more value bits of a certain pixel have been encoded, for example Figure 3 the sixth to the eighth columns shown, then it is determined that the pixel is fully significant, and the corresponding pixel category is 3.
[0120] After determining the pixel category corresponding to the pixel, according to the pixel category corresponding to the pixel and the number of significant pixels among adjacent pixels, the context information corresponding to the pixel is determined.
[0121] Optionally, there are 17 types of context information. Among them, contexts 0 - 8 are for the bits of category 0 pixels; contexts 9 - 10 are for the bits of category 1 pixels; context 11 is for the bits of category 2 pixels; contexts 12 - 16 are for sign bits; since the bits of category 3 are almost incompressible, they are not encoded.
[0122] In this embodiment, for the pixel bits of category 0, if the subband is not the HL subband, the context is determined according to the combination of h, v, and d, where h is the number of significant pixels among horizontally adjacent pixels, v is the number of significant pixels among vertically adjacent pixels, and d is the number of significant pixels among diagonal adjacent pixels. As shown in Table 1.
[0123] Table 1:
[0124]
[0125] For example, as shown in Table 1, if a pixel has a corresponding pixel level of 0 and this pixel does not belong to the HL subband, the number of significant pixels among horizontally adjacent pixels is 0, the number of significant pixels among vertically adjacent pixels is 0, and the number of significant pixels among diagonal adjacent pixels is 0, then it is determined that the context corresponding to this pixel is 0.
[0126] For the pixel bits of category 0 and the pixels belonging to the HL subband, the roles of h and v are interchanged. The context is determined according to the combination of h, v, and d, as shown in Table 2 below.
[0127] Table 2:
[0128]
[0129] For example, as shown in Table 2, if the pixel level corresponding to a pixel is 0, and the pixel belongs to the HL sub-band, the number of significant pixels among the diagonal neighboring pixels is 0, and the number of significant pixels among the horizontal neighboring pixels is 2, then the context corresponding to this pixel is determined to be 2.
[0130] For the pixel bits of category 1, if there are no significant pixels among the horizontal and vertical neighboring pixels, the context is 9, otherwise it is 10.
[0131] For the pixel bits of category 2, the context is 11.
[0132] In this embodiment, for the sign bit, the sign bit prediction and the determination of the sign bit context can be performed through two horizontal neighboring pixels and two vertical neighboring pixels. If the sub-band to which the sign bit belongs is not the HL sub-band, let h1 and h2 represent the signs and significances of two horizontal neighboring pixels, taking values of 1 (positive), -1 (negative), and 0 (not significant). Similarly, let v1 and v2 represent the signs and significances of two vertical neighboring pixels. For the HL sub-band, the roles of h and v are interchanged. The sign bit prediction and context modeling are shown in Table 3.
[0133] Table 3:
[0134] <![CDATA[v1+v2]]> <![CDATA[h1 + h2 < 0]]> <![CDATA[h1 + h2 = 0]]> <![CDATA[h1+h1>0]]> <![CDATA[v1 + v2 < 0]]> -,16 +,13 +,14 <![CDATA[v1 + v2 = 0]]> -,15 +,12 +,15 <![CDATA[v1 + v2 > 0]]> -,14 -,13 +,16
[0135] Optionally, the performing of the staggered entropy coding on each coding segment according to the estimated value corresponding to each pixel includes:
[0136] For each pixel, determining a probability threshold value associated with the estimated value corresponding to the pixel;
[0137] Based on the probability threshold value, determining the coding method corresponding to the pixel;
[0138] Performing staggered entropy coding on the pixel through the coding method.
[0139] In this embodiment, 17 coding methods are set, as well as 17 probability threshold values corresponding to these 17 coding methods one by one. After obtaining the estimated value corresponding to each pixel, determining the probability threshold value associated with the estimated value corresponding to the pixel. The above-mentioned association relationship can be understood as a proximity relationship. That is to say, for a pixel, the probability threshold value among the 17 probability threshold values that is close to the estimated value corresponding to the pixel is determined as the probability threshold value associated with the pixel. Then, staggered entropy coding is performed on the pixel through the coding method corresponding to the probability threshold value.
[0140] Specifically, please refer to Table 4:
[0141]
[0142]
[0143] As shown in Table 4, for the first 8 sub - intervals, a shorthand notation is used to select the appropriate component source code; for sub - intervals 9 to 17, for the bits with higher probability estimates, Golomb coding is used for compression.
[0144] To facilitate the understanding of the technical solution of the interleaved entropy coding provided in this embodiment, please refer to Figure 4 . Whenever the source bit bi arrives, it is assigned to the corresponding sub - interval according to its probability estimate pi. If there is already a partial input codeword in the current sub - interval, the source bit bi is added to this partial codeword. If there is no partial codeword in the sub - interval, the new bit bi starts a new codeword and is added to the end of the encoder's list. If the front - most codeword in the list is complete, that is, it is a complete input codeword, the encoder generates the corresponding output bit and removes this codeword from the list. In this way, space in the list is freed up for new codewords. If the list is full, to ensure the smooth progress of decoding, "flushbits" are added to the front - most input codeword to make it a complete codeword. When the input bit sequence ends, all remaining partial input codewords are also completed through "flushbits".
[0145] Please refer to Figure 5 , Figure 5 which is the flowchart of the anti - interference image compression decoding method provided by the embodiment of the present application. As Figure 5 shown, the anti - interference image compression decoding method provided by the embodiment of the present application includes:
[0146] S210, input the target code stream.
[0147] S220, perform entropy decoding on the target code stream according to the preset decoding order to obtain multiple sub - bands.
[0148] In this embodiment, the target code stream is input, and entropy decoding is performed on the target code stream according to the preset decoding order. Optionally, the above - mentioned entropy decoding method is interleaved entropy decoding. Among them, the definition of the decoding order is the same as the definition of the above - mentioned encoding order, and will not be repeated here.
[0149] To facilitate the understanding of the technical solution of the interleaved entropy decoding provided in this embodiment, please refer to Figure 6 .
[0150] The decoder determines the sub-interval to which each source bit \(b_i\) belongs based on the probability estimate \(p_i\) (using the same method as the encoder) and decodes the source bits in sequence. Whenever a bit \(b_i\) is decoded, the decoder checks whether there is a partial codeword in the corresponding sub-interval. If there is, it removes the leading bit from the partial codeword and uses it as the decoded bit \(b_i\). If there is no corresponding partial codeword, the decoder needs to parse the complete output codeword from the encoded bit stream and convert it back to the corresponding input codeword. The decoder first decodes the first bit and remembers the remaining codeword suffix. The decoder needs to ensure that it does not misdecode "flushbits" as source bits. To this end, the decoder records the number of each reconstructed codeword and compares it with the encoder buffer size. If, when the decoder recovers the codeword, the difference between the number of reconstructed codewords and the number of corresponding codewords in the cache exceeds the buffer size, the remaining bits are determined to be "flushbits" and discarded.
[0151] S230, perform an inverse deep learning wavelet transform on the multiple sub-bands to obtain an image.
[0152] In this step, the inverse deep learning wavelet transform of the multiple sub-bands can be implemented through an updater and a predictor.
[0153] Please refer to Figure 7 , as Figure 7 shown, the high-frequency information in the sub-band is processed by the updater, and an even part is obtained based on the high-frequency information processed by the updater and the low-frequency information in the sub-band; the low-frequency information in the sub-band is processed by the predictor, and an odd part is obtained based on the low-frequency information processed by the predictor and the high-frequency information in the sub-band; and then the above even part and odd part are combined to obtain an image.
[0154] S240, perform image enhancement processing on the image to obtain a target image.
[0155] In this step, after the image is obtained, image enhancement processing is performed to improve the image quality and obtain a target image. For specific implementation manners, please refer to the subsequent embodiments.
[0156] Optionally, the performing image enhancement processing on the image to obtain a target image includes:
[0157] Performing feature extraction on the image through a primary feature extraction layer to obtain initial features;
[0158] Obtaining multi-scale features through a deep residual nested module based on the initial features;
[0159] Performing scaling processing on the feature map through an upsampling module, where the feature map is generated based on the multi-scale features;
[0160] The target image is obtained by the reconstruction module based on the feature map.
[0161] In this embodiment, the primary feature extraction layer performs an initial feature mapping on the input image using a single-layer convolution operation, and its mathematical representation is:
[0162] F0 = H SF (I LR )
[0163] where I LR is the input low-quality image, H SF is the shallow feature extraction function, and F0 is the extracted initial feature.
[0164] The deep residual nested module extracts multi-scale features of the image through multi-level residual learning and channel attention mechanism. The upsampling module uses sub-pixel convolution to enlarge the size of the feature map to the target size, and finally the reconstruction module generates the final enhanced image, and its mathematical representation is:
[0165] I SR = H REC (H UP (F DF )) = H RCAN-ECA (I LR )
[0166] where, I SR is the output image, I LR is the input image, H UP represents upsampling, H REC respectively represent the reconstruction function, and H RCAN-ECA represents the mapping function of the entire network.
[0167] Please refer to Figure 8 , Figure 8 which is the structural schematic diagram of the deep residual nested module provided by the embodiment of the present application. As Figure 8 shown, the deep residual nested module is composed of G cascaded residual groups (Residual Group, RG), and realizes cross-level feature fusion by introducing a long skip connection (Long Skip Connection, LSC). The inside of the residual group adopts a modular design, and B improved residual channel attention blocks (Residual Channel Attention Block with Efficient Channel Attention, RCAB-ECA) are integrated in the basic operation unit. This structured nested structure realizes the step-by-step extraction of multi-scale features of the image through a step-by-step feature refinement mechanism. Specifically, the feature evolution process of the g-th residual group can be expressed as:
[0168] F g= H g (F g-1 ) = H g (H g-1 (…H1(F0)…))
[0169] where H g represents the G-th RG function, F g-1 and F g represent the input and output of the g-th RG, and H g-1 represents the (G - 1)-th RG function, and F0 represents the initial input.
[0170] The long-range skip connections in the RIR structure enable the direct transmission of low-frequency information, allowing the backbone network to focus on learning high-frequency detailed information. At the same time, the short-range skip connections within each RG also facilitate the backpropagation of gradients, effectively alleviating the training difficulty of deep networks. The design of this residual nested structure is based on the following considerations: Image super-resolution can be regarded as a processing process, that is, it is necessary to recover as much high-frequency information as possible. The input image contains the most low-frequency information, which can be directly transmitted to the final output image without much calculation. Through the design of LSC and SSC, a large amount of low-frequency information can be directly transmitted through skip connections, enabling the main network to focus on learning more valuable high-frequency information.
[0171] Please refer to Figure 9 , Figure 9 which is the structural schematic diagram of the RCAB module provided by the embodiment of the present application.
[0172] In this embodiment, the ECA attention mechanism is used to replace and optimize the original attention mechanism. The traditional SE (Squeeze-and-Excitation) attention mechanism module compresses and restores the channel dimension through a fully connected layer. This operation of dimension transformation may introduce high computational complexity and may also cause loss of feature information. The improved RCAB module with the ECA mechanism uses a one-dimensional convolution kernel to achieve cross-channel interaction and directly models the inter-channel dependence relationship by adaptively selecting the local receptive field, optimizing the channel attention modeling while avoiding the original dimension transformation operation of the SE mechanism. The calculation process of the RCAB-ECA module is simplified in this improvement scheme as follows: After the feature map is globally averaged and pooled, channel feature interaction is performed through a one-dimensional convolution kernel with parameter k, and finally, the channel attention weight is generated through the Sigmoid function. This structural optimization significantly improves the parameter efficiency of the model, saving computational resources while maintaining the lightweight characteristics of the module.
[0173] Compared with the global compression excitation method adopted by the traditional SE module, the ECA module adopts non-dimensionality reduction processing, effectively avoiding the loss of feature information during the channel compression process. At the same time, the ECA module realizes the dynamic calibration of channel weights within the limited neighborhood of the 1D convolution kernel by constructing a local cross-channel interaction strategy, and realizes the establishment of long-range dependencies between channels in a more efficient computational manner.
[0174] In addition, the RIR module adopts a series of optimization strategies in the design of the network structure to balance performance and efficiency. The backbone part is designed with a modular architecture, and the core computing module consists of 10 cascaded residual units. 20 improved RCAB-ECA composite modules are integrated in each residual unit, and the feature channel dimension is uniformly configured as 64.
[0175] To ensure that the receptive field of the attention mechanism can be appropriately expanded as the number of channels increases, the adjustable parameter γ = 2 and the bias term b = 1 are introduced into the ECA module, ensuring the attention to local features while not ignoring global information.
[0176] In the design of the convolutional layer, except for the one-dimensional convolution in the ECA module, the standard convolutional layers in the network all use 3×3 convolution kernels. An activation function is connected after all convolutional layers except the last reconstruction layer. The upsampling module uses the sub-pixel convolution method, which has better performance compared to the transposed convolution layer and the nearest neighbor interpolation upsampling.
[0177] In the design of the loss function, since the L2 loss will generate too large a gradient when dealing with large errors, while the L1 loss can maintain a stable gradient, and compared with the L2 loss, the L1 loss is less sensitive to outliers and can produce a clearer reconstruction result. Therefore, the L1 loss is used to optimize the network parameters.
[0178] In this embodiment, the image is subjected to feature extraction through the primary feature extraction layer and the deep residual nested module, the feature map is scaled through the upsampling module, and the image is reconstructed through the reconstruction module, thereby obtaining a high-quality target image.
[0179] Please refer to Figure 10 , as Figure 10 shown, the embodiment of the present application also provides an anti-interference image compression and coding device 300, and the anti-interference image compression and coding device 300 includes:
[0180] A first input module 310, configured to input an image;
[0181] A transformation module 320, configured to perform deep learning wavelet transform processing on the image to obtain a plurality of subbands;
[0182] A determination module 330, configured to determine the coding order corresponding to each coding segment included in each subband;
[0183] An encoding module 340, configured to perform entropy encoding on each of the encoding segments according to the encoding order to obtain a target bitstream.
[0184] Optionally, the transform module 320 is specifically configured to:
[0185] Split the image in the row direction to obtain a high-frequency subband and a low-frequency subband;
[0186] Split the low-frequency subband in the column direction to obtain a first subband and a second subband; the first subband represents the low-frequency part in the low-frequency subband, and the second subband represents the high-frequency part in the low-frequency subband;
[0187] Split the high-frequency subband in the column direction to obtain a third subband and a fourth subband; the third subband represents the low-frequency part in the high-frequency subband, and the fourth subband represents the high-frequency part in the high-frequency subband.
[0188] Optionally, the determination module 330 is specifically configured to:
[0189] Perform quantization processing on each of the subbands;
[0190] Obtain the decomposition level, subband type, and the bit plane to which each of the subbands after quantization processing belongs;
[0191] Determine the encoding order corresponding to each of the encoding segments included in each of the subbands according to the decomposition level, the subband type, and the bit plane to which it belongs.
[0192] Optionally, the encoding module 340 is specifically configured to:
[0193] Perform context modeling on each of the encoding segments to determine the context information corresponding to each pixel in each of the encoding segments;
[0194] Perform probability estimation based on the context information corresponding to each pixel to obtain an estimated value corresponding to each pixel;
[0195] Perform interleaved entropy encoding on each of the encoding segments according to the estimated value corresponding to each pixel to obtain a target bitstream.
[0196] Optionally, the encoding module 340 is further specifically configured to:
[0197] For each pixel, determine the pixel category corresponding to the pixel based on the encoding situation of adjacent pixels;
[0198] Determine the context information corresponding to the pixel according to the pixel category corresponding to the pixel and the number of significant pixels among the adjacent pixels.
[0199] Optionally, the encoding module 340 is further specifically configured to:
[0200] For each pixel, determine a probability threshold associated with the estimated value corresponding to the pixel;
[0201] Based on the probability threshold, determine the encoding method corresponding to the pixel;
[0202] Perform cross-entropy encoding on the pixel through the encoding method.
[0203] Please refer to Figure 11 , such as Figure 11 shown, an anti-interference image compression and decoding device 400 is further provided in an embodiment of the present application. The anti-interference image compression and decoding device 400 includes:
[0204] A second input module 410 for inputting a target bitstream;
[0205] A decoding module 420 for performing entropy decoding on the target bitstream in a preset decoding order to obtain a plurality of subbands;
[0206] An inverse transformation module 430 for performing an inverse deep learning wavelet transformation on the plurality of subbands to obtain an image;
[0207] An enhancement module 440 for performing image enhancement processing on the image to obtain a target image.
[0208] Optionally, the enhancement module 440 is specifically configured to:
[0209] Extract features from the image through a primary feature extraction layer to obtain initial features;
[0210] Obtain multi-scale features through a deep residual nesting module based on the initial features;
[0211] Scale the feature map through an upsampling module, where the feature map is generated based on the multi-scale features;
[0212] Obtain a target image through a reconstruction module based on the feature map.
[0213] To solve the above technical problems, an embodiment of the present application also provides a computer device. Specifically, please refer to Figure 12 , Figure 12 which is the basic structural block diagram of the computer device in the embodiment of the present application.
[0214] The computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 5 with components 51-53 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0215] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.
[0216] The memory 51 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 51 may be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 may also be an external storage device of the computer device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 5. Of course, the memory 51 may also include both the internal storage unit of the computer device 5 and its external storage device. In this embodiment, the memory 51 is generally used to store the operating system and various application software installed on the computer device 5, such as the program codes of the anti-interference image compression encoding method and the anti-interference image compression decoding method, etc. In addition, the memory 51 can also be used to temporarily store various data that have been output or will be output.
[0217] In some embodiments, the processor 52 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 52 is generally used to control the overall operation of the computer device 5. In this embodiment, the processor 52 is used to run the program code stored in the memory 51 or process data, such as running the program code of the anti-interference image compression encoding method and the program code of the anti-interference image compression decoding method.
[0218] The network interface 53 may include a wireless network interface or a wired network interface, and the network interface 53 is generally used to establish a communication connection between the computer device 5 and other electronic devices.
[0219] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing the program code of the anti-interference image compression encoding method and the program code of the anti-interference image compression decoding method. The program code of the anti-interference image compression encoding method and the program code of the anti-interference image compression decoding method can be executed by at least one processor, so that the at least one processor executes the steps of the anti-interference image compression encoding method and the anti-interference image compression decoding method as described above.
[0220] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware online platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0221] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0222] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are given in the accompanying drawings, but do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structures directly or indirectly using the content of the specification and drawings of this application in other related technical fields are equally within the scope of the patent protection of this application.
Claims
1. An anti-interference image compression and encoding method, characterized in that Comprising: Input image; Performing deep learning wavelet transform processing on the said image to obtain multiple sub-bands; Determining the coding order corresponding to each coding segment included in each sub-band; Performing entropy coding on each coding segment according to the said coding order to obtain a target bitstream.
2. The method according to claim 1, wherein The performing deep learning wavelet transform processing on the said image to obtain multiple sub-bands includes: Performing row-direction splitting on the said image to obtain a high-frequency sub-band and a low-frequency sub-band; Performing column-direction splitting on the low-frequency sub-band to obtain a first sub-band and a second sub-band; the first sub-band represents the low-frequency part in the low-frequency sub-band, and the second sub-band represents the high-frequency part in the low-frequency sub-band; Performing column-direction splitting on the high-frequency sub-band to obtain a third sub-band and a fourth sub-band; the third sub-band represents the low-frequency part in the high-frequency sub-band, and the fourth sub-band represents the high-frequency part in the high-frequency sub-band.
3. The method according to claim 1, wherein The determining the coding order corresponding to each coding segment included in each sub-band includes: Performing quantization processing on each sub-band; Obtaining the decomposition level, sub-band type, and the bit plane to which each sub-band belongs after quantization processing; Determining the coding order corresponding to each coding segment included in each sub-band according to the decomposition level, the sub-band type, and the bit plane to which it belongs.
4. The method according to claim 1, wherein The performing entropy coding on each coding segment to obtain a target bitstream includes: Performing context modeling on each coding segment to determine the context information corresponding to each pixel in each coding segment; Performing probability estimation based on the context information corresponding to each pixel to obtain an estimated value corresponding to each pixel; Performing interleaved entropy coding on each coding segment according to the estimated value corresponding to each pixel to obtain a target bitstream.
5. The method according to claim 4, characterized in that, The performing context modeling on each coding segment to determine the context information corresponding to each pixel in each coding segment includes: For each pixel, determining the pixel category corresponding to the pixel based on the coding situation of adjacent pixels; Determining the context information corresponding to the pixel according to the pixel category corresponding to the pixel and the number of significant pixels among the adjacent pixels.
6. The method according to claim 4, wherein The performing interleaved entropy coding on each coding segment according to the estimated value corresponding to each pixel includes: For each pixel, determining a probability threshold value associated with the estimated value corresponding to the pixel; Determining the coding method corresponding to the pixel based on the probability threshold value; Performing interleaved entropy coding on the pixel through the coding method.
7. An anti-interference image compression and decoding method, characterized in that Comprising: Inputting the target bitstream; Performing entropy decoding on the target bitstream according to a preset decoding order to obtain multiple sub-bands; Performing inverse deep learning wavelet transform on the multiple sub-bands to obtain an image; Performing image enhancement processing on the said image to obtain a target image.
8. The method according to claim 7, wherein The performing image enhancement processing on the said image to obtain a target image includes: Performing feature extraction on the said image through a primary feature extraction layer to obtain initial features; Obtaining multi-scale features through a deep residual nested module based on the initial features; Performing scaling processing on the feature map through an upsampling module, the feature map being generated based on the multi-scale features; Obtaining a target image through a reconstruction module based on the feature map.
9. An anti-interference image compression and encoding device, characterized in that The anti-interference image compression and encoding device includes: A first input module for inputting an image; A transformation module for performing deep learning wavelet transformation processing on the image to obtain multiple subbands; A determination module for determining the encoding order corresponding to each coding segment included in each subband; An encoding module for performing entropy encoding on each coding segment according to the encoding order to obtain a target bitstream.
10. An anti-interference image compression and decoding device, characterized in that, The anti-interference image compression and decoding device includes: A second input module for inputting a target bitstream; A decoding module for performing entropy decoding on the target bitstream according to a preset decoding order to obtain multiple subbands; An inverse transformation module for performing deep learning inverse wavelet transformation on the multiple subbands to obtain an image; An enhancement module for performing image enhancement processing on the image to obtain a target image.