Three-dimensional biomedical image compression method, system, device and storage medium
By adopting training-based three-dimensional affine wavelet transformation and affine pattern methods in three-dimensional biomedical image compression technology, the problem that the prior art cannot adaptively process different texture images is solved, and more efficient image coding performance is achieved.
Patent Information
- Application Number
- CN202210099132.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-01-27
AI Technical Summary
The existing three-dimensional biomedical image compression technology cannot implement adaptive processing when processing images with different textures, which affects the performance of image encoding.
The training-based three-dimensional affine wavelet transformation is adopted to introduce affine patterns, so that the prediction and update network can adjust the wavelet basis function according to the different contents of the input after training, thereby improving the encoding performance of three-dimensional biomedical images.
Through the adjustment of affine patterns, images with different textures can be better adapted to images, improving the encoding performance of three-dimensional biomedical images.
Smart Images

Figure CN116563396B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image compression coding technology, and in particular to a three-dimensional biomedical image compression method, system, device and storage medium. Background Technology
[0002] Three-dimensional biomedical images mainly include three-dimensional biological images and three-dimensional medical images. With the development of artificial intelligence, three-dimensional biomedical images have broader application prospects and demands. The enormous size of three-dimensional biomedical images poses significant challenges to storage and transmission. Therefore, three-dimensional image compression technology is a key technology enabling the widespread use of these images. Currently, the mainstream technology for three-dimensional biomedical images is based on traditional wavelet transform schemes, with the most widely used technique being the JP3D standard (Extensions for Three-Dimensional Data in JPEG-2000).
[0003] Traditional wavelet transform is a local transform method that inherits and develops the localization idea of short-time Fourier transform while overcoming the shortcomings such as the window size not changing with frequency. It provides a "time-frequency" window that changes with frequency, making it an ideal tool for signal time-frequency analysis and processing. It can perform localized, multi-scale analysis of images and is therefore often used in image coding tasks. Traditional wavelet transform has two implementations, referred to as first-generation wavelets and second-generation wavelets. First-generation wavelet transform decomposes the signal using basis functions, resulting in complex matrix operations that are not conducive to hardware implementation. Second-generation wavelet transform, known as the wavelet lifting scheme, decomposes the first-generation wavelet into a lifting structure, which is easier to design manually and its structure is suitable for hardware implementation.
[0004] Taking a 3D image as an example, the process of second-generation wavelet transform decomposing a signal is as follows: Figure 1As shown, the input 3D image (input signal) is first split along a certain direction, separating its odd-numbered frames (odd signals) and even-numbered frames (even signals). The odd-numbered frames then undergo a prediction step, implemented in traditional wavelets using a linear prediction filter. The filtered result is then subtracted from the even-numbered frames; this difference is considered to contain the detail components of the original signal, i.e., the high-frequency signal obtained from the decomposition. The high-frequency signal is then filtered by an update filter, and the resulting signal is added to the odd-numbered frames to obtain an approximate component of the original signal, i.e., the low-frequency signal obtained from the decomposition. This completes one lifting step. It's worth noting that the number of lifting steps can be freely chosen, not limited to one. In subsequent lifting steps, the previously obtained low-frequency signal is used as the odd signal, and the high-frequency signal as the even signal, and the prediction and update steps are repeated to obtain new low-frequency and high-frequency signals. Multiple lifting steps can further yield more accurate high-frequency and low-frequency components. After one or more enhancement steps, the decomposition in the corresponding direction is completed, resulting in the decomposed high-frequency and low-frequency signals. The remaining two directions are then decomposed from the decomposed high-frequency and low-frequency signals to complete a full decomposition. The number of full decompositions, N, is greater than or equal to 1. Ultimately, the 3D image, after one decomposition process, yields 8 sub-bands. The three directions of the 3D image are denoted as x, y, and z, respectively. Assume that a decomposition is first performed in the z-direction (or multiple times), yielding the low-frequency signal L and the high-frequency signal H. L and H are then decomposed in the y-direction, where L is decomposed into LL and HL; H is decomposed into LH and HH. These four sub-bands are then decomposed in the x-direction, yielding 8 sub-bands: LLL, HLL, LHL, HHL, LLH, HLH, LHH, and HHH. LLL is the lowest frequency sub-band, containing the most important low-frequency details of the original signal, and can be used for the next full decomposition. N decompositions will yield 7N+1 sub-bands. The subbands obtained after wavelet transform have more concentrated energy than the original signal. Therefore, discrete wavelet transform is often used as the transform part of image coding. These subbands are further quantized and entropy encoded to obtain the final bitstream. Commonly used subband entropy coding methods include EZW, SPIHT, and EBCOT.
[0005] like Figure 2The diagram illustrates the JPEG-2000 encoding process, a common encoding method based on wavelet transform, and a standard developed by JPEG for image encoding. The original image data undergoes preprocessing and discrete wavelet transform (i.e., the traditional wavelet transform mentioned earlier) to obtain a series of subband coefficients with more concentrated energy. These subband coefficients are floating-point numbers, which is not conducive to saving bitrate during encoding. A quantization step converts these floating-point numbers to integers, and then entropy coding is used to organize the subband coefficients into a bitstream, resulting in compressed image data. At the decoding end, after obtaining these transmitted coefficients, the original image is reconstructed through entropy decoding, inverse quantization, and inverse discrete wavelet transform. The process is as follows: Figure 3 As shown.
[0006] Furthermore, a 3D image coding method based on trained wavelet transform has been proposed. The trained wavelet transform starts from image texture features, designs targeted wavelet transforms, and applies them to image coding. The main idea is to use a deep learning network to replace the prediction and update filters in traditional wavelet transforms. Based on a large number of training samples, a wavelet transform specific to the texture distribution is trained. For example... Figure 4 The diagram shows the steps involved in training a wavelet.
[0007] In image encoding, a deep network-based wavelet transform is applied to decompose the image and obtain wavelet coefficients. Then, a sub-band coding method is applied to obtain a compressed bitstream. In image decoding, a sub-band decoding method is applied to decode the wavelet coefficients from the compressed bitstream. Finally, a deep network-based inverse wavelet transform is applied to reconstruct the image.
[0008] Image coding methods based on trained wavelets can effectively capture non-directional features, allowing wavelet transforms to adapt to different textures to a large extent. However, it still has a problem: after training, the wavelet transform coefficients of the trained wavelet are determined. When processing images with different textures, the same set of coefficients is used, which makes the trained wavelet transform unable to adapt to different images, affecting the performance of image coding. Summary of the Invention
[0009] The purpose of this invention is to provide a three-dimensional biomedical image compression method, system, device, and storage medium that can improve encoding performance.
[0010] The objective of this invention is achieved through the following technical solution:
[0011] A three-dimensional biomedical image compression method, comprising:
[0012] In the encoding stage, a training-based 3D affine wavelet transform is used to decompose the input 3D biomedical image. The steps include: splitting the input 3D biomedical image into odd-numbered frames and even-numbered frames in the current direction; inputting the odd-numbered frames into a prediction network based on a deep network to obtain a prediction result and a first affine map; scaling the prediction result using the first affine map and subtracting it from the even-numbered frames, with the difference being a high-frequency signal; inputting the high-frequency signal into an update network based on a deep network to obtain an update result and a second affine map; scaling the update result using the second affine map and adding it to the odd-numbered frames, with the addition result being a low-frequency signal, thus completing one decomposition process; executing several decomposition processes to complete the decomposition in the current direction; decomposing the low-frequency and high-frequency signals obtained from the current direction into other directions; and using the decomposition results from all directions for entropy coding.
[0013] In the decoding stage, the process is reversed from that in the encoding stage to reconstruct the three-dimensional biomedical image.
[0014] A three-dimensional biomedical image compression system, comprising:
[0015] An encoder is used in the encoding stage. In this stage, a training-based 3D affine wavelet transform is employed to decompose the input 3D biomedical image. The steps include: splitting the input 3D biomedical image into odd-numbered frames and even-numbered frames in the current direction; inputting the odd-numbered frames into a deep network-based prediction network; scaling the prediction result using an affine graph and then subtracting it from the even-numbered frames; the difference result is a high-frequency signal; inputting the high-frequency signal into a deep network-based update network; scaling the update result using an affine graph and then adding it to the odd-numbered frames; the sum result is a low-frequency signal. This completes one decomposition process. Several decomposition processes are executed to complete the decomposition in the current direction; the low-frequency signal and high-frequency signal obtained from the current direction decomposition are then used to decompose the image in other directions; and entropy coding is performed using the decomposition results from all directions.
[0016] A decoder is used in the decoding stage; the decoding stage employs a process that is the reverse of the encoding stage to reconstruct a three-dimensional biomedical image.
[0017] A processing device includes: one or more processors; and a memory for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0019] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0020] As can be seen from the technical solution provided by the present invention, a training-based three-dimensional affine wavelet transform is designed, and an affine graph is introduced. After training, although the prediction and update network is fixed, the affine graph will change according to different inputs. Therefore, it is equivalent to the wavelet basis function being adjusted according to different input contents, thereby improving the coding performance of three-dimensional biomedical images. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A schematic diagram of the second-generation wavelet transform lifting structure is provided for the background art of this invention;
[0023] Figure 2 The JPEG-2000 encoding flowchart provided as a background feature of this invention;
[0024] Figure 3 The JPEG-2000 decoding flowchart provided in the background of this invention;
[0025] Figure 4 A flowchart of the wavelet training process steps is provided for the background technology of this invention;
[0026] Figure 5 A flowchart of a three-dimensional biomedical image compression method provided in an embodiment of the present invention;
[0027] Figure 6 A flowchart of the lifting steps of training-based three-dimensional affine wavelet transform provided for embodiments of the present invention;
[0028] Figure 7 This is a schematic diagram of the prediction network and update network structure provided in an embodiment of the present invention;
[0029] Figure 8 A flowchart of context modeling based on deep networks provided for embodiments of the present invention;
[0030] Figure 9 This is a schematic diagram of a three-dimensional inter-subband context depth network structure provided in an embodiment of the present invention;
[0031] Figure 10 This is a schematic diagram of a three-dimensional sub-band context depth network structure provided in an embodiment of the present invention;
[0032] Figure 11This is a schematic diagram of a three-dimensional context fusion deep network structure provided in an embodiment of the present invention;
[0033] Figure 12 This is a schematic diagram of subband coefficient entropy encoding provided in an embodiment of the present invention;
[0034] Figure 13 A flowchart of the affine wavelet inverse transform based on a deep network provided in an embodiment of the present invention;
[0035] Figure 14 The diagram illustrates the beneficial results of the three-dimensional affine wavelet transform provided in the embodiments of the present invention.
[0036] Figure 15 A flowchart of lossy image coding based on training-based three-dimensional affine wavelet transform provided for embodiments of the present invention;
[0037] Figure 16 A flowchart of lossy image decoding based on training-based three-dimensional affine wavelet transform provided in an embodiment of the present invention;
[0038] Figure 17 A flowchart of lossless image coding based on training-based three-dimensional affine wavelet transform provided in an embodiment of the present invention;
[0039] Figure 18 A flowchart of lossless image decoding based on training-based three-dimensional affine wavelet transform provided in an embodiment of the present invention;
[0040] Figure 19 This is a flowchart of an image coding method combining affine wavelet transform and parameter sharing strategy provided in an embodiment of the present invention.
[0041] Figure 20 This is a flowchart of an image decoding method combining affine wavelet transform and parameter sharing strategy provided in an embodiment of the present invention.
[0042] Figure 21 This is a flowchart of an image coding method combining affine wavelet transform and trained three-dimensional entropy coding, provided in an embodiment of the present invention.
[0043] Figure 22 This is a flowchart of an image decoding method combining affine wavelet transform and trained three-dimensional entropy coding, provided in an embodiment of the present invention.
[0044] Figure 23 A flowchart of an image coding method combining affine wavelet transform based on a parameter sharing strategy and trained three-dimensional entropy coding provided in an embodiment of the present invention;
[0045] Figure 24 A flowchart of an image decoding method combining affine wavelet transform based on a parameter sharing strategy and training a 3D encoder / decoder, provided in an embodiment of the present invention;
[0046] Figure 25 A schematic diagram of a three-dimensional biomedical image compression system provided in an embodiment of the present invention;
[0047] Figure 26 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0049] First, the following explanations are provided for the terms that may be used in this article:
[0050] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions shall be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” shall be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements known in the art that are not expressly listed.
[0051] The three-dimensional biomedical image compression scheme provided by this invention will be described in detail below. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they should be performed according to conventional conditions in the art or conditions recommended by the manufacturer.
[0052] Example 1
[0053] like Figure 1 As shown, a three-dimensional biomedical image compression method mainly includes the following steps:
[0054] In the encoding stage, a training-based 3D affine wavelet transform is used to decompose the input 3D biomedical image. The steps include: splitting the input 3D biomedical image into odd-numbered frames and even-numbered frames in the current direction; inputting the odd-numbered frames into a prediction network based on a deep network to obtain a prediction result and a first affine map; scaling the prediction result using the first affine map and subtracting it from the even-numbered frames, with the difference being a high-frequency signal; inputting the high-frequency signal into an update network based on a deep network to obtain an update result and a second affine map; scaling the update result using the second affine map and adding it to the odd-numbered frames, with the addition result being a low-frequency signal, thus completing one decomposition process; executing several decomposition processes to complete the decomposition in the current direction; decomposing the low-frequency and high-frequency signals obtained from the current direction into other directions; and using the decomposition results from all directions for entropy coding.
[0055] In the decoding stage, the process is reversed from that in the encoding stage to reconstruct the three-dimensional biomedical image.
[0056] To facilitate understanding, the following section provides a detailed explanation of each stage of the encoding and decoding process.
[0057] I. Encoding stage.
[0058] 1. Training-based 3D affine wavelet transform.
[0059] 1) Improvement steps based on training of 3D affine wavelet transform.
[0060] Figure 6 This paper demonstrates the enhancement steps of a training-based 3D affine wavelet transform. The deep network-based affine wavelet transform incorporates a deep network for prediction and update steps. In addition, the training-based 3D affine wavelet transform update and prediction networks, besides learning update and prediction filters, also output an affine graph. The interaction between the affine graph and the prediction and update results is used to scale the output size at each spatial location. After training, although the prediction and update networks are fixed, the affine graph changes according to different inputs, thus equivalent to the wavelet basis functions adjusting according to different input content.
[0061] In this embodiment of the invention, training is performed based on an image compression objective function, which is: R + λD; where R represents the encoding bitrate, D represents the L2 loss between the reconstructed image and the input image, and λ is a coefficient. During training, the training parameters of the prediction network and the update network are optimized by minimizing the objective function.
[0062] In this embodiment of the invention, the value of each value in the affine graph is generally between 0 and 1. The affine graph can be implemented in any of the following ways: 1) The first affine graph is a tensor with the same size as the prediction result, and the second affine graph is a tensor with the same size as the update result. The prediction result and the update result each have their own affine graphs calculated using the sigmoid function. The outputs of the update network and the prediction network are different (i.e., the prediction result and the update result are different), therefore the affine graphs applied to the prediction result and the update result are also different. Each of the two affine graphs provides a separate scaling value for each spatial location of the output results of the prediction network and the update network, similar to a spatial attention mechanism. 2) The affine graph is a scalar. The same scaling value is provided for each spatial location of the prediction result and the update result, similar to a channel attention mechanism. During the learning process, the affine graph is set as a learnable variable. Applying this variable to the output results of the prediction network and the update network is equivalent to applying the same scaling to each position of the output result.
[0063] In this embodiment of the invention, the prediction network and the update network adopt the same network structure, such as... Figure 7 As shown, it mainly includes five three-dimensional convolutional layers connected in sequence. The output of the first three-dimensional convolutional layer is connected to the output of the fifth three-dimensional convolutional layer, the output of the second three-dimensional convolutional layer is connected to the output of the fourth three-dimensional convolutional layer, and the third and fourth three-dimensional convolutional layers first use the tanh activation function before performing convolution operations. Figure 7 In this context, 3×3×3×1 indicates that a 3D convolutional layer uses a 3×3×3 convolutional kernel to generate one feature map. Similarly, 3×3×3×16 indicates that a 3D convolutional layer uses a 3×3×3 convolutional kernel to generate 16 feature maps. Of course, the convolutional kernel parameters and the number of feature maps provided here are only examples and do not constitute a limitation. In practical applications, users can make corresponding adjustments according to the situation or experience.
[0064] Based on the above principles, the embodiments of the present invention can realize both lossy and lossless encoding.
[0065] a) In lossy encoding, the first affine map is multiplied by the prediction result, and the second affine map is multiplied by the update result to scale the output size for each spatial location in the prediction and update results. For example... Figure 6 An example of lossy encoding is shown.
[0066] b) Lossless encoding.
[0067] Although the forward and inverse transformations of the affine wavelet transform are completely reversible in terms of algorithm flow, the presence of multiplication and division operations (in the decoding stage) in the affine wavelet transform makes complete reversibility impossible in practical implementation. To address this issue, this invention designs an efficient lossless form of the affine wavelet transform.
[0068] As is well known, bitwise left or right shifts do not incur precision loss in practice. Therefore, by quantizing the value at each position in the first and second affine graphs, implementing multiplication through a right shift operation, and scaling the output size of each spatial position in the prediction and update results using the right shift operation, precision loss is concisely avoided. For example, the affine graph... Figure 4 Rounding quantization rounds to the nearest power of 2; for example, 0.5 is quantized to 2. -1 .
[0069] It should be noted that when decomposing in each direction, the number of times the decomposition process is executed can be set by those skilled in the art based on the actual situation or experience. Similar to the prior art, in the next decomposition, the low-frequency signal and high-frequency signal obtained from the previous decomposition are used as input (i.e., no splitting is required). After the current direction is decomposed, all the low-frequency signals and high-frequency signals obtained from the decomposition are decomposed in other directions (the number of decompositions is the same as the number of decompositions in the first direction). After all three directions are decomposed, a complete decomposition is completed. Similarly, the number of complete decompositions can be one or more.
[0070] 2) Parameter sharing strategy based on training-based 3D wavelet transform.
[0071] As mentioned in the background section, the 3D image coding method based on trained wavelet transform shares the parameters of the wavelet transform in the three dimensions of the image. However, this is not reasonable for 3D images with different properties in the three directions. Therefore, sharing the parameters in all directions will affect the performance of 3D image coding based on trained wavelet transform.
[0072] Three-dimensional biomedical images exhibit different characteristics in the three directions. Due to limitations in imaging equipment, three-dimensional images can be categorized into isotropic and anisotropic images based on their axial resolution. Isotropic images refer to images where the axial resolution is the same as the row and column resolutions, resulting in consistent properties across all three dimensions. Anisotropic images, on the other hand, have different axial resolutions than the row and column resolutions (generally with a smaller axial resolution), leading to axial properties that differ from the other two directions.
[0073] To address the aforementioned characteristics, this invention proposes a parameter-sharing strategy based on trained 3D wavelet transform. For isotropic images, the prediction and update networks in the trained 3D affine wavelet transform use the same structure (e.g., ...). Figure 7 As shown), and share all parameters, which reduces the complexity of the network and saves parameters; for anisotropic images, in order to better adapt to different axial properties, the prediction network and the update network in the trained 3D affine wavelet transform use the same structure (as shown). Figure 7 As shown, the axial direction uses a separate set of parameters, while all other directions share all parameters.
[0074] 2. Three-dimensional entropy coding.
[0075] Based on the processing described in 1), the high-frequency and low-frequency signals obtained from the decomposition in all directions are referred to as three-dimensional sub-bands; all sub-bands are entropy encoded sequentially according to a set order. If a lossy encoding method is used, quantization is required before entropy encoding; if a lossless encoding method is used, entropy encoding can be performed directly.
[0076] Currently, commonly used subband entropy coding methods include EZW, EBCOT, etc.
[0077] EZW, or Embedded Zerotree Wavelet Encoding, refers to a mathematical structure based on image wavelet transform. In the EZW algorithm, the embedded bitstream is implemented using a zerotree structure combined with successive approximation quantization. The purpose of the zerotree structure is to efficiently represent the positions of non-zero values (effective value mapping) in the wavelet transform coefficient matrix.
[0078] EBCOT stands for Embedded Block Coding with Optimized Truncation, an algorithm published in 1999. JPEG2000 uses EBCOT for wavelet coefficient quantization, a core feature of the JPEG2000 standard and a method for embedded bit-level coding of wavelet coefficients. EBCOT coding consists of two parts: Part 1 (Tier 1) divides each subband into independent coding blocks, performs embedded coding scans on each block independently, performs bit-level coding on each block, and finally performs MQ arithmetic coding on the scan results to obtain the embedded bitstream; Part 2 (Tier 2) combines the embedded bitstreams of each coding block according to the output bitrate requirements, and performs optimization, truncation, sorting, and packing on the coding streams of all blocks to obtain the JPEG2000 bitstream.
[0079] While the commonly used 3D versions of subband entropy coding can be used for 3D image coding, they are not data-driven, meaning they cannot adjust to the texture of the 3D image. Therefore, these methods do not offer good coding performance.
[0080] To address this, this invention proposes a three-dimensional entropy coding method based on deep networks. For a given three-dimensional subband to be encoded, a three-dimensional context coarse extraction network is used to extract context from already encoded three-dimensional subbands. This context is then input into an inter-subband context extraction network to obtain inter-subband context. The given three-dimensional subband to be encoded is then input into an intra-subband context extraction network to obtain intra-subband context. The inter-subband context and the intra-subband context are then input into a three-dimensional context fusion deep network to obtain entropy coding parameters for the given three-dimensional subband. Finally, these entropy coding parameters are used to encode the given three-dimensional subband. The following provides a detailed description of the three-dimensional entropy coding method based on deep networks.
[0081] As previously mentioned, 7N+1 three-dimensional subbands can be obtained through training-based three-dimensional affine wavelet transform. Taking N=2 as an example, the first complete decomposition yields 8 subbands, of which the lowest frequency subband LLL can be further decomposed into 8 subbands, resulting in 15 subbands after two complete decompositions. When encoding the 15 three-dimensional subbands, each subband is encoded sequentially in the order {LLL,HLL,LHL,HHL,LLH,HLH,LHH,HHH}. During encoding, two aspects of correlation are emphasized: (1) inter-subband correlation; (2) intra-subband correlation. The coefficients of each subband are encoded sequentially from left to right and from top to bottom. For each coefficient, a deep network is first used for context modeling to give the probability distribution of the coefficient; then, based on the probability distribution, an arithmetic encoder is used for entropy encoding. Figure 8 The diagram illustrates the main process of context modeling, which includes:
[0082] 1) A 3D context coarse extraction network extracts coarse context C. t .
[0083] Encoding the current three-dimensional subband S t At this point, the already encoded 3D subband contains the context of the current 3D subband, which is extracted using a deep neural network. First, the encoded 3D subband is reconstructed using an inverse transform to match the scale of the current 3D subband, and then merged along the channel dimension. Next, it passes through a coarse 3D context extraction network, which consists of a single convolutional neural network layer using 3×3×3 convolutional kernels. The network outputs the coarsely extracted context C. t .
[0084] 2) 3D subband context extraction.
[0085] a) Input the context into the 3D sub-band context extraction network to obtain the 3D sub-band context C_b t .
[0086] like Figure 9 As shown, the three-dimensional subband context extraction network mainly includes: a first convolutional unit and a second convolutional unit arranged in sequence, each convolutional unit including two three-dimensional convolutional layers arranged in sequence, and a ReLU function between the two three-dimensional convolutional layers; the input and output of each convolutional unit are connected.
[0087] b) Input the current 3D subband to be encoded into the 3D subband context extraction network to obtain the context C_w within the 3D subband. t .
[0088] like Figure 10 As shown, the three-dimensional sub-band context extraction network mainly includes: a masked three-dimensional convolutional layer, a third convolutional unit, and a fourth convolutional unit arranged sequentially; each convolutional unit includes two masked three-dimensional convolutional layers arranged sequentially, and a ReLU function is provided between the two masked three-dimensional convolutional layers; the input of the third convolutional unit includes the output of the masked three-dimensional convolutional layer in front of it and the input of the three-dimensional sub-band context extraction network, and the input of the third convolutional unit is also connected to its output; the input and output of the fourth convolutional unit are connected.
[0089] 3) After the aforementioned steps, the three-dimensional subband S to be encoded is obtained. t The intra-subband context and inter-subband context are fused by a 3D context fusion deep network to output the parameters of the cumulative distribution probability function of the coefficients of the 3D subband to be encoded (i.e., entropy coding parameters).
[0090] like Figure 11 As shown, the three-dimensional context fusion deep network mainly includes: a masked three-dimensional convolutional layer and three three-dimensional convolutional layers arranged sequentially. ReLU functions are provided between the masked three-dimensional convolutional layers and between adjacent three-dimensional convolutional layers.
[0091] Overall, for each 3D subband to be encoded, each coefficient is encoded sequentially from left to right, top to bottom, and front to back. For each coefficient, first use... Figure 10The network shown performs context modeling to obtain the probability distribution of the 3D subband coefficients; then, based on the probability distribution, an arithmetic encoder is used for entropy encoding. A 3D PixelCNN (Pixel Convolutional Neural Network) is used for context modeling to obtain the probability distribution of the coefficients to be encoded. PixelCNN takes the current 3D subband as input and outputs the parameters of the cumulative probability function of the coefficients to be encoded. Figure 12 For example, the center point represents the coefficient to be encoded, the light-colored part represents the coefficients that have already been encoded, and the dark-colored part represents the coefficients that have not yet been encoded. By ensuring that the encoded coefficients are usable and the uncoded coefficients are unusable, the correctness of the decoding logic is guaranteed.
[0092] II. Decoding stage.
[0093] The decoding stage employs the reverse process of the encoding stage.
[0094] First, a training-based 3D entropy decoding method is used: each sub-band coefficient is decoded sequentially from left to right and from top to bottom. For each 3D sub-band coefficient, a deep network is used for context modeling to obtain the probability distribution of the 3D sub-band coefficient; then, based on the probability distribution, an arithmetic decoder is called to perform entropy decoding to obtain the reconstructed 3D sub-band coefficient.
[0095] Next, a training-based inverse 3D affine wavelet transform is performed, which is the inverse process of the training-based 3D affine wavelet transform. The 3D biomedical image is then reconstructed through this training-based inverse 3D affine wavelet transform, such as... Figure 13 As shown. The networks of the inverse transform generally share parameters with the forward transform (i.e., the training-based three-dimensional affine wavelet transform).
[0096] Similarly, in accordance with the lossy and lossless encoding methods described in the encoding stage, the decoding stage employs corresponding lossy and lossless decoding methods. Specifically: during lossy decoding, the prediction result is divided by the first affine map, and the update result is divided by the second affine map, used to scale the output size at each spatial location in the prediction and update results. Figure 13 This paper demonstrates a lossy decoding scheme based on training-based inverse 3D affine wavelet transform. In lossless decoding, the values at each position in the first and second affine maps are quantized, and division is implemented using a left shift operation. This left shift operation scales the output size of each spatial position in the prediction and update results. Similarly, the lossy decoding method requires inverse quantization before performing the training-based inverse 3D affine wavelet transform; the lossless decoding method directly performs the training-based inverse 3D affine wavelet transform on the entropy decoding result.
[0097] The above solutions of the present invention mainly include three aspects of improvement:
[0098] 1. A training-based three-dimensional affine wavelet transform (inverse transform) is proposed, which includes lossy coding and lossless coding (lossy decoding and lossless decoding).
[0099] 2. A parameter sharing strategy based on training of three-dimensional affine wavelet transform (inverse transform) designed for the characteristics of three-dimensional biomedical images.
[0100] 3. A 3D entropy encoding (entropy decoding) method based on deep networks designed for 3D biomedical images.
[0101] The most basic improvement scheme in the first aspect above can improve coding performance compared to the existing scheme; based on this, combining the second aspect, or the third aspect, or combining the second and third aspects at the same time can further improve coding performance.
[0102] like Figure 14 As shown, the beneficial effects of training-based 3D affine wavelet transform (i.e., using only the improvement scheme of aspect 1) are demonstrated. The training-based 3D affine wavelet transform is universal on various datasets. Here, the beneficial effects of affine wavelet transform (referred to as affinewavelet in the figure) compared to the traditional 9 / 7 wavelet transform (referred to as 9 / 7 wavelet in the figure) and the training wavelet transform (referred to as additive wavelet in the figure) that only replaces the update and prediction network are demonstrated on the 3D image dataset, namely the electron microscopy image FAFB dataset, are shown. Figure 24 The three line segments from top to bottom correspond to the affine wavelet, the additive wavelet, and the 9 / 7 wavelet, respectively.
[0103] Below are some relevant solution examples based on the above three aspects of improvement.
[0104] Example 1: Training-based 3D affine wavelet transform.
[0105] This section introduces a scheme for 3D biomedical image compression based on trained 3D affine wavelet transform, where the subsequent entropy coding uses an existing method.
[0106] 1. Three-dimensional affine wavelet transform for lossy image coding.
[0107] 1) Encoding stage.
[0108] like Figure 15 As shown, the input is a 3D biomedical image to be encoded, and the output is a compressed bitstream, which can be sent to the decoding end for decoding. The main steps are described below:
[0109] Step 1: Apply a training-based 3D affine wavelet transform (lossy form) to the input 3D biomedical image to be encoded, performing N affine wavelet transforms to obtain wavelet coefficients, consisting of 7N+1 subbands. The value of N is prior; for example, in this example, N is set to 4.
[0110] Step 2: Use subband coding to quantize and entropy code the wavelet coefficients to obtain a compressed bitstream.
[0111] Subband coding methods for wavelet coefficients involve two steps: quantization and entropy coding. Common subband coding methods include 3D EZW, SPIHT, and EBCOT, which can be selected at this step based on specific requirements.
[0112] 2. Decoding stage.
[0113] Corresponding to the coding phase scheme architecture and coding operations, such as Figure 16 As shown, the input is the received compressed bitstream, and the output is the reconstructed 3D biomedical image. The main steps are described below:
[0114] Step 1: Perform entropy decoding and dequantization on the compressed bitstream using the subband decoding method to obtain the reconstructed wavelet coefficients. The subband decoding method is compatible with the subband coding method.
[0115] Step 2: Apply the inverse transform (lossy form) of the training-based 3D affine wavelet transform to the reconstructed wavelet coefficients to obtain the reconstructed 3D biomedical image.
[0116] 2. Three-dimensional affine wavelet transform for lossless image coding.
[0117] 1) Encoding stage.
[0118] like Figure 17 As shown, the input is a 3D biomedical image to be encoded, and the output is a compressed bitstream, which can be sent to the decoding end for decoding. The main steps are described below:
[0119] Step 1: Apply a training-based 3D affine wavelet transform (lossless form) to the input 3D biomedical image to be encoded, performing N affine wavelet transforms to obtain wavelet coefficients, consisting of 7N+1 subbands. The value of N is prior; for example, in this example, N is set to 4.
[0120] Step 2: Use subband coding to entropy encode the wavelet coefficients to obtain the compressed bitstream.
[0121] Entropy coding is performed on the subbands of the wavelet coefficients. Common subband coding methods include 3D EZW, SPIHT, and EBCOT, which can be selected at this step based on specific requirements.
[0122] 2. Decoding stage.
[0123] Corresponding to the coding phase scheme architecture and coding operations, such as Figure 18 As shown, the input is the received compressed bitstream, and the output is the reconstructed 3D biomedical image. The main steps are described below:
[0124] Step 1: Perform entropy decoding on the compressed bitstream using the subband decoding method to obtain the reconstructed wavelet coefficients. The subband decoding method is compatible with the subband coding method.
[0125] Step 2: Apply the inverse transform (lossless form) of the training-based 3D affine wavelet transform to the reconstructed wavelet coefficients to obtain the reconstructed 3D biomedical image.
[0126] Example 2: Image coding method combining affine wavelet transform and parameter sharing strategy.
[0127] This section introduces a scheme for 3D biomedical image compression that combines training-based 3D affine wavelet transform with a parameter sharing strategy. This is a compression scheme that combines aspects 1 and 2 mentioned above, and is referred to below as affine wavelet transform (inverse transform) based on a parameter sharing strategy. The entropy coding used in the following section is an existing method.
[0128] 1. Encoding stage.
[0129] like Figure 19 As shown, the input is a 3D biomedical image to be encoded, and the output is a compressed bitstream, which can be sent to the decoding end for decoding. The main steps are described below:
[0130] Step 1: Apply an affine wavelet transform based on a parameter-sharing strategy to the input image to be encoded, performing N affine wavelet transforms to obtain wavelet coefficients, which consist of 7N+1 subbands. The value of N is prior; for example, in this example, N is set to 4.
[0131] Step 2: Use subband coding to quantize and entropy code the wavelet coefficients to obtain a compressed bitstream.
[0132] Subband coding methods for wavelet coefficients involve two steps: quantization and entropy coding. Common subband coding methods include 3D EZW, SPIHT, and EBCOT, which can be selected at this step based on specific requirements.
[0133] 2. Decoding stage.
[0134] Corresponding to the coding phase scheme architecture and coding operations, such as Figure 20 As shown, the input is the received compressed bitstream, and the output is the reconstructed 3D biomedical image. The main steps are described below:
[0135] Step 1: Perform entropy decoding and dequantization on the compressed bitstream using the subband decoding method to obtain the reconstructed wavelet coefficients. The subband decoding method is compatible with the subband coding method.
[0136] Step 2: Apply the inverse affine wavelet transform based on the parameter sharing strategy to the reconstructed wavelet coefficients to obtain the reconstructed three-dimensional biomedical image.
[0137] In this section, for lossless encoding, simply remove the quantization in the encoding stage and the sum and dequantization in the decoding stage; the other processes are the same, so they will not be elaborated upon.
[0138] Example 3: An image coding method combining affine wavelet transform and 3D entropy coding based on deep networks.
[0139] This section introduces a scheme for 3D biomedical image compression that combines training-based 3D affine wavelet transform with a deep network-based 3D entropy coding method, namely, a compression scheme that combines the aforementioned first and third aspects.
[0140] 1. Encoding stage.
[0141] like Figure 21 As shown, the input is a 3D biomedical image to be encoded, and the output is a compressed bitstream, which can be sent to the decoding end for decoding. The main steps are described below:
[0142] Step 1: Apply a training-based affine wavelet transform to the input image to be encoded, performing N affine wavelet transforms to obtain wavelet coefficients, which consist of 7N+1 subbands. The value of N is prior; for example, in this example, N is set to 4.
[0143] Step 2: After quantizing the wavelet coefficients, entropy coding is performed using a three-dimensional entropy coding method based on deep networks to obtain a compressed bitstream.
[0144] 2. Decoding stage.
[0145] Corresponding to the coding phase scheme architecture and coding operations, such as Figure 22 As shown, the input is the received compressed bitstream, and the output is the reconstructed 3D biomedical image. The main steps are described below:
[0146] Step 1: Use a deep network-based 3D entropy decoding method to perform entropy decoding, and then perform inverse quantization to obtain the reconstructed wavelet coefficients.
[0147] Step 2: Apply training-based affine wavelet inverse transform to the reconstructed wavelet coefficients to obtain the reconstructed 3D biomedical image.
[0148] In this section, for lossless encoding, simply remove the quantization in the encoding stage and the sum and dequantization in the decoding stage; the other processes are the same, so they will not be elaborated upon.
[0149] Example 4: An image coding method combining affine wavelet transform based on parameter sharing strategy and 3D entropy coding based on deep networks.
[0150] This section introduces a scheme for 3D biomedical image compression that combines training-based 3D affine wavelet transform, parameter sharing strategy, and deep network-based 3D entropy coding method. Specifically, it combines the compression schemes mentioned in aspects 1, 2, and 3 above. The scheme that combines training-based 3D affine wavelet transform (inverse transform) with parameter sharing strategy will be referred to as affine wavelet transform (inverse transform) based on parameter sharing strategy.
[0151] 1. Encoding stage.
[0152] like Figure 23 As shown, the input is a 3D biomedical image to be encoded, and the output is a compressed bitstream, which can be sent to the decoding end for decoding. The main steps are described below:
[0153] Step 1: Apply an affine wavelet transform based on a parameter-sharing strategy to the input image to be encoded, performing N affine wavelet transforms to obtain wavelet coefficients, which consist of 7N+1 subbands. The value of N is prior; for example, in this example, N is set to 4.
[0154] Step 2: After quantizing the wavelet coefficients, entropy coding is performed using a three-dimensional entropy coding method based on deep networks to obtain a compressed bitstream.
[0155] 2. Decoding stage.
[0156] Corresponding to the coding phase scheme architecture and coding operations, such as Figure 24 As shown, the input is the received compressed bitstream, and the output is the reconstructed 3D biomedical image. The main steps are described below:
[0157] Step 1: Use a deep network-based 3D entropy decoding method to perform entropy decoding, and then perform inverse quantization to obtain the reconstructed wavelet coefficients.
[0158] Step 2: Apply the inverse affine wavelet transform based on the parameter sharing strategy to the reconstructed wavelet coefficients to obtain the reconstructed three-dimensional biomedical image.
[0159] In this section, for lossless encoding, simply remove the quantization in the encoding stage and the sum and dequantization in the decoding stage; the other processes are the same, so they will not be elaborated upon.
[0160] The relevant calculation processes involved in the above four examples have been described in detail above, therefore, they will not be repeated here. The superior performance of this invention is demonstrated below through comparative experiments.
[0161] Example 2
[0162] This invention also provides a three-dimensional biomedical image compression system, which is mainly based on the method provided in Embodiment 1 above, such as... Figure 25 As shown, the system mainly includes:
[0163] An encoder is used in the encoding stage. In this stage, a training-based 3D affine wavelet transform is employed to decompose the input 3D biomedical image. The steps include: splitting the input 3D biomedical image into odd-numbered frames and even-numbered frames in the current direction; inputting the odd-numbered frames into a deep network-based prediction network; scaling the prediction result using an affine graph and then subtracting it from the even-numbered frames; the difference result is a high-frequency signal; inputting the high-frequency signal into a deep network-based update network; scaling the update result using an affine graph and then adding it to the odd-numbered frames; the sum result is a low-frequency signal. This completes one decomposition process. Several decomposition processes are executed to complete the decomposition in the current direction; the low-frequency signal and high-frequency signal obtained from the current direction decomposition are then used to decompose the image in other directions; and entropy coding is performed using the decomposition results from all directions.
[0164] A decoder is used in the decoding stage; the decoding stage employs a process that is the reverse of the encoding stage to reconstruct a three-dimensional biomedical image.
[0165] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0166] It should be noted that the main technical details involved in the above encoding and decoding stages, as well as the combination of specific improvement schemes involved in the encoding and decoding stages, can also refer to the method in Implementation 1, so they will not be repeated here.
[0167] Example 3
[0168] The present invention also provides a processing device, such as Figure 26 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0169] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0170] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example:
[0171] Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc.
[0172] The output device can be a display terminal;
[0173] The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.
[0174] Example 4
[0175] The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.
[0176] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0177] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A three-dimensional biomedical image compression method, characterized in that: include: In the encoding stage, a training-based three-dimensional affine wavelet transform is used to decompose the input three-dimensional biomedical image, and the steps include: splitting the current direction of the input three-dimensional biomedical image into odd frames and even frames, inputting the odd frames into a prediction network based on a deep network, obtaining a prediction result and a first affine image, subtracting the prediction result from the even frame after scaling the first affine image, and the subtraction result is a high-frequency signal, and inputting the high-frequency signal into an update network based on a deep network to obtain an update result and a second affine image, and adding the update result to the odd frame after scaling the second affine image, and the addition result is a low-frequency signal, completing a decomposition process, executing the decomposition process several times, and completing the decomposition in the current direction; decomposing the low-frequency signal and the high-frequency signal obtained by the decomposition of the current direction in other directions; and performing entropy coding using the decomposition results of all directions; In the decoding stage, a process opposite to that of the encoding stage is adopted to reconstruct the three-dimensional biomedical image.
2. A three-dimensional biomedical image compression method according to claim 1, characterized in that: The first affine map is a tensor of the same size as the prediction result, and the second affine map is a tensor of the same size as the update result. The prediction result and the update result are each calculated using a sigmoid function to obtain a corresponding affine map, and the corresponding affine map provides a scaling value for each spatial position of the prediction result and the update result. Alternatively, the affine map is a scalar, which is set as a learnable variable, and the same scaling value is provided for each spatial position of the prediction result and the update result through the affine map.
3. A three-dimensional biomedical image compression method according to claim 1 or 2, characterized in that: The prediction result is scaled by the first affine map, and the update result is scaled by the second affine map, including: In the encoding stage, the first affine map is multiplied by the prediction result, and the second affine map is multiplied by the update result to scale the output size of each spatial position in the prediction result and the update result; or, the values of each position of the first affine map and the second affine map are quantized, and the multiplication operation is implemented by a right shift operation, and the output size of each spatial position in the prediction result and the update result is scaled by the right shift operation; In the decoding stage, the prediction result is divided by the first affine map, and the update result is divided by the second affine map to scale the output size of each spatial position in the prediction result and the update result; or, the numerical value of each position of the first affine map and the second affine map is quantized, and the division operation is implemented by a left shift operation, and the output size of each spatial position in the prediction result and the update result is scaled by the left shift operation.
4. A three-dimensional biomedical image compression method according to claim 1, characterized in that: The training-based three-dimensional affine wavelet transform uses a parameter sharing strategy; The three-dimensional biomedical images are divided into isotropic images and anisotropic images according to different axial resolutions; for isotropic images, the prediction network and the update network based on the trained three-dimensional affine wavelet transform use the same structure and share all parameters; for anisotropic images, the prediction network and the update network based on the trained three-dimensional affine wavelet transform use the same structure, a separate set of parameters is used in the axial direction, and all parameters are shared in the other directions.
5. A three-dimensional biomedical image compression method according to claim 1 or 4, characterized in that: The structure of the prediction network and the update network includes: five three-dimensional convolutional layers connected in sequence, wherein the output of the first three-dimensional convolutional layer is connected to the output of the fifth three-dimensional convolutional layer, the output of the second three-dimensional convolutional layer is connected to the output of the fourth three-dimensional convolutional layer, and the third three-dimensional convolutional layer and the fourth three-dimensional convolutional layer first use the tanh activation function and then perform the convolution operation.
6. A three-dimensional biomedical image compression method according to claim 1, characterized in that: The encoding stage uses a three-dimensional entropy coding method based on a deep network. The high-frequency signals and low-frequency signals obtained by decomposition in all directions are called three-dimensional sub-bands. All sub-bands are entropy coded in sequence according to the set order. For the current 3D subband to be encoded, a context is extracted from the coded 3D subband using a 3D context coarse extraction network; the context is input into a 3D inter-subband context extraction network to obtain an inter-3D subband context; Inputting the current 3D sub-band to be encoded into a 3D sub-band context extraction network to obtain a context within the 3D sub-band; The context between the three-dimensional sub-bands and the context within the three-dimensional sub-band are input into a three-dimensional context fusion deep network to obtain entropy coding parameters of the three-dimensional sub-band to be encoded currently; and the three-dimensional sub-band to be encoded currently is entropy encoded using the entropy coding parameters.
7. A three-dimensional biomedical image compression method according to claim 6, characterized in that: The three-dimensional inter-subband context extraction network includes: a first convolution unit and a second convolution unit arranged in sequence, each convolution unit includes two three-dimensional convolution layers arranged in sequence, and a Relu function is arranged between the two three-dimensional convolution layers; the input and output of each convolution unit are connected; The three-dimensional sub-band context extraction network includes: a masked three-dimensional convolution layer, a third convolution unit and a fourth convolution unit arranged in sequence; each convolution unit includes two masked three-dimensional convolution layers arranged in sequence, and a Relu function is provided between the two masked three-dimensional convolution layers; the input of the third convolution unit includes the output of the masked three-dimensional convolution layer in front of it and the input of the three-dimensional sub-band context extraction network, and at the same time, the input of the third convolution unit is also connected to its output; the input of the fourth convolution unit is connected to the output; The three-dimensional context fusion deep network includes: a masked three-dimensional convolutional layer and three three-dimensional convolutional layers arranged in sequence, and Relu functions are arranged between the masked three-dimensional convolutional layer and the three-dimensional convolutional layer, and between adjacent three-dimensional convolutional layers.
8. A three-dimensional biomedical image compression system, characterized in that: The method according to any one of claims 1 to 7 is implemented, and the system comprises: The encoder is applied to the encoding stage; in the encoding stage, the input three-dimensional biomedical image is decomposed by using a training-based three-dimensional affine wavelet transform, and the steps include: splitting the current direction of the input three-dimensional biomedical image into odd frames and even frames, inputting the odd frames into a prediction network based on a deep network, performing difference between the prediction results and the even frames after scaling by an affine image, the difference result is a high-frequency signal, the high-frequency signal is input into an update network based on a deep network, the update results are scaled by an affine image and added to the odd frames, the addition result is a low-frequency signal, completing a decomposition process, executing the decomposition process several times, and completing the decomposition in the current direction; performing decomposition in other directions on the low-frequency signal and the high-frequency signal obtained by decomposition in the current direction; performing entropy coding using the decomposition results in all directions; The decoder is applied to the decoding stage; the decoding stage adopts the opposite process of the encoding stage to reconstruct the three-dimensional biomedical image.
9. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image dimension reduction and reconstruction method based on deep neural network
CN111009018A
AUPP091197A0