Three-dimensional sonar data compression method and system based on adaptive quantization
Through adaptive quantization and entropy coding methods, the problem of large amount of three-dimensional sonar data is solved, efficient data compression and image quality improvement is achieved, suitable for marine applications and reduce computing power consumption.
Patent Information
- Application Number
- CN202510092146.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The data volume of three-dimensional sonar is large, and traditional compression algorithms are not suitable for marine applications. The computing power consumes a lot of power based on deep learning, making it difficult to implement embedded implementation.
The three-dimensional sonar data compression method based on adaptive quantization is adopted, and the frequency domain information distribution model is obtained through 2D-DCT transformation, combined with the visual characteristics of the human eye, block activity is divided into flat areas, texture areas and edge areas, and quantization tables with different quantization factors are used for adaptive quantization, and finally a compressed code stream is generated through entropy coding.
It significantly improves image compression magnification and image quality, reduces computing power consumption, and is suitable for embedded implementation.
Smart Images

Figure CN120014077A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of three-dimensional sonar data compression methods, and in particular relates to a three-dimensional sonar data compression method and system based on adaptive quantization. Background Art
[0002] Three-dimensional image sonar can perform beamforming in space and obtain three-dimensional information of the target, providing more powerful support for ocean exploration and development. As human beings continue to deepen their research on the ocean, traditional two-dimensional image sonar can no longer meet the increasingly complex application needs, and three-dimensional image sonar has unprecedented development opportunities. Three-dimensional sonar uses a planar transducer array to receive acoustic echoes in space. When performing beamforming calculations, the amount of data in the electronic system of three-dimensional image sonar is huge because three-dimensional sonar has an additional imaging dimension compared to two-dimensional sonar, which makes it difficult to transmit data to the host computer in real time.
[0003] In order to relieve the pressure of real-time transmission of image data to the host computer through the network protocol, a compression algorithm is needed to process the image data. However, the quantization table proposed by the traditional standard protocol is based on the statistical results of a large number of natural images and is not suitable for the marine application background of 3D sonar. The adaptive data compression method based on deep learning requires a large amount of computing power, which is not conducive to the embedded implementation of 3D sonar electronic systems. Summary of the invention
[0004] The present invention aims to solve the technical problems existing in the background technology and provides a three-dimensional sonar data compression method and system based on adaptive quantization. The data compression method can obtain a frequency domain information distribution model of a pixel block through 2D-DCT transformation, and classify the pixel blocks according to the frequency domain information distribution model, complete adaptive quantization of different types of pixel blocks, and perform entropy coding on the quantized data to obtain a final compressed code stream, thereby improving the image compression ratio and image quality.
[0005] In order to solve the technical problem, the technical solution of the present invention is:
[0006] A three-dimensional sonar data compression method based on adaptive quantization, the method comprising:
[0007] S1: Use the wavelet transform algorithm to separate the high and low frequency sub-bands of the normalized three-dimensional sonar image, compress the data dynamic range of the high and low frequency sub-bands respectively, obtain the reconstructed image and perform image segmentation;
[0008] S2: Perform 2D-DCT transformation on the segmented image to obtain the information distribution model in the image frequency domain;
[0009] S3: Based on the information distribution model in the image frequency domain, the block activity of each pixel block combined with the visual characteristics of the human eye is counted, and the upper and lower thresholds of block classification are determined according to the average activity of all blocks, and the pixel blocks are divided into flat areas, texture areas and edge areas;
[0010] S4: using quantization tables with different quantization factors for different blocks to complete the adaptive quantization function;
[0011] S5: Perform entropy coding on the adaptively quantized image. First, differentially encode the DC coefficient to obtain an intermediate result. Then, run-length encode the AC coefficient to obtain an intermediate result. Then, entropy code the two intermediate results and concatenate them to obtain the final compressed bit stream result.
[0012] Further, in step S1, the data dynamic range compression includes:
[0013] S101: Perform global normalization calculation on image data:
[0014] U norml =(UU min ) / (U max -U min )
[0015] Among them, U represents the pixel value of each point in a frame of 3D image, U max and U min are the maximum and minimum values of all pixel values in a frame of 3D image, and the result of pixel value normalization is U norml ;
[0016] S102: Processing U using wavelet transform norml , and get the high and low frequency subbands:
[0017] DWT(U normal )=[LL,HB]
[0018] Among them, LL is a low-frequency subband, LH, HL and HH are all high-frequency subbands, representing horizontal, vertical and diagonal details respectively. For convenience of representation, HB = {LH, HL, HH} represents the set of all high-frequency subbands;
[0019] S103: A linear mapping algorithm is used for the low-frequency sub-band, a gamma transform algorithm is used for the high-frequency sub-band, and an inverse wavelet transform is used to splice the transformed high- and low-frequency sub-bands back into the image data:
[0020]
[0021] U result =IDWT(LL enh ,HB enh )
[0022] Among them, LL enh is the result of contrast enhancement transformation of low-frequency subband, α is the gain factor, μ LL is the mean of the low-frequency subband, ensuring that the overall brightness of the image is maintained after the transformation. γ is the gamma factor. The processed different subbands are subjected to IDWT inverse wavelet transform to obtain the reconstructed image U after the dynamic range of the data is compressed. result .
[0023] Further, in step S2, the 2D-DCT transformation includes:
[0024] Use the 2D-DCT algorithm to get the frequency domain information distribution model of the pixel block:
[0025]
[0026] Among them, i, j represent the horizontal and vertical coordinates of a point in the image data, f(i, j) represents the pixel value of the point in the image data, u, v represent the horizontal and vertical coordinates of a point after 2D-DCT transformation, and F(u, v) represents the value of the matrix after 2D-DCT transformation.
[0027] Furthermore, in step S3, the block activity statistics include:
[0028] Statistical AC coefficient energy expression T m,n (u,v) is:
[0029] T m,n (u,v)=F m,n (u,v) 2 (u,v)≠0
[0030] Set En (m,n) It represents the block activity coefficient after combining the visual characteristics of the human eye. The calculation formula is:
[0031]
[0032] Where ρ = 1 / F m,n (0,0) indicates unit weighting, H -1 (ω) represents the weighting function of the AC coefficient, and its calculation formula is:
[0033]
[0034] The calculation formula of the multiplication function A(ω) is as follows, σ is a constant,
[0035]
[0036] Furthermore, in step S4, the quantization table using different quantization factors for different blocks includes:
[0037] S401: Determine the small threshold for classification as T L , the classification threshold is T H , the classification criteria are as follows:
[0038]
[0039] Determine the average activity of an image Gate limit T H Select as Small threshold T L Select as
[0040] S402: The larger the quality factor QF is, the higher the image visual quality is and the lower the compression rate is. f The relationship between (u,v) and the quality factor QF is:
[0041]
[0042] in, represents the rounding down operation, Q(u,v) represents the classic quantization chart provided by the JPEG Joint Photographic Experts Group, and the quantization strategy is: for the texture area, a lower quality factor is used, for the edge area, a medium quality factor is used, and for the flat area, a higher quality factor is used, and the frequency domain information distribution model after adaptive quantization is obtained.
[0043]
[0044] Further, in step S5, entropy coding the adaptively quantized image includes:
[0045] S501: Perform differential pulse coding (DPCM) on the DC coefficient, and use the difference Diff as the data to be encoded, Diff = DC i -DC i-1 , where DC i is the DC component in the current image matrix, DC i-1 For the DC component of the previous image matrix, when Huffman coding the DC coefficient, first convert Diff into an intermediate state represented by the symbol (SSSS, AMP), where SSSS represents the binary digit of the Diff value, and AMP represents the data size, represented by the VLI variable-length integer coding table. In the VLI code, the Diff value is divided into 12 categories, and each category corresponds to a VLI code, and then the calculated intermediate state (SSSS, AMP) is Huffman coded;
[0046] S502: Run-length encoding RLE is performed on the AC coefficients. The encoding result is represented by symbols. The format of the symbols is as follows: RRRR, SSSS, AMP, where RRRR and SSSS each occupy 4 bits, forming the upper 8 bits. RRRR indicates how many consecutive 0 values there are before the non-zero AC component; SSSS indicates the number of binary bits occupied by the data of the non-zero AC component; AMP occupies the lower 8 bits, indicating the value of the AC component; when there are 16 consecutive 0-value AC components, the value of the high byte (RRRR, SSSS) is (15, 0), indicating zero run length (ZRL). In the 8*8 data matrix, if the subsequent AC components are all 0 values during encoding, the run-length encoding ends and is marked as EOB (0, 0);
[0047] S503: The coding rules of the DC coefficient and the AC coefficient are that the Huffman codeword occupies the high bit and the VLI codeword occupies the low bit, and the Huffman code streams of the DC coefficient and the AC coefficient are spliced together to obtain the final compressed code stream.
[0048] A three-dimensional sonar data compression system based on adaptive quantization, the system is applied to any of the above methods, the system comprising:
[0049] Wavelet transform and dynamic range compression module: used to separate the high and low frequency sub-bands of the normalized three-dimensional sonar image using the wavelet transform algorithm, and perform data dynamic range compression on the high and low frequency sub-bands respectively to obtain the reconstructed image and perform image segmentation;
[0050] Two-dimensional discrete cosine transform module: used to perform two-dimensional discrete cosine transform 2D-DCT on the segmented image to obtain the information distribution model of the image frequency domain;
[0051] Block activity estimation and classification module: It is used to count the block activity of each pixel block based on the information distribution model in the image frequency domain after combining the visual characteristics of the human eye, and determine the upper and lower thresholds of block classification according to the average activity of all blocks, and divide the pixel blocks into flat areas, texture areas and edge areas;
[0052] Adaptive quantization module: used to use quantization tables with different quantization factors for different blocks to complete the adaptive quantization function;
[0053] Entropy coding module: used to perform entropy coding on the adaptively quantized image. First, the DC coefficient is differentially coded to obtain an intermediate result, then the AC coefficient is run-length coded to obtain an intermediate result, and then the two intermediate results are entropy coded and spliced to obtain the final compressed bit stream result.
[0054] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any one of the above-mentioned three-dimensional sonar data compression methods based on adaptive quantization is implemented.
[0055] A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the three-dimensional sonar data compression method based on adaptive quantization described in any one of the above is implemented.
[0056] Compared with the prior art, the advantages of the present invention are:
[0057] The adaptive three-dimensional sonar data compression method of the present invention achieves a better compression effect by classifying different pixel blocks, and significantly improves the image compression ratio. In the adaptive data compression method of the present invention, different pixel blocks are classified in combination with the human eye data characteristic function, and the compressed image quality is more suitable for human eye viewing. Compared with other compression algorithms based on deep learning, the adaptive three-dimensional sonar data compression method of the present invention requires less computing power and is easy to implement in embedded systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 , a flowchart of the adaptive quantized three-dimensional sonar data compression method in a specific implementation manner.
[0059] FIG2 is a comparison diagram of data before and after data compression in a specific implementation manner, wherein (a) is the original image and (b) is the reconstructed image.
[0060] Figure 3 , a schematic diagram of 2D-DCT transformation in a specific implementation method, wherein (a) is the original image and (b) is a schematic diagram of the frequency domain information distribution of the image.
[0061] Fig. 4 is a schematic diagram of the results of activity classification in a specific implementation manner, wherein (a) is the original image, (b) is the flat area, (c) is the texture area, and (d) is the edge area.
[0062] Figure 5 , a schematic diagram of entropy coding processing in a specific implementation method.
[0063] FIG6 is a visual comparison diagram of an image before and after adaptive quantization in a specific implementation manner, wherein (a) is a sonar sector diagram before compression, and (b) is a sonar sector diagram after adaptive quantization. DETAILED DESCRIPTION
[0064] The specific implementation mode of the present invention is described below in conjunction with embodiments:
[0065] It should be noted that the structures, proportions, sizes, etc. shown in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them, and are not used to limit the conditions under which the present invention can be implemented. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0066] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" cited in this specification are only for the convenience of description and are not used to limit the scope of implementation of the present invention. Changes or adjustments to their relative relationships should be regarded as the scope of implementation of the present invention without substantially changing the technical content.
[0067] Embodiment 1:
[0068] See attached Figure 1 In some specific implementations, the adaptive three-dimensional sonar data compression method proposed in the present invention comprises the following steps:
[0069] S1. Data normalization and wavelet transform
[0070] Read the 3D sonar image data and normalize each pixel value so that its range is mapped between 0 and 1. Use wavelet transform on the normalized image to decompose the image into a low-frequency part (mainly containing the overall outline information of the image) and a high-frequency part (containing the detail information in the horizontal, vertical and diagonal directions). Linear mapping is performed on the low-frequency part to enhance the contrast; gamma transform is used to enhance the details of the high-frequency part. The enhanced low-frequency and high-frequency parts are synthesized into a new image through inverse wavelet transform to complete dynamic range compression.
[0071] S2.2D-DCT Transform
[0072] The compressed image is divided into 8×8 pixel blocks. A two-dimensional discrete cosine transform (2D-DCT) is used for each pixel block to convert the pixel value from the spatial domain to the frequency domain, and the frequency information distribution of each pixel block is obtained.
[0073] S3. Block activity statistics and classification
[0074] The energy value in the frequency information of each pixel block is counted to reflect the richness of the details in the block. The activity of each pixel block is weighted according to the visual characteristics. The pixel blocks are divided into three categories: flat area (less details), texture area (medium details), and edge area (edges and changes are more obvious). The classification is based on the size of the activity value, and three thresholds are set: low, medium, and high.
[0075] S4. Adaptive Quantization
[0076] According to the classification results of each pixel block, different quantization strategies are used for different types of areas: a high quality factor is used in the flat area to keep more details; a low quality factor is used in the texture area to improve the compression rate; a medium quality factor is used in the edge area to balance details and compression rate.
[0077] Use the JPEG standard quantization table for adjustment and divide the frequency information by the corresponding quantization factor and round it off.
[0078] S5. Entropy Coding
[0079] The frequency information of each pixel block is encoded. First, the low-frequency coefficient (ie, the DC component) is encoded using a differential encoding method, and only the difference between the current and previous low-frequency coefficients is stored.
[0080] The run-length coding method is used for high-frequency coefficients (i.e., AC components), and continuous zero values are counted and non-zero values are encoded.
[0081] The low-frequency and high-frequency encoding results are compressed using Huffman coding.
[0082] The final bit streams of low frequency and high frequency are stitched together to generate compressed image data.
[0083] Summary: The entire compression process starts from the original three-dimensional sonar image, and goes through normalization, wavelet transform, blocking, frequency domain conversion, activity calculation, adaptive quantization and entropy coding in sequence, and finally generates an efficient compressed bit stream while maintaining the image quality perceived by the human eye as much as possible.
[0084] The specific steps include:
[0085] S1: Traverse and scan all the pixels in the complete 3D sonar image, obtain the maximum and minimum values, and then perform normalization calculations. Then perform image segmentation on each frame slice to obtain 8*8 pixel blocks. Use wavelet transform to separate the high and low frequency sub-bands of the pixel blocks, use linear mapping algorithm for low frequency sub-bands, and use gamma transform algorithm for high frequency sub-bands. Use inverse wavelet transform to splice the transformed high and low frequency sub-bands back to image data:
[0086]
[0087] U result =IDWT(LL enh ,HB enh )
[0088] Among them, U represents the pixel value of each point in a frame of 3D image, U max and U min are the maximum and minimum values of all pixel values in a frame of 3D image, and the result of pixel value normalization is Unorml . norml The DWT wavelet transform is decomposed into four sub-bands, LL is the low-frequency sub-band, LH, HL and HH are all high-frequency sub-bands, representing horizontal, vertical and diagonal details respectively. For convenience, HB = {LH, HL, HH} is written to represent the set of all high-frequency sub-bands. enh is the result of contrast enhancement transformation of low-frequency subband, α is the gain factor, μ LL HB is the mean of the low-frequency subband, ensuring that the overall brightness of the image is maintained after the transformation. enh ={LH enh ,HL enh ,HH enh}, represents the result set after gamma transformation of high-frequency sub-bands, and γ is the gamma factor. The processed different sub-bands are subjected to IDWT inverse wavelet transform to obtain the reconstructed image U result The schematic diagram of the processing effect of this part is shown in Figure 2.
[0089] S2: After the data dynamic range compression model is processed, the frequency domain information distribution model of the pixel block is obtained using the 2D-DCT algorithm:
[0090]
[0091] Among them, i, j represent the horizontal and vertical coordinates of a point in the image data, f(i, j) represents the pixel value of the point in the image data, u, v represent the horizontal and vertical coordinates of a point after 2D-DCT transformation, and F(u, v) represents the value of the matrix after 2D-DCT transformation. The effect comparison diagram of this part is as follows Figure 3 shown.
[0092] S3: Classify different pixel blocks according to the frequency domain information distribution model of the pixel block. Statistical AC coefficient energy expression T m,n (u,v) is:
[0093] T m,n (u,v)=F m,n (u,v) 2 (u,v)≠0
[0094] Set En (m,n) It represents the block activity coefficient after combining the visual characteristics of the human eye. The calculation formula is:
[0095]
[0096] Where ρ = 1 / F m,n (0,0) indicates unit weighting, H -1 (ω) represents the weighting function of the AC coefficient, and its calculation formula is:
[0097]
[0098] The calculation formula of the multiplication function A(ω) is as follows, σ is a constant,
[0099] Assume that the minimum threshold for classification is T L , the classification threshold is T H , the classification criteria are as follows:
[0100]
[0101] Assuming the average activity of the image The gate limit is selected as The small threshold is selected as The effect of block classification on an image is shown in Figure 4.
[0102] S4: Different quality factors are used for different types of image blocks. The larger the quality factor QF, the higher the image visual quality and the lower the compression rate. f The relationship between (u,v) and the quality factor QF is:
[0103]
[0104] in, represents the rounding down operation, Q(u,v) represents the classic quantization chart provided by the JPEG Joint Photographic Experts Group. The quantization strategy is: for texture areas, a lower quality factor is used, for edge areas, a medium quality factor is used, and for flat areas, a higher quality factor is used. This strategy can better ensure that the image is not distorted in the human eye and improve the compression rate as much as possible. The frequency domain information distribution model after adaptive quantization is obtained
[0105]
[0106] S5: Perform entropy coding according to the frequency domain information distribution model after adaptive quantization. The preprocessing of DC coefficient is differential pulse coding (DPCM). Use the difference Diff as the data to be coded, Diff = DC i -DC i-1 , where DC i is the DC component in the current image matrix, DC i-1 is the DC component of the previous image matrix.
[0107] S51: When Huffman coding the DC coefficient, first convert Diff into an intermediate state and represent it with the symbol (SSSS, AMP). Among them, SSSS represents the binary digit of the Diff value, and AMP represents the data size, which is represented by VLI (variable length integer coding table). In the VLI code, the Diff value is divided into 12 categories, and each category (except for category 0) corresponds to a VLI code. Then the calculated intermediate state (SSSS, AMP) is Huffman coded.
[0108] S52: Run-length encoding (RLE) is used for preprocessing of AC coefficients. The encoding result is represented by symbols, and the format of the symbols is as follows (RRRR, SSSS, AMP). Among them, RRRR and SSSS each occupy 4 bits, forming the upper 8 bits, and RRRR indicates how many consecutive 0 values there are before the non-zero AC component. SSSS indicates the number of binary bits occupied by the data of the non-zero AC component. AMP occupies the lower 8 bits and indicates the value of the AC component. When there are 16 consecutive AC components with zero values, the value of the high byte (RRRR, SSSS) is (15,0), indicating the zero run length (ZRL). In the 8*8 data matrix, since there are only 63 AC components in 64 coefficients, the run-length encoding of the AC component has a maximum of three ZRLs. If the subsequent AC components are all 0 values during encoding, the run-length encoding ends and is marked as EOB (0,0). The processing diagram of this part is shown as follows. Figure 5 shown.
[0109] The coding rules of DC coefficient and AC coefficient are that Huffman codeword occupies high bit and VLI codeword occupies low bit. The Huffman code streams of DC coefficient and AC coefficient are concatenated together to obtain the final compressed code stream.
[0110] Embodiment 2:
[0111] The image is adaptively quantized and compressed through the entire process of the specific implementation method. The test data used is a sonar image of a real cubic target in a pool. The final image obtained after compression and the original image before compression are shown in Figure 6. Figure 6 (b) is the final compression result. It can be seen that the algorithm in this paper has no distortion in human vision, and has good detail retention in the texture without obvious block effect.
[0112] The adaptive compression algorithm is evaluated using the following evaluation metrics: compression ratio (CR), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM).
[0113]
[0114] Among them, S ori is the size of the original image, S comis the size of the compressed image, both in bytes. MAX is the maximum pixel value that may appear in the image. For an 8-bit grayscale image, this value is 255. MSE is the mean square error of the image. The larger the PSNR, the better the image quality. 30dB is considered to be a good image quality. μ x and μ y Represent the pixel sample means of images x and y respectively, and represents the variance of images x and y, σ xy Represents the covariance of two images. C1 and C2 are constants. The closer the SSIM value is to 1, the higher the structural similarity between the compressed image and the original image, and the better the image quality.
[0115] According to the above evaluation indicators, two static quantization algorithms: distortion-free quantization and joint image expert group quantization, and two dynamic quantization algorithms: a segmented quantization table based on noise statistics and a quantization table of the present invention are used for a total of four comparative experiments. The results are shown in Table 1 below, where ARACATI2017 and UATD are both open source datasets.
[0116] Table 1 - Performance comparison of different quantization methods
[0117]
[0118] Compared with the other three algorithms, the present invention has more advantages in terms of structural similarity and increases the compression ratio by 12.9% under the condition that the PSNR remains almost unchanged.
[0119] Embodiment 3:
[0120] This embodiment provides a terminal device, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions; the processor described in the embodiment of the present invention can be used for the operation of a three-dimensional sonar data compression method based on adaptive quantization, including the following steps:
[0121] S1: Use the wavelet transform algorithm to separate the high and low frequency sub-bands of the normalized three-dimensional sonar image, compress the data dynamic range of the high and low frequency sub-bands respectively, obtain the reconstructed image and perform image segmentation;
[0122] S2: Perform 2D-DCT transformation on the segmented image to obtain the information distribution model in the image frequency domain;
[0123] S3: Based on the information distribution model in the image frequency domain, the block activity of each pixel block combined with the visual characteristics of the human eye is counted, and the upper and lower thresholds of block classification are determined according to the average activity of all blocks, and the pixel blocks are divided into flat areas, texture areas and edge areas;
[0124] S4: using quantization tables with different quantization factors for different blocks to complete the adaptive quantization function;
[0125] S5: Perform entropy coding on the adaptively quantized image. First, differentially encode the DC coefficient to obtain an intermediate result. Then, run-length encode the AC coefficient to obtain an intermediate result. Then, entropy code the two intermediate results and concatenate them to obtain the final compressed bit stream result.
[0126] Embodiment 4:
[0127] This embodiment provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.
[0128] One or more instructions stored in a computer-readable storage medium may be loaded and executed by a processor to implement the corresponding steps of a three-dimensional sonar data compression method based on adaptive quantization in the above embodiment; one or more instructions in the computer-readable storage medium may be loaded and executed by a processor as follows:
[0129] S1: Use the wavelet transform algorithm to separate the high and low frequency sub-bands of the normalized three-dimensional sonar image, compress the data dynamic range of the high and low frequency sub-bands respectively, obtain the reconstructed image and perform image segmentation;
[0130] S2: Perform 2D-DCT transformation on the segmented image to obtain the information distribution model in the image frequency domain;
[0131] S3: Based on the information distribution model in the image frequency domain, the block activity of each pixel block combined with the visual characteristics of the human eye is counted, and the upper and lower thresholds of block classification are determined according to the average activity of all blocks, and the pixel blocks are divided into flat areas, texture areas and edge areas;
[0132] S4: using quantization tables with different quantization factors for different blocks to complete the adaptive quantization function;
[0133] S5: Perform entropy coding on the adaptively quantized image. First, differentially encode the DC coefficient to obtain an intermediate result. Then, run-length encode the AC coefficient to obtain an intermediate result. Then, entropy code the two intermediate results and concatenate them to obtain the final compressed bit stream result.
[0134] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0136] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0138] The preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
[0139] Many other changes and modifications may be made without departing from the concept and scope of the present invention.It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.
Claims
1. A three-dimensional sonar data compression method based on adaptive quantization, characterized in that: The method comprises: S1: Use the wavelet transform algorithm to separate the high and low frequency sub-bands of the normalized three-dimensional sonar image, compress the data dynamic range of the high and low frequency sub-bands respectively, obtain the reconstructed image and perform image segmentation; S2: Perform 2D-DCT transformation on the segmented image to obtain the information distribution model in the image frequency domain; S3: Based on the information distribution model in the image frequency domain, the block activity of each pixel block combined with the visual characteristics of the human eye is counted, and the upper and lower thresholds of block classification are determined according to the average activity of all blocks, and the pixel blocks are divided into flat areas, texture areas and edge areas; S4: using quantization tables with different quantization factors for different blocks to complete the adaptive quantization function; S5: Perform entropy coding on the adaptively quantized image. First, differentially encode the DC coefficient to obtain an intermediate result. Then, run-length encode the AC coefficient to obtain an intermediate result. Then, entropy code the two intermediate results and concatenate them to obtain the final compressed bit stream result.
2. The three-dimensional sonar data compression method based on adaptive quantization according to claim 1, characterized in that: In the step S1, the data dynamic range compression includes: S101: Perform global normalization calculation on image data: IN norml =(UU min ) / (IN max -IN min ) Among them, U represents the pixel value of each point in a frame of 3D image, U max and U min They are the maximum and minimum values of all pixel values in a frame of 3D image, and the result of pixel value normalization is U norml ; S102: Processing U using wavelet transform norml , and get the high and low frequency subbands: DWT(U normal )=[LL,HB] Among them, LL is a low-frequency subband, LH, HL and HH are all high-frequency subbands, representing horizontal, vertical and diagonal details respectively. For convenience of representation, HB = {LH, HL, HH} represents the set of all high-frequency subbands; S103: A linear mapping algorithm is used for the low-frequency sub-band, a gamma transform algorithm is used for the high-frequency sub-band, and an inverse wavelet transform is used to splice the transformed high- and low-frequency sub-bands back into the image data: U result =IDWT(LL enh ,HB enh ) Among them, LL enh is the result of contrast enhancement transformation of low-frequency subband, α is the gain factor, μ LL is the mean of the low-frequency subband, ensuring that the overall brightness of the image is maintained after the transformation. γ is the gamma factor. The processed different subbands are subjected to IDWT inverse wavelet transform to obtain the reconstructed image U after the dynamic range of the data is compressed. result .
3. The three-dimensional sonar data compression method based on adaptive quantization according to claim 1, characterized in that: In step S2, the 2D-DCT transformation includes: Use the 2D-DCT algorithm to get the frequency domain information distribution model of the pixel block: Among them, i, j represent the horizontal and vertical coordinates of a point in the image data, f(i, j) represents the pixel value of the point in the image data, u, v represent the horizontal and vertical coordinates of a point after 2D-DCT transformation, and F(u, v) represents the value of the matrix after 2D-DCT transformation.
4. The three-dimensional sonar data compression method based on adaptive quantization according to claim 1, characterized in that: In step S3, the block activity statistics include: Statistical AC coefficient energy expression T m,n (u,v) is: T m,n (u,v)=F m,n (u,v) 2 (u,v)≠0 Set En (m,n) It represents the block activity coefficient after combining the visual characteristics of the human eye. The calculation formula is: Where ρ = 1 / F m,n (0,0) indicates unit weighting, H -1 (ω) represents the weighting function of the AC coefficient, and its calculation formula is: The calculation formula of the multiplication function A(ω) is as follows, σ is a constant, 5. The three-dimensional sonar data compression method based on adaptive quantization according to claim 1, characterized in that: In step S4, the quantization table using different quantization factors for different blocks includes: S401: Determine the small threshold for classification as T L , the classification threshold is T H , the classification criteria are as follows: Determine the average activity of an image Gate limit T H Select as Small threshold T L Select as S402: The larger the quality factor QF is, the higher the image visual quality is and the lower the compression rate is. f The relationship between (u,v) and the quality factor QF is: in, represents the rounding down operation, Q(u,v) represents the classic quantization chart provided by the JPEG Joint Photographic Experts Group, and the quantization strategy is: for the texture area, a lower quality factor is used, for the edge area, a medium quality factor is used, and for the flat area, a higher quality factor is used, and the frequency domain information distribution model after adaptive quantization is obtained.
6. The three-dimensional sonar data compression method based on adaptive quantization according to claim 1, characterized in that: In step S5, entropy coding the adaptively quantized image includes: S501: Perform differential pulse coding (DPCM) on the DC coefficient, and use the difference Diff as the data to be encoded, Diff = DC i -DC i-1 , where DC i is the DC component in the current image matrix, DC i-1 For the DC component of the previous image matrix, when performing Huffman coding of the DC coefficient, Diff is first converted into an intermediate state represented by the symbol (SSSS, AMP), where SSSS represents the binary digit of the Diff value, and AMP represents the data size, represented by the VLI variable-length integer coding table. In the VLI code, the Diff value is divided into 12 categories, and each category corresponds to a VLI code, and then the calculated intermediate state (SSSS, AMP) is processed by Huffman coding; S502: Run-length encoding RLE is performed on the AC coefficients. The encoding result is represented by symbols. The format of the symbols is as follows: RRRR, SSSS, AMP, where RRRR and SSSS each occupy 4 bits, forming the upper 8 bits. RRRR indicates how many consecutive 0 values there are before the non-zero AC component; SSSS indicates the number of binary bits occupied by the data of the non-zero AC component; AMP occupies the lower 8 bits, indicating the value of the AC component; when there are 16 consecutive 0-value AC components, the value of the high byte (RRRR, SSSS) is (15, 0), indicating zero run length (ZRL). In the 8*8 data matrix, if the subsequent AC components are all 0 values during encoding, the run-length encoding ends and is marked as EOB (0, 0); S503: The coding rules of the DC coefficient and the AC coefficient are that the Huffman codeword occupies the high bit and the VLI codeword occupies the low bit, and the Huffman code streams of the DC coefficient and the AC coefficient are spliced together to obtain the final compressed code stream.
7. A three-dimensional sonar data compression system based on adaptive quantization, characterized in that: The system is applied to the method described in any one of claims 1 to 6, and the system comprises: Wavelet transform and dynamic range compression module: used to separate the high and low frequency sub-bands of the normalized three-dimensional sonar image using the wavelet transform algorithm, and perform data dynamic range compression on the high and low frequency sub-bands respectively to obtain the reconstructed image and perform image segmentation; Two-dimensional discrete cosine transform module: used to perform two-dimensional discrete cosine transform 2D-DCT on the segmented image to obtain the information distribution model of the image frequency domain; Block activity estimation and classification module: It is used to count the block activity of each pixel block based on the information distribution model in the image frequency domain after combining the visual characteristics of the human eye, and determine the upper and lower thresholds of block classification according to the average activity of all blocks, and divide the pixel blocks into flat areas, texture areas and edge areas; Adaptive quantization module: used to use quantization tables with different quantization factors for different blocks to complete the adaptive quantization function; Entropy coding module: used to perform entropy coding on the adaptively quantized image. First, the DC coefficient is differentially coded to obtain an intermediate result, then the AC coefficient is run-length coded to obtain an intermediate result, and then the two intermediate results are entropy coded and spliced to obtain the final compressed bit stream result.
8. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a three-dimensional sonar data compression method based on adaptive quantization according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements a three-dimensional sonar data compression method based on adaptive quantization according to any one of claims 1 to 6.