Data compression method, device, and storage medium
Through a binary compression method, the hexadecimal data of the medical image file is converted into a set of index numbers by using the training data list, and the index numbers are allocated by replacing the maximum index number and probability analysis, the problems of insufficient compression ratio and high computing resource consumption in the prior art are solved, and efficient medical image file compression is achieved.
Patent Information
- Application Number
- CN202111191101.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-10-12
AI Technical Summary
The existing medical image file compression method sacrifices a huge compression ratio margin to maintain bit-to-bit data after decompression, resulting in the inability to effectively save disk space and requires a large amount of computing resources during the compression process, resulting in low compression speed and high energy consumption.
Through a binary compression-based method, the hexadecimal data of medical image files is converted into a set of index numbers using a training data inventory, and the allocation of index numbers is improved by replacing the maximum index number to reduce storage space, combined with probability analysis, to improve compression efficiency.
It achieves a low compression ratio without losing visual quality, enabling the ability to compress medical image files to about 6% of the original file size, and significantly reduces the consumption of computing resources.
Smart Images

Figure CN114036323B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data compression method and device, and a storage medium. Background Art
[0002] Medical image files (typically Digital Imaging and Communications in Medicine (DICOM) and Neuroimaging Informatics Technology Initiative (NIFTI)) are structured as raw image files (RAW), which are usually compressed using a mathematical (totally) lossless compression method (TLC).
[0003] Most TLC methods provide maximum performance results that do not exceed about 25% compression ratio. For example, JPEGLS and JPEG2K (JPEG2000) provide an average compression ratio of 30% to 25%.
[0004] Despite these techniques, as well as fully lossless compression techniques with message integrity check (MIC), obtaining an exact bit-for-bit equivalent image file of the original medical image file after decompression, binary data is usually useless to healthcare workers, such as doctors, because their ultimate requirement is the same medical image that is visually equivalent to the original medical image, rather than the binary data in the original medical image.
[0005] Current medical image file compression methods sacrifice a huge compression ratio margin in order to obtain bit-for-bit data after decompression, making it impossible for compression software to save more advantageous disk space through compression.
[0006] In addition, since current compression methods require a lot of computation to compress binary data, a lot of computing resources, such as central processing unit (CPU) and random access memory (RAM), are consumed during the compression process. This results in lower compression speed and higher energy consumption in medical image data centers. Summary of the invention
[0007] The present application provides a data compression method and device, and a storage medium, which provide a lower compression ratio without losing any visual quality.
[0008] In a first aspect, a data compression method is provided, the method comprising:
[0009] According to the training data list, one byte of hexadecimal data of the first image file is saved as a first index number set and a second index number set;
[0010] Replace the second index number with a value corresponding to a difference between any second index number in the second index number set and the minimum index number in the second index number set;
[0011] The first index number set and the replaced second index number set are saved.
[0012] Optionally, the first index number in the first index number set is at least a two-digit index number, and the second index number in the second index number set is a one-digit index number.
[0013] Optionally, the method further comprises:
[0014] Saving one byte of binary data of the second image file into two hexadecimal files;
[0015] Get the type of the hexadecimal value in each of the two hexadecimal files;
[0016] Get the probability of each type of hexadecimal value appearing in the two hexadecimal files;
[0017] According to the probability, assigning a corresponding index number to each hexadecimal value in the two hexadecimal files;
[0018] The second image file is saved as a third index number set and a fourth index number set.
[0019] Optionally, the method further comprises:
[0020] Acquire binary data of the second image file;
[0021] extracting a header of the second image file;
[0022] The step of saving one byte of binary data of the second image file into two hexadecimal files comprises:
[0023] The binary data of the message body of the second image file except the header is saved as the two hexadecimal files.
[0024] Optionally, the method further comprises:
[0025] Training is performed on the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
[0026] In a second aspect, a data compression device is provided, the device comprising:
[0027] A first saving unit, used for saving one byte of hexadecimal data of the first image file as a first index number set and a second index number set according to the training data list;
[0028] a replacement unit, configured to replace the second index number with a value corresponding to a difference between any second index number in the second index number set and the minimum index number in the second index number set;
[0029] The second storage unit is used to store the first index number set and the replaced second index number set.
[0030] Optionally, the first index number in the first index number set is an index number of at least two digits, and the second index number in the second index number set is an index number of one digit.
[0031] Optionally, the device further comprises:
[0032] A third saving unit, used for saving one byte of binary data of the second image file into two hexadecimal files;
[0033] A first acquiring unit, configured to acquire a type of a hexadecimal value in each of the two hexadecimal files;
[0034] A second obtaining unit, used for obtaining the probability of each type of hexadecimal value appearing in the two hexadecimal files;
[0035] an allocating unit, configured to allocate a corresponding index number to each hexadecimal value in the two hexadecimal files according to the probability;
[0036] The fourth saving unit is used to save the second image file as a third index number set and a fourth index number set.
[0037] Optionally, the device further comprises:
[0038] A third acquiring unit, configured to acquire binary data of the second image file;
[0039] an extraction unit, configured to extract a header of the second image file;
[0040] The third saving unit is used to save the binary data of the message body of the second image file except the header as the two hexadecimal files.
[0041] Optionally, the device further comprises:
[0042] A training unit is used to train the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
[0043] In a third aspect, a data compression device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in the first aspect or any one of the implementations of the first aspect is implemented.
[0044] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect or any one of the implementations of the first aspect is implemented.
[0045] The data compression scheme of the present application has the following beneficial effects:
[0046] A binary compression method is used to provide a high-end compression ratio, which can ensure the exact quality of the original medical image after compression, and a very low compression ratio, which can compress medical image files to about 6% of the original file size without any loss of visual quality after decompression. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 A flowchart of a data compression method provided in an embodiment of the present application;
[0049] Figure 2 A flowchart of another data compression method provided in an embodiment of the present application;
[0050] Figure 3 A schematic diagram of a data compression training example of an embodiment of the present application;
[0051] Figure 4 A schematic diagram of data compression according to an example of an embodiment of the present application;
[0052] Figure 5 A schematic diagram of the structure of a data compression device provided in an embodiment of the present application;
[0053] Figure 6 A schematic diagram of the structure of another data compression device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0055] The present application provides a different method to compress medical image files, namely, using a binary compression method to provide a high-end compression ratio, which can not only ensure the accurate quality of the original medical image after compression, but also provide a very low compression ratio, which can compress the medical image file to about 6% of the original file size without losing any visual quality after decompression.
[0056] Medical image files, which can be in the digital imaging and communication in medicine (DICOM) or NIFTI (neuroimaging informatics technology initiative) format, will display accurate pixel-to-pixel visual images, just as they were in the original medical image files before compression. Through a variety of image analysis methods such as structural similarity (SSIM) comparison, the decompressed medical images can be visualized and analyzed, proving that the decompressed medical image files have exactly the same visual quality as the corresponding original medical image files.
[0057] This application describes a solution to compress medical image files to approximately 6% of the original file size, which after decompression will produce exactly the same visual quality as the original file.
[0058] This method uses binary compression technology and does not analyze the pixel and grayscale data of DICOM files as visual units. Instead, the pixel data and grayscale data are exported as binary values and analyzed using the most efficient calculations to save a lot of computing power.
[0059] This approach, referred to here as medical image replication (MIR), is primarily based on data models trained by artificial intelligence (AI) algorithms.
[0060] like Figure 1 FIG. 1 is a flow chart of a data compression method provided in an embodiment of the present application, the method comprising the following steps:
[0061] S101. Save the hexadecimal data of one byte of the first image file as the first index number set and the second index number set according to the training data list.
[0062] S102. Replace the second index number in any of the second index number sets with the value corresponding to the difference between the second index number and the smallest index number in the second index number set.
[0063] S103. Save the first index number set and the replaced second index number set.
[0064] Store the hexadecimal values of one byte one by one into a new file, which is actually a compressed file (the output result of the compression process):
[0065] Step S101: Given the hexadecimal values in the screenshot as Figure 4 shown, save the hexadecimal value of one byte as an index, either a one - digit or two - digit index number:
[0066] DD => index 85 00 => index 5
[0067] E4 => index 80 00 => index 5
[0068] EF => index 85 00 => index 5
[0069] E7 => index 81 00 => index 5
[0070] D3 => index 55 00 => index 5
[0071] C1 => index 50 00 => index 5
[0072] BF => index 55 00 => index 5
[0073] C8 => index 56 00 => index 5
[0074] Step S102: Replace the largest index number with the value corresponding to the difference between the largest index number and the smallest index number. For example, if there are 85 index numbers (DD, E4, EF, E7, D3, C1, BF, C8,....) in file 1 and the smallest index number is 50, then 85 - 50 = 35.
[0075] In this way, the largest index number (85) is replaced by 35, and the index numbers are sorted in such a way that the smallest index number is 0.
[0076] This enables saving more disk space by storing smaller index numbers.
[0077] Step S103: Each hexadecimal value of the first byte is replaced by the modified index number representing them, and the storage method is as follows:
[0078] DD=>35
[0079] 00=>5
[0080] E4=>30
[0081] 00=>5
[0082] EF=>35
[0083] 00=>5
[0084] E7=>31
[0085] 00=>5
[0086] D3=>20
[0087] 00=>5
[0088] C1=>15
[0089] 00=>5
[0090] BF=>20
[0091] 00=>5
[0092] C8=>21
[0093] 00=>5
[0094] The largest index value in the first byte (DD, E4, EF, ....) is 35, which is a 6-digit number. The largest index value in the second byte (00, 00, 00, 00, ....) is 5, which is a 3-digit number.
[0095] This way, the first 1-byte hexadecimal value will be saved with only 6 bits, and the second 1-byte hexadecimal value will be saved with only 3 bits, which means that the 16-bit data is saved in a data block of only 9 bits. This will automatically save about 45% of the disk space, which means the compression ratio will be about 55%.
[0096] Now that the original data has been compressed, but its redundancy has not been significantly reduced, another compression algorithm (patents have been obtained / submitted) will be used to compress the file that still has a lot of redundancy. This method includes:
[0097] Ⅰ) Redundancy Generator
[0098] II) Probabilistic Predictor
[0099] III) Batch of data analyzers of the same type
[0100] Since the input file has been halved, the compression process can be started using the three methods described above, compressing data to approximately 12% of the input data. This enables the creation of a compression ratio of approximately 6% over the original medical image file.
[0101] Unzip:
[0102] Referring to the data collected by TDI, changing step 4 of the compression engine to step 1 will obtain the original image. The binary data lost during the compression process includes:
[0103] Extra digits that are not considered to be an integer. For example, in step 8 of TDI, we have:
[0104] FE=>1.125% (index: 2)
[0105] Replace the probability number 1.125% with 1.1%. The remaining numbers (2 and 5) are lost.
[0106] This will result in no quality issues in the final image, since the probability quantities are always distributed in the MIFs and the corresponding values are distributed in the binary data of the MIFs and will not be affected by simplifying the quantity to only one number.
[0107] According to the flow chart of a data compression method provided in the present application, a binary compression method is used to provide a high-end compression ratio, which can not only ensure the accurate quality of the compressed original medical image, but also provide a very low compression ratio, which can compress the medical image file to about 6% of the original file size without losing any visual quality after decompression.
[0108] like Figure 2 FIG. 1 is a flow chart of a data compression method provided in an embodiment of the present application, the method comprising the following steps:
[0109] S201. Obtain binary data of the second image file.
[0110] S202. Extract the header of the second image file.
[0111] S203. Save the binary data of the message body of the second image file except the header as the two hexadecimal files.
[0112] S204. Save one byte of binary data of the second image file as two hexadecimal files.
[0113] S205. Obtain the type of the hexadecimal value in each of the two hexadecimal files.
[0114] S206. Obtain the probability of each type of hexadecimal value appearing in the two hexadecimal files.
[0115] S207. According to the probability, assign a corresponding index number to each hexadecimal value in the two hexadecimal files.
[0116] S208. Save the second image file as a third index number set and a fourth index number set.
[0117] S209. Train the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
[0118] S210. According to the training data list, save one byte of hexadecimal data of the first image file as a first index number set and a second index number set.
[0119] S211. Replace the second index number with a value corresponding to the difference between any second index number in the second index number set and the minimum index number in the second index number set.
[0120] S212. Save the first index number set and the replaced second index number set.
[0121] The existing trained data inventories (TDIs) that underlie this approach minimize the computational power of any new medical image file.
[0122] MIR is divided into two algorithms:
[0123] I) Trained data inventory (TDI)
[0124] II) Compression Engine
[0125] Method Description:
[0126] I) TDI
[0127] Introduction:
[0128] The training data inventory is responsible for collecting analysis data obtained from thousands of DICOM or NIFTI files, here called medical image files (MIFs).
[0129] The working of TDI is as follows:
[0130] Receive a single MIF.
[0131] S201: Read binary data of MIF. The binary data can be read as hexadecimal (HEX) equivalent data of the binary data.
[0132] S202: Extract the header of the MIF. The header of the MIF contains metadata of the MIF, including patient information and instruments used by the MIF. The header of the file is usually in the range of several kilobytes in size, and this range varies from file to file.
[0133] Save the extracted MIF header as a separate binary file.
[0134] S203: Divide the remaining binary data (the MIF message body) into two parts, each part represents one byte, and obtain two independent binary files: one composed of single-byte data with odd indexes, and the other composed of single-byte data with even indexes.
[0135] This is achieved by storing the first, third, fifth, seventh, .... byte of data in one binary file, and storing the second, fourth, sixth, eighth, tenth, .... byte of data in another binary file.
[0136] Continue this step for all remaining bytes.
[0137] Example: Keep the HEX values in a rectangular box in one file. The rest of the HEX values in another file.
[0138] S204: Perform the following analysis and research on these two files respectively:
[0139] Determine how many types of values there are per file. For example, considering a file that displays Figure 4 The hexadecimal values of the files in the rectangular box in the diagram of the compressed data training are shown. There are 5 types of hexadecimal values:
[0140] (00,01,02,03,04) in the selected data:
[0141] 00,00,00,00,00,00,00,00,00,00,00,01,01,……,02,03,03,04,....
[0142] Step 205: Generate probability lists for the following two files:
[0143] File 1: [ie 1.bin]
[0144] 00=>20%
[0145] 01=>12.5%
[0146] 02=>8%
[0147] 03=>5%
[0148] 04=>2%
[0149] …=>…%
[0150] A1=>3.5%
[0151] A2=>4.5%
[0152] A3=>5.5%
[0153] A4=>19%
[0154] A5=>0.5%
[0155] …=>…%
[0156] B1=>2%
[0157] B2=>3%
[0158] B3=>3.5%
[0159] B3=>4.5%
[0160] B4=>5.5%
[0161] B5=>11.5%
[0162] …=>…%
[0163] File 2: [ie 2.bin]
[0164] 00=>25%
[0165] 01=>5%
[0166] 02=>12.5%
[0167] 03=>20%
[0168] 04=>30%
[0169] 05=>1%
[0170] …=>…%
[0171] Considering that file 2 is the representation of the second 1-byte hexadecimal value (highlighted by the rectangle), the odd thing is that the probability distribution between the hexadecimal values corresponding to this file does not show such a large difference, but only a collection of a small number of hexadecimal values. Unlike file 1, it has a large number of hexadecimal values, each with a corresponding probability value.
[0172] S206: Create a separate index file (.bin file) for each of the two files.
[0173] The index file contains each hexadecimal value and its corresponding index number.
[0174] Index numbers are assigned in hexadecimal values from the highest digit belonging to the most probable hexadecimal value to the least probable hexadecimal value.
[0175] For example, if we assume that file 2 consists only of Figure 3 The index file corresponding to file 2 is as follows:
[0176] File 2: [ie 2.bin]
[0177] 00=>53.125% (index: 5)
[0178] 01=>40.625% (index: 4)
[0179] 02=>1.5625% (index: 2)
[0180] 03=>3.125% (index: 3)
[0181] 04=>1.5625% (index: 1)
[0182] Note: If two hexadecimal values (02 and 04 in this example) have equal probability, the software will randomly assign two consecutive index numbers between them.
[0183] Repeat steps S201 to S206 for the next MIFs. By training TDI with more medical image files, the accuracy of TDI will become higher and higher.
[0184] Step 208: Generate an overall analysis data (as a .bin file) containing the sorted probabilities of all 1-byte hexadecimal values.
[0185] The .bin file should contain results similar to those generated in step S208, except that the index numbers are generated from a set of MIFs files rather than a single medical image file.
[0186] For example, analyzing 100,000 DICOM files might return the following values:
[0187] [Index2.bin]
[0188] 00=>53.125% (index: 85)
[0189] ....=>...%(index:...)
[0190] FD=>1.5625% (index: 3)
[0191] FE=>1.125% (index: 2)
[0192] FF=>1.0625% (index: 1)
[0193] (II) Compression Engine
[0194] From the data collected from TDI, it is known how much information is usually available for each 1-byte hexadecimal value in the MIFs.
[0195] This enables it to replace 1-byte hex values, which would normally consume 8 bits of disk space per hex value, with only the single digit of the second 1-byte hex value and the two-digit number of the first 1-byte hex value (this is the result of the analysis of data training algorithms using thousands of medical image files collected).
[0196] Store the hexadecimal value of a byte one by one into a new file, which is actually the compressed file (the output of the compression process):
[0197] Step S209: Given Figure 4 The hexadecimal values in the screenshot of the data compression diagram shown are stored as a byte hexadecimal value as an index, or a one- and two-digit index number:
[0198] DD=>index 85 00=>index 5
[0199] E4=>index 80 00=>index 5
[0200] EF=>index 85 00=>index 5
[0201] E7=>index 81 00=>index 5
[0202] D3=>index 55 00=>index 5
[0203] C1=>index 50 00=>index 5
[0204] BF=>index 55 00=>index 5
[0205] C8=>index 56 00=>index 5
[0206] Step S210: Replace the maximum index number with the value corresponding to the difference between the maximum index number and the minimum index number. For example, if there are 85 index numbers (DD, E4, EF, E7, D3, C1, BF, C8, ...) in file 1 and the minimum index number is 50, then 85-50=35.
[0207] Thus, the largest index number (85) is replaced by 35, and the index numbers are sorted in such a way that the smallest index number is 0.
[0208] This enables more disk space to be saved by storing smaller index numbers.
[0209] Step S211: The hexadecimal value of each first byte is replaced by its modified index number and stored as follows:
[0210] DD => 35
[0211] 00 => 5
[0212] E4 => 30
[0213] 00 => 5
[0214] EF => 35
[0215] 00 => 5
[0216] E7 => 31
[0217] 00 => 5
[0218] D3 => 20
[0219] 00 => 5
[0220] C1 => 15
[0221] 00 => 5
[0222] BF => 20
[0223] 00 => 5
[0224] C8 => 21
[0225] 00 => 5
[0226] The largest index value in the first byte (DD, E4, EF,....) is 35, which is a 6-digit number. The largest index value in the second 1-byte (00, 00, 00, 00,....) is 5, which is a 3-digit number.
[0227] In this way, the hexadecimal value of the first 1-byte that saves only 6 bits, and the hexadecimal value of the second 1-byte that saves only 3 bits can be stored, which means that 16-bit data is stored in a data block of only 9 bits. This will automatically save approximately 45% of the disk space, which means the compression ratio will be approximately 55%.
[0228] The original data has now been compressed, but its redundancy has not been significantly reduced. Other compression algorithms (already patented / submitted for patent) will be used to compress the file that still has a large amount of redundancy. The method includes:
[0229] Ⅰ) Redundancy generator
[0230] Ⅱ) Probability predictor
[0231] III) Batch of data analyzers of the same type
[0232] Since the input file has been halved, the compression process can be started using the three methods described above, compressing data to approximately 12% of the input data. This enables the creation of a compression ratio of approximately 6% over the original medical image file.
[0233] Unzip:
[0234] Referring to the data collected by TDI, changing step 4 of the compression engine to step 1 will obtain the original image. The binary data lost during the compression process includes:
[0235] Extra digits that are not considered to be an integer. For example, in step 8 of TDI, we have:
[0236] FE=>1.125% (index: 2)
[0237] Replace the probability number 1.125% with 1.1%. The remaining numbers (2 and 5) are lost.
[0238] This will result in no quality issues in the final image, since the probability quantities are always distributed in the MIFs and the corresponding values are distributed in the binary data of the MIFs and will not be affected by simplifying the quantity to only one number.
[0239] According to the flow chart of a data compression method provided in the present application, a binary compression method is used to provide a high-end compression ratio, which can not only ensure the accurate quality of the compressed original medical image, but also provide a very low compression ratio, which can compress the medical image file to about 6% of the original file size without losing any visual quality after decompression.
[0240] It is understandable that, in order to implement the functions in the above-mentioned embodiments, the data compression device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0241] like Figure 5 FIG. 5 is a schematic diagram of a data compression device provided by the present application. The device 500 may include:
[0242] A first saving unit 501 is used to save one byte of hexadecimal data of the first image file as a first index number set and a second index number set according to the training data list;
[0243] A replacement unit 502 is used to replace the second index number with a value corresponding to a difference between any second index number in the second index number set and the minimum index number in the second index number set;
[0244] The second storage unit 503 is used to store the first index number set and the replaced second index number set.
[0245] Optionally, the first index number in the first index number set is an index number of at least two digits, and the second index number in the second index number set is an index number of one digit.
[0246] Optionally, the device further includes (indicated by dotted lines in the figure):
[0247] A third saving unit 504 is used to save one byte of binary data of the second image file into two hexadecimal files;
[0248] A first acquiring unit 505 is used to acquire the type of the hexadecimal value in each of the two hexadecimal files;
[0249] A second obtaining unit 506 is used to obtain the probability of each type of hexadecimal value appearing in the two hexadecimal files;
[0250] An allocating unit 507, configured to allocate a corresponding index number to each hexadecimal value in the two hexadecimal files according to the probability;
[0251] The fourth saving unit 508 is configured to save the second image file as a third index number set and a fourth index number set.
[0252] Optionally, the device further includes (indicated by dotted lines in the figure):
[0253] A third acquiring unit 509 is used to acquire binary data of the second image file;
[0254] An extraction unit 510, configured to extract a header of the second image file;
[0255] The third saving unit 504 is used to save the binary data of the message body of the second image file except the header as the two hexadecimal files.
[0256] Optionally, the device further includes (indicated by dotted lines in the figure):
[0257] The training unit 511 is used to train the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
[0258] It should be noted that the above units or one or more of the units can be implemented by software, hardware or a combination of the two. When any of the above units or units is implemented by software, the software exists in the form of computer program instructions and is stored in a memory, and the processor can be used to execute the program instructions and implement the above method flow. The processor can be built into a system on chip (SoC) or ASIC, or it can be an independent semiconductor chip. In addition to the core used to execute software instructions for calculation or processing in the processor, it can also further include necessary hardware accelerators, such as field programmable gate arrays (FPGA), programmable logic devices (PLD), or logic circuits that implement dedicated logic operations.
[0259] When the above units or units are implemented in hardware, the hardware can be any one or any combination of a CPU, a microprocessor, a digital signal processing (DSP) chip, a microcontroller unit (MCU), an artificial intelligence processor, an ASIC, a SoC, an FPGA, a PLD, a dedicated digital circuit, a hardware accelerator or a non-integrated discrete device, which can run the necessary software or not rely on the software to execute the above method flow.
[0260] According to a data compression device provided in an embodiment of the present application, a binary compression method is used to provide a high-end compression ratio, which can not only ensure the accurate quality of the compressed original medical image, but also provide a very low compression ratio. The medical image file can be compressed to about 6% of the original file size without losing any visual quality after decompression.
[0261] like Figure 6 FIG. 6 is a schematic diagram of another data compression device provided by the present application. The device 600 may include:
[0262] An input device 61, an output device 62, a memory 63 and a processor 64 (the number of processors 64 in the device can be one or more, Figure 6 In some embodiments of the present application, the input device 61, the output device 62, the memory 63 and the processor 64 may be connected via a bus or other means, wherein: Figure 6 The example of connecting through bus is taken in the following.
[0263] The processor 64 is used to perform the following steps:
[0264] According to the training data list, one byte of hexadecimal data of the first image file is saved as a first index number set and a second index number set;
[0265] Replace the second index number with a value corresponding to a difference between any second index number in the second index number set and the minimum index number in the second index number set;
[0266] The first index number set and the replaced second index number set are saved.
[0267] Optionally, the first index number in the first index number set is an index number of at least two digits, and the second index number in the second index number set is an index number of one digit.
[0268] Optionally, the processor 64 is further configured to perform the following steps:
[0269] Saving one byte of binary data of the second image file into two hexadecimal files;
[0270] Get the type of the hexadecimal value in each of the two hexadecimal files;
[0271] Get the probability of each type of hexadecimal value appearing in the two hexadecimal files;
[0272] According to the probability, assigning a corresponding index number to each hexadecimal value in the two hexadecimal files;
[0273] The second image file is saved as a third index number set and a fourth index number set.
[0274] Optionally, the processor 64 is configured to perform the following steps:
[0275] Acquire binary data of the second image file;
[0276] extracting a header of the second image file;
[0277] The step of saving one byte of binary data of the second image file into two hexadecimal files comprises:
[0278] The binary data of the message body of the second image file except the header is saved as the two hexadecimal files.
[0279] Optionally, the processor 64 is further configured to perform the following steps:
[0280] Training is performed on the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
[0281] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0282] According to a data compression device provided in an embodiment of the present application, a binary compression method is used to provide a high-end compression ratio, which can not only ensure the accurate quality of the compressed original medical image, but also provide a very low compression ratio. The medical image file can be compressed to about 6% of the original file size without losing any visual quality after decompression.
[0283] The method steps in the embodiments of the present application can be implemented by hardware, or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a data compression device. Of course, the processor and the storage medium can also be present in the data compression device as discrete components.
[0284] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the embodiment of the present application is executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a base station, a user device or other programmable device. The computer program or instruction may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instruction may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, for example, a floppy disk, a hard disk, a tape; it may also be an optical medium, for example, a digital video disc; it may also be a semiconductor medium, for example, a solid-state hard disk.
[0285] In the various embodiments of the present application, unless otherwise specified or provided for in any logical conflict, the terms and / or descriptions between the different embodiments are consistent and may be referenced to each other, and the technical features in the different embodiments may be combined to form new embodiments according to their inherent logical relationships.
[0286] It should be understood that in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; wherein A and B can be singular or plural. Also, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be necessarily different. Meanwhile, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0287] It is understood that the various numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic.
Claims
1. A data compression method, It is characterized in that The method comprises: According to the training data list, one byte of hexadecimal data of the first image file is saved as a first index number set and a second index number set, specifically: binary data of the first image file is read, a header of the binary data is extracted, and the remaining binary data is divided into two parts, each part represents a byte, one of the two parts is composed of single-byte data with an odd index and the other is composed of single-byte data with an even index, and the first index number set and the second index number set are determined according to the two parts; the training data list includes medical image files, and the medical image files include analysis data obtained from thousands of DICOM or NIFTI files; Replace the second index number with a value corresponding to a difference between any second index number in the second index number set and the minimum index number in the second index number set; The first index number set and the replaced second index number set are saved.
2. The method according to claim 1, It is characterized in that The first index number in the first index number set is at least a two-digit index number, and the second index number in the second index number set is a one-digit index number.
3. The method according to claim 1 or 2, It is characterized in that The method further comprises: Saving one byte of binary data of the second image file into two hexadecimal files; Get the type of the hexadecimal value in each of the two hexadecimal files; Get the probability of each type of hexadecimal value appearing in the two hexadecimal files; According to the probability, assigning a corresponding index number to each hexadecimal value in the two hexadecimal files; The second image file is saved as a third index number set and a fourth index number set.
4. The method according to claim 3, It is characterized in that The method further comprises: Acquire binary data of the second image file; extracting a header of the second image file; The step of saving one byte of binary data of the second image file into two hexadecimal files comprises: The binary data of the message body of the second image file except the header is saved as the two hexadecimal files.
5. The method according to claim 3, It is characterized in that The method further comprises: Training is performed on the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
6. A data compression device, It is characterized in that The device comprises: A first saving unit is used to save one byte of hexadecimal data of a first image file as a first index number set and a second index number set according to a training data list, specifically: reading binary data of the first image file, extracting a header of the binary data, dividing the remaining binary data into two parts, each part representing a byte, one of the two parts respectively consisting of single-byte data of odd indexes and the other consisting of single-byte data of even indexes, and determining the first index number set and the second index number set according to the two parts; the training data list includes medical image files, and the medical image files include analysis data obtained from thousands of DICOM or NIFTI files; a replacement unit, configured to replace the second index number with a value corresponding to a difference between any second index number in the second index number set and the minimum index number in the second index number set; The second storage unit is used to store the first index number set and the replaced second index number set.
7. The device according to claim 6, It is characterized in that The device also includes: A third saving unit, used for saving one byte of binary data of the second image file into two hexadecimal files; A first acquiring unit, configured to acquire a type of a hexadecimal value in each of the two hexadecimal files; A second obtaining unit, used for obtaining the probability of each type of hexadecimal value appearing in the two hexadecimal files; an allocating unit, configured to allocate a corresponding index number to each hexadecimal value in the two hexadecimal files according to the probability; The fourth saving unit is used to save the second image file as a third index number set and a fourth index number set.
8. The device according to claim 7, It is characterized in that The device also includes: A training unit is used to train the plurality of second image files in the training data list to obtain the sorting probabilities of all 1-byte hexadecimal values.
9. A data compression device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data compression method, device and equipment, and data decompression method, device and equipment
CN110958212A
Resource file packaging method and device, storage medium and computer equipment
CN111443942A