Data Compression System and Data Compression Method
The data compression system addresses the issue of suboptimal compression ratios by employing multiple irreversible compression methods selectively for different data parts, resulting in enhanced storage and transfer efficiency for large-scale IoT data.
Patent Information
- Application Number
- JP2022023881
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2042-02-18
AI Technical Summary
Existing data compression technologies often fail to achieve optimal compression ratios when a single compression method is used across different data types or regions, leading to suboptimal storage and transfer efficiency for large-scale data from IoT devices.
A data compression system that employs multiple irreversible compression methods selectively for different parts of the data, involving a first irreversible compression method for initial data processing, followed by decompression, extraction of difference information, and subsequent compression using a second irreversible method, thereby optimizing compression ratios.
This approach allows for improved compression ratios by selectively using appropriate compression techniques for each data part, outperforming single-method compression systems in terms of storage and transfer efficiency.
Smart Images

Figure 0007691383000002 
Figure 0007691383000003 
Figure 0007691383000004
Abstract
Description
Technical Field
[0001] The present invention relates to reduction of data volume.
Background Art
[0002] A storage system for reducing data volume is known (for example, Patent Document 1). Such a storage system generally reduces the data volume by compression. As one of the existing compression methods, a method is known in which a string having a high appearance frequency within a predetermined block unit is dictionary-ized and replaced with a code of a smaller size, such as the run-length method.
[0003] As a technique for reducing the data volume rather than reversible compression such as the run-length method, an irreversible compression technique is known. For example, for video data, High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC), which are standardized compression techniques, are known (hereinafter, standard codecs).
[0004] In addition, as a technique (DeepVideo Compression) for reducing the data volume of a video by a compressor and an expander configured by a Deep Neural Network (DNN), for example, there is Non-Patent Document 2.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non-Patent Documents
[0006]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] From the perspective of reducing the costs required for data storage, transfer, etc., it is considered that irreversible compression with a high compression ratio is required for the storage, transfer, etc. of large-scale data generated by IoT (Internet-of-Things) devices and the like.
[0008] However, since the optimal irreversible compression technology varies for each part of the data, there is a problem that the compression ratio is not optimal when only a single compression technology is used. For example, in intra-frame coding for video compression, the compression ratio may differ between the standard codec and Deep Video Compression for each spatial region of each frame.
[0009] This problem is not limited to the standard codec and Deep Video Compression in video data, but can occur in two or more compression technologies for various types of data.
Means for Solving the Problems
[0010] A data compression system according to an aspect of the present invention includes one or more processors and one or more storage devices. The one or more processors compress original data by a first irreversible compression method to generate first compressed data, decompress the first compressed data to generate first decompressed data, extract difference information between the original data and the first decompressed data, compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data, and store the first compressed data and the second compressed data in the one or more storage devices.
Effects of the Invention
[0011] According to one aspect of the present invention, since an appropriate compression technique can be selectively used for each part of the data, the compression ratio is improved compared to the case where only a single compression technique is used.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Modes for Carrying Out the Invention
[0013] Next, examples of this specification will be described with reference to the drawings. Note that the present invention is not limited to the examples described below.
Examples
[0014] (1-1) Outline First, the outline of Example 1 will be described with reference to FIG. 1. FIG. 1 shows the logical configuration of the system. The system includes a data generation source 100, a client 101, a compression unit 102, a decompression unit 103, a storage / communication unit 104, and a storage 105.
[0015] The data generation source 100 is the entity that generates the data to be compressed. For example, it is an image sensor that generates video data. In this example, the case where the data generation source 100 is an image sensor that generates video data will be described as an example. However, the data generation source 100 and the data it generates are not limited to this. For example, it may be an image sensor that generates still image data, a vibration sensor that generates one-dimensional time-series data, etc.
[0016] Also, the data generation source 100 is not limited to sensors and may be software such as Computer Graphics that generates video data or still image data. Further, the data generation source 100 may be data obtained by processing data generated by sensors, software, etc., such as a Segmentation Map obtained by Semantic Segmentation of each frame of video data. Also, there may be a plurality of data generation sources 100.
[0017] The compression unit 102 is a module that compresses the data generated by the data generation source 100. When the compression unit 102 receives the video data to be compressed, it inputs the frame of the data (hereinafter, the original frame) to the compressor 120 (hereinafter, compressor A) of the first compression technique, and obtains the compressed data Main stream 121 obtained as its output. The first compression technique is an irreversible compression technique. At this time, the Main stream is generated so that the bit consumption is less than when compressed to a desired image quality using only the first compression technique. The bit consumption represents the size of the compressed data, and the smaller the value, the higher the compression ratio.
[0018] The bit consumption may be reduced by any method. For example, the bit consumption may be reduced by uniformly increasing the quantization parameter (QP) throughout the entire frame. Alternatively, the compression ratios of the first compression technique and the second compression technique may be compared for each region within the frame, and in regions where the compression ratio of the second compression technique is good, the bit consumption may be reduced by increasing the QP. The second compression technique is a lossy compression technique.
[0019] The compression parameter setter 128 is a block that determines the parameters of compressor A 120, or compressor B 125, or both in the compression unit 102. For example, the compression parameter setter 128 can reduce the bit consumption by compressor A 120 throughout the entire frame by setting, as the parameter of compressor A 120, a QP obtained by adding a constant to the QP specified by the user.
[0020] Alternatively, for example, the compression parameter setter 128 can set the parameter of compressor A 120 such that the QP of compressor A 120 is increased in regions where the compression ratio of compressor B 125 is good, based on the compression ratios measured by actually compressing each patch obtained by dividing the original frame into tiles using compressor A 120 and compressor B 125. Alternatively, for example, the compression parameter setter 128 may output parameters that result in the image quality specified by the user, based on the relationship between the bit consumption and the image quality measured in advance.
[0021] In addition, for example, when compressor B 125 described later is a compressor configured by a neural network trained for each parameter of compressor A 120, the compression parameter setter 128 may set the learned parameters of the neural network of compressor B 125 corresponding to the parameters of compressor A 120.
[0022] However, the compression parameter setter 128 is not limited to these. Also, when the parameters of compressor A 120 and compressor B 125 are fixed values or when the user is allowed to specify the parameters, the compression parameter setter 128 may not be necessary.
[0023] Next, the compression unit 102 inputs the Main stream 121 into the expander 122 of the first compression technique (hereinafter referred to as expander A) to obtain an expansion frame (hereinafter referred to as the first expansion frame). Next, the compression unit 102 inputs the original frame and the first expansion frame into the second compression unit 123 to obtain the Side stream 126 which is the compressed data obtained as its output.
[0024] At this time, for the first compression technique, the bit consumption of the Side stream is controlled so as to improve the image quality in the region where the compression rate of the second compression technique is good. The control method may be any method, for example, as described later, it can be controlled by a Deep Neural Network (DNN). Or, as described above, it can be controlled by the compression parameter setter 128.
[0025] The second compression unit 123 includes an image quality improvement information extractor 124 and a compressor 125 of the second compression technique (hereinafter referred to as compressor B). The image quality improvement information extractor 124 takes the original frame and the first expansion frame as inputs and outputs data in a format that can be compressed by the compressor B125. The image quality improvement information extractor 124 outputs a new frame representing the residual between the two by, for example, subtracting the first expansion frame from the original frame element by element. The output of the image quality improvement information extractor 124 is compressed into the Side stream by the compressor B125.
[0026] However, the image quality improvement information extractor 124 is not limited to this. For example, it may be a block that outputs a frame obtained by dividing the original frame by the first expansion frame element by element, or may be a block composed of any computable process. Also, the second compression unit 123 does not necessarily need to be composed of an independent image quality improvement information extractor 124 and compressor B125, and may be a single functional block that includes the functions of both. For example, as described later, the second compression unit 123 may be a set of DNNs that takes the original frame and the first expansion frame as inputs and outputs the Side stream 126.
[0027] Also, the blocks included in the second compression unit 123 are not limited to the image quality improvement information extractor 124 and the compressor B125, and other functional blocks may be included. For example, a block that outputs the setting information of the compressor B125 from the original frame and the first decompressed frame may be included.
[0028] Finally, the compression unit 102 associates the Main stream 121 and the Side stream 126 with each other by means of the compressed data management table 127. The compression unit 102 transmits this as the final compressed data to the storage / communication unit 104.
[0029] The storage / communication unit 104 is a module that stores the data received from the compression unit 102 in the storage 105, transfers it to the decompression unit 103, or, in response to a request from the decompression unit 103, responds to the decompression unit 103 with the compressed data stored in the storage 105. The decompression unit 103 is a module that decompresses and responds to the compressed data obtained from the storage / communication unit 104 in response to a request from the client 101.
[0030] The client 101 may be a computer different from the computer that processes the decompression unit 103, or may be software such as video display software or video analysis software that operates on the same computer as the decompression unit 103, or may be any other hardware and software that consumes the decompressed data. The client 101 may request data for each frame from the decompression unit 103, may request data for each video, may request that the data generated by the data generation source 100 be transmitted at any time, or may request data under any other conditions.
[0031] When the extension unit 103 receives the compressed data from the storage / communication unit 104, it acquires the Main stream 130 and the Side stream 132 from the compressed data management table 136 that constitutes the data. Next, it inputs the Main stream into the decompressor A 122 to obtain the first decompressed frame. Next, it inputs the Side stream 132 and the first decompressed frame into the second decompression unit 133 to obtain the final decompressed frame (hereinafter, the final decompressed frame), and responds to the client 101.
[0032] The second decompression unit 133 includes a decompressor 134 for the second compression technique (hereinafter, decompressor B) and a frame generator 135. The frame generator 135 is a block that obtains the final decompressed frame by taking the output of the compressor B 134 and the first decompressed frame as inputs. For example, when the image quality improvement information extractor 124 outputs the difference between the original frame and the first decompressed frame, the corresponding frame generator 135 can perform a process of adding the first decompressed frame and the output of the decompressor B 134.
[0033] However, the frame generator 135 is not limited to this, and it may be a block composed of any computable process. Also, the frame generator 135 is not limited to the inverse conversion process of the image quality improvement information extractor 124.
[0034] Also, the second decompression unit 133 does not necessarily need to be composed of an independent decompressor B 134 and a frame generator 135, and it may be a single functional block that includes the functions of both. For example, as will be described later, the second decompression unit 133 may be a set of DNNs that takes the first decompressed frame and the Side stream 132 as inputs and outputs the final decompressed frame. Also, the blocks included in the second decompression unit 133 are not limited to the decompressor B 134 and the frame generator 135, and other functional blocks may be included.
[0035] The processes of the compression unit 102 and the extension unit 103 described above may be performed for each frame of the video, or may be performed for each unit that groups a plurality of frames. For a plurality of framesFor each summarized unit When processing, the first compression technique, or the second compression technique, or both may perform encoding that takes into account temporal redundancy, such as inter-frame encoding in video compression.
[0036] (1-2) System Configuration The system configuration of the first embodiment will be described with reference to FIG. 2. The compression unit 102, the decompression unit 103, and the storage / communication unit 104 are, for example, computers equipped with hardware resources such as a processor, memory, and network interface, and software resources such as an Operating System, middleware, a data compression program, and a data decompression program. The switch 206 interconnects the compression unit 102, the decompression unit 103, and the storage / communication unit 104.
[0037] The compression unit 102 includes a Front-end Interface 220, a processor 221, a RAM 223, a Back-end Interface 226, and a switch 222. The Front-end Interface 220 is an interface for connecting the compression unit 102 to the data generation source 100. The processor 221 controls the entire compression unit 102 based on the program 224 stored in the RAM 223 and the management information (metadata) 225 via the switch 222. The Back-end Interface 226 connects the compression unit 102 to the storage / communication unit 104.
[0038] The decompression unit 103 includes a Front-end Interface 230, a processor 231, a RAM 233, a Back-end Interface 236, and a switch 232. The Front-end Interface 230 is an interface for connecting the decompression unit 103 to the client 101. The processor 231 controls the entire decompression unit 103 based on the program 234 stored in the RAM 233 and the management information (metadata) 235 via the switch 232. The Back-end Interface 236 connects the decompression unit 103 to the storage / communication unit 104.
[0039] In FIG. 2, although the detailed configuration of the storage / communication unit 104 is omitted, for example, it can have the same configuration as the compression unit 102 or the decompression unit 103.
[0040] The processors 221 and 231 may be general-purpose arithmetic processors such as a CPU (Central Processing Unit), or may be accelerators such as a GPU (Graphical Processing Unit) or an FPGA (Field Programmable Gate Array), or may be hardware encoders / decoders of standard codecs such as HEVC, or may be a combination thereof.
[0041] The storage 105 may be a block device composed of a Hard Disk Drive (HDD) or a Solid State Drive (SSD), or may be a file storage, or may be a content storage, or may be a volume constructed on a storage system, or may be realized by any other method of accumulating data using one or more storage devices.
[0042] The compression unit 102, the decompression unit 103, and the storage / communication unit 104 may be configured by interconnecting hardware such as an IC (Integrated Circuit) that implements the components described above, or some of them may be implemented by one semiconductor element as an ASIC (Application Specific Integrated Circuit) or an FPGA, or may be a VM (Virtual Machine) implemented software-wise. Also, components other than those shown here may be added.
[0043] Also, the data generation source 100, the client 101, the compression unit 102, the decompression unit 103, and the storage / communication unit 104 may be different hardware devices, or different VMs operating on the same computer, or different containers operating on the same OS (Operating System), or different applications operating on the same OS, or each may be composed of a plurality of computers, or a combination of these.
[0044] For example, the data generation source 100 is an image sensor, the compression unit 102 is an edge device composed of a CPU and a GPU connected to the image sensor, the client 101 and the decompression unit 103 are programs operating on the same PC, and the storage / communication unit 104 may be a program operating on a Hyper Converged Infrastructure.
[0045] (1-3) RAM configuration FIG. 3 shows a configuration 300 of data stored in the RAM 223 of the compression unit 102 and the RAM 233 of the decompression unit 103. The RAM stores a program 310 executed by the processor and management information 320 used in the program.
[0046] The program 310 、pressure includes a compression program 311, a data decompression program 312, and a data learning program 313. The management information 320 includes compressed data 321 . Note that the data decompression program 312 may not be included in the program 224 of the compression unit 102, and the data compression program 311 may not be included in the program 234 of the decompression unit 103.
[0047] Also, when the DNN learning is executed on a third computer not included in the system shown in FIG. 2, the learning program 313 may not be included in the compression unit 102 or the decompression unit 103. However, in that case, the learning program 313 is expanded onto the RAM of the third computer. Note that the RAM may include data other than the above-described programs and configuration information.
[0048] The data compression program 311 is a program that compresses data in the compression unit 102. The data decompression program 312 is a program that decompresses the compressed data in the decompression unit 103. The learning program 313 is a program that executes the learning when the DNN is included in the compression unit 102 and the decompression unit 103.
[0049] The compressed data 321 is a memory area for storing the compressed data and has a data structure including a Main stream and a Side stream.
[0050] (1-4) Table Configuration FIG. 4 shows a compressed data management table 400, which is a data structure constituting the compressed data 321. Note that the representation method of the compressed data 321 is not limited to the format of the compressed data management table 400 and may be represented by a data structure other than a table, such as XML (Extensible Markup Language), YAML (YAML Ain’t a Markup Language), a hash table, or a tree structure.
[0051] The data name column 401 of the compressed data management table 400 is a field that stores an identifier representing the data source 100. The identifier may be a character string named by a user for the data source 100, a Media Access Control (MAC) address or an Internet Protocol (IP) address assigned to the data source 100, or any other code that can identify the data source 100. Furthermore, if the data source 100 is self-evident, the data name column 401 may not be present.
[0052] The main stream column 402 is a field for storing the main stream 121 obtained by compressing data received from the data generation source 100 by the compressor A 120. The side stream column 403 is a field for storing the side stream 126 which is the output of the second compression unit 123.
[0053] The model ID column 404 is a field that stores information for identifying a model used to generate a side stream when, for example, the second compression technique is Deep Video Compression and multiple models are prepared for each target image quality. However, the model ID column 404 is optional and does not need to be included in the compressed data management table 400. In addition, fields other than those described above, such as setting information for the first compression technique and a timestamp, may be included in the compressed data management table 400.
[0054] (1-5) Data compression and decompression 5 is a flow diagram of the data compression program 311. The processor 221 of the compression unit 102 starts the data compression program 311 upon receipt of video data generated by the data generation source 100 (S500).
[0055] S501 is the step in which the compressor 102 receives one or more frames of the video received from the data generation source 100, and the processor 221 acquires them from the Front-end Interface 220.
[0056] S502 is the step in which the processor 221 compresses the frames acquired in S501 by the compressor A120 to generate the Main stream 121.
[0057] S503 is the step in which the processor 221 inputs the Main stream 121 generated in S501 to the decompressor A122 to generate the first decompressed frame. S504 is the step in which the processor 221 uses the frames acquired in S501 and the first decompressed frame generated in S502 as the input to the second compression unit 123 to generate the Side stream 126.
[0058] S505 is the step of storing the Main stream 121 generated in S502 and the Side stream 126 generated in S504 in the compressed data management table 400 within the compressed data 321. Information such as the data name column 401 and the model ID column 404, etc. are also set in this step if necessary.
[0059] S506 is the step of transmitting the information of the compressed data management table 400 created in S505 to the storage / communication unit 104 through the Back-end Interface 226. After that, the data compression program 311 ends (S507).
[0060] Figure 6 is a flowchart of the data decompression program 312. When the processor 231 of the decompression unit 103 receives the compressed data from the storage / communication unit 104, it starts the data decompression program 312 (S600).
[0061] S601 is a step in which the processor 231 obtains, from the Back-end Interface 236, the compressed data received by the extension part 103 from the accumulation / communication part 104 and stores the compressed data in the compressed data management table 400 in the form of the compressed data in the RAM 233. tag 3 21.
[0062] S602 is a step in which the processor 231 obtains the Main stream 130 from the compressed data management table 400 within the compressed data 21. S603 is a step in which the processor 231 extends the Main stream 130 obtained in S602 into a first extended frame by the extender A122. tag 3
[0063] S604 is a step in which the processor 231 obtains the Side stream 132 from the compressed data management table 400.
[0064] S605 is a step in which the processor 231 inputs the first extended frame generated in S603 and the Side stream obtained in S604 into the second extension unit 133 to generate a final extended frame.
[0065] S606 is a step in which the processor 231 transmits the final extended frame generated in S605 to the client 101 through the Front-end Interface 230. Then, the data extension program 312 ends (S607).
[0066] The flow of the data compression program 311 and the data extension program 312 has been described above. In the following, when the first compression technology is the standard codec and the second compression technology is Deep Video Compression, three more specific examples of the flow will be described. However, the data compression program 311 and the data extension program 312 are not limited to the examples described below.
[0067] Further, two or more of the following examples may be combined and used. For example, while encoding key frames at regular intervals within a frame, the frames in between may be encoded by inter-frame encoding. The frequency of performing intra-frame encoding is, for example, for every predetermined number of frames, but may be any frequency, such as performing it at a variable frequency.
[0068] Also, all frames may be encoded by intra-frame encoding. Also, inter-frame encoding is not limited to being based on the frame one time step before in time. For example, it may be based on frames two or more time steps before in time, or may be a frame that is in the future in time but has already been extended, or may be a combination of these. Also, intra-frame encoding and inter-frame encoding may be synchronized between the first compression technique and the second compression technique, or each frame may be encoded by an independent method.
[0069] FIG. 7 shows an example of intra-frame encoding of a video. Compression process 700 is a block diagram representing the compression process of intra-frame encoding. The original frame 701 generated by data generation source 100 is input to compressor A120, which is a compressor of a standard codec, and is compressed into Main stream 121. At this time, the bit consumption of Main stream 121 is made less than the amount necessary to achieve a desired image quality using only the standard codec. For example, as described above, the bit consumption may be reduced by increasing the QP of the standard codec for the entire frame, or the bit consumption may be selectively reduced in regions where the compression rate of the standard codec is worse than the compression rate of Deep Video Compression.
[0070] Next, Main stream 121 is input to expander A122, which is an expander of a standard codec, to obtain the first expanded frame 702. Then, the original frame 701 and the first expanded frame 702 are input to encoder 703 configured by a DNN.
[0071] The encoder 703 is a DNN composed of a convolutional layer or a pooling layer that outputs a three-dimensional tensor, taking as input a tensor of 6×Height×Width obtained by concatenating tensors of the original frame and the first extended frame 702, both of size 3×Height×Width and represented in, for example, the RGB format, in the channel axis direction.
[0072] The coder 704 encodes the tensor output by the encoder 703 into a bit stream and outputs Side stream 126. The coder 704 may simply serialize the bit stream of the floating-point numbers representing the tensor output by the encoder 703, or, in order to improve the compression rate, may estimate the probability of occurrence of the value of each element of the tensor using an entropy estimator such as an Auto Regressive Model or a Hyper Prior Network composed of a DNN, and perform entropy coding such as Range Coder based on the result, or any other arbitrary means may be used.
[0073] Note that the DNNs included in the encoder 703 and the coder 704 may be trained to allocate particularly more bits in a region where the compression rate of Deep Video Compression is better than that of the standard codec. An example of the learning process will be described later. The encoder 703 and the coder 704 constitute the second compression unit 123. The function of the image quality improvement information extractor 124 is included in the encoder 703, and the coder 704 has a compression function.
[0074] The extension process 710 is a block diagram representing the extension process of in-frame encoding. The main stream 130 is input into the extender A122, which is an extender of the standard codec, and is extended into the first extended frame 711. The side stream 132 is input into the decoder 712 and is decoded from the bit stream into a form such as a tensor. Note that the decoder 712 is, for example, the inverse transformation of the encoder 704. When entropy encoding is performed by the encoder 704, the decoder 712 performs decoding using the same entropy model as that used by the encoder 704.
[0075] The decoded tensor and the first extended frame 711 are input into the decoder 713, and the final extended frame 714 is generated. The decoder 713 is, for example, a DNN composed of a transpose convolution layer or the like that outputs the final extended frame 714 with a size of 3×Height×Width by taking as input the tensor of the first extended frame 711 with a size of 3×Height×Width represented in the RGB format and the tensor output by the decoder 712. decryption Encoder 7 12 And the decoder 713 constitutes the second extension unit 133. decryption Encoder 7 12 Has an extension function. The function of the frame generator 135 is included in the decoder 713.
[0076] Figure 8 shows the first example of inter-frame encoding of a video. Among the arrow lines connecting the blocks, the thick line represents the path required during extension. During compression, both the thin line and the thick line paths are used.
[0077] The original frame 801 generated by the data generation source 100 is compressed by the compressor A120, which is a compressor of the standard codec, and is converted into the main stream 121. Then, it is converted into the first extended frame 802 by the extender A122, which is an extender of the standard codec. At this time, similar to in-frame encoding, the bit consumption of the main stream 121 is suppressed.
[0078] Next, the original frame 801 and the first extended frame 802 are converted by the image quality improvement information extractor 803 into Feature 804 expressed in a form such as a tensor. The image quality improvement information extractor 803 is, for example, a DNN composed of a convolutional layer, a pooling layer, and the like. The image quality improvement information extractor 803 takes, as input, a tensor of the original frame 801 with a size of 3×Height×Width and a tensor of the first extended frame 802 with the same size, which are expressed in the RGB format, and are concatenated in the channel axis direction to form a tensor with a size of 6×Height×Width, and outputs the 3D tensor Feature 804.
[0079] Next, the first extended frame 805 of the frame one frame before the original frame 801 in time and the final extended frame 806 are input to the image quality improvement information extractor 807 to extract Feature 808 (hereinafter referred to as the forward Feature) in the previous frame.
[0080] Note that the image quality improvement information extractor 807 used at this time may be the same as or different from the image quality improvement information extractor 803. Also, the image quality improvement information extractors 803 and 807 do not necessarily use a DNN, and may be, for example, a process of obtaining the difference between two input frames.
[0081] Also, the first extended frame 805 and the final extended frame 806 are not limited to the frame one frame ahead in time of the original frame 801, and may be frames two or more frames before, or may be frames that are behind in time but have already been extended. The forward Feature can be extracted from these frames by the image quality improvement information extractor 807.
[0082] Next, Feature 804 and the forward Feature 808 are input into motion extraction 809 to extract the information required in subsequent motion compensation 812. Motion extraction 809 may be, for example, a pre-trained DNN that estimates Optical Flow, or a DNN that is end-to-end trained together with other DNNs included in FIG. 8, or a motion vector predictor used in a standard codec, etc., or any other process.
[0083] Motion compression 810 compresses the output of motion extraction 809 into a bit string. Motion compression 810, for example, converts the tensor output by motion extraction 809 by a DNN including a convolutional layer, and encodes the resulting tensor using an entropy estimator such as an Auto Regressive model composed of a DNN with a Range Coder or the like. Note that the method of motion compression 810 is not limited to this.
[0084] The output of motion compression 810 is decompressed by motion decompression 811 and then input into motion compensation 812 together with the forward Feature 808. Motion compensation 812 is a process of correcting the forward Feature 808 based on the information output by motion decompression 811. Motion compensation 812 is, for example, a block that warps the forward Feature 808, which is a three-dimensional tensor, with offset information having the same width and height as the forward Feature 808 and a channel number of 2 output by motion decompression 811, but is not limited to this.
[0085] Next, the residual extractor 813 subtracts the tensor obtained as the result of motion compensation 812 from Feature 804 element by element and outputs residual information. However, the residual extractor 813 is not limited to this and may be a DNN or the like. residual The information is compressed into a bit string by residual compression 814. Residual compression 814 may use the same technology as motion compression 810, or any other compression technology may be used.
[0086] The side stream 126 has a data structure including the bit sequence generated by the motion compression 810 and the bit sequence generated by the residual compression 814. The bit sequence generated by the residual compression 814 is extended by the residual extension 815 and then input to the residual compensator 816 together with the output of the motion compensation 812. The residual compensator 816 is, for example, a process of outputting a tensor 817 (hereinafter, extended Feature) obtained by adding the output of the residual extension 815 to the output of the motion compensation 812 element by element, but is not limited thereto.
[0087] Finally, the first extended frame 802 and the extended Feature 817 are input to the frame generator 818 to obtain the final extended frame 819. The frame generator 818 is a DNN composed of, for example, a transposed convolutional layer, etc., but is not limited thereto.
[0088] Note that, in the above, an example of performing the motion extraction 809 and the motion compensation 812 using the extended first extended frame 805 and the final extended frame 806 has been shown, but it is not limited thereto. For example, the extended extended Feature 817 may be buffered and used as the forward Feature 808.
[0089] FIG. 9 shows a second example of the inter-frame encoding. First, the compression process will be described. The original frame 901 composed of a plurality of frames is input to the compressor A120 which is a compressor of the standard codec, and after obtaining the Main stream 902, the first extended frame 903 composed of a plurality of frames is obtained by its expander A122.
[0090] Next, the original frame 901 and the first extended frame 903 are input to the encoder 904 simultaneously for a plurality of frames. The encoder 904 is, for example, a DNN composed of a two-dimensional convolutional layer, etc., which takes as input a tensor of 6N×Height×Width obtained by concatenating the original frame 901 and the first extended frame for N frames, each of which is represented in the RGB format and has a size of 3×Height×Width, in the channel axis direction, and outputs a three-dimensional tensor.
[0091] Also, the encoder 904 may perform a process of converting a tensor of 6×N×Height×Width, which is obtained by concatenating the original frames 901 of N frames and the first extended frame, both of which are represented in the RGB format and have a size of 3×Height×Width, in the channel axis direction and the frame axis direction, into a tensor by a DNN composed of a 3D convolutional layer or the like, or any other arbitrary process.
[0092] The coder 905 converts data such as the tensor generated by the encoder 904 into a bit string to generate the Side stream 906. The coder 905 is, for example, a process of encoding the tensor output by the encoder 904 using an entropy estimator of an Auto Regressive model composed of a DNN with a Range Coder or the like, but is not limited thereto.
[0093] Next, the extension process will be described. The extender A122 outputs the first extended frame 903 from the Main stream 902. Also, the decoder 907 decodes the Side stream 906 into data such as a tensor. Finally, the first extended frame 903 and the output of the decoder 907 are input to the decoder 908 to obtain the final extended frame 909 for a plurality of frames.
[0094] The decoder 908 is, for example, a DNN composed of a 2D Transposed convolutional layer or the like that outputs the final extended frame 909 for N frames by outputting a tensor with a size of 3N×Height×Width. Also, the decoder 908 may be a DNN composed of a 3D convolutional layer that inputs a plurality of 3D tensors and outputs a plurality of tensors with a size of 3×Height×Width, or any other arbitrary process.
[0095] (1-6) DNN learning process Figure 10 shows an overview of the DNN learning program 313. In the following, the overview of learning is shown using the intra-frame encoding shown in Figure 7 as an example, but the DNN can also be learned in the same way for the inter-frame encoding shown in Figures 8 and 9. Note that the DNN learning method is not limited to that described below, and any learning data, Optimizer, Loss function, etc. may be used.
[0096] The learning dataset 1000 is data used for learning the DNN. The original frame 1001 is data consisting of frames of the video before compression. The first decompressed frame 1002 is a frame obtained by compressing and decompressing the original frame 1001 by intra-frame encoding with a standard codec.
[0097] The DNN learning process will be described. First, from the learning dataset 1000, the original frames 1001 for the batch size used for learning and the corresponding first decompressed frames 1002 are obtained. Next, the original frame 1001 and the first decompressed frame 1002 are input to the encoder 703 to output a Feature 1010 such as a tensor.
[0098] In the output of the encoder 703, if the process of quantizing the values of the Feature 1010 into integers, etc. is included, changes such as adding noise to the tensor may be made during learning so that the backpropagation method becomes possible. In addition, generally known quantization approximation methods that enable the backpropagation method may be used. Next, the Feature 1010 and the first decompressed frame 1002 are input to the decoder 713 to obtain the final decompressed frame 1011.
[0099] Next, the image quality between the obtained final extension frame 1011 and the original frame 1001 is quantified using, for example, the Mean Squared Error (MSE) 1014. Note that the index of the image quality is not limited to MSE, and any index such as the L1 norm or Multi-scale Structural Similarity may be used. When an encoder 704 that entropy-encodes Feature 1010 is used, the occurrence probability of the value of each element of Feature 1010 is estimated by an entropy estimator 1012 such as an Auto Regressive Model composed of a DNN.
[0100] Next, based on the estimation result of the entropy estimator 1012, the bit consumption amount after encoding of Feature 1010 is calculated by a bit-per-pixel (bpp) calculator 1013. Note that bpp is an index representing the bit consumption amount per pixel. bpp calculator The bpp calculated by 1013 and MSE the MSE calculated by 1014 are input to a Loss function 1015, and the Loss value for learning is calculated.
[0101] Thereafter, based on the value of the Loss function, the learning parameters of the DNNs included in the encoder 703, decoder 713, entropy estimator 1012, etc. are updated using, for example, the backpropagation method. Note that the input to the Loss function 1015 is not limited to the calculated bpp and MSE, and regularization such as Weight Decay may be reflected in the learning with the learning parameters of the DNN as the input. Also, when the entropy estimator 1012 is a Hyper Prior Network, the bpp of the Hyper Prior may be estimated in the same way and used as the input to the Loss function 1015.
[0102] The Loss function 1015 is, for example, a function that linearly combines bpp and MSE by a hyperparameter a (L = MSE + a × bpp). The hyperparameter a is a parameter that adjusts the bit consumption amount of the Side stream 126.
[0103] Also, as the Loss function 1015, the following formula (1) may be used.
[0104] [Number]
[0105] By using formula (1), without adjusting the hyperparameter a, the DNN can be trained to maximize the reduction rate of the bit consumption of this embodiment with respect to the standard codec. Formula (1) is a formula that represents, in terms of the image quality of the final extended frame 1011, the ratio of the bit consumption to the bit consumption of the standard codec as a percentage.
[0106] Using FIG. 11, formula (1) will be described. Curve 1100 represents the rate-distortion curve of the standard codec in the learning batch x. The function rate_x(mse) is a function representing curve 1100, and is a function that returns the bit consumption of the Main stream 126 when the learning batch x is compressed and decompressed by the standard codec so that the image quality becomes mse. This function can be obtained by interpolation with a quartic function or the like from the measured values of the image quality and the bit consumption when the original frame 1001 of the learning batch x is compressed at a plurality of QPs, but is not limited thereto.
[0107] Also, the measured values of the image quality and the bit consumption for each QP required for the interpolation process may be included in the learning dataset 1000. Point 1101 is the point when the original frame 1001 is compressed and decompressed into the first extended frame 1002 by the standard codec, and its bpp is bpp_main. Point 1102 is the point when the original frame 1001 is compressed and decompressed according to this embodiment, and its image quality is mse_xhat.
[0108] Assuming that the bit consumption of the side stream 126 is bpp_side, the bit consumption of this embodiment is bpp_main + bpp_side, which corresponds to the numerator of Equation (1). When decompressing and expanding the original frame 1001 by the standard codec and the image quality is mse_xhat, the bit consumption can be estimated as rate_x(mse_xhat), which corresponds to the denominator of Equation (1).
[0109] That is, by using Equation (1) as the Loss function 1015, the DNN can be trained so that the bit consumption of the side stream 126 is such that the ratio of the bit consumption of this embodiment to that of the standard codec is minimized when the image quality is the same. Note that the Loss function 1015 is not limited to the functions described above and may be other functions.
[0110] Note that the present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail for easy understanding of the present invention and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Also, it is possible to add, delete, or replace other configurations for a part of the configuration of each embodiment.
[0111] In addition, each of the above-described configurations, functions, processing units, etc. may be realized in hardware by designing a part or all of them, for example, by an integrated circuit. Also, each of the above-described configurations, functions, etc. may be realized in software by a processor interpreting and executing a program that realizes each function. Information such as programs, tables, files, etc. that realize each function can be stored in a memory, a recording device such as a hard disk or an SSD (Solid State Drive), or a recording medium such as an IC card or an SD card.
[0112] In addition, the control lines and information lines show those considered necessary for explanation, and not necessarily all control lines and information lines are shown on the product. In reality, it may be considered that almost all components are interconnected.
Explanation of Signs
[0113] 100 Data generation source 102 Compression unit 103 Decompression unit 104 Accumulation / communication unit 105 Storage 120, 125 Compressor 122 Decompressor 123 Second compression unit 124 Image quality improvement information extractor 128 Compression parameter setter 221, 231 Processor 223, 233 RAM
Claims
1. A data compression system, comprising: one or more processors; and one or more storage devices, wherein the one or more processors compress original data by a first irreversible compression method to generate first compressed data, decompress the first compressed data to generate first decompressed data, extract difference information between the original data and the first decompressed data, compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data, store the first compressed data and the second compressed data in the one or more storage devices, wherein one of the first irreversible compression method or the second irreversible compression method performs compression using a neural network, and the other of the first irreversible compression method or the second irreversible compression method performs compression without using a neural network. A data compression system.
2. A data compression system, comprising: one or more processors; and one or more storage devices, wherein the one or more processors compress original data by a first irreversible compression method to generate first compressed data, decompress the first compressed data to generate first decompressed data, extract difference information between the original data and the first decompressed data, compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data, store the first compressed data and the second compressed data in the one or more storage devices, wherein the one or more processors compare the compression ratios measured by compressing the divided parts of the original data by the first irreversible compression method and the second irreversible compression method, and reduce the bit consumption by the first irreversible compression method in a part where the compression ratio of the second irreversible compression method is higher than that of the first irreversible compression method. A data compression system.
3. A data compression system, comprising: one or more processors; and one or more storage devices, wherein the one or more processors compress original data by a first irreversible compression method to generate first compressed data, decompress the first compressed data to generate first decompressed data, extract difference information between the original data and the first decompressed data, compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data, store the first compressed data and the second compressed data in the one or more storage devices, wherein the second irreversible compression method uses a neural network. The neural network is a data compression system that is trained to increase the reduction rate of the bit consumption for the first irreversible compression method. **Claim 4**: A data compression system, comprising: One or more processors; One or more storage devices, wherein the one or more processors: Compress original data by a first irreversible compression method to generate first compressed data; Decompress the first compressed data to generate first decompressed data; Extract difference information between the original data and the first decompressed data; Compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data; Store the first compressed data and the second compressed data in the one or more storage devices; and extract the difference information by a neural network. A data compression system. **Claim 5**: A data compression system, comprising: One or more processors; One or more storage devices, wherein the one or more processors: Compress original data by a first irreversible compression method to generate first compressed data; Decompress the first compressed data to generate first decompressed data; Extract difference information between the original data and the first decompressed data; Compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data; Store the first compressed data and the second compressed data in the one or more storage devices; wherein the original data is video data; and the second irreversible compression method executes at least one of intra-frame coding and inter-frame coding using a neural network. A data compression system. **Claim 6**: A data compression method by a data compression system, comprising: Compress original data by a first irreversible compression method to generate first compressed data; Decompress the first compressed data to generate first decompressed data; Extract difference information between the original data and the first decompressed data; Compress the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data; Store the first compressed data and the second compressed data in a storage; wherein one of the first irreversible compression method or the second irreversible compression method executes compression using a neural network, and the other of the first irreversible compression method or the second irreversible compression method executes compression without using a neural network. A data compression method.
7. A data compression method by a data compression system, comprising: Compressing original data by a first irreversible compression method to generate first compressed data; Decompressing the first compressed data to generate first decompressed data; Extracting difference information between the original data and the first decompressed data; Compressing the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data; Storing the first compressed data and the second compressed data in a storage; Comparing the compression ratios measured by compressing the divided parts of the original data by the first irreversible compression method and the second irreversible compression method, and reducing the bit consumption by the first irreversible compression method in a part where the compression ratio of the second irreversible compression method is higher than that of the first irreversible compression method. A data compression method.
8. A data compression method by a data compression system, comprising: Compressing original data by a first irreversible compression method to generate first compressed data; Decompressing the first compressed data to generate first decompressed data; Extracting difference information between the original data and the first decompressed data; Compressing the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data; Storing the first compressed data and the second compressed data in a storage; The second irreversible compression method uses a neural network; The neural network is trained to increase the reduction rate of bit consumption with respect to the first irreversible compression method. A data compression method.
9. A data compression method by a data compression system, comprising: Compressing original data by a first irreversible compression method to generate first compressed data; Decompressing the first compressed data to generate first decompressed data; Extracting difference information between the original data and the first decompressed data; Compressing the difference information by a second irreversible compression method different from the first irreversible compression method to generate second compressed data; Storing the first compressed data and the second compressed data in a storage; Extracting the difference information by a neural network. A data compression method.
10. A data compression method by a data compression system, comprising: Compressing original data by a first irreversible compression method to generate first compressed data; Decompressing the first compressed data to generate first decompressed data; Extracting difference information between the original data and the first decompressed data; Compress the differential information with a second irreversible compression method different from the first irreversible compression method to generate second compressed data. Store the first compressed data and the second compressed data in a storage. The original data is video data. The second irreversible compression method is a data compression method that performs at least one of intra-frame encoding and inter-frame encoding using a neural network.
Citation Information
Patent Citations
Picture quality evaluation system
JP1995184062A
Extended image coding method, extended image decoding method, extended image coder, extended image decoder and extended image recording medium
JP2002369220A
Image encoder and image decoder
JP2004015226A
Storage system and storage control apparatus
JP2007199891A
Scalable encoding method and device, scalable decoding method and device, their programs and recording medium on which the programs are recorded
WO2006038607A1