Image Encoding Device, Image Decoding Device, and Program
Patent Information
- Application Number
- JP2024551157
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-16
- Estimated Expiration
- 2042-10-20
AI Technical Summary
Learning-based image compression technologies face challenges in being applied to devices with limited computational resources due to high calculation loads, making it difficult to restore original images effectively from compressed data, especially when the filter group selection is not optimal during the encoding and decoding stages.
An image encoding and decoding system that employs multiple machine learning models to extract context information, determine parameters, compress data, and encode/decode it, ensuring that the amount of data processed is reduced while maintaining high image quality by using a context extraction unit, parameter determination units, and encoding/decoding units within the system.
The system effectively reduces processing requirements while ensuring that high-quality images approximating the original can be restored from encoded data without significantly increasing processing loads, even on devices with limited resources.
Abstract
Description
Image encoding device, image decoding device, image processing system, model learning device, image encoding method, image decoding method, and computer-readable storage medium
[0001] The present invention relates to an image encoding device, an image decoding device, an image processing system, a model learning device, an image encoding method, an image decoding method, and a computer-readable storage medium.
[0002] Learning-based image compression technology is an image compression technology that uses a machine learning model to convert an image into compressed data with less information so that the original image can be restored. Learning-based image compression technology can achieve higher compression performance than general image compression technology. Examples of indicators of compression performance include the compression rate from the original image to compressed data, the image quality of the restored data restored from the compressed data, and the recognition rate in image recognition of the restored data.
[0003] Image compression technology is applied to a wide range of applications and fields, including remote monitoring, communication, and education. Generally, the computational load required for machine learning models is large. Therefore, it can be difficult to apply learning-based image compression technology to devices with limited computing resources. Examples of small-scale devices with limited computing resources include low-power edge devices. Reducing the amount of computation is expected to promote the widespread use of learning-based image compression technology.
[0004] In this regard, Non-Patent Document 1 proposes Omni-dimensional Dynamic Convolution (ODConv). ODConv is a method that utilizes a multi-dimensional attention mechanism that enables parallel learning of complementary attentions of convolution kernels along all four dimensions of the kernel space in any convolution layer. That is, ODConv dynamically determines attention for each kernel constituting a convolution filter group in each convolution layer using the input data. This method allows dynamic selection of filter group parameters required in the convolution layer. Although the filter group parameters may vary significantly depending on the input image data, calculations for unnecessary filters in each layer are not required. This is expected to reduce the amount of calculations.
[0005] Chao Li, Aojun Zhou, Anbang Yao, “OMNI-DIMENSIONAL DYNAMIC CONVOLUTION, “Generative Adversarial Networks for Extreme Learned image Compression”, International Conference on Learning Representations (ICLR 2022), 1-20, April 25-29, 2022
[0006] However, in image coding, the compressed data obtained by compressing the information volume of image data is expressed as a code through coding, and this code heavily depends on the filter group selected. In model learning for image coding, the goal is to ensure that the data restored from compressed data is as close as possible to the original image data. Therefore, dynamic filter group selection does not work effectively, and there is a risk that the same filter group will continue to be selected. The compressed data to be coded depends on the filters selected in the coding stage. If the filter group is not selected appropriately in the decoding stage, it may not be possible to restore the original image from the compressed data.
[0007] An object of the present disclosure is to provide an image encoding device, an image decoding device, an image processing system, a model learning device, an image encoding method, an image decoding method, and a computer-readable storage medium that solve the above-mentioned problems.
[0008] According to a first aspect of the present disclosure, an image encoding device includes a context extraction unit that uses a first machine learning model on image data to extract context information that indicates characteristics of the image data; a parameter determination unit that uses a second machine learning model on the context information to determine parameters of a third machine learning model; a compression unit that uses the third machine learning model on the image data to generate compressed data with a smaller amount of data than the image data; and an encoding unit that encodes the context information and the compressed data and generates a code sequence.
[0009] According to a second aspect of the present disclosure, an image decoding device includes a decoding unit that decodes context information and compressed data from a code sequence, a parameter determination unit that determines parameters of a fifth machine learning model using a fourth machine learning model on the context information, and a restoration unit that generates restored data using the fifth machine learning model on the compressed data.
[0010] According to a third aspect of the present disclosure, an image processing system includes an image encoding device and an image decoding device, wherein the image encoding device includes a context extraction unit that uses a first machine learning model on image data to extract context information indicating characteristics of the image data, a parameter determination unit that uses a second machine learning model on the context information to determine parameters of a third machine learning model, a compression unit that uses the third machine learning model on the image data to generate compressed data with a smaller amount of data than the image data, and an encoding unit that encodes the context information and the compressed data and generates a code sequence, and the image decoding device includes a decoding unit that decodes the context information and the compressed data from the code sequence, a parameter determination unit that uses a fourth machine learning model on the context information to determine parameters of a fifth machine learning model, and a restoration unit that uses the fifth machine learning model on the compressed data to generate restored data.
[0011] According to a fourth aspect of the present disclosure, a model learning device includes a context extraction unit that uses a first machine learning model on image data to extract context information indicating characteristics of the image data; a parameter determination unit that uses a second machine learning model on the context information to determine parameters of a third machine learning model; a compression unit that uses a third machine learning model on the image data to generate compressed data with a smaller amount of data than the image data; a parameter determination unit that uses a fourth machine learning model on the context information to determine parameters of a fifth machine learning model; a restoration unit that uses the fifth machine learning model on the compressed data to generate restored data; and a model learning unit that determines parameters of the first machine learning model, the second machine learning model, the third machine learning model, the fourth machine learning model, and the fifth machine learning model so that the difference between the image data and the restored data is reduced.
[0012] According to a fifth aspect of the present disclosure, there is provided an image encoding method in an image encoding device, wherein the image encoding device executes a context extraction step of extracting context information indicating characteristics of image data using a first machine learning model for the image data, a parameter determination step of determining parameters of a third machine learning model for the context information using a second machine learning model for the context information, a compression step of generating compressed data having a smaller amount of data than the image data using the third machine learning model for the image data, and an encoding step of encoding the context information and the compressed data to generate a code sequence.
[0013] According to a sixth aspect of the present disclosure, there is provided an image decoding method in an image decoding device, wherein the image decoding device executes a decoding step of decoding context information and compressed data from a code sequence, a parameter determination step of determining parameters of a fifth machine learning model using a fourth machine learning model on the context information, and a restoration step of generating restored data using the fifth machine learning model on the compressed data.
[0014] According to the present disclosure, it is possible to restore a high-quality image that is close to the original image from a code obtained by encoding the original image without significantly increasing the amount of processing.
[0015] FIG. 1 is a schematic block diagram showing an example of the functional configuration of an image processing system according to the present embodiment. FIG. 2 is a schematic block diagram showing an example of the functional configuration of a model learning device according to the present embodiment. FIG. 3 is a schematic block diagram showing a second example of the functional configuration of an image encoding device according to the present embodiment. FIG. 4 is a schematic block diagram showing an example of the functional configuration of a compression layer according to the present embodiment. FIG. 5 is a schematic block diagram showing a second example of the functional configuration of an image decoding device according to the present embodiment. FIG. 6 is a schematic block diagram showing an example of the functional configuration of a restoration layer according to the present embodiment. FIG. 7 is an explanatory diagram for explaining an example of a first machine learning model according to the present embodiment. FIG. 8 is an explanatory diagram for explaining another example of the first machine learning model according to the present embodiment. FIG. 9 is a schematic block diagram showing an implementation example in an image processing system according to the present embodiment. FIG. 10 is a schematic block diagram showing an example of the configuration of a statistical value calculation unit according to the present embodiment. FIG. 11 is a schematic block diagram showing another example of the configuration of an encoding unit and a decoding unit according to the present embodiment. FIG. 12 is a diagram showing an example of processing delay. FIG. 13 is a diagram showing an example of MS-SSIM. FIG. 14 is a schematic block diagram showing an example of the minimum configuration of an image encoding device. FIG. 15 is a schematic block diagram showing an example of the minimum configuration of an image decoding device. FIG. 16 is a schematic block diagram showing an example of the minimum configuration of an image processing system. FIG. 17 is a schematic block diagram showing an example of the minimum configuration of a model learning device. FIG. 18 is a schematic block diagram showing an example of the hardware configuration of an image encoding device according to the present embodiment.
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. <First Embodiment> The first embodiment will be described. FIG. 1 is a schematic block diagram showing an example of the functional configuration of an image processing system 1 according to this embodiment. The image processing system 1 includes an image encoding device 10 and an image decoding device 20. The image encoding device 10 and the image decoding device 20 are connected using, for example, a transmission path that enables the transmission of various types of data. The image encoding device 10 and the image decoding device 20 may be synchronously or asynchronously accessible to a common storage medium that enables the storage of various types of data.
[0017] The image encoding device 10 acquires image data and extracts context information indicating characteristics of the image data using a first machine learning model on the acquired image data. The image encoding device 10 determines parameters of a third machine learning model on the extracted context information using a second machine learning model. The image encoding device 10 generates compressed data with a smaller amount of data using the third machine learning model on the acquired image data. The image encoding device 10 encodes the generated context information and compressed data to generate a code sequence (i.e., a data stream). The image encoding device 10 outputs the generated code sequence.
[0018] The image decoding device 20 acquires a code sequence sent from the image encoding device 10 and decodes the context information and compressed data indicated by the acquired code sequence. The image decoding device 20 determines parameters of a fifth machine learning model using a fourth machine learning model for the decoded context information. The image decoding device 20 generates reconstructed data using the fifth machine learning model for the decoded compressed data. The image decoding device 20 may output the generated reconstructed data to an external device or may temporarily or continuously store the reconstructed data in a storage medium of its own device.
[0019] In the image processing system 1 illustrated in Fig. 1, the number of image encoding devices 10 and the number of image decoding devices 20 are one each, but this is not limited to this. The number of image encoding devices 10 may be two or more, and the number of image decoding devices 20 may be two or more. The image processing system 1 is applied to, for example, a distributed processing system. A distributed processing system includes multiple devices, and the individual devices are distributed and arranged at different spatial locations.
[0020] The image processing system 1 includes one or more data centers (not shown), and each data center may be connected to one or more edge devices (not shown). Generally, each edge device is installed near a source of information to be processed and may provide computing resources for that information. On the other hand, the data center performs processing related to the entire distributed processing system using various information provided by the edge devices. The data center may be installed in a location spatially separated from the edge devices. Each edge device includes an image encoding device 10 and a capturing unit (camera) (not shown). The capturing unit is an example of an information source that captures images within its field of view and generates image data representing the captured images. The capturing unit provides the generated image data to the image encoding device 10.
[0021] The data center may be configured as a single device, but is not limited to this. The data center may include multiple devices and be configured as a cloud that can transmit and receive data between them (not shown). The data center may be configured to include, for example, a server device and a model learning device. The server device may include, for example, an image decoding device 20 and an image recognition unit (not shown). Code sequences from individual image encoding devices 10 may be input to the image decoding device 20 via a network as a transmission path. The image decoding device 20 outputs restored data obtained by decoding the code sequence to the image recognition unit. The image recognition unit performs image recognition processing on the restored image indicated in the restored data input from the image decoding device 20. The data center uses the recognition results from the image recognition processing to monitor the field of view of the imaging unit.
[0022] Next, an example of the functional configuration of the image encoding device 10 according to this embodiment will be described. The image encoding device 10 includes a context extraction unit 102, a parameter determination unit 104, a compression unit 106, and an encoding unit .
[0023] The context extraction unit 102 acquires image data. The context extraction unit 102 extracts context information by using a first machine learning model for image data representing an image of each frame. The first machine learning model is a mathematical model for calculating, as an output value, an element value representing context information for each pixel represented in the image data as an input value. The context extraction unit 102 outputs the extracted context information to the parameter determination unit 104 and the encoding unit 108.
[0024] The parameter determination unit 104 determines parameters of a third machine learning model by using the second machine learning model on the context information input from the context extraction unit 102. The second machine learning model is a mathematical model for calculating parameters of the third machine learning model as output values for element values of the context information as input values. The parameter determination unit 104 sets the determined parameters of the third machine learning model in the compression unit 106.
[0025] The compression unit 106 uses a third machine learning model on the acquired image data to generate compressed data. The amount of information in the generated compressed data is compressed so that it is less than the amount of data in the acquired image data (data compression). The third machine learning model is a mathematical model for calculating element values that indicate compressed data as output values for signal values for each pixel indicated in the image data as input values. When calculating the output values, the compression unit 106 uses parameters set by the parameter determination unit 104 as parameters of the third machine learning model. The compression unit 106 outputs the generated compressed data to the encoding unit 108.
[0026] The encoding unit 108 includes a context encoding unit 108a and a compressed data encoding unit 108b. The context encoding unit 108a performs entropy coding on the context information input from the context extraction unit 102 to generate an entropy code sequence as a first code sequence. The context encoding unit 108a may also perform entropy coding on quantized values obtained by quantizing the input context information. The context encoding unit 108a and the compressed data encoding unit 108b may use any entropy coding method, such as Huffman coding or arithmetic coding.
[0027] The compressed data encoding unit 108b performs entropy encoding on the compressed data input from the compression unit 106, and generates an entropy code sequence as a second code sequence. The compressed data encoding unit 108b may also perform entropy encoding on quantized values obtained by quantizing the input compressed data. The encoding unit 108 outputs a code sequence including the generated first code sequence and second code sequence.
[0028] Next, an example of the functional configuration of the image decoding device 20 according to this embodiment will be described. The image decoding device 20 includes a decoding unit 202, a parameter determining unit 204, and a restoring unit 206.
[0029] The decoding unit 202 acquires a code sequence sent from the image encoding device 10. The decoding unit 202 includes a context decoding unit 202a and a compressed data decoding unit 202b. The context decoding unit 202a performs entropy decoding on a first code sequence included in the acquired code sequence to generate context information. The context decoding unit 202a uses an entropy decoding method corresponding to the entropy encoding method used by the context encoding unit 108a. The context decoding unit 202a outputs the generated context information to the parameter determination unit 204. The compressed data decoding unit 202b performs entropy decoding on a second code sequence included in the acquired code sequence to generate compressed data. The compressed data decoding unit 202b uses an entropy decoding method corresponding to the entropy encoding method used by the compressed data encoding unit 108b. The compressed data decoding unit 202b outputs the generated compressed data to the restoration unit 206.
[0030] The parameter determination unit 204 determines parameters of a fifth machine learning model by using a fourth machine learning model on the context information input from the decoding unit 202. The fourth machine learning model is a mathematical model for calculating parameters of the fifth machine learning model as output values for element values of the context information as input values. The parameter determination unit 204 sets the determined parameters of the fifth machine learning model in the restoration unit 206.
[0031] The restoration unit 206 uses a fifth machine learning model on the acquired compressed data input from the decoding unit to generate restored data. The amount of data in the generated restored data is extended so that it is greater than the amount of data in the compressed data (data restoration). The fifth machine learning model is a mathematical model for calculating, as an output value, an element value representing the compressed data for an element value of the compressed data as an input value. The fifth machine learning model can be a mathematical model equivalent to an inverse function corresponding to the third machine learning model. When calculating the output value, the restoration unit 206 uses the parameters set by the parameter determination unit 204 as parameters of the fifth machine learning model. The restoration unit 206 outputs the generated restored data.
[0032] Next, an example of the functional configuration of the model learning device according to this embodiment will be described. Fig. 2 is a schematic block diagram showing an example of the functional configuration of a model learning device 30 according to this embodiment. The model learning device 30 includes a context extraction unit 312, a parameter determination unit 314, a compression unit 316, an encoding unit 318, a decoding unit 322, a parameter determination unit 324, a restoration unit 326, and a model learning unit 330. The encoding unit includes a context encoding unit 318a and a compressed data encoding unit 318b. The decoding unit 322 includes a context decoding unit 322a and a compressed data decoding unit 322b.
[0033] The context extraction unit 312, the parameter determination unit 314, the compression unit 316, the context encoding unit 318a, and the compressed data encoding unit 318b are capable of performing the same processes and have the same configurations as the context extraction unit 102, the parameter determination unit 104, the compression unit 106, the context encoding unit 108a, and the compressed data encoding unit 108b of the image encoding device 10, respectively. The description of the image encoding device 10 is used for the processes and configurations. The context decoding unit 322a, the compressed data decoding unit 322b, the parameter determination unit 324, and the restoration unit 326 are capable of performing the same processes and have the same configurations as the context decoding unit 202a, the compressed data decoding unit 202b, the parameter determination unit 204, and the restoration unit 206 of the image decoding device 20, respectively. The description of the image decoding device 20 is used for the processes and configurations.
[0034] The model learning unit 330 determines parameters for each of the first machine learning model, the second machine learning model, the third machine learning model, the fourth machine learning model, and the fifth machine learning model so as to reduce the difference between the image data acquired by the context extraction unit 312 and the parameter determination unit 314 and the restored data output from the restoration unit 326. The model learning unit 330 sequentially sets the parameters for each of the first machine learning model, the second machine learning model, the third machine learning model, the fourth machine learning model, and the fifth machine learning model in the context extraction unit 312, the parameter determination unit 314, the compression unit 316, the parameter determination unit 324, and the restoration unit 326. As a result, these parameters are recurrently updated and can be adaptively learned.
[0035] The loss function indicating the magnitude of the difference between the image data and the restored data can be any of the L1 norm, L2 norm, cross entropy between the image data and the restored data, or a weighted sum thereof. The model training unit 330 can use, for example, a gradient method to update the parameters of each machine learning model. The model training unit 330 sequentially calculates the update amount for the parameters of each machine learning model for each update step, and obtains the updated parameter by adding the calculated update amount to the corresponding parameter. The model training unit 330 can determine whether the parameter change has converged based on whether the magnitude of the update amount is equal to or less than a predetermined threshold. The model training unit 330 terminates the learning of the machine learning model when it is determined that convergence has occurred. The model training unit 330 may also repeat the update of the parameters of the machine learning model a predetermined number of times and then terminate the learning of the machine learning model.
[0036] After completing the learning, the model learning unit 330 sets the parameters of the first machine learning model, the second machine learning model, and the third machine learning model in the context extraction unit 102, the parameter determination unit 104, and the compression unit 106 of the image encoding device 10, respectively. The model learning unit 330 sets the parameters of the fourth machine learning model and the fifth machine learning model in the parameter determination unit 204 and the restoration unit 206 of the image decoding device 20, respectively.
[0037] Note that the learning of the machine learning model in the model learning unit 330 may be performed independently of the operation of the image encoding device 10 and the image decoding device 20 (offline learning), or may be performed in parallel with the operation of the image encoding device 10 and the image decoding device 20. In this case, the model learning unit 330 may use image data acquired by the context extraction unit 102 or the compression unit 106 of the image encoding device 10 and restored data acquired from the restoration unit 206 of the image decoding device 20 (online learning). In this case, the model learning unit 330 may be provided in one or both of the image encoding device 10 and the image decoding device 20. In this case, the dedicated context extraction unit 312, parameter determination unit 314, compression unit 316, encoding unit 318, decoding unit 322, parameter determination unit 324, and restoration unit 326 may be omitted.
[0038] The context extraction unit 102, parameter determination unit 104, and compression unit 106 of the image encoding device 10 may be set with parameters of a first machine learning model, a second machine learning model, and a third machine learning model obtained by training in a device separate from the model learning device 30. Furthermore, the parameter determination unit 204 and restoration unit 206 of the image decoding device 20 may be set with parameters of a fourth machine learning model and a fifth machine learning model obtained by training in a device separate from the model learning device 30.
[0039] The image encoding process in the image encoding device 10 and the image decoding process in the image decoding device 20 may each have two or more layers (layering). In the example of FIG. 3, the image encoding device 10 includes N layers (N is a predetermined integer equal to or greater than two) of compression layers 110-1 to 110-N and an encoding unit 108. The nth layer (n is an integer between 1 and N) of the compression layer 110-n receives output data from the immediately preceding layer, the n-1th layer of the compression layer 110-n-1, as input data, and outputs nth layer context information obtained for the input data to the context encoding unit 108a. The compression layer 110-n outputs compressed data obtained based on the nth layer context information for the input data input from the compression layer 110-n-1 to the immediately succeeding layer, the n+1th layer of the compression layer 110n+1. However, image data acquired by the image encoding device 10 is input as input data to the first layer of the compression layer 110-1. The Nth compression layer 110-N outputs the compressed data to be output data to the compressed data encoding unit 108b.
[0040] Next, an example of the functional configuration of the compression layer 110-n of the nth layer will be described. FIG. 4 is a schematic block diagram showing an example of the functional configuration of the compression layer 110-n according to this embodiment. The compression layer 110-n includes a context extraction unit 102-n, a parameter determination unit 104-n, and a compression unit 106-n. In FIG. 4, dashed lines indicate optional functional units. First, the description will be made assuming that the compression layer 110-n does not include a merge unit 112-n.
[0041] The context extraction unit 102-n receives the output data from the compression layer 110-n-1 as input data, and extracts n-th layer context information using the first machine learning model of the n-th layer (hereinafter referred to as the "n-th layer first machine learning model" or the like) as input data. The context extraction unit 102-n outputs the extracted n-th layer context information to the parameter determination unit 104-n.
[0042] The parameter determination unit 104-n determines parameters of the n-th layer third machine learning model using the n-th layer second machine learning model for the n-th layer context information. The parameter determination unit 104-n sets the determined parameters of the n-th layer third machine learning model in the compression unit 106-n. The compression unit 106-n uses output data input from the compression layer 110-n-1 as input data, and generates output data with a smaller amount of data by using the n-th layer third machine learning model on the input data. The compression unit 106-n outputs the generated output data to the compression layer 110-n+1.
[0043] Next, a case where the compression layer 110-n includes a merger unit 112-n will be described. The merger unit 112-n receives the nth layer context information from the context extraction unit 102-n and the n-1th layer accumulated context information from the compression layer 110-n-1. The merger unit 112-n merges the nth layer context information with the n-1th layer accumulated context information to generate nth layer accumulated context information. Here, "merge" includes the meaning of "concatenation." The merger unit 112-n outputs the nth layer accumulated context information to the parameter determination unit 104-n and the compression layer 110-n+1.
[0044] The parameter determination unit 104-n calculates parameters of the n-th layer third machine learning model using the n-th layer second machine learning model for the n-th layer accumulated context information instead of the n-th layer context information. Therefore, by merging context information from a certain layer and providing it to the next layer, parameter determination in the next layer can be conditionalized. Note that in the first layer (n=1), the merger unit 112-1 is omitted because there is no layer 0 accumulated context information as input information. In the N-th layer (n=N), there is no compression layer 110-N+1 to which the merger unit 112-N outputs the N-th layer accumulated context information.
[0045] In the example of FIG. 5 , the image decoding device 20 includes a decoding unit 202 and N layers of reconstruction layers 210-1 to 210-N. The reconstruction layers 210-1 to 210-N are layers corresponding to the compressed layers 110-1 to 110-N, respectively. The nth layer reconstruction layer 210-n receives context information for the nth layer from the context decoding unit 202a, and receives output data from the immediately preceding layer, the n+1th layer reconstruction layer 210-n+1, as input data. The reconstruction layer 210-n outputs reconstructed data obtained from the input data based on the nth layer context information to the immediately succeeding layer, the n-1th layer reconstruction layer 210-n-1. However, the Nth layer reconstruction layer 210-N receives compressed data as input data from the compressed data decoding unit 202b. The first layer reconstruction layer 210-1 outputs the reconstructed data to be output data to the outside of the image decoding device 20.
[0046] Next, an example of the functional configuration of the n-th restoration layer 210-n will be described. Fig. 6 is a schematic block diagram showing an example of the functional configuration of the restoration layer 210-n according to this embodiment. The restoration layer 210-n includes a parameter determination unit 204-n and a restoration unit 206-n. First, the description will be made assuming that the restoration layer 210-n does not include a merge unit 212-n.
[0047] The parameter determination unit 104-n receives n-th layer context information from the context decoding unit 202a. The parameter determination unit 104-n determines parameters of the n-th layer fifth machine learning model using the n-th layer fourth machine learning model for the n-th layer context information. The parameter determination unit 104-n sets the determined parameters of the n-th layer fifth machine learning model in the reconstruction unit 206-n. The reconstruction unit 206-n receives output data output from the reconstruction layer 210-n+1 as input data. The reconstruction unit 206-n uses the n-th layer fifth machine learning model on the input data to generate output data with a larger amount of data than the input data. The reconstruction unit 206-n outputs the generated output data to the reconstruction layer 210-n-1.
[0048] Next, a case where the reconstruction layer 210-n includes a merger 212-n will be described. The merger 212-n receives the layer n context information from the context decoding unit 202a and the layer n-1 accumulated context information from the reconstruction layer 210-n-1. The merger 212-n merges the layer n context information with the layer n-1 accumulated context information to construct layer n accumulated context information. The merger 212-n outputs the constructed layer n accumulated context information to the reconstruction unit 206-n and the reconstruction layer 210-n+1.
[0049] The parameter determination unit 204-n calculates parameters of the n-th layer fifth machine learning model using the n-th layer fourth machine learning model for the n-th layer cumulative context information instead of the n-th layer context information. Note that in the first layer (n=1), the merger unit 212-1 is omitted because there is no layer 0 cumulative context information as input information. In the N-th layer (n=N), the reconstruction layer 210-N+1 does not exist as an output destination for the layer N cumulative context information from the merger unit 212-N.
[0050] Next, a first machine learning model according to this embodiment will be described. The first machine learning model extracts context information from image data as a function of the context extraction unit 102. In this application, context information refers to the physical or technical characteristics or meaning of the image data to be processed, or signs (signs, tokens, etc.) or other identifiers (identifiers, etc.) for distinguishing them. Image recognition is one mode of extracting context information. FIG. 7 shows, as an example of the first machine learning model, a mathematical model that, when image data is input, identifies, as context information, an identifier representing a "bicycle" as the object represented by the image. FIG. 8 shows, as another example of the first machine learning model, a mathematical model that, when image data is input, identifies, as context information, an identifier representing a "passenger car."
[0051] The first machine learning model may be a mathematical model that provides individual context information candidates for input image data. The first machine learning model may be, for example, a neural network, a decision tree, a random forest, or the like. The neural network may be, for example, a convolutional neural network (CNN) having a convolution section (convolution layer). The activation function constituting the neural network may be, for example, a rectified linear unit (ReLU), a sigmoid function, a soft sign, or the like. In the present application, it is not necessary to set a specific object or its state in advance as a candidate for context information. In model learning, it is sufficient to simultaneously calculate the parameters of the second, third, fourth, and fifth machine learning models.
[0052] The second machine learning model may be a mathematical model that can provide parameters of the third machine learning model as an output corresponding to context information as an input. The context information may be represented by discrete values that do not themselves represent technical or physical characteristics. Similarly, the fourth machine learning model may be a mathematical model that can provide parameters of the fifth machine learning model as an output corresponding to context information as an input. The fourth machine learning model may be a mathematical model that exhibits the same computational procedure as the second machine learning model.
[0053] The second machine learning model and the fourth machine learning model may each be a mathematical model that is linearly separable with respect to the input, or may each be a mathematical model that is linearly inseparable with respect to the input. For example, a multi-layer perceptron (MLP), which is a type of neural network, may be applied as the second machine learning model and the fourth machine learning model.
[0054] The third machine learning model may be a mathematical model that can restore technical or physical characteristics of input data and derive output data expressed with a smaller amount of information from the input data. The fifth machine learning model may be a mathematical model that performs an operation corresponding to the inverse operation of the third machine learning model. The third and fifth machine learning models may be, for example, a CNN, a recurrent neural network (RNN), or the like.
[0055] Next, an implementation example of a machine learning model according to this embodiment will be described. FIG. 9 is a schematic block diagram showing an implementation example in an image processing system 1 according to this embodiment. In the example of FIG. 9, the image encoding device 10 includes N-layer compression layers 110-1 to 110-N and an encoding unit 108. In the compression layer 110-n, the third machine learning model functioning as the compression unit 106-n is a CNN having two convolutional layers 1 and 2, two ReLU computation units 1 and 2, and a merging unit. The two convolutional layers 1 and 2 and the two ReLU computation units 1 and 2 perform their respective operations in parallel. The output value from one convolutional layer 1 is supplied as an input value to one ReLU computation unit 1, and the output value from the other convolutional layer 2 is supplied as an input value to the other ReLU computation unit 2.
[0056] In the convolutional layer 1, weight values (corresponding to attention) for each input value constituting a plurality of sample data belonging to each kernel and a bias value (bias) for that kernel are set as one of the parameters obtained by the parameter determination unit 104-n. The convolutional layer 1 performs a convolution operation on the input values constituting the data samples of that input data for each kernel, and outputs the obtained operation value as an output value to the ReLU operation unit 1. The operation value for the convolution operation is obtained by adding the bias value for that kernel to the sum in the kernel of the multiplied values obtained by multiplying the input value by the corresponding weight value. The ReLU operation unit 1 uses the operation value for each kernel as an input value, calculates a ReLU function value for the input value, and outputs the obtained function value as an output value to the merge unit in the compression layer 110-n.
[0057] In the convolution layer 2, parameters for the convolution operation for each kernel are set as the other parameters obtained by the parameter determination unit 104-n. The convolution operation is performed for each kernel on input values, which are data samples that form the input data to the convolution layer 2, and the calculated values obtained are output as output values to the other ReLU operation unit 2. The ReLU operation unit 2 uses the calculated values for each kernel as input values, calculates a ReLU function value for the input values, and outputs the obtained function value as an output value to a merge unit in the compression layer 110-n. The merge unit outputs the input values for each kernel from the ReLU operation unit 1 and output data indicating the input values for each kernel from the ReLU operation unit 2 to the compression layer 110-n+1.
[0058] In the compression layer 110-n, the first machine learning model functioning as the context extraction unit 102-n is a CNN including a convolutional layer 2, a ReLU calculation unit 2, a statistical value calculation unit, an MLP, and a quantization unit. The convolutional layer 2 and the ReLU calculation unit 2 are also shared as part of the third machine learning model. The ReLU calculation unit 2 also outputs the output value for each kernel to the statistical value calculation unit. The statistical value calculation unit is an example of a pooling layer that performs pooling on the output value for each kernel and calculates statistics with a smaller number of elements as a representative value for that layer. The statistical value calculation unit outputs the calculated statistical value as an output value to the MLP1. An example of the statistical value calculation unit will be described later.
[0059] The MLP functions as an encoder. The output values from the statistical value calculation unit are input to the MLP as input values, and the calculated values for the input values are output to the quantization unit as output values. The quantization unit outputs n-th layer context information indicating, as elements, quantized values obtained by quantizing the output values input from the MPL to the merge unit 112-n.
[0060] In the compression layer 110-n, the second machine learning model functioning as the parameter determination unit 104-n is a neural network including two MLPs 1 and 2. The two MLPs 1 and 2 function as either decoder. One of the MLPs, MLP1, receives n-th layer accumulated context information from the merge unit 112-n. The other MLP1 uses the element values indicated in the n-th layer accumulated context information as input values, and sets the calculated values for the input values as parameters for the convolutional layer 1 of the third machine learning model in the compression unit 106-n. The other MLP2 uses the element values indicated in the (n-1)-th layer context information as input values, and sets the calculated values for the input values as parameters for the convolutional layer 2 of the third machine learning model in the compression unit 106-n.
[0061] However, in the first layer, MLP2 is omitted from the second machine learning model of the parameter determination unit 104-1. Parameters of convolutional layer 2 of the third machine learning model of the compression unit 106-1 are set in advance by model learning. Furthermore, the merging unit 112-1 is omitted from the compression layer 110-1. The quantization unit of the context extraction unit 102-1 outputs first-layer context information to the parameter determination unit 104-1 and the compression layer 110-2. The first-layer context information corresponds to first-layer accumulated context information.
[0062] The encoding unit 108 includes a context encoding unit 108a, a compressed data encoding unit 108b, and a quantization unit 108d. The context encoding unit 108a receives N-th layer accumulated context information from the compression layer 110-N. The N-th layer accumulated context information includes and combines the first to N-th layer context information. The context encoding unit 108a performs entropy encoding on the N-th layer accumulated context information to generate a first code sequence.
[0063] The quantization unit 108d receives compressed data from the compression layer 110-N, quantizes element values indicated in the compressed data, and outputs quantized compressed data indicating the quantized element values to the compressed data encoding unit 108b. The compressed data encoding unit 108b performs entropy encoding on the quantized compressed data received from the quantization unit 108d to generate a second code sequence. The encoding unit 108 outputs a code sequence including the first code sequence and the second code sequence to the image decoding device 20.
[0064] Next, we will explain an example of the functional configuration of the image decoding device 20. The image decoding device 20 includes a decoding unit 202 and N-level restoration layers 210-1 to 210-N. The decoding unit 202 decodes the code sequence input from the image encoding device 10. The decoding unit 202 includes a context decoding unit 202a and a compressed data decoding unit 202b.
[0065] The context decoding unit 202a performs entropy decoding on a first code sequence included in the input code sequence to obtain first-layer context information to N-th layer context information. The context decoding unit 202a outputs the obtained first-layer context information to N-th layer context information to the reconstruction layers 210-1 to 210-N, respectively. The compressed data decoding unit 202b performs entropy decoding on a second code sequence included in the input code sequence to convert it into compressed data. The compressed data decoding unit 202b outputs the converted compressed data to the reconstruction layer 210-N.
[0066] In the reconstruction layer 210-n, the fourth machine learning model functioning as the parameter determination unit 204-n is a neural network having an MLP, and functions as a decoder. The MLP receives n-th layer context information from the context decoding unit 202a. The MLP uses element values indicated by the n-th layer context information as input values, and sets values calculated on the input values as parameters of the deconvolution layer of the fifth machine learning model in the reconstruction unit 206-n.
[0067] In the restoration layer 210-n, the fifth machine learning model functioning as the restoration unit 206-n has a ReLU calculation unit and a deconvolution layer. The ReLU calculation unit receives the output data from the restoration layer 210-n+1 as input data, calculates a ReLU function value for the input value that constitutes the input data, and inputs the obtained function value as an output value to the deconvolution layer. The deconvolution layer receives the output value from the ReLU calculation unit as an input value for each kernel, and outputs output data indicating the calculated value for each sample obtained by performing a deconvolution operation on the input value to the restoration layer 210-n-1. In the deconvolution operation, multiple calculation values are obtained for each input value. A set of multiple samples corresponds to one kernel, and for each kernel, a weight value for each sample and one bias value become parameters of the deconvolution layer. In the deconvolution operation, the restoration layer 210-n calculates the sum of the bias value and the product of the weight value and the input value for each sample as the output value of that sample.
[0068] The ReLU calculation unit in the reconstruction layer 210-N receives compressed data as input data from the compressed data decoding unit 202b in the decoding unit 202. Reconstructed data is sent as output data from the deconvolution layer in the reconstruction layer 210-1.
[0069] Next, an example configuration of the statistical value calculation unit 122-n constituting the context extraction unit 102-n of the n-th compression layer 110-n will be described. FIG. 10 is a schematic block diagram showing an example configuration of the statistical value calculation unit 122-n according to this embodiment. The statistical value calculation unit 122-n includes at least a global average pooling (GAP) unit 1221-n. The input value of the input data input to the GAP unit 1221-n is calculated by outputting the calculated value for each kernel from the ReLU calculation unit 2 as a data sample. When the image data input to the image encoding device 10 represents, for example, a color image, the channels correspond to the primary colors that represent the color image. When the color image is expressed in the RGB color system, the channels represent any of red (R), green (G), and blue (B). The input value Y cxy are data samples for each sample arranged in a three-dimensional space stretched in mutually orthogonal horizontal (x-direction), vertical (y-direction), and color (c-direction) directions. Each sample in the first layer corresponds to a pixel, and the input value corresponds to a color signal value. The GAP unit 1221-n calculates an input value Y as a data sample acquired for each of a plurality of kernels arranged in a two-dimensional plane stretched in the horizontal and vertical directions for each channel. cxy Average value Y c The statistical value calculation unit 122-n outputs the calculated statistical value to the MLP as an output value.
[0070] The statistical value calculation unit 122-n may further include a cross product calculation unit 1222-n, a GAP unit 1223-n, a triangulation and flattening unit 1224-n, and a merging unit 1225-n. The cross product calculation unit 1222-n calculates the input value Y cxy The cross product Y between the channels cxy *Y c’xyis calculated for each channel pair (c, c') and each kernel (x, y) arranged in a two-dimensional plane. The cross product calculation unit 1222-n outputs the calculated cross product to the GAP unit 1223-n. Here, in a channel pair, channel c and channel c' may be equal. The GAP unit 1223-n calculates the cross product Y cxy *Y c’xy The average value Z between the kernels cc’ is calculated between pairs of channels. The GAP unit 1223-n calculates the calculated average value Z cc’ is output to the triangulation / flattening unit 1224-n. cc’ The set of channels c and c' is the row and column, respectively, and the average value Z cc’ can be expressed as a matrix with elements
[0071] The triangulation and flattening unit 1224-n calculates the average value Z cc’ Among them, the element value Z where channel c is equal to or less than channel c' cc’ {c, c'∈c≦c'} is adopted (triangularization). The adopted element value Z cc’ can be expressed as a triangular matrix. The triangularization and flattening unit 1224-n calculates the adopted element value Z cc’ are aligned (flattened) and the mean vector W d The triangularization and flattening unit 1224-n constructs the mean vector W d The merge unit 1225-n outputs the statistical value Y c The mean vector W d The vector value obtained by combining these is output to the MLP as a new statistical value.
[0072] Next, another example configuration of the encoding unit 108 of the image encoding device 10 and the decoding unit 202 of the image decoding device 20 will be described. Fig. 11 is a schematic block diagram showing another example configuration of the encoding unit 108 and the decoding unit 202 according to this embodiment. The encoding unit 108 includes a context encoding unit 108a, a compressed data encoding unit 108b, a quantization unit 108d, and a distribution estimation unit 108c.
[0073] The distribution estimation unit 108c estimates a probability distribution of element values representing compressed data that will be input values to the quantization unit 108d using a sixth machine learning model for the N-th layer accumulated context information input from the compression layer 110-N. The distribution estimation unit 108c can calculate the probability distribution using, for example, a Gaussian Mixture Model (GMM).
[0074] The Gaussian mixture model is a mathematical model that uses a predetermined number of normal distributions (Gaussian functions) as basis functions and expresses a continuous probability distribution as a linear combination of these basis functions. The output values from the sixth machine learning model include parameters representing the probability distribution, i.e., weighting coefficients (weight), mean, and variance, which are parameters of each normal distribution. The distribution estimation unit 108c sets the estimated probability distribution in the compressed data encoding unit 108b.
[0075] The sixth machine learning model may be, for example, an MLP. The distribution estimation unit 108c may use each element value of the N-th layer cumulative context information as an input value to the sixth machine learning model and calculate parameters of a probability distribution as an output value. The distribution estimation unit 108c sets the probability distribution represented by the calculated parameters to the quantization unit 108d.
[0076] The compressed data encoding unit 108b uses the probability distribution set by the distribution estimation unit 108c to perform entropy encoding on the element values of the quantized compressed data input from the quantization unit 108d. By entropy encoding, a larger amount of information is assigned to compressed data with less entropy estimated from the probability distribution, thereby reducing the overall amount of information in the code sequence generated by encoding.
[0077] The decoding unit 202 includes a context decoding unit 202a, a compressed data decoding unit 202b, and a distribution estimation unit 202c. The distribution estimation unit 202c estimates a probability distribution of element values representing compressed data decoded from the second code sequence using a seventh machine learning model, similar to the method used by the distribution estimation unit 108c in the encoding unit 108, for the N-th layer accumulated context information input from the context decoding unit 202a. As the seventh machine learning model, for example, MLP can be used. The distribution estimation unit 202c sets the estimated probability distribution in the compressed data decoding unit 202b. The compressed data decoding unit 202b performs entropy decoding on the second code sequence using the probability distribution set by the distribution estimation unit 202c, and decodes the compressed data.
[0078] In addition, as the entropy encoding method in the compressed data encoding unit 108b and the entropy decoding method in the compressed data decoding unit 202b, for example, the encoding method and decoding method described in International Publication No. WO2022 / 130477 can be applied.
[0079] 3-6, 9, and 10, even when the encoding process and the decoding process are layered into multiple layers, the model training unit 330 can determine the parameters of the first machine learning model, the second machine learning model, the third machine learning model, the fourth machine learning model, and the fifth machine learning model for each layer so as to minimize the difference between the image data to be encoded and the restored data obtained by decoding. Furthermore, even when a probability distribution is used for entropy encoding and entropy decoding, as illustrated in FIG. 11, the model training unit 330 can determine the parameters of the sixth machine learning model and the seventh machine learning model by including the entropy encoding and entropy decoding processes in obtaining the restored data.
[0080] The encoding unit 108 may combine the generated first and second code sequences and transmit them as a single code sequence, or may transmit them separately without combining them. When combining and transmitting the first and second code sequences, the first and second code sequences may be assigned to different predetermined timings. When transmitting the first and second code sequences separately, the first and second code sequences may be transmitted to different transmission paths or storage areas.
[0081] When an integrated code sequence is transmitted, the decoding unit 202 extracts a first code sequence and a second code sequence from the acquired code sequence and provides the extracted first code sequence and second code sequence to the compressed data decoding unit 202b and the context decoding unit 202a, respectively. When the first code sequence and the second code sequence are assigned to different timings, the decoding unit 202 may extract the first code sequence and the second code sequence from the code sequence according to the respective timings. When the first code sequence and the second decoded sequence are transmitted individually, the first code sequence and the second decoded sequence may be provided to the compressed data decoding unit 202b and the context decoding unit 202a, respectively.
[0082] <Experimental Example> Next, a description will be given of a simulation experiment conducted by the applicant to verify the effectiveness of the image processing system 1 according to this embodiment. As indicators of effectiveness, processing delay and MS-SSIM (Multi-Scale Structural Similarity) index were measured. To verify the effectiveness of this embodiment, the results were compared with indicators obtained by executing other methods on a computer system equipped with the same GPU (Graphic Processing Unit).
[0083] The processing delay was measured as the time from input of image data to the image encoding device 10 to output of restored data from the image decoding device 20. MS-SSIM is an example of an index value for image quality. A larger value indicates higher image quality. However, in this experiment, the configuration examples shown in Figures 9 to 11 were adopted as this embodiment. In convolution layers 1 and 2 in each compression layer, the kernel size of each kernel was set to 3 pixels horizontally and 3 pixels vertically, with a stride of 2 pixels in both the horizontal and vertical directions. The stride corresponds to the interval at which calculations are applied. In the deconvolution layer in each restoration layer, the kernel size was set to 3 pixels horizontally and 3 pixels vertically, with a stride of 2 pixels. Furthermore, for each method, the image size was set to 1920 pixels horizontally and 1080 pixels vertically (1080p).
[0084] FIG. 12 is a diagram showing an example of processing delay. The vertical axis of FIG. 12 indicates processing delay (unit: ms (milliseconds)), and the horizontal axis indicates the method. Methods (1) to (5) are all comparative examples, and method (6) is the present embodiment. Methods (1) to (4) are all "no parameter selection" methods that do not use context information extracted from image data to select parameters to be used for data compression or data decompression. Methods (1) to (4) differ in the amount of processing. N=1 to 4 is an index indicating the level of processing amount. N is approximately proportional to the number of parameters. Method (5) is a method "no parameter selection, no encoding / decoding" that selects parameters corresponding to context information, but does not involve encoding or decoding of the context information.
[0085] According to methods (1) to (4), the processing delay increases as the processing volume increases. For example, the processing delay in method (1) (N=1) was 40 ms, while the processing delay in method (4) (N=4) was 130 ms. In methods (5) and (6), the processing delay was 45 ms and 50 ms, respectively. The difference in processing delay between methods (1), (5), and (6) is relatively small.
[0086] FIG. 13 is a diagram showing an example of MS-SSIM. The vertical axis of FIG. 13 represents MS-SSIM, and the horizontal axis represents bpp (bits per pixel). bpp is a unit of information amount transmitted per pixel. A smaller bpp indicates higher coding efficiency. Generally, each technique has a relationship in which the greater the amount of information, the higher the quality. FIG. 13 shows that the quality increases in the order of techniques (1), (5), (2), (7), (3), (6), and (4). Technique (7) is a technique that selects parameters corresponding to context information as "parameter selection, GAP encoding / decoding," but involves encoding and decoding of statistical values that are part of the context information (i.e., output values from GAP). However, it does not involve encoding and decoding of other parts of the context information, such as the filter coefficients of each kernel (i.e., weight values and bias values for each sample).
[0087] 13 shows that, under the same amount of processing, the selection of parameters corresponding to context information, the encoding and decoding of statistical values, and the encoding and decoding of the entire context information each contribute to improving the quality of the restored data. Furthermore, there is no significant difference in the quality of the restored data among methods (3), (6), and (4). This indicates that while an increase in processing volume beyond a certain level limits the improvement in the quality of the restored data, encoding and decoding of context information can improve the quality of the restored data without a significant increase in processing volume. Note that method (6) according to this embodiment can improve the quality of the restored data more than methods such as JPEG (Joint Photographic Experts Group), which is commonly used for general still image compression, and ITU-T H.264 (AVC: Advanced Video Coding) and ITU-T H.265 (HEVC: High Efficiency Video Coding), which are commonly used for video compression.
[0088] <Modifications> Next, a first modification of this embodiment will be described. Generally, a series of moving images is represented by switching between still images of each frame at regular time intervals. If an image does not change over time, the context information indicating its characteristics also does not change. Furthermore, if the image changes little, or if the context information changes little, the context information also changes little. Therefore, the context extraction unit 102 of the image encoding device 10 may stop extracting context information if the amount of variation between the image data of the frame from which context information was last extracted (hereinafter sometimes referred to as the "reference frame") and the image data of the current frame is within a predetermined reference value of variation. In this case, new context information is not provided to the parameter determination unit 104. When compressing the image data of the current frame, the compression unit 106 continues to use the parameters of the third machine learning model corresponding to the last context information of the reference frame.
[0089] Because the encoding unit 108 does not perform encoding on the new context information, the first code sequence for the current frame is not provided to the image decoding device 20. Consequently, no new context information is provided to the parameter determination unit 204 of the image decoding device 20. The reconstruction unit 206 continues to use the parameters of the fifth machine learning model corresponding to the last context information for the reference frame in reconstruction of the reconstructed data for the current frame.
[0090] The context extraction unit 102 determines, for each frame, whether the amount of variation from the image data of the reference frame to the image data of the current frame is within a predetermined reference value of the amount of variation. In this case, the context extraction unit 102 can use, for example, the sum of squared differences (SSD) or the sum of absolute differences (SAD) of signal values as an index of the amount of variation from the image data of the reference frame to the image data of the current frame. Furthermore, instead of the inter-frame difference in signal values for each pixel, the context extraction unit 102 can use the norm (distance) between vectors representing context information related to the reference frame and the current frame as an index value of the amount of variation.
[0091] If the image encoding device 10 is layered, i.e., if it has multiple compression layers, each compression layer may determine whether to stop extracting context information based on whether the amount of variation is within a predetermined reference value. Also, at least one compression layer (e.g., the first compression layer 110-1) may determine whether to stop extracting context information for all layers based on whether the amount of variation from a reference frame to a current frame is within a predetermined reference value.
[0092] Next, a second modified example of this embodiment will be described. Information on a target information amount (bit rate) or target image quality may be set as setting information in the context extraction unit 102 of the image encoding device 10, and the set setting information may be used in calculations as part of input values for the first machine learning model to extract context information. The target information amount is a target value for the information amount of the entire code sequence including a first code sequence indicating the context information and a second code sequence indicating the compressed data. The target information amount may be defined in any unit, such as the number of bits per pixel, the number of bits per frame, or the number of bits per second (bit rate). The target image quality is a target value for the image quality of the context information obtained by decoding the code sequence and the restored data restored from the compressed data. The target image quality may be defined using any index, such as MS-SSIM or SNR (Signal-to-Noise Ratio).
[0093] When a target information amount is set in the context extraction unit 102, the model training unit 330 determines parameters for each of the first to fifth machine learning models so that the difference between the information amount of the code sequence and the target information amount is reduced while the target information amount is input to the first machine learning model. In model training, the model training unit 330 of the model training device 30 determines parameters for each of the first to fifth machine learning models so that a loss function including a first factor indicating the magnitude of the difference between the image data and the restored data and a second factor indicating the magnitude of the difference between the information amount of the code sequence and the target information amount is reduced. The model training unit 330 determines the information amount of the code sequence input from the encoding unit 108 of the image encoding device 10 or the encoding unit 318 of its own device, and uses the determined information amount for model training. Prior to model training, the model training unit 330 previously acquires the target information amount set in the context extraction unit 102 of the image encoding device 10 or the context extraction unit 312 of its own device.
[0094] When setting a target image quality in the context extraction unit 102, the model training unit 330 determines parameters for each of the first to fifth machine learning models in model training so that the target image quality is input to the first machine learning model and the difference between the information amount of the code sequence and the target information amount is reduced. In model training, the model training unit 330 of the model training device 30 determines parameters for each of the first to fifth machine learning models so that a loss function including a first factor indicating the magnitude of the difference between the image data and the restored data and a second factor indicating the magnitude of the difference between the image quality of the restored image indicated by the restored data and the target image quality is reduced. The model training unit 330 determines the image quality of the restored image indicated by the restored data input from the restoration unit 206 of the image decoding device 20 or the restoration unit 326 of its own device, and uses the determined image quality for model training. Prior to model training, the model training unit 330 previously acquires the target image quality to be set in the context extraction unit 102 of the image encoding device 10 or the context extraction unit 312 of its own device.
[0095] As described above, the image encoding device 10 according to this embodiment includes a context extraction unit 102 that extracts context information indicating characteristics of the image data using a first machine learning model on the image data, a parameter determination unit 104 that determines parameters of a third machine learning model on the context information using a second machine learning model on the context information, a compression unit 106 that uses the third machine learning model on the image data to generate compressed data with a smaller amount of data than the image data, and an encoding unit 108 that encodes the context information and the compressed data to generate a code sequence. Furthermore, the image decoding device 20 according to this embodiment includes a decoding unit 202 that decodes the context information and compressed data from the code sequence, a parameter determination unit 204 that determines parameters of a fifth machine learning model on the context information using a fourth machine learning model, and a restoration unit 206 that generates restored data using the fifth machine learning model on the compressed data.
[0096] According to this configuration, parameters of a third machine learning model are determined using the second machine learning model for context information indicating characteristics of the image data extracted from the image data using the first machine learning model, and parameters of a fifth machine learning model are determined using the fourth machine learning model for the context information for generating decompressed data from the compressed data.
[0097] Therefore, parameters of the third machine learning model are determined according to the characteristics of the image data, and these parameters are estimated based on the context information decoded by the fourth machine learning model. By effectively using dynamic parameter selection, an image that approximates the original image data can be restored from the code obtained by encoding without increasing the amount of processing.
[0098] In the image encoding device 10, the third machine learning model may include a convolution layer that performs a convolution operation for each kernel having multiple data samples, and the parameters of the third machine learning model may include at least a weighting coefficient for the data sample. In the image decoding device 20, the fifth machine learning model may include a deconvolution layer that performs a deconvolution operation for each kernel having one or more data samples, and the parameters of the fifth machine learning model may include at least a deconvolution coefficient for the data sample. With this configuration, the weighting coefficients used in the convolution operation for the data samples are estimated using the third machine learning model, and the deconvolution coefficients used in the deconvolution operation for the data samples are estimated using the fifth machine learning model. Therefore, the estimated weighting coefficients and deconvolution coefficients enable efficient compression of information volume through convolution operation according to image features and restoration through deconvolution operation.
[0099] Furthermore, in the image encoding device 10, the first machine learning model may include a pooling layer that determines a representative value based on the output value of each kernel. This configuration allows the parameters of the third machine learning model to be determined by analyzing global characteristics and the strength of those characteristics rather than kernels consisting of multiple data samples. This can contribute to efficient compression of information volume through convolution operations based on the global characteristics and the strength of those characteristics.
[0100] Furthermore, the image encoding device 10 may receive image data for each frame, and the context extraction unit 102 may stop extracting context information when the amount of variation between the image data of the frame from which context information was last extracted and the image data of the current frame is within a predetermined reference value. According to this configuration, when the variation in the image data is small, context information is not extracted. Therefore, the context information is not encoded, and compressed data is generated using parameters of a third machine learning model corresponding to the last extracted context information. Furthermore, restored data is generated using parameters of a fifth machine learning model corresponding to that context information. Therefore, it is possible to reduce the amount of information in the coded sequence resulting from encoding while minimizing degradation in image quality due to the restored data.
[0101] The context extraction unit 102 may also determine context information based on the image data and a target information amount or target image quality. The target information amount is a target value for the information amount of a code sequence including the context information and the compressed data code, and the target image quality is a target value for the image quality of restored data restored from the context information and the compressed data. This configuration makes it possible to obtain a code sequence having the set target information amount or restored data having the set target image quality. Therefore, it is possible to compress and restore image data according to the required information amount or image quality.
[0102] Furthermore, the image encoding device 10 may have a first machine learning model, a second machine learning model, and a third machine learning model in N (N is an integer of 2 or more) layers (e.g., compression layers 110-1 to 110-N), wherein output data from the third machine learning model in the (n-1)th layer is input to the third machine learning model in the n-th (n is an integer of 2 or more and N or less) layer, and output data from the first machine learning model in the n-th layer is input to the second machine learning model in the n-th layer, and the encoding unit 108 may encode the output data from the third machine learning model in the N-th layer as compressed data and encode the output data from the first machine learning models from the first to N layers as context information. Furthermore, the image decoding device 20 may have a fourth machine learning model and a fifth machine learning model in N layers (e.g., restoration layers 210-1 to 210-N), wherein output data from the fifth machine learning model in the n-1th layer is input to the fifth machine learning model in the n-th layer, and the context information of the n-th layer is input to the fourth machine learning model in the n-th layer. According to this configuration, context information from the first machine learning model in each layer is accumulated and encoded, and compressed data with a high compression ratio is obtained from the image data by the third machine learning model in each layer. Then, parameters of the fifth machine learning model are obtained from the context information decoded by the fourth machine learning model in each layer. Therefore, restored data that approximates the original image data can be obtained from the compressed data with a high compression ratio. This allows for improved encoding efficiency without degrading the image quality of the restored data.
[0103] Furthermore, the model learning device 30 according to this embodiment includes a context extraction unit 312 that extracts context information indicating characteristics of the image data using a first machine learning model from the image data, a parameter determination unit 314 that determines parameters of a third machine learning model using a second machine learning model from the context information, a compression unit 316 that uses the third machine learning model from the image data to generate compressed data with a smaller amount of information than the image data, a parameter determination unit 324 that determines parameters of a fifth machine learning model using a fourth machine learning model from the context information, a restoration unit 326 that generates restored data using the fifth machine learning model from the compressed data, and a model learning unit 330 that determines parameters of the first machine learning model, the second machine learning model, the third machine learning model, the fourth machine learning model, and the fifth machine learning model so as to reduce the difference between the image data and the restored data. This configuration makes it possible to obtain the parameters of the first machine learning model, the second machine learning model, and the third machine learning model used in the image encoding device 10, and the parameters of the fourth machine learning model and the fifth machine learning model used in the image decoding device 20. This can contribute to the restoration of an image that approximates the original image data from the code obtained by encoding, without increasing the amount of processing.
[0104] (Minimum Configuration) Next, the minimum configuration of the above embodiment will be described. Fig. 14 is a schematic block diagram showing an example of the minimum configuration of an image encoding device 10 of the present application. The image encoding device 10 includes a context extraction unit 102 that extracts context information indicating characteristics of the image data using a first machine learning model on the image data, a parameter determination unit 104 that determines parameters of a third machine learning model using a second machine learning model on the context information, a compression unit 106 that uses the third machine learning model on the image data to generate compressed data with a smaller amount of data than the image data, and an encoding unit 108 that encodes the context information and the compressed data and generates a code sequence.
[0105] 15 is a schematic block diagram showing an example of a minimum configuration of the image decoding device 20 of the present application. The image decoding device 20 includes a decoding unit 202 that decodes context information and compressed data from a code sequence, a parameter determination unit 204 that determines parameters of a fifth machine learning model using a fourth machine learning model on the context information, and a restoration unit 206 that generates restored data using the fifth machine learning model on the compressed data.
[0106] 16 is a schematic block diagram showing an example of a minimum configuration of an image processing system 1 of the present application. The image processing system 1 includes an image encoding device 10 and an image decoding device 20. The image encoding device 10 includes a context extraction unit 102 that extracts context information indicating characteristics of image data using a first machine learning model on the image data, a parameter determination unit 104 that determines parameters of a third machine learning model on the context information using a second machine learning model on the context information, a compression unit 106 that generates compressed data with a smaller data volume than the image data using the third machine learning model on the image data, and an encoding unit 108 that encodes the context information and the compressed data and generates a code sequence. The image decoding device 20 includes a decoding unit 202 that decodes the context information and the compressed data from the code sequence, a parameter determination unit 204 that determines parameters of a fifth machine learning model on the context information using a fourth machine learning model, and a restoration unit 206 that generates restored data using the fifth machine learning model on the compressed data.
[0107] 17 is a schematic block diagram showing an example of a minimum configuration of a model learning device 30 of the present application. The model learning device 30 includes a context extraction unit 312 that extracts context information indicating characteristics of image data using a first machine learning model, a parameter determination unit 314 that determines parameters of a third machine learning model using a second machine learning model on the context information, a compression unit 316 that generates compressed data with a smaller data volume than the image data using the third machine learning model on the image data, a parameter determination unit 324 that determines parameters of a fifth machine learning model using a fourth machine learning model on the context information, a restoration unit 326 that generates restored data using the fifth machine learning model on the compressed data, and a model learning unit 330 that determines parameters of the first, second, third, fourth, and fifth machine learning models so as to reduce the difference between the image data and the restored data.
[0108] The above-mentioned devices, such as the image encoding device 10, the image decoding device 20, the model learning device 30, and edge devices and server devices including any of them, may each include a computer system. The computer system includes one or more processors, such as a CPU (Central Processing Unit). The processors included in the computer system may include one or more GPUs. The processing steps of each of the above-mentioned components are stored in a computer-readable storage medium in the form of a program for each device or apparatus. The computer reads the instructions written in the program and executes the processing indicated by the instructions, thereby achieving the functions of the components. The program may include the above-mentioned machine learning model. The computer system includes software such as an operating system (OS), device drivers, and utility programs, as well as hardware such as a processor, storage media, and peripheral devices. Furthermore, the term "computer-readable storage medium" refers to portable media such as magnetic disks, magneto-optical disks, read-only memories (ROMs), and semiconductor memories, as well as storage devices such as hard disks built into the computer system. Furthermore, the term "computer-readable recording medium" may also include a medium that dynamically stores a program for a short period of time, such as a communication line used when transmitting a program over a network such as the Internet or a communication line such as a telephone line, or a medium that stores a program for a certain period of time, such as volatile memory within a computer system that serves as a server or client in such a case. The program may also be one that realizes part of the above-mentioned functions, or one that can realize the above-mentioned functions in combination with a program already recorded in the computer system, such as a so-called differential file (differential program).
[0109] For example, the image encoding device 10 may have a hardware configuration including a processor 152, a drive unit 156, an input / output unit 158, a ROM 162, and a RAM (Random Access Memory) 164, as illustrated in FIG. 18 . The processor 152 controls processes for enabling the image encoding device 10 to function and the functions of each unit constituting the image encoding device 10. The drive unit 156 includes an auxiliary storage device that reads various data stored in the storage medium 154 or stores various data in the storage medium 154. The drive unit 156 may be, for example, a solid-state drive (SSD) or a hard disk drive (HDD). The storage medium 154 is, for example, a non-volatile memory such as a flash memory. The drive unit 156 may be configured so that the storage medium 154 is detachable. The input / output unit 158 inputs or outputs various data to or from other devices wirelessly or via a wired connection. The input / output unit 158 may be connected to other devices via a communication network so that various types of data can be input and output. The input / output unit 158 may include, for example, an input / output interface, a communication interface, or a combination thereof.
[0110] The ROM 162 continuously stores programs in which instructions for various processes to be executed by each unit of the image encoding device 10 are written, various data such as parameters for the execution of the programs, and various data acquired by each unit of the image encoding device 10. The RAM 164 is mainly used as a work area (main storage area) for the processor 152. The processor 152 records the programs and parameters stored in the ROM 162 in the RAM 164 upon startup. The processor then temporarily records the calculation results obtained by the execution of the programs and the acquired data in the RAM 164. Note that the image decoding device 20, the model learning device 30, and part or all of the edge device and server device including either of them may also have the hardware configuration illustrated in FIG. 18.
[0111] Furthermore, some or all of the devices or apparatuses in the above-described embodiments may be realized as integrated circuits such as LSIs (Large Scale Integration). Each functional block, each unit, and each step of each device or apparatus may be individually implemented as a processor, or may be integrated in part or in whole as a processor, or may be configured as a module. Furthermore, the integrated circuit implementation method is not limited to LSIs, and may be implemented using dedicated circuits or general-purpose processors. Furthermore, if an integrated circuit implementation technology that can replace LSIs emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.
[0112] The above embodiment may be realized as follows: (Supplementary Note 1) An image encoding device including: a context extraction unit that extracts context information indicating characteristics of image data using a first machine learning model for the image data; a parameter determination unit that determines parameters of a third machine learning model using a second machine learning model for the context information; a compression unit that uses the third machine learning model for the image data to generate compressed data with a smaller amount of data than the image data; and an encoding unit that encodes the context information and the compressed data to generate a code sequence.
[0113] (Supplementary Note 2) In the image encoding device of Supplementary Note 1, the third machine learning model includes a convolution layer that performs a convolution operation for each kernel having a plurality of data samples, and the parameters include at least weighting coefficients for the data samples.
[0114] (Supplementary Note 3) In the image encoding device of Supplementary Note 2, the first machine learning model includes a pooling layer that determines a representative value based on the output value for each kernel.
[0115] (Supplementary Note 4) In the image encoding device of Supplementary Note 1, the image data is input for each frame, and the context extraction unit stops extracting the context information when the amount of variation between the image data of the frame from which context information was last extracted and the image data of the current frame is within a predetermined reference value.
[0116] (Supplementary Note 5) In the image encoding device of Supplementary Note 1, the context extraction unit determines the context information based on the image data and further on a target information amount or a target image quality, the target information amount being a target value of the information amount of a code sequence including the context information and the code of the compressed data, and the target image quality being a target value of the image quality of the restored data restored from the context information and the compressed data.
[0117] (Supplementary Note 6) An image encoding device according to Supplementary Note 1, comprising the first machine learning model, the second machine learning model, and the third machine learning model in N (N is an integer greater than or equal to 2) layers, wherein output data from the third machine learning model in the n-1th layer is input to the third machine learning model in the nth (n is an integer greater than or equal to 2 and less than or equal to N) layer, and output data from the first machine learning model in the nth layer is input to the second machine learning model in the nth layer, and the encoding unit encodes the output data from the third machine learning model in the Nth layer as the compressed data and encodes the output data from the first to Nth layers as the context information.
[0118] (Supplementary Note 7) An image decoding device comprising: a decoding unit that decodes context information and compressed data from a code sequence; a parameter determination unit that determines parameters of a fifth machine learning model using a fourth machine learning model on the context information; and a restoration unit that generates restored data using the fifth machine learning model on the compressed data.
[0119] (Supplementary Note 8) An image decoding device according to Supplementary Note 7, wherein the fifth machine learning model comprises a deconvolution layer that performs a deconvolution operation for each kernel having one or more data samples, and the parameters include at least deconvolution coefficients of the data samples.
[0120] (Supplementary Note 9) The image decoding device of Supplementary Note 7 has the fourth machine learning model and the fifth machine learning model in N (N is an integer greater than or equal to 2) layers, and the fifth machine learning model in the n-1 (n is an integer greater than or equal to 2 and less than or equal to N) layer receives output data from the fifth machine learning model in the nth layer, and the fourth machine learning model in the nth layer receives the context information in the nth layer.
[0121] (Supplementary Note 10) A computer-readable storage medium storing a program for causing a computer to function as the image encoding device of Supplementary Note 1 or the image decoding device of Supplementary Note 7.
[0122] (Supplementary Note 11) An image processing system comprising an image encoding device and an image decoding device, wherein the image encoding device comprises: a context extraction unit that extracts context information indicating characteristics of image data using a first machine learning model on the image data; a parameter determination unit that determines parameters of a third machine learning model on the context information using a second machine learning model on the context information; a compression unit that uses the third machine learning model on the image data to generate compressed data having less information than the image data; and an encoding unit that encodes the context information and the compressed data and generates a code sequence; and the image decoding device comprises: a decoding unit that decodes the context information and the compressed data from the code sequence; a parameter determination unit that determines parameters of a fifth machine learning model on the context information using a fourth machine learning model; and a restoration unit that generates restored data using the fifth machine learning model on the compressed data.
[0123] (Supplementary Note 12) The image processing system of Supplementary Note 11, wherein the encoding unit separately outputs a first code sequence generated by encoding the context information and a second code sequence generated by encoding the compressed data, and the decoding unit decodes the context information from the first code sequence and decodes the compressed data from the second code sequence.
[0124] (Supplementary Note 13) A model learning device comprising: a context extraction unit that extracts context information indicating characteristics of image data using a first machine learning model from the image data; a parameter determination unit that determines parameters of a third machine learning model using a second machine learning model from the context information; a compression unit that uses a third machine learning model from the image data to generate compressed data having less information than the image data; a parameter determination unit that determines parameters of a fifth machine learning model using a fourth machine learning model from the context information; a restoration unit that generates restored data using the fifth machine learning model from the compressed data; and a model learning unit that determines parameters of the first machine learning model, the second machine learning model, the third machine learning model, the fourth machine learning model, and the fifth machine learning model so that a difference between the image data and the restored data is reduced.
[0125] (Supplementary Note 14) An image encoding method in an image encoding device, wherein the image encoding device executes a context extraction step of extracting context information indicating characteristics of image data using a first machine learning model for the image data, a parameter determination step of determining parameters of a third machine learning model for the context information using a second machine learning model for the context information, a compression step of generating compressed data having less information than the image data using the third machine learning model for the image data, and an encoding step of encoding the context information and the compressed data to generate a code sequence.
[0126] (Supplementary Note 15) An image decoding method in an image decoding device, wherein the image decoding device executes a decoding step of decoding context information and compressed data from a code sequence, a parameter determination step of determining parameters of a fifth machine learning model using a fourth machine learning model on the context information, and a restoration step of generating restored data using the fifth machine learning model on the compressed data.
[0127] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments and their variations. Addition, omission, substitution, and other modifications of the configuration are possible without departing from the spirit of the present invention. The direction of arrows shown in block diagrams and other drawings is for the convenience of explanation, and the disclosure of this application does not limit the direction of the flow of information, data, signals, etc. during implementation. Furthermore, the present invention is not limited by the above description, but is limited only by the appended claims.
[0128] The image encoding device, image decoding device, image processing system, model learning device, image decoding method, and computer-readable storage medium of each of the above aspects can be used for encoding, compressing, transmitting, decoding, etc., of various types of data including one or both of still images and moving images.
[0129] 1...image processing system, 10...image encoding device, 20...image decoding device, 30...model learning device, 102, 102-n...context extraction unit, 104, 104n...parameter determination unit, 106, 106-n...compression unit, 108...encoding unit, 108a...context encoding unit, 108b...compressed data encoding unit, 108d...quantization unit, 110-1 to 110-N...compression layer, 112-n...merge unit, 152...processor, 156...drive unit, 158...input / output unit, 162...ROM, 164...RAM, 202...decoding unit, 202 a...context decoding unit, 202b...compressed data decoding unit, 204, 204-n...parameter determination unit, 206, 206-n...reconstruction unit, 210-1 to 210-N...reconstruction layer, 212-n...merging unit, 312...context extraction unit, 314...parameter determination unit, 316...compression unit, 318...encoding unit, 318a...context encoding unit, 318b...compressed data encoding unit, 322...decoding unit, 322a...context decoding unit, 322b...compressed data decoding unit, 324...parameter determination unit, 326...reconstruction unit, 330...model learning unit
Claims
1. a context extraction unit that extracts context information indicating characteristics of the image data by using a first machine learning model for the image data; a parameter determination unit that determines parameters of a third machine learning model using the second machine learning model for the context information; a compression unit that generates compressed data having a smaller amount of data than the image data by using a third machine learning model on the image data; an encoding unit that encodes the context information and the compressed data to generate a code sequence; Image encoding device.
2. The third machine learning model includes a convolution layer that performs a convolution operation for each kernel having a plurality of data samples; The parameters include at least weighting factors for the data samples. The image encoding device according to claim 1 .
3. The first machine learning model includes a pooling layer that determines a representative value based on the output value of each kernel. The image encoding device according to claim 2 .
4. The image data is input for each frame, The context extraction unit stops extracting the context information when an amount of change between image data of a frame from which context information was last extracted and the image data of a current frame is within a predetermined reference value. The image encoding device according to claim 1 .
5. The context extraction unit determining the context information based on the image data and further based on a target information amount or a target image quality; the target information amount is a target value of an information amount of a code sequence including the context information and the code of the compressed data, The target image quality is a target value of the image quality of the restored data restored from the context information and the compressed data. The image encoding device according to claim 1 .
6. The first machine learning model, the second machine learning model, and the third machine learning model have N layers (N is an integer equal to or greater than 2); The third machine learning model in the n-th layer (n is an integer equal to or greater than 2 and equal to or less than N) receives output data from the third machine learning model in the n-1-th layer, The second machine learning model in the nth layer receives output data from the first machine learning model in the nth layer, The encoding unit encodes output data from the third machine learning model in an N-th layer as the compressed data, and encodes output data from the first machine learning model from the first layer to the N-th layer as the context information. The image encoding device according to claim 1 .
7. a decoding unit for decoding the context information and the compressed data from the code sequence; a parameter determination unit that determines parameters of a fifth machine learning model using a fourth machine learning model for the context information; A restoration unit that generates restored data by using the fifth machine learning model on the compressed data. Image decoding device.
8. The fifth machine learning model includes a deconvolution layer that performs a deconvolution operation for each kernel having one or more data samples; The parameters include at least the deconvolution coefficients of the data samples. The image decoding device according to claim 7.
9. The fourth machine learning model and the fifth machine learning model have N layers (N is an integer equal to or greater than 2), The fifth machine learning model in the n-1th layer (n is an integer between 2 and N) receives output data from the fifth machine learning model in the nth layer, The fourth machine learning model of the nth layer is input with the context information of the nth layer. The image decoding device according to claim 7.
10. A program for causing a computer to function as the image encoding device according to claim 1 or the image decoding device according to claim 7.