Image Encoding Method, Image Decoding Method, Image Processing Method, Image Encoding Apparatus, and Image Decoding Apparatus
The image encoding method improves task processing accuracy by applying region-specific encoding processes and transmitting relevant parameters, addressing redundant processing in existing systems.
Patent Information
- Application Number
- JP2023516421
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-23
- Filing Date
- 2022-03-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Existing image processing systems redundantly execute the same task processing on both the encoder and decoder sides, leading to inefficiencies and reduced accuracy in task processing.
An image encoding method that differentiates between regions in an image, applying distinct encoding processes for regions with and without objects based on parameters from a task processing apparatus, and transmits these parameters with the encoded bitstream to improve decoding accuracy.
Enhances the accuracy of task processing by ensuring higher image quality and resolution in regions containing objects, thereby improving the overall performance of subsequent neural network tasks.
Smart Images

Figure 0007704842000001 
Figure 0007704842000002 
Figure 0007704842000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding method, an image decoding method, an image processing method, an image encoding apparatus, and an image decoding apparatus.
Background Art
[0002] A neural network is a series of algorithms that attempts to recognize the underlying relationships in a dataset through a process that mimics the way the human brain processes information. In this sense, a neural network essentially refers to a system of organic or artificial neurons. Different types of neural networks in deep learning, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and artificial neural networks (ANNs), are changing the way we interact with the world. These different types of neural networks are at the core of power applications such as the deep learning revolution, drones, self-driving cars, and speech recognition. A CNN, consisting of multiple stacked layers, is the most commonly applied class of deep neural networks for visual image analysis.
[0003] Edge artificial intelligence (Edge AI) is a system that uses machine learning algorithms to process data generated by hardware sensors at the local level. In order to process such data and make decisions in real time, there is no need to connect the system to the Internet. In other words, Edge AI leads data and its processing to the closest contact point to the user, whether it is a computer, an IoT device, or an edge server. For example, in the setup of a surveillance camera system, the camera system can be deployed using Edge AI consisting of a simple neural network. Edge AI may be deployed to perform simple task analysis that requires real-time results. The camera system typically also includes a video / audio codec. The video / audio codec compresses video / audio for efficient transmission to a recording server. Another neural network may be deployed to a cloud or server to perform complex task analysis.
[0004] The image encoding system architecture according to the background art is disclosed in, for example, Patent Documents 1 and 2.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
[0006] The present disclosure aims to improve the accuracy of task processing.
[0007] An image encoding method according to an aspect of the present disclosure is such that an image encoding apparatus encodes a first region including an object in the image by a first encoding process based on one or more parameters related to the object included in the image and input from a first processing apparatus that executes a predetermined task process based on the image, encodes a second region not including the object in the image by a second encoding process based on the one or more parameters, generates a bitstream by encoding the first region and the second region, and transmits the generated bitstream to an image decoding apparatus.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Embodiments for Carrying Out the Invention
[0009] (Knowledge on which the present disclosure is based) FIG. 3 is a diagram showing a configuration example of an image processing system 1100 according to the background art. The image encoding apparatus 1102 inputs an image captured by a camera and outputs a compressed bitstream by encoding the image. The bitstream is transmitted from the image encoding apparatus 1102 to the image decoding apparatus 1103 via a communication network.
[0010] On the encoder side, the image captured by the camera is also input to the task processing unit 1101. The task processing unit 1101 executes a predetermined neural network task based on the input image. The task processing unit 1101 executes, for example, a face detection process for detecting the face of a person included in the image. The processing result R1 (detected face bounding box) of the task processing by the task processing unit 1101 is fed back to the image capture device.
[0011] The image decoder 1103 receives the bitstream transmitted from the image encoder 1102, decodes the bitstream, and inputs the decompressed image to the task processing unit 1104. The task processing unit 1104 executes the same neural network task (in this example, face detection processing) as the task processing unit 1101 based on the input image.
[0012] The processing result R1 (detected face bounding box) of the task processing by the task processing unit 1104 is input to the task processing unit 1105. The task processing unit 1105 executes a predetermined neural network task based on the input processing result R1. The task processing unit 1105 executes, for example, a person identification process for identifying the person based on the face of the person included in the image, and outputs the processing result R2.
[0013] The problem of the background art shown in FIG. 3 is that the same task processing is redundantly executed by the task processing unit 1101 on the encoder side and the task processing unit 1104 on the decoder side.
[0014] To solve such a problem, the inventor introduced a new method of signaling the output of the previous neural network. The concept is to utilize the information from the previous neural network in order to retain important data of the image transmitted to the subsequent neural network. By using this information, the size of the encoded bitstream can be further reduced, and the accuracy of determination in the task processing of the subsequent neural network can be improved.
[0015] Next, each aspect of the present disclosure will be described.
[0016] An image encoding method according to an aspect of the present disclosure is such that an image encoding apparatus encodes a first region including an object in the image by a first encoding process based on one or more parameters regarding the object included in the image, which are input from a first processing apparatus that executes a predetermined task process based on the image, encodes a second region not including the object in the image by a second encoding process based on the one or more parameters, generates a bitstream by encoding the first region and the second region, and transmits the generated bitstream toward an image decoding apparatus.
[0017] According to this aspect, the image encoding apparatus encodes a first region including an object by a first encoding process based on one or more parameters input from the first processing apparatus, and encodes a second region not including the object by a second encoding process. As a result, the image encoding apparatus can set the first encoding process and the second encoding process so that the first region has higher image quality or higher resolution than the second region, and the image decoding apparatus can output an image in which the first region has higher image quality or higher resolution than the second region toward the second processing apparatus. As a result, it becomes possible to improve the accuracy of the task process in the second processing apparatus.
[0018] In the above aspect, the image encoding apparatus adds the one or more parameters to the bitstream and transmits the bitstream toward the image decoding apparatus.
[0019] According to this aspect, the image decoding apparatus can output the one or more parameters received from the image encoding apparatus toward the second processing apparatus. As a result, by the second processing apparatus executing a predetermined task process based on the one or more parameters input from the image decoding apparatus, it becomes possible to further improve the accuracy of the task process in the second processing apparatus.
[0020] In the above aspect, the image encoding device further adds control information indicating whether or not the one or more parameters are added to the bitstream to the bitstream, and transmits the bitstream to the image decoding device.
[0021] According to this aspect, the image decoding device can easily determine whether or not one or more parameters are added to the received bitstream by checking whether control information is added to the received bitstream.
[0022] In the above aspect, the image decoding device receives the bitstream from the image encoding device, obtains the one or more parameters from the received bitstream, and based on the obtained one or more parameters, decodes the first area by first decoding processing and decodes the second area by second decoding processing.
[0023] According to this aspect, the image decoding device decodes the first area including the object by first decoding processing and decodes the second area not including the object by second decoding processing based on one or more parameters obtained from the received bitstream. Thereby, the image decoding device can set the first decoding processing and the second decoding processing so that the first area has higher image quality or higher resolution than the second area, and can output an image in which the first area has higher image quality or higher resolution than the second area to the second processing device. As a result, it is possible to improve the accuracy of task processing in the second processing device.
[0024] In the above aspect, the image decoding device outputs the obtained one or more parameters to a second processing device that executes a predetermined task process.
[0025] According to this aspect, it is possible to further improve the accuracy of task processing in the second processing device by having the second processing device execute a predetermined task process based on one or more parameters input from the image decoding device.
[0026] In the above aspect, the first encoding process and the second encoding process include at least one of a quantization process, a filtering process, an intra prediction process, an inter prediction process, and an arithmetic coding process.
[0027] According to this aspect, since the first encoding process and the second encoding process include at least one of a quantization process, a filtering process, an intra prediction process, an inter prediction process, and an arithmetic coding process, it becomes possible to set the first encoding process and the second encoding process so that the first region has higher image quality or higher resolution than the second region.
[0028] In the above aspect, the image encoding device sets the first encoding process and the second encoding process so that the first region has higher image quality or higher resolution than the second region.
[0029] According to this aspect, the image encoding device can transmit an image in which the first region has higher image quality or higher resolution than the second region to the image decoding device, and the image decoding device can output an image in which the first region has higher image quality or higher resolution than the second region to the second processing device. As a result, it becomes possible to improve the accuracy of task processing in the second processing device.
[0030] In the above aspect, the first encoding process and the second encoding process include a quantization process, and the image encoding device sets the value of a first quantization parameter related to the first encoding process to be smaller than the value of a second quantization parameter related to the second encoding process.
[0031] According to this aspect, by setting the value of the first quantization parameter related to the first encoding process to be smaller than the value of the second quantization parameter related to the second encoding process, it becomes possible to obtain an image in which the first region has higher image quality or higher resolution than the second region.
[0032] In the above aspect, the one or more parameters include at least one of a confidence level value of a neural network task which is the predetermined task processing, a counter value indicating the number of the objects included in the image, category information indicating the attributes of the objects included in the image, feature information indicating the features of the objects included in the image, and boundary information indicating a boundary surrounding the objects included in the image.
[0033] According to this aspect, since one or more parameters include at least one of a confidence level value of a neural network task, a counter value, category information, feature information, and boundary information regarding the objects included in the image, it is possible to improve the accuracy of task processing by the second processing device executing predetermined task processing based on the one or more parameters.
[0034] An image decoding method according to an aspect of the present disclosure is such that an image decoding device receives a bit stream from an image encoding device, obtains one or more parameters regarding an object included in the image from the bit stream, and based on the obtained one or more parameters, decodes a first region including the object in the image by a first decoding process, and decodes a second region not including the object in the image by a second decoding process.
[0035] According to this aspect, the image decoding device decodes a first region including an object by a first decoding process and decodes a second region not including the object by a second decoding process based on one or more parameters obtained from the received bit stream. Thereby, the image decoding device can set the first decoding process and the second decoding process so that the first region has higher image quality or higher resolution than the second region, and can output an image in which the first region has higher image quality or higher resolution than the second region to the second processing device. As a result, it is possible to improve the accuracy of task processing in the second processing device.
[0036] An image processing method according to an aspect of the present disclosure includes an image decoding device receiving, from an image encoding device, a bitstream including an encoded image and one or more parameters related to an object included in the image, obtaining the one or more parameters from the received bitstream, decoding, based on the obtained one or more parameters, a first region including the object in the image by a first decoding process, and decoding, by a second decoding process, a second region not including the object in the image.
[0037] According to this aspect, the image decoding device decodes, by a first decoding process, a first region including an object based on one or more parameters obtained from the received bitstream, and decodes, by a second decoding process, a second region not including the object. Thereby, the image decoding device can set the first decoding process and the second decoding process so that the first region has higher image quality or higher resolution than the second region, and can output, to a second processing device, an image in which the first region has higher image quality or higher resolution than the second region. As a result, it is possible to improve the accuracy of task processing in the second processing device.
[0038] An image encoding device according to an aspect of the present disclosure encodes, by a first encoding process, a first region including an object in the image based on one or more parameters related to the object included in the image input from a first processing device that executes a predetermined task process based on the image, encodes, by a second encoding process, a second region not including the object in the image based on the one or more parameters, generates a bitstream by encoding the first region and the second region, and transmits the generated bitstream to an image decoding device.
[0039] According to this aspect, an image encoding device encodes a first region including an object by a first encoding process and encodes a second region not including the object by a second encoding process based on one or more parameters input from a first processing device. As a result, the image encoding device can set the first encoding process and the second encoding process so that the first region has higher image quality or higher resolution than the second region, and the image decoding device can output an image in which the first region has higher image quality or higher resolution than the second region to the second processing device. As a result, it is possible to improve the accuracy of task processing in the second processing device.
[0040] An image decoding device according to an aspect of the present disclosure receives a bitstream from an image encoding device, obtains one or more parameters related to an object included in the image from the bitstream, and based on the obtained one or more parameters, decodes a first region including the object in the image by a first decoding process and decodes a second region not including the object in the image by a second decoding process.
[0041] According to this aspect, an image decoding device decodes a first region including an object by a first decoding process and decodes a second region not including the object by a second decoding process based on one or more parameters obtained from the received bitstream. As a result, the image decoding device can set the first decoding process and the second decoding process so that the first region has higher image quality or higher resolution than the second region, and can output an image in which the first region has higher image quality or higher resolution than the second region to the second processing device. As a result, it is possible to improve the accuracy of task processing in the second processing device.
[0042] (Embodiments of the present disclosure) Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that elements denoted by the same reference numerals in different drawings indicate the same or corresponding elements.
[0043] Note that all the embodiments described below show specific examples of the present disclosure. The numerical values, shapes, components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Also, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components. Further, in all embodiments, the respective contents can be combined with each other.
[0044] FIG. 4 is a diagram showing a configuration example of an image processing system 1200 according to an embodiment of the present disclosure. Also, FIG. 2 is a flowchart showing a processing procedure 2000 of an image encoding method according to an embodiment of the present disclosure.
[0045] The image encoding device 1202 inputs an image captured by a camera and outputs a compressed bitstream by encoding the image. The image captured by the camera is also input to the task processing unit 1201. The task processing unit 1201 executes a predetermined task process such as a neural network task based on the input image. The task processing unit 1201, for example, executes a face detection process for detecting the face of a person included in the image. The processing result R1 (detected face bounding box) of the task processing by the task processing unit 1201 is fed back to the image capture device and input to the image encoding device 1202.
[0046] The processing result R1 includes one or more parameters regarding the object included in the image. The one or more parameters include at least one of a confidence level value of a neural network task, a counter value indicating the number of objects included in the image, category information indicating the attributes of the objects included in the image, feature information indicating the features of the objects included in the image, and boundary information indicating a figure (bounding box) of the boundary surrounding the objects included in the image.
[0047] FIG. 13 is a diagram showing an example of attribute table information for setting category information. The attribute table information describes various attributes related to objects, such as person, bicycle, car, etc. The task processing unit 1201 selects the attributes of the objects included in the image from the attribute table information and outputs them as category information related to the objects.
[0048] FIG. 14 is a diagram showing an example of feature table information for setting feature information. The feature table information describes various features related to objects, such as colour, size, shape, etc. The task processing unit 1201 sets the features of the objects included in the image based on the feature table information and outputs them as feature information related to the objects.
[0049] In step S2001, the image encoding device 1202 adds one or more parameters to the bitstream. The one or more parameters may be added to the bitstream by being encoded, or may be added to the bitstream by being stored in the header of the bitstream. The header may be a VPS, SPS, PPS, PH, SH, or SEI. Further, the image encoding device 1202 may further add control information such as flag information indicating whether one or more parameters are added to the bitstream to the header of the bitstream or the like.
[0050] Also, the image encoding device 1202 encodes the image based on one or more parameters included in the processing result R1. In step S2002, the image encoding device 1202 encodes the first region including the object in the image by the first encoding process, and in step S2003, encodes the second region not including the object in the image by the second encoding process.
[0051] Figs. 15 and 16 are diagrams showing examples of the first region and the second region. Referring to Fig. 15, the image encoding device 1202 sets, as the first region, a region including objects (persons and animals in this example) set by one or more parameters in the input image, and sets, as the second region, a region not including the objects. Referring to Fig. 16, the image encoding device 1202 sets, as the first region, a region including objects (a car and a person in this example) set by one or more parameters in the input image, and sets, as the second region, a region not including the objects. Thus, by defining a plurality of regions in the same image, it becomes possible to switch the encoding process for each region.
[0052] The first encoding process and the second encoding process include at least one of quantization processing, filtering processing, intra prediction processing, inter prediction processing, and arithmetic encoding processing. The image encoding device 1202 sets the first encoding process and the second encoding process so that the first region has higher image quality or higher resolution than the second region. Thereby, it is possible to perform encoding so that the first region including the object has higher image quality or higher resolution, and to appropriately reduce the processing amount in the encoding of the second region.
[0053] Fig. 6 is a block diagram showing a configuration example 2300 of the image encoding device 1202 according to an embodiment of the present disclosure. The image encoding device 1202 is configured to encode an input image in block units and output an encoded bit stream. As shown in Fig. 6, the image encoding device 1202 includes a conversion unit 2301, a quantization unit 2302, an inverse quantization unit 2303, an inverse conversion unit 2304, a filter processing unit 2305, a block memory 2306, an intra prediction unit 2307, a picture memory 2308, a block memory 2309, a motion vector prediction unit 2310, an interpolation unit 2311, an inter prediction unit 2312, and an entropy encoding unit 2313. In the configuration example 2300, one or more parameters are input to the quantization unit 2302 and the entropy encoding unit 2313.
[0054] Next, an exemplary operation flow will be described. The input image and the predicted image are input to the adder, and the addition value corresponding to the difference image between the input image and the predicted image is input from the adder to the conversion unit 2301. The conversion unit 2301 inputs the frequency coefficients obtained by converting the addition value to the quantization unit 2302.
[0055] The quantization unit 2302 quantizes the input frequency coefficients and inputs the quantized frequency coefficients to the inverse quantization unit 2303 and the entropy encoding unit 2313. For example, when the reliability level of the object included in the first region is equal to or higher than a predetermined threshold, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when the reliability level of the object included in the first region is less than a predetermined threshold, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when a bounding box exists, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. As another example, when no bounding box exists, the quantization unit 2302 sets a quantization parameter having a value different from the value of the quantization parameter for the second region in the first region. Examples of setting the quantization parameter include adding a parameter indicating the set value of the quantization parameter together with the quantization parameter in the header and selecting the quantization parameter.
[0056] For example, the quantization unit 2302 sets the value of the quantization parameter for the first region to be smaller than the value of the quantization parameter for the second region based on one or more input parameters. That is, the value of the first quantization parameter for the first encoding process is set to be smaller than the value of the second quantization parameter for the second encoding process.
[0057] FIG. 17 is a diagram showing an example of setting quantization parameters. The quantization unit 2302 sets the quantization parameter QP1 for blocks corresponding to the second region that does not contain an object in the image, and sets a quantization parameter QP2 smaller than the quantization parameter QP1 for blocks corresponding to the first region that contains an object in the image.
[0058] The entropy encoding unit 2313 generates a bitstream by entropy encoding the quantized frequency coefficients. Further, the entropy encoding unit 2313 adds one or more input parameters to the bitstream by entropy encoding them together with the quantized frequency coefficients, or by storing them in the header of the bitstream. Further, the entropy encoding unit 2313 may further add control information such as flag information indicating whether one or more parameters are added to the bitstream to the header of the bitstream or the like.
[0059] The inverse quantization unit 2303 inverse-quantizes the frequency coefficients input from the quantization unit 2302 and inputs the inverse-quantized frequency coefficients to the inverse transformation unit 2304. The inverse transformation unit 2304 generates a difference image by inverse-transforming the frequency coefficients and inputs the difference image to an adder. The adder adds the difference image input from the inverse transformation unit 2304 and the prediction image input from the intra prediction unit 2307 or the inter prediction unit 2312. The adder inputs the addition value 2320 corresponding to the input image to the filter processing unit 2305. The filter processing unit 2305 performs a predetermined filter process on the addition value 2320 and inputs the value after the filter process to the block memory 2306 and the picture memory 2308 for further prediction.
[0060] The intra prediction unit 2307 and the inter prediction unit 2312 search for an image region most similar to the input image for prediction within the reconstructed image stored in the block memory 2306 or the picture memory 2308. The block memory 2309 fetches blocks of the reconstructed image from the picture memory 2308 using the motion vectors input from the motion vector prediction unit 2310. The block memory 2309 inputs the blocks of the reconstructed image to the interpolation unit 2311 for interpolation processing. The interpolated image is input from the interpolation unit 2311 to the inter prediction unit 2312 for the inter prediction process.
[0061] The image encoding device 1202 transmits a bitstream with one or more parameters added thereto to the image decoding device 1203 via a communication network.
[0062] FIG. 8 is a block diagram showing a configuration example 2400 of the image encoding device 1202 according to an embodiment of the present disclosure. As shown in FIG. 8, the image encoding device 1202 includes a conversion unit 2401, a quantization unit 2402, an inverse quantization unit 2403, an inverse conversion unit 2404, a filter processing unit 2405, a block memory 2406, an intra prediction unit 2407, a picture memory 2408, a block memory 2409, a motion vector prediction unit 2410, an interpolation unit 2411, an inter prediction unit 2412, and an entropy encoding unit 2413. In the configuration example 2400, one or more parameters are input to the filter processing unit 2405 and the entropy encoding unit 2413.
[0063] FIG. 18 is a diagram showing a setting example of filter processing. Examples of filter processing are a deblocking filter, an adaptive loop filter (ALF), a cross-component adaptive loop filter (CCALF), a sample adaptive offset (SAO), or a luma mapping by chroma scaling (LMCS). The filter processing unit 2405 sets the filter strength, the filter length, or the activation / inactivation of filter processing based on the one or more input parameters.
[0064] The filter processing unit 2405 sets the filter intensity A for blocks corresponding to the second region that does not contain an object in the image, and sets the filter intensity B for blocks corresponding to the first region that contains an object in the image. Alternatively, the filter processing unit 2405 sets the filter length A for blocks corresponding to the second region, and sets the filter length B for blocks corresponding to the first region. Alternatively, the filter processing unit 2405 invalidates the filter processing for blocks corresponding to the second region, and validates the filter processing for blocks corresponding to the first region. For example, in the case of an adaptive loop filter where a larger filter length has a higher effect of reducing the difference from the original image of the decoded image, the filter length of the first region is increased for the second region. Also, in the case of a deblocking filter where a larger filter length has a higher effect of suppressing block distortion, the filter length of the first region may be increased for the second region.
[0065] For example, when the trust level of an object included in the first region is equal to or higher than a predetermined threshold, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when the trust level of an object included in the first region is less than a predetermined threshold, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when the category or feature of an object is the same as a predetermined category or feature, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when the category or feature of an object is different from a predetermined category or feature, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when a bounding box exists, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region. As another example, when a bounding box does not exist, the filter processing unit 2405 applies a setting different from the filter strength, filter length, or activation / deactivation setting of the filter processing for the second region to the first region.
[0066] FIG. 24 is a diagram showing an example of filter strength, and FIG. 25 is a diagram showing an example of filter tap length. The filter processing unit 2405 can set the filter strength according to the filter coefficient and can set the filter length according to the number of filter taps.
[0067] FIG. 10 is a block diagram showing a configuration example 2500 of the image encoding apparatus 1202 according to an embodiment of the present disclosure. As shown in FIG. 10, the image encoding apparatus 1202 includes a conversion unit 2501, a quantization unit 2502, an inverse quantization unit 2503, an inverse conversion unit 2504, a filter processing unit 2505, a block memory 2506, an intra prediction unit 2507, a picture memory 2508, a block memory 2509, a motion vector prediction unit 2510, an interpolation unit 2511, an inter prediction unit 2512, and an entropy encoding unit 2513. In the configuration example 2500, one or more parameters are input to the intra prediction unit 2507 and the entropy encoding unit 2413.
[0068] FIG. 19 is a diagram showing a setting example of prediction processing. The intra prediction unit 2507 does not perform prediction processing on blocks corresponding to a second region that does not include an object in the image, and performs intra prediction processing on blocks corresponding to a first region that includes an object in the image based on one or more input parameters.
[0069] For example, when the reliability level of the object included in the first region is equal to or higher than a predetermined threshold, the intra prediction unit 2507 performs intra prediction processing on the block corresponding to the first region. As another example, when the reliability level of the object included in the first region is less than a predetermined threshold, the intra prediction unit 2507 does not perform intra prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the intra prediction unit 2507 performs intra prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the intra prediction unit 2507 does not perform intra prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the intra prediction unit 2507 performs intra prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the intra prediction unit 2507 does not perform intra prediction processing on the block corresponding to the first region. As another example, when a bounding box exists, the intra prediction unit 2507 performs intra prediction processing on the block corresponding to the first region. As another example, when a bounding box does not exist, the intra prediction unit 2507 does not perform intra prediction processing on the block corresponding to the first region. In the above examples, it is also possible to select an encoding-specific intra prediction mode based on one or more parameters. For example, the intra prediction mode can be replaced with the DC intra mode.
[0070] FIG. 12 is a block diagram showing a configuration example 2600 of an image encoding apparatus 1202 according to an embodiment of the present disclosure. As shown in FIG. 12, the image encoding apparatus 1202 includes a conversion unit 2601, a quantization unit 2602, an inverse quantization unit 2603, an inverse conversion unit 2604, a filter processing unit 2605, a block memory 2606, an intra prediction unit 2607, a picture memory 2608, a block memory 2609, a motion vector prediction unit 2610, an interpolation unit 2611, an inter prediction unit 2612, and an entropy encoding unit 2613. In the configuration example 2600, one or more parameters are input to the inter prediction unit 2612 and the entropy encoding unit 2413.
[0071] FIG. 20 is a diagram showing a setting example of prediction processing. The inter prediction unit 2612 does not perform prediction processing on blocks corresponding to a second region that does not include an object in the image, and performs inter prediction processing on blocks corresponding to a first region that includes an object in the image, based on one or more input parameters.
[0072] For example, when the reliability level of the object included in the first region is equal to or higher than a predetermined threshold, the inter prediction unit 2612 executes inter prediction processing on the block corresponding to the first region. As another example, when the reliability level of the object included in the first region is less than a predetermined threshold, the inter prediction unit 2612 does not execute inter prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the inter prediction unit 2612 executes inter prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the inter prediction unit 2612 does not execute inter prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the inter prediction unit 2612 executes inter prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the inter prediction unit 2612 does not execute inter prediction processing on the block corresponding to the first region. As another example, when a bounding box exists, the inter prediction unit 2612 executes inter prediction processing on the block corresponding to the first region. As another example, when a bounding box does not exist, the inter prediction unit 2612 does not execute inter prediction processing on the block corresponding to the first region. In the above examples, it is also possible to select an encoding-specific inter prediction mode based on one or more parameters. For example, the inter prediction mode can be replaced only with the skip prediction mode, or cannot be replaced with the skip prediction mode.
[0073] For example, referring to FIG. 6, when the first encoding process and the second encoding process include arithmetic encoding processes, the entropy encoding unit 2313 can set different context models for the first region and the second region based on one or more input parameters.
[0074] FIG. 21 is a diagram showing an example of setting a context model. The entropy encoding unit 2313 sets context model A for blocks corresponding to the second region that does not include an object in the image, and sets context model B for blocks corresponding to the first region that includes an object in the image.
[0075] For example, when the reliability level of the object included in the first region is equal to or higher than a predetermined threshold, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when the reliability level of the object included in the first region is less than a predetermined threshold, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when a bounding box exists, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region. As another example, when a bounding box does not exist, the entropy encoding unit 2313 sets a context model different from the context model regarding the second region for the first region.
[0076] Referring to FIG. 2, in step S2004, the image encoding device 1202 may generate pixel samples of the encoded image and output a signal including the pixel samples of the image and the above one or more parameters toward the first processing device. The first processing device may be the task processing unit 1201.
[0077] The first processing device uses the pixel samples of the image and one or more parameters included in the input signal to execute a predetermined task process such as a neural network task. In the neural network task, at least one determination process may be executed. An example of the neural network is a convolutional neural network. Examples of neural network tasks are object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, human-machine hybrid vision, or any combination thereof.
[0078] FIG. 22 is a diagram showing object detection and object segmentation as an example of a neural network task. In object detection, the attributes of the objects (in this example, a TV and a person) included in the input image are detected. In addition to the attributes of the objects included in the input image, the position and number of the objects in the input image may be detected. Thereby, for example, the position of the object to be recognized may be narrowed down, or the objects other than the object to be recognized may be excluded. Specific applications include, for example, face detection in a camera and detection of pedestrians in autonomous driving. In object segmentation, the pixels of the region corresponding to the object are segmented (i.e., separated). Thereby, for example, applications such as separating obstacles and roads in autonomous driving to assist safe driving of a vehicle, detecting product defects in a factory, and identifying terrain in a satellite image are considered.
[0079] FIG. 23 shows object tracking, action recognition, and pose estimation as an example of neural network tasks. In object tracking, the movement of an object included in an input image is tracked. Applications include, for example, counting the number of users in facilities such as stores and analyzing the movements of sports players. If the processing is further accelerated, real-time object tracking becomes possible, and applications to camera processing such as autofocus also become possible. In action recognition, the type of an object's action (in this example, "riding a bicycle" and "walking") is detected. For example, it can be applied to uses such as preventing and detecting criminal acts such as robbery and shoplifting and preventing work mistakes in factories by using it in security cameras. In pose estimation, the pose of an object is detected by detecting key points and joints. For example, it can be considered for use in industrial fields such as improving work efficiency in factories, security fields such as detecting abnormal behavior, and fields such as healthcare and sports.
[0080] The first processing device outputs a signal indicating the execution result of the neural network task. The signal may include at least one of the number of detected objects, the confidence level of the detected objects, the boundary information or position information of the detected objects, and the classification category of the detected objects. The signal may be input from the first processing device to the image encoding device 1202.
[0081] FIG. 1 is a flowchart showing a processing procedure 1000 of an image decoding method according to an embodiment of the present disclosure. Referring to FIGS. 1 and 4, an image decoding apparatus 1203 receives a bitstream transmitted from an image encoding apparatus 1202 and obtains one or more parameters from the bitstream (step S1001). Based on the one or more obtained parameters, the image decoding apparatus 1203 decodes a first region including an object in the image by a first decoding process (step S1002), and decodes a second region not including the object in the image by a second decoding process (step S1003). Further, the image decoding apparatus 1203 inputs the extended image and a processing result R1 including one or more parameters to a task processing unit 1204 as a second processing apparatus. Further, the image decoding apparatus 1203 outputs the decoded image to a display device, and the display device displays the image.
[0082] Based on the input image and the processing result R1, the task processing unit 1204 executes a predetermined task process such as a neural network task (step S1004). For example, the task processing unit 1204 executes a person identification process for identifying a person based on the face of the person included in the image, and outputs a processing result R2.
[0083] FIGS. 15 and 16 are diagrams showing examples of the first region and the second region. The image decoding apparatus 1203 sets a region including an object set by one or more parameters as a first region, and sets a region not including the object as a second region.
[0084] The first decoding process and the second decoding process include at least one of a quantization process, a filtering process, an intra prediction process, an inter prediction process, and an arithmetic coding process. The image decoding apparatus 1203 sets the first decoding process and the second decoding process so that the first region has higher image quality or higher resolution than the second region.
[0085] FIG. 5 is a block diagram showing a configuration example 1300 of an image decoding apparatus 1203 according to an embodiment of the present disclosure. The image decoding apparatus 1203 is configured to decode an input bitstream in block units and output a decoded image. As shown in FIG. 5, the image decoding apparatus 1203 includes an entropy decoding unit 1301, an inverse quantization unit 1302, an inverse transform unit 1303, a filter processing unit 1304, a block memory 1305, an intra prediction unit 1306, a picture memory 1307, a block memory 1308, an interpolation unit 1309, an inter prediction unit 1310, an analysis unit 1311, and a motion vector prediction unit 1312.
[0086] Next, an exemplary operation flow will be described. The encoded bitstream input to the image decoding apparatus 1203 is input to the entropy decoding unit 1301. The entropy decoding unit 1301 decodes the input bitstream and inputs the frequency coefficients, which are decoded values, to the inverse quantization unit 1302. Also, the entropy decoding unit 1301 acquires one or more parameters from the bitstream and inputs the acquired one or more parameters to the inverse quantization unit 1302. The inverse quantization unit 1302 inverse quantizes the frequency coefficients input from the entropy decoding unit 1301 and inputs the inverse quantized frequency coefficients to the inverse transform unit 1303. The inverse transform unit 1303 generates a differential image by inverse transforming the frequency coefficients and inputs the differential image to an adder. The adder adds the differential image input from the inverse transform unit 1303 and the predicted image input from the intra prediction unit 1306 or the inter prediction unit 1310. The adder inputs an addition value 1320 corresponding to the input image to the filter processing unit 1304. The filter processing unit 1304 performs a predetermined filter process on the addition value 2320 and inputs the filtered image to the block memory 1305 and the picture memory 1307 for further prediction. Also, the filter processing unit 1304 inputs the filtered image to a display device, and the display device displays the image.
[0087] The parsing unit 1311 inputs several pieces of prediction information, such as a block of residual samples, a reference index indicating the reference picture used, and a delta motion vector, to the motion vector prediction unit 1312 by parsing the input bitstream. The motion vector prediction unit 1312 predicts the motion vector of the current block based on the prediction information input from the parsing unit 1311. The motion vector prediction unit 1312 inputs a signal indicating the predicted motion vector to the block memory 1308.
[0088] The intra prediction unit 1306 and the inter prediction unit 1310 search for the image region in the reconstructed image stored in the block memory 1305 or the picture memory 1307 that is most similar to the input image for prediction. The block memory 1308 fetches the block of the reconstructed image from the picture memory 1307 using the motion vector input from the motion vector prediction unit 1312. The block memory 1308 inputs the block of the reconstructed image to the interpolation unit 1309 for interpolation processing. The interpolated image is input from the interpolation unit 1309 to the inter prediction unit 1310 for the inter prediction process.
[0089] In the configuration example 1300, one or more parameters are input to the inverse quantization unit 1302.
[0090] For example, when the reliability level of an object included in the first region is equal to or higher than a predetermined threshold value, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when the reliability level of an object included in the first region is less than a predetermined threshold value, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when a counter value indicating the number of objects is equal to or higher than a predetermined threshold value, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when a counter value indicating the number of objects is less than a predetermined threshold value, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when the category or feature of an object is the same as a predetermined category or feature, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when the category or feature of an object is different from a predetermined category or feature, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when a bounding box exists, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region. As another example, when a bounding box does not exist, the inverse quantization unit 1302 sets, for the first region, a quantization parameter having a value different from the value of the quantization parameter for the second region.
[0091] FIG. 7 is a block diagram showing a configuration example 1400 of an image decoding apparatus 1203 according to an embodiment of the present disclosure. As shown in FIG. 7, the image decoding apparatus 1203 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, a filter processing unit 1404, a block memory 1405, an intra prediction unit 1406, a picture memory 1407, a block memory 1408, an interpolation unit 1409, an inter prediction unit 1410, an analysis unit 1411, and a motion vector prediction unit 1412.
[0092] In Configuration Example 1400, one or more parameters are input to the filter processing unit 1404.
[0093] For example, when the trust level of the object included in the first region is equal to or higher than a predetermined threshold value, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when the trust level of the object included in the first region is less than a predetermined threshold value, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold value, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold value, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when a bounding box exists, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region. As another example, when no bounding box exists, the filter processing unit 1404 applies a setting different from the filter strength, filter length, or activation / inactivation setting of the filter processing for the second region to the first region.
[0094] FIG. 9 is a block diagram showing a configuration example 1500 of an image decoding apparatus 1203 according to an embodiment of the present disclosure. As shown in FIG. 9, the image decoding apparatus 1203 includes an entropy decoding unit 1501, an inverse quantization unit 1502, an inverse transform unit 1503, a filter processing unit 1504, a block memory 1505, an intra prediction unit 1506, a picture memory 1507, a block memory 1508, an interpolation unit 1509, an inter prediction unit 1510, an analysis unit 1511, and a motion vector prediction unit 1512.
[0095] In the configuration example 1500, one or more parameters are input to the intra prediction unit 1506.
[0096] For example, when the reliability level of the object included in the first region is equal to or higher than a predetermined threshold, the intra prediction unit 1506 executes intra prediction processing on the block corresponding to the first region. As another example, when the reliability level of the object included in the first region is less than a predetermined threshold, the intra prediction unit 1506 does not execute intra prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the intra prediction unit 1506 executes intra prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the intra prediction unit 1506 does not execute intra prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the intra prediction unit 1506 executes intra prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the intra prediction unit 1506 does not execute intra prediction processing on the block corresponding to the first region. As another example, when a bounding box exists, the intra prediction unit 1506 executes intra prediction processing on the block corresponding to the first region. As another example, when a bounding box does not exist, the intra prediction unit 1506 does not execute intra prediction processing on the block corresponding to the first region.
[0097] FIG. 11 is a block diagram showing a configuration example 1600 of an image decoding apparatus 1203 according to an embodiment of the present disclosure. As shown in FIG. 11, the image decoding apparatus 1203 includes an entropy decoding unit 1601, an inverse quantization unit 1602, an inverse transformation unit 1603, a filter processing unit 1604, a block memory 1605, an intra prediction unit 1606, a picture memory 1607, a block memory 1608, an interpolation unit 1609, an inter prediction unit 1610, an analysis unit 1611, and a motion vector prediction unit 1612.
[0098] In the configuration example 1600, one or more parameters are input to the inter prediction unit 1610.
[0099] For example, when the reliability level of the object included in the first region is equal to or higher than a predetermined threshold, the inter prediction unit 1610 executes inter prediction processing on the block corresponding to the first region. As another example, when the reliability level of the object included in the first region is less than a predetermined threshold, the inter prediction unit 1610 does not execute inter prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the inter prediction unit 1610 executes inter prediction processing on the block corresponding to the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the inter prediction unit 1610 does not execute inter prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the inter prediction unit 1610 executes inter prediction processing on the block corresponding to the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the inter prediction unit 1610 does not execute inter prediction processing on the block corresponding to the first region. As another example, when a bounding box exists, the inter prediction unit 1610 executes inter prediction processing on the block corresponding to the first region. As another example, when a bounding box does not exist, the inter prediction unit 1610 does not execute inter prediction processing on the block corresponding to the first region.
[0100] For example, referring to FIG. 5, when the first decoding process and the second decoding process include an arithmetic coding process, the entropy decoding unit 1301 can set different context models for the first region and the second region based on one or more parameters obtained from the received bit stream.
[0101] For example, when the trust level of the object included in the first region is equal to or higher than a predetermined threshold, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when the trust level of the object included in the first region is less than a predetermined threshold, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when the counter value indicating the number of objects is equal to or higher than a predetermined threshold, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when the counter value indicating the number of objects is less than a predetermined threshold, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when the category or feature of the object is the same as a predetermined category or feature, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when the category or feature of the object is different from a predetermined category or feature, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when a bounding box exists, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region. As another example, when no bounding box exists, the entropy decoding unit 1301 sets a context model different from the context model for the second region for the first region.
[0102] Referring to FIG. 1, in step S1004, the image decoding apparatus 1203 may generate pixel samples of the decoded image and output the pixel samples of the image and the processing result R1 including the above one or more parameters to the task processing unit 1204 which is a second processing apparatus.
[0103] The task processing unit 1204 executes a predetermined task process such as a neural network task by using the pixel samples of the image and one or more parameters included in the input signal. In the neural network task, at least one determination process may be executed. An example of the neural network is a convolutional neural network. Examples of the neural network task are object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, machine and human hybrid vision, or any combination thereof.
[0104] According to the present embodiment, the image encoding apparatus 1202 encodes a first region including an object by a first encoding process and encodes a second region not including the object by a second encoding process based on one or more parameters input from the task processing unit 1201. Thereby, the image encoding apparatus 1202 can set the first encoding process and the second encoding process so that the first region has higher image quality or higher resolution than the second region, and the image decoding apparatus 1203 can output an image in which the first region has higher image quality or higher resolution than the second region to the task processing unit 1204. As a result, it is possible to improve the accuracy of the task processing in the task processing unit 1204.
Industrial Applicability
[0105] The present disclosure is particularly useful for application to an image processing system including an encoder that transmits an image and a decoder that receives the image.
Claims
1. An image decoding apparatus, receives a bit stream from an image encoding apparatus, obtains one or more parameters regarding an object included in the image from the bit stream, switches a decoding process between a first region including the object in the image and a second region not including the object in the image based on the obtained one or more parameters, and outputs the obtained one or more parameters to a second processing apparatus that executes a predetermined task process. An image decoding method.
2. An image decoding apparatus, receives a bit stream including an encoded image and one or more parameters regarding an object included in the image from an image encoding apparatus, obtains the one or more parameters from the received bit stream, switches a decoding process between a first region including the object in the image and a second region not including the object in the image based on the obtained one or more parameters, and outputs the obtained one or more parameters to a second processing apparatus that executes a predetermined task process. An image processing method.
3. An image decoding apparatus that receives a bit stream from an image encoding apparatus, obtains one or more parameters regarding an object included in the image from the bit stream, switches a decoding process between a first region including the object in the image and a second region not including the object in the image based on the obtained one or more parameters, and outputs the obtained one or more parameters to a second processing apparatus that executes a predetermined task process.
Citation Information
Patent Citations
Image processing apparatus, image processing method and program
JP2009049976A
Tiling in video decoding and encoding
US20100046635A1
Utilizing a neural network having a two-stream encoder architecture to generate composite digital images
US20210027470A1