Image processing method
By encoding environmental parameters in the bitstream, the method improves machine accuracy in image processing tasks by enabling adaptive neural network adjustments.
Patent Information
- Application Number
- JP2025151115
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-04-23
- Filing Date
- 2025-09-11
- Publication Date
- 2025-12-16
AI Technical Summary
Machines trained for image processing deteriorate in accuracy when environmental conditions change, leading to poor detection and judgment.
An image encoding device adds parameters such as camera characteristics, object size, and depth to the bitstream, which are used by a decoding device to adjust neural network models and improve task processing accuracy.
Enhances the accuracy of task processing by allowing neural networks to adapt to environmental changes through informed model adjustments and parameter adjustments.
Smart Images

Figure 2025183340000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing method. [Background technology]
[0002] For example, as shown in Patent Documents 1 and 2, a conventional image coding system architecture includes a camera or sensor for capturing an image, an encoder for encoding the captured image into a bitstream, a decoder for decoding the image from the bitstream, and a display device for displaying the image for human evaluation. Since the advent of machine learning or neural network-based applications, machines are rapidly replacing humans in image evaluation because they are superior to humans in terms of scalability, efficiency, and accuracy. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] US Patent Application Publication No. 2010 / 0046635 [Patent Document 2] U.S. Patent Application Publication No. 2021 / 0027470 Summary of the Invention [Problem to be solved by the invention]
[0004] Machines tend to only work well in the situations they are trained in. If some of the environmental information changes on the camera side, the machine's performance will deteriorate, resulting in poor detection accuracy and poor judgment. If we could teach the machine environmental information, perhaps we could customize it to adapt to changes to achieve better detection accuracy.
[0005] The present disclosure aims to improve the accuracy of task processing. [Means for solving the problem]
[0006] In an image processing method according to one embodiment of the present disclosure, an image encoding device generates a bitstream by encoding an image, adds one or more parameters not used in encoding the image to the bitstream by encoding, transmits the bitstream with the one or more parameters added to it to an image decoding device, and outputs the image and the one or more parameters to a first processing device that performs a predetermined task processing, the predetermined task processing including at least one of object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, and hybrid vision, and the image decoding device outputs the image and the one or more parameters to a second processing device that performs the predetermined task processing. [Effects of the Invention]
[0007] According to the present disclosure, the accuracy of task processing can be improved. [Brief explanation of the drawings]
[0008] [Figure 1] 3 is a flowchart showing a processing procedure of an image encoding method according to the first embodiment of the present disclosure. [Figure 2] 10 is a flowchart showing a processing procedure of an image decoding method according to the first embodiment of the present disclosure. [Figure 3] 3 is a flowchart showing a processing procedure of an image encoding method according to the first embodiment of the present disclosure. [Figure 4] 10 is a flowchart showing a processing procedure of an image decoding method according to the first embodiment of the present disclosure. [Figure 5] FIG. 1 is a block diagram showing a configuration of an encoder according to a first embodiment of the present disclosure. [Figure 6] FIG. 2 is a block diagram showing the configuration of a decoder according to the first embodiment of the present disclosure. [Figure 7] 1 is a block diagram illustrating an example configuration of an image encoding device according to a first embodiment of the present disclosure. [Figure 8]FIG. 2 is a block diagram illustrating an example configuration of an image decoding device according to a first embodiment of the present disclosure. [Figure 9] FIG. 1 is a diagram illustrating an example of the configuration of an image processing system according to the background art. [Figure 10] 1 is a diagram illustrating a first exemplary configuration of an image processing system according to the present disclosure. [Figure 11] FIG. 10 is a diagram illustrating a second exemplary configuration of an image processing system according to the present disclosure. [Figure 12] FIG. 10 is a diagram illustrating an example of camera characteristics related to the installation position of a fixed camera. [Figure 13] FIG. 10 is a diagram illustrating an example of camera characteristics related to the installation position of a fixed camera. [Figure 14] FIG. 1 illustrates an example of a neural network task. [Figure 15] FIG. 1 illustrates an example of a neural network task. [Figure 16] 10 is a flowchart illustrating an exemplary procedure for determining the size of an object. [Figure 17] 1 is a flowchart illustrating an exemplary procedure for determining the depth of an object. [Figure 18] FIG. 10 is a diagram illustrating an example of calculation of the depth and size of an object. [Figure 19] 10 is a flowchart showing a processing procedure of a first example of using one or more parameters. [Figure 20] 10 is a flowchart showing a processing procedure of a second example of using one or more parameters. [Figure 21] 10 is a flowchart showing a processing procedure of a third example of using one or more parameters. [Figure 22] 10 is a flowchart showing a processing procedure of a fourth example of using one or more parameters. [Figure 23] 10 is a flowchart showing a processing procedure of a fifth example of using one or more parameters. [Figure 24] 10 is a flowchart showing a processing procedure of a sixth example of using one or more parameters. [Figure 25]13 is a flowchart showing a processing procedure of a seventh example of using one or more parameters. [Figure 26] 13 is a flowchart showing a processing procedure of an eighth example of using one or more parameters. [Figure 27] FIG. 10 is a diagram illustrating an example of camera characteristics related to a camera mounted on a moving object. [Figure 28] FIG. 10 is a diagram illustrating an example of camera characteristics related to a camera mounted on a moving object. [Figure 29] 10 is a flowchart showing a processing procedure of an image decoding method according to a second embodiment of the present disclosure. [Figure 30] 10 is a flowchart showing a processing procedure of an image encoding method according to a second embodiment of the present disclosure. [Figure 31] FIG. 10 is a block diagram showing an example configuration of a decoder according to a second embodiment of the present disclosure. [Figure 32] FIG. 10 is a block diagram showing an example configuration of an encoder according to a second embodiment of the present disclosure. [Figure 33] FIG. 10 is a diagram showing a comparison of output images from a normal camera and a camera with large distortion. [Figure 34] FIG. 10 is a diagram illustrating an example of boundary information. [Figure 35] FIG. 10 is a diagram illustrating an example of boundary information. [Figure 36] FIG. 10 is a diagram illustrating an example of boundary information. [Figure 37] FIG. 10 is a diagram illustrating an example of boundary information. DETAILED DESCRIPTION OF THE INVENTION
[0009] (Findings that formed the basis of this disclosure) 9 is a diagram showing an example of the configuration of an image processing system 3000 according to the background art. An encoder 3002 receives an image or feature signal from a camera or sensor 3001 and encodes the signal to output a compressed bitstream. The compressed bitstream is transmitted from the encoder 3002 to a decoder 3004 via a communication network 3003. The decoder 3004 receives the compressed bitstream and decodes the bitstream to input a decompressed image or feature signal to a task processing unit 3005. In the background art, information on camera characteristics, object size, and object depth is not transmitted from the encoder 3002 to the decoder 3004.
[0010] A problem with the background art described above is that the encoder 3002 does not transmit information necessary to improve the accuracy of task processing to the decoder 3004. By transmitting this information to the decoder 3004, the encoder 3002 can provide the task processor 3005 with important data related to the application environment, etc., which can be used to improve the accuracy of task processing. This information can include camera characteristics, the size of an object included in an image, or the depth of an object included in an image. The camera characteristics can include the camera installation height, the camera tilt angle, the distance from the camera to a region of interest (ROI), the camera field of view, or any combination thereof. The size of an object can be calculated from the width and height of the object in the image or estimated by executing a computer vision algorithm. The size of the object can be used to estimate the distance between the object and the camera. The depth of the object can be obtained by using a stereo camera or by executing a computer vision algorithm. The depth of the object can be used to estimate the distance between the object and the camera.
[0011] To solve the problems of the background art, the present inventors have introduced a new method for signaling camera characteristics, the size of an object contained in an image, the depth of an object contained in an image, or any combination thereof. The concept is to convey important information to a neural network so that the neural network can adapt to the environment from which the image or feature originates. One or more parameters indicative of this important information are added to the bitstream by being coded with the image or by being stored in the bitstream header. The header may be a VPS, SPS, PPS, PH, SH, or SEI. The one or more parameters may also be signaled at a system layer in the bitstream. An important aspect of this solution is that the conveyed information is intended to improve the accuracy of decisions, etc., in task processing involving a neural network.
[0012] 10 is a diagram illustrating a first exemplary configuration of an image processing system 3100 according to the present disclosure. An encoder 3102 (image encoding device) receives an image or feature signal from a camera or sensor 3101 and encodes the signal to generate a compressed bitstream. The encoder 3102 also receives one or more parameters from the camera or sensor 3101 and adds the one or more parameters to the bitstream. The compressed bitstream with the one or more parameters added is transmitted from the encoder 3102 to a decoder 3104 (image decoding device) via a communication network 3103. The decoder 3104 receives the compressed bitstream and decodes the bitstream to input the decompressed image or feature signal and the one or more parameters to a task processing unit 3105 that executes predetermined task processing.
[0013] FIG. 11 is a diagram illustrating a second exemplary configuration of an image processing system 3200 according to the present disclosure. The preprocessing unit 3202 receives an image or feature signal from a camera or sensor 3201 and outputs the preprocessed image or feature signal and one or more parameters. The encoder 3203 (image encoding device) receives the image or feature signal from the preprocessing unit 3202 and encodes the signal to generate a compressed bitstream. The encoder 3203 also receives one or more parameters from the preprocessing unit 3202 and adds the one or more parameters to the bitstream. The compressed bitstream with the one or more parameters added is transmitted from the encoder 3203 to a decoder 3205 (image decoding device) via a communication network 3204. The decoder 3205 receives the compressed bitstream and decodes the bitstream to input a decompressed image or feature signal to a postprocessing unit 3206 and input one or more parameters to a task processing unit 3207 that executes a predetermined task process. The post-processing unit 3206 inputs the post-processed decompressed image or feature signal to the task processing unit 3207 .
[0014] In the task processing units 3105 and 3207, the information signaled as one or more parameters can be used to change the neural network model being used. For example, a complex or simple neural network model can be selected depending on the size of the object or the installation height of the camera. Task processing may be performed using the selected neural network model.
[0015] The information signaled as one or more parameters can be used to change parameters used to adjust the estimated output of the neural network. For example, the signaled information can be used to set a detection threshold used for the estimation. Task processing can be performed using the new detection threshold for the neural network estimation.
[0016] The information signaled as one or more parameters can be used to adjust the scaling of an image input to the task processors 3105 and 3207. For example, the signaled information is used to set a scaling size. The image input to the task processors 3105 and 3207 is scaled to the set scaling size before the task processors 3105 and 3207 execute task processing.
[0017] Next, each aspect of the present disclosure will be described.
[0018] In an image encoding method according to one embodiment of the present disclosure, an image encoding device generates a bitstream by encoding an image, adds one or more parameters not used in encoding the image to the bitstream, transmits the bitstream with the one or more parameters added to it to an image decoding device, and outputs the image and the one or more parameters to a first processing device that performs a specified task processing.
[0019] According to this aspect, the image encoding device transmits, to the image decoding device, one or more parameters to be output to the first processing device for executing a predetermined task processing. This allows the image decoding device to output the one or more parameters received from the image encoding device to the second processing device that executes a task processing similar to the predetermined task processing. As a result, the second processing device executes the predetermined task processing based on the one or more parameters input from the image decoding device, thereby making it possible to improve the accuracy of the task processing in the second processing device.
[0020] In the above aspect, the image decoding device receives the bitstream from the image encoding device and outputs the image and the one or more parameters to a second processing device that executes a task processing similar to the specified task processing.
[0021] According to this aspect, the second processing device executes a predetermined task process based on one or more parameters input from the image decoding device, thereby making it possible to improve the accuracy of the task process in the second processing device.
[0022] In the above aspect, when executing the specified task processing, the first processing device and the second processing device switch at least one of a machine learning model, a detection threshold, a scaling value, and a post-processing method based on the one or more parameters.
[0023] According to this aspect, it is possible to improve the accuracy of task processing in the first processing device and the second processing device by switching at least one of the machine learning model, the detection threshold, the scaling value, and the post-processing method based on one or more parameters.
[0024] In the above aspect, the predetermined task processing includes at least one of object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, and hybrid vision.
[0025] According to this aspect, it is possible to improve the accuracy of each of these processes.
[0026] In the above aspect, the predetermined task processing includes image processing for improving the image quality or resolution of the image.
[0027] According to this aspect, it is possible to improve the accuracy of image processing for improving the image quality or resolution of an image.
[0028] In the above aspect, the image processing includes at least one of morphological transformation and edge enhancement processing for enhancing an object included in the image.
[0029] According to this aspect, it is possible to improve the accuracy of each of these processes.
[0030] In the above aspect, the one or more parameters include at least one of an installation height, a tilt angle, a distance to a region of interest, and a field of view for a camera that outputs the image.
[0031] According to this aspect, by including this information in one or more parameters, it is possible to improve the accuracy of task processing.
[0032] In the above aspect, the one or more parameters include at least one of depth and size of an object included in the image.
[0033] According to this aspect, by including this information in one or more parameters, it is possible to improve the accuracy of task processing.
[0034] In the above aspect, the one or more parameters include boundary information indicating a boundary surrounding an object included in the image, and distortion information indicating whether or not there is distortion in the image.
[0035] According to this aspect, by including this information in one or more parameters, it is possible to improve the accuracy of task processing.
[0036] In the above aspect, the boundary information includes position coordinates of a plurality of vertices of a graphic that defines the boundary.
[0037] According to this aspect, even if the image is distorted, it is possible to accurately define the boundary surrounding the object.
[0038] In the above aspect, the boundary information includes center coordinates, width information, height information, and tilt information related to a graphic that defines the boundary.
[0039] According to this aspect, even if the image is distorted, it is possible to accurately define the boundary surrounding the object.
[0040] In the above aspect, the distortion information includes additional information indicating that the image is an image captured by a fisheye camera, an ultra-wide-angle camera, or an omnidirectional camera.
[0041] According to this aspect, it is possible to easily determine whether a fisheye camera, an ultra-wide-angle camera, or an omnidirectional camera is being used, depending on whether additional information is included in one or more parameters.
[0042] In an image decoding method according to one embodiment of the present disclosure, an image decoding device receives a bitstream from an image encoding device, decodes an image from the bitstream, obtains one or more parameters from the bitstream that are not used in decoding the image, and outputs the image and the one or more parameters to a processing device that executes a specified task processing.
[0043] According to this aspect, the image decoding device outputs one or more parameters received from the image encoding device to a processing device that executes a predetermined task processing. As a result, the processing device executes the predetermined task processing based on the one or more parameters input from the image decoding device, thereby improving the accuracy of the task processing in the processing device.
[0044] In an image processing method according to one embodiment of the present disclosure, an image decoding device receives a bitstream from an image encoding device, the bitstream including an encoded image and one or more parameters not used in encoding the image, acquires the one or more parameters from the received bitstream, and outputs the acquired one or more parameters to a processing device that executes a specified task processing.
[0045] According to this aspect, the image decoding device outputs one or more parameters acquired from a bitstream received from the image encoding device to a processing device that executes a predetermined task processing. As a result, the processing device executes the predetermined task processing based on the one or more parameters input from the image decoding device, thereby improving the accuracy of the task processing in the processing device.
[0046] An image encoding device according to one embodiment of the present disclosure generates a bitstream by encoding an image, adds one or more parameters not used in encoding the image to the bitstream, transmits the bitstream with the one or more parameters added to it to an image decoding device, and outputs the image and the one or more parameters to a first processing device that performs a specified task processing.
[0047] According to this aspect, the image encoding device transmits, to the image decoding device, one or more parameters to be output to the first processing device for executing a predetermined task processing. This allows the image decoding device to output the one or more parameters received from the image encoding device to the second processing device that executes a task processing similar to the predetermined task processing. As a result, the second processing device executes the predetermined task processing based on the one or more parameters input from the image decoding device, thereby making it possible to improve the accuracy of the task processing in the second processing device.
[0048] An image decoding device according to one embodiment of the present disclosure receives a bitstream from an image encoding device, decodes an image from the bitstream, obtains one or more parameters from the bitstream that are not used in decoding the image, and outputs the image and the one or more parameters to a processing device that executes a specified task processing.
[0049] According to this aspect, the image decoding device outputs one or more parameters received from the image encoding device to a processing device that executes a predetermined task processing. As a result, the processing device executes the predetermined task processing based on the one or more parameters input from the image decoding device, thereby improving the accuracy of the task processing in the processing device.
[0050] (Embodiments of the present disclosure) Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Elements with the same reference numerals in different drawings indicate the same or corresponding elements.
[0051] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all embodiments, the respective contents can be combined.
[0052] (First embodiment) Fig. 5 is a block diagram showing the configuration of an encoder 1100A according to the first embodiment of the present disclosure. The encoder 1100A corresponds to the encoder 3102 shown in Fig. 10 or the encoder 3203 shown in Fig. 11. The encoder 1100A includes an image encoding device 1101A and a first processing device 1102A. However, the first processing device 1102A may be implemented within the image encoding device 1101A as part of the functions of the image encoding device 1101A.
[0053] 6 is a block diagram showing the configuration of a decoder 2100A according to the first embodiment of the present disclosure. The decoder 2100A includes an image decoding device 2101A and a second processing device 2102A. However, the second processing device 2102A may be implemented within the image decoding device 2101A as part of the functions of the image decoding device 2101A. The image decoding device 2101A corresponds to the decoder 3104 shown in FIG. 10 or the decoder 3205 shown in FIG. 11. The second processing device 2102A corresponds to the task processing unit 3105 shown in FIG. 10 or the task processing unit 3207 shown in FIG. 11.
[0054] The image encoding device 1101A generates a bitstream by encoding an input image on a block-by-block basis. The image encoding device 1101A also adds one or more input parameters to the bitstream. The one or more parameters are not used in encoding the image. The image encoding device 1101A also transmits the bitstream to which the one or more parameters have been added to the image decoding device 2101A. The image encoding device 1101A also generates pixel samples of the image and outputs a signal 1120A including the pixel samples of the image and one or more parameters to the first processing device 1102A. The first processing device 1102A executes a predetermined task processing, such as a neural network task, based on the signal 1120A input from the image encoding device 1101A. The first processing device 1102A may also input a signal 1121A obtained as a result of executing the predetermined task processing to the image encoding device 1101A.
[0055] The image decoding device 2101A receives a bitstream from the image encoding device 1101A. The image decoding device 2101A decodes an image from the received bitstream and outputs the decoded image to a display device. The display device displays the image. The image decoding device 2101A obtains one or more parameters from the received bitstream. The one or more parameters are not used in decoding the image. The image decoding device 2101A generates pixel samples of the image and outputs a signal 2120A including the pixel samples of the image and the one or more parameters to the second processing device 2102A. The second processing device 2102A performs predetermined task processing similar to that of the first processing device 1102A, based on the signal 2120A input from the image decoding device 2101A. The second processing device 2102A may input a signal 2121A obtained as a result of executing the predetermined task processing to the image decoding device 2101A.
[0056] (Encoder side processing) FIG. 1 is a flowchart showing a processing procedure 1000A of an image encoding method according to a first embodiment of the present disclosure. In a first step S1001A, an image encoding device 1101A encodes one or more parameters into a bitstream. Examples of the one or more parameters are parameters indicating camera characteristics. The camera characteristics include, but are not limited to, the camera installation height, the camera's perspective angle, the distance from the camera to the region of interest, the camera's tilt angle, the camera's field of view, the camera's orthographic projection size, the camera's near / far clipping planes, and the camera's image quality. The one or more parameters may be added to the bitstream by being encoded or by being stored in the bitstream's header. The header may be a VPS, SPS, PPS, PH, SH, or SEI. The one or more parameters may be added to the system layer of the bitstream.
[0057] 12 and 13 are diagrams showing examples of camera characteristics related to the installation position of a fixed camera. The camera characteristics may be predefined for the camera. Fig. 12 shows a side view 3300 and a top view 3400 of a wall-mounted camera. Fig. 13 shows a side view 3500 and a top view 3600 of a ceiling-mounted camera.
[0058] As shown in FIG. 12, camera installation height 3301 is the vertical distance from the ground to the camera. Camera tilt angle 3302 is the inclination angle of the camera's optical axis relative to the vertical direction. The distance from the camera to a region of interest (ROI) 3306 includes at least one of distance 3303 and distance 3304. Distance 3303 is the horizontal distance from the camera to the region of interest 3306. Distance 3304 is the distance from the camera to the region of interest 3306 in the direction of the optical axis. Camera field of view 3305 is the vertical angle of view centered on the optical axis toward the region of interest 3306. As shown in FIG. 12, camera field of view 3401 is the horizontal angle of view centered on the optical axis toward the region of interest 3402.
[0059] As shown in Figure 13, camera installation height 3501 is the vertical distance from the ground to the camera. Camera field of view 3502 is the vertical angle of view centered on the optical axis directed toward the region of interest. As shown in Figure 13, camera field of view 3601 is the horizontal angle of view centered on the optical axis directed toward the region of interest.
[0060] 27 and 28 are diagrams showing examples of camera characteristics for a camera mounted on a moving object. FIG. 27 shows a side view and a top view of a camera mounted on a vehicle or robot. FIG. 28 shows a side view and a top view of a camera mounted on an air vehicle. The camera can be mounted on a vehicle, robot, or air vehicle. For example, the camera can be mounted on a car, bus, truck, wheeled robot, legged robot, robotic arm, drone, or unmanned aerial vehicle.
[0061] As shown in Figure 27, the camera installation height is the vertical distance from the ground to the camera. The distance from the camera to the region of interest is the distance from the camera to the region of interest in the direction of the optical axis. The camera's field of view is the vertical and horizontal angles of view centered on the optical axis directed toward the region of interest.
[0062] As shown in Figure 28, the camera installation height is the vertical distance from the ground to the camera. The distance from the camera to the region of interest is the distance from the camera to the region of interest in the direction of the optical axis. The camera's field of view is the vertical and horizontal angles of view centered on the optical axis directed toward the region of interest.
[0063] The camera characteristics may be dynamically updated via other sensors mounted on the moving object. In the case of a camera mounted on a vehicle, the distance from the camera to the region of interest may be changed depending on the driving situation, such as driving on a highway or driving in a city. For example, braking distances differ between driving on a highway and driving in a city due to differences in vehicle speed. Specifically, high-speed driving on a highway requires longer braking distances, so it is necessary to be able to detect objects that are farther away. In contrast, normal-speed driving in a city requires shorter braking distances, so it is sufficient to be able to detect relatively closer objects. In practice, the distance from the camera to the ROI is changed by switching the focal length. For example, increasing the focal length increases the distance from the camera to the ROI. In the case of a camera mounted on an aircraft, the installation height of the camera may be changed based on the flight altitude of the aircraft. In the case of a camera mounted on a robotic arm, the distance from the camera to the region of interest may be changed according to the movement of the robotic arm.
[0064] As another example, the one or more parameters include at least one of a depth and a size for an object contained in the image.
[0065] 18 is a diagram showing an example of calculation of the depth and size of an object. In a side view 4200, an object 4204 is located at a location physically distant from a camera 4201 and is included in a field of view 4202 of the camera 4201. The separation distance between the camera 4201 and the object 4204, i.e., the depth, corresponds to a depth 4203 of the object 4204. An image 4300 captured by the camera 4201 includes an object 4301 corresponding to the object 4204. The image 4300 has a horizontal width 4302 and a vertical height 4303, and the object 4301 included in the image 4300 has a horizontal width 4304 and a vertical height 4305.
[0066] 16 is a flowchart showing an exemplary processing procedure S4000 for determining the size of an object. In step S4001, an image 4300 is read from a camera 4201. In step S4002, the size (e.g., horizontal width and vertical height) of an object 4204 is calculated based on the width 4304 and height 4305 of an object 4301 included in the image 4300. Alternatively, the size of the object 4204 may be estimated by running a computer vision algorithm on the image 4300. The size of the object 4204 may be used to estimate the distance between the object 4204 and the camera 4201. In step S4003, the size of the object 4204 is written to a bitstream encoding the image 4300 as one of one or more parameters related to the object 4301 included in the image 4300.
[0067] 17 is a flowchart showing an exemplary processing procedure S4100 for determining the depth of an object. In step S4101, an image 4300 is read from a camera 4201. In step S4102, a depth 4203 of an object 4204 is determined by using a stereo camera or by running a computer vision algorithm on the image 4300. Based on the depth 4203 of the object 4204, the distance between the object 4204 and the camera 4201 can be estimated. In step S4103, the depth 4203 of the object 4204 is written into a bitstream encoding the image 4300 as one of one or more parameters related to the object 4301 included in the image 4300.
[0068] 1, next, in step S1002A, the image encoding device 1101A generates a bitstream by encoding an image to generate pixel samples of the image. Here, one or more parameters are not used in encoding the image. The image encoding device 1101A adds one or more parameters to the bitstream and transmits the bitstream with the one or more parameters added to it to the image decoding device 2101A.
[0069] In a final step S1003A, the image coding device 1101A outputs a signal 1120A containing pixel samples of the image and one or more parameters to the first processing device 1102A.
[0070] The first processing unit 1102A performs a predetermined task process, such as a neural network task, using pixel samples of an image included in the input signal 1120A and one or more parameters. The neural network task may include at least one decision process. An example of a neural network is a convolutional neural network. Examples of neural network tasks include object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, machine-human hybrid vision, or any combination thereof.
[0071] FIG. 14 illustrates object detection and object segmentation as examples of neural network tasks. In object detection, attributes of objects (in this example, a television and a person) included in an input image are detected. In addition to the attributes of objects included in the input image, the positions and number of objects in the input image may also be detected. This may, for example, narrow down the positions of objects to be recognized or eliminate objects other than those to be recognized. Specific applications include face detection in cameras and pedestrian detection in autonomous driving. In object segmentation, pixels in areas corresponding to objects are segmented (i.e., separated). This may be used, for example, to separate obstacles from roads in autonomous driving to assist safe driving of automobiles, to detect product defects in factories, and to identify terrain in satellite images.
[0072] Figure 15 illustrates object tracking, action recognition, and pose estimation as examples of neural network tasks. Object tracking tracks the movement of objects in an input image. Potential applications include counting the number of users in a facility such as a store or analyzing the movements of athletes. Further increasing the processing speed enables real-time object tracking, enabling applications to camera processing such as autofocus. Action recognition detects the type of object movement (in this example, "riding a bicycle" or "walking"). For example, application to security cameras can prevent and detect criminal behavior such as robbery and shoplifting, or prevent factory workers from forgetting to perform tasks. Pose estimation detects the posture of an object by detecting key points and joints. Potential applications include industrial applications such as improving work efficiency in factories, security applications such as detecting abnormal behavior, and healthcare and sports.
[0073] The first processing device 1102A outputs a signal 1121A indicating the execution result of the neural network task. The signal 1121A may include at least one of the number of detected objects, the confidence level of the detected objects, boundary information or position information of the detected objects, and the classification category of the detected objects. The signal 1121A may be input from the first processing device 1102A to the image encoding device 1101A.
[0074] An example of how the first processing device 1102A utilizes one or more parameters will now be described.
[0075] FIG. 19 is a flowchart showing a processing procedure S5000 of a first example of use of one or more parameters. In step S5001, one or more parameters are acquired from a bitstream. In step S5002, the first processing device 1102A determines whether the values of the one or more parameters are less than a predetermined value. If it is determined that the values of the one or more parameters are less than the predetermined value (S5002: Yes), the first processing device 1102A selects machine learning model A in step S5003. If it is determined that the values of the one or more parameters are greater than or equal to the predetermined value (S5002: No), the first processing device 1102A selects machine learning model B in step S5004. In step S5005, the first processing device 1102A executes a neural network task using the selected machine learning model. Machine learning model A and machine learning model B may be models trained using different datasets or may include different neural network layer designs.
[0076] FIG. 20 is a flowchart showing processing procedure S5100 of a second example of use of one or more parameters. In step S5101, one or more parameters are obtained from a bitstream. In step S5102, the first processing device 1102A checks the values of the one or more parameters. If the values of the one or more parameters are less than a predetermined value A, the first processing device 1102A selects machine learning model A in step S5103. If the values of the one or more parameters are greater than a predetermined value B, the first processing device 1102A selects machine learning model B in step S5105. If the values of the one or more parameters are greater than or equal to the predetermined value A and less than or equal to the predetermined value B, the first processing device 1102A selects machine learning model C in step S5104. In step S5106, the first processing device 1102A executes a neural network task using the selected machine learning model.
[0077] FIG. 21 is a flowchart showing a processing procedure S5200 of a third example of using one or more parameters. In step S5201, one or more parameters are obtained from a bitstream. In step S5202, the first processing unit 1102A determines whether the values of the one or more parameters are less than a predetermined value. If it is determined that the values of the one or more parameters are less than the predetermined value (S5202: Yes), the first processing unit 1102A sets a detection threshold A in step S5203. If it is determined that the values of the one or more parameters are greater than or equal to the predetermined value (S5202: No), the first processing unit 1102A sets a detection threshold B in step S5204. In step S5205, the first processing unit 1102A executes a neural network task using the set detection threshold. The detection threshold may be used to control the estimated output of the neural network. For example, the detection threshold is used to compare with the confidence level of a detected object. If the confidence level of a detected object exceeds the detection threshold, the neural network outputs the confidence level.
[0078] 22 is a flowchart showing processing procedure S5300 of a fourth example of use of one or more parameters. In step S5301, one or more parameters are obtained from the bitstream. In step S5302, the first processing unit 1102A checks the values of the one or more parameters. If the values of the one or more parameters are less than a predetermined value A, the first processing unit 1102A sets a detection threshold A in step S5303. If the values of the one or more parameters are greater than a predetermined value B, the first processing unit 1102A sets a detection threshold B in step S5305. If the values of the one or more parameters are greater than or equal to the predetermined value A and less than or equal to the predetermined value B, the first processing unit 1102A sets a detection threshold C in step S5304. In step S5306, the first processing unit 1102A executes a neural network task using the set detection thresholds.
[0079] FIG. 23 is a flowchart showing a processing procedure S5400 of a fifth example of using one or more parameters. In step S5401, one or more parameters are obtained from the bitstream. In step S5402, the first processing unit 1102A determines whether the values of the one or more parameters are less than a predetermined value. If it is determined that the values of the one or more parameters are less than the predetermined value (S5402: Yes), in step S5403, the first processing unit 1102A sets a scaling value A. If it is determined that the values of the one or more parameters are greater than or equal to the predetermined value (S5402: No), in step S5404, the first processing unit 1102A sets a scaling value B. In step S5405, the first processing unit 1102A scales the input image based on the set scaling value. As an example, the input image is upscaled or downscaled based on the set scaling value. In step S5406, the first processing unit 1102A executes a neural network task using the scaled input image.
[0080] FIG. 24 is a flowchart showing processing procedure S5500 of a sixth example of using one or more parameters. In step S5501, one or more parameters are obtained from the bitstream. In step S5502, the first processing unit 1102A checks the values of the one or more parameters. If the values of the one or more parameters are less than a predetermined value A, the first processing unit 1102A sets a scaling value A in step S5503. If the values of the one or more parameters are greater than a predetermined value B, the first processing unit 1102A sets a scaling value B in step S5505. If the values of the one or more parameters are greater than or equal to the predetermined value A and less than or equal to the predetermined value B, the first processing unit 1102A sets a scaling value C in step S5504. In step S5506, the first processing unit 1102A scales the input image based on the set scaling value. In step S5507, the first processing unit 1102A executes a neural network task using the scaled input image.
[0081] FIG. 25 is a flowchart showing processing procedure S5600 of a seventh example of use of one or more parameters. In step S5601, one or more parameters are obtained from the bitstream. In step S5602, the first processing unit 1102A determines whether the values of the one or more parameters are less than a predetermined value. If it is determined that the values of the one or more parameters are less than the predetermined value (S5602: Yes), the first processing unit 1102A selects post-processing method A in step S5603. If it is determined that the values of the one or more parameters are greater than or equal to the predetermined value (S5602: No), the first processing unit 1102A selects post-processing method B in step S5604. In step S5605, the first processing unit 1102A performs filtering of the input image using the selected post-processing method. The post-processing method may be sharpening, blurring, morphological transformation, unsharp masking, or any combination of image processing methods. In step S5606, the first processing unit 1102A executes a neural network task using the filtered input image.
[0082] FIG. 26 is a flowchart showing processing procedure S5700 of an eighth example of using one or more parameters. In step S5701, one or more parameters are obtained from the bitstream. In step S5702, the first processing unit 1102A determines whether the values of the one or more parameters are less than a predetermined value. If it is determined that the values of the one or more parameters are less than the predetermined value (S5702: Yes), in step S5703, the first processing unit 1102A performs filtering of the input image using a predetermined post-processing method. If it is determined that the values of the one or more parameters are greater than or equal to the predetermined value (S5702: No), the first processing unit 1102A does not perform the filtering. In step S5704, the first processing unit 1102A performs a neural network task using the filtered input image or the unfiltered input image.
[0083] 7 is a block diagram showing an example configuration of an image encoding device 1101A according to the first embodiment of the present disclosure. The image encoding device 1101A is configured to encode an input image on a block-by-block basis and output an encoded bitstream. As shown in FIG. 7, the image encoding device 1101A includes a transform unit 1301, a quantization unit 1302, an inverse quantization unit 1303, an inverse transform unit 1304, a block memory 1306, an intra prediction unit 1307, a picture memory 1308, a block memory 1309, a motion vector prediction unit 1310, an interpolation unit 1311, an inter prediction unit 1312, and an entropy encoding unit 1313.
[0084] Next, an exemplary operation flow will be described. An input image and a predicted image are input to an adder, and an added value corresponding to a difference image between the input image and the predicted image is input from the adder to a transform unit 1301. The transform unit 1301 converts the added value and inputs the frequency coefficients obtained by converting the added value to a quantizer 1302. The quantizer 1302 quantizes the input frequency coefficients and inputs the quantized frequency coefficients to an inverse quantizer 1303 and an entropy encoder 1313. One or more parameters including the depth and size of an object are input to the entropy encoder 1313. The entropy encoder 1313 entropy encodes the quantized frequency coefficients to generate a bitstream. The entropy encoder 1313 adds one or more parameters including the depth and size of the object to the bitstream by entropy encoding them together with the quantized frequency coefficients or by storing them in the bitstream header.
[0085] The inverse quantization unit 1303 inversely quantizes the frequency coefficients input from the quantization unit 1302 and inputs the inversely quantized frequency coefficients to the inverse transform unit 1304. The inverse transform unit 1304 inversely transforms the frequency coefficients to generate a difference image and inputs the difference image to the adder. The adder adds the difference image input from the inverse transform unit 1304 to the predicted image input from the intra prediction unit 1307 or the inter prediction unit 1312. The adder inputs an addend 1320 (corresponding to the pixel sample described above) corresponding to the input image to the first processing unit 1102A, the block memory 1306, and the picture memory 1308. The addend 1320 is used for further prediction.
[0086] The first processing unit 1102A performs at least one of morphological transformation and edge enhancement processing such as unsharp masking on the sum 1320 based on at least one of the object's depth and size, thereby emphasizing the features of the object included in the input image corresponding to the sum 1320. The first processing unit 1102A performs object tracking, which involves at least a determination process, using the object-enhanced sum 1320 and at least one of the object's depth and size. The object's depth and size improve the accuracy and speed performance of object tracking. Here, the first processing unit 1102A may perform object tracking using position information indicating the position of the object included in the image (e.g., boundary information indicating the boundary surrounding the object) in addition to at least one of the object's depth and size. This further improves the accuracy of object tracking. In this case, the entropy encoding unit 1313 includes the object's position information in addition to the object's depth and size in the bitstream. The determination result 1321 is input from the first processing unit 1102A to the picture memory 1308 and used for further prediction. For example, by performing processing to emphasize an object based on the determination result 1321 in an input image corresponding to the added value 1320 stored in the picture memory 1308, it is possible to improve the accuracy of subsequent inter prediction. However, input of the determination result 1321 to the picture memory 1308 may be omitted.
[0087] The intra predictor 1307 and the inter predictor 1312 search for image regions that are most similar to the input image for prediction within the reconstructed image stored in the block memory 1306 or the picture memory 1308. The block memory 1309 fetches blocks of the reconstructed image from the picture memory 1308 using motion vectors input from the motion vector predictor 1310. The block memory 1309 inputs the blocks of the reconstructed image to the interpolator 1311 for interpolation processing. The interpolated image is input from the interpolator 1311 to the inter predictor 1312 for the inter prediction process.
[0088] 3 is a flowchart showing a processing procedure 1200A of the image coding method according to the first embodiment of the present disclosure. In the first step S1201A, the entropy coding unit 1313 encodes the depth and size of an object into a bitstream. The depth and size of the object may be added to the bitstream by being entropy coded, or may be added to the bitstream by being stored in the bitstream header.
[0089] Next, in step S1202A, the entropy encoder 1313 generates a bitstream by entropy encoding the image, generating pixel samples of the image. Here, the depth and size of the object are not used in the entropy encoding of the image. The entropy encoder 1313 adds the depth and size of the object to the bitstream, and transmits the bitstream with the added depth and size of the object to the image decoding device 2101A.
[0090] Next, in step S1203A, the first processing unit 1102A performs a combination of morphological transformation and edge enhancement processing such as unsharp masking on pixel samples of the image based on the depth and size of the object to enhance features of at least one object contained in the image. The object enhancement processing in step S1203A improves the accuracy of the neural network task performed by the first processing unit 1102A in the next step S1204A.
[0091] In the final step S1204A, the first processing unit 1102A performs object tracking, which involves at least a decision process, based on the pixel samples of the image and the depth and size of the object, where the depth and size of the object improve the accuracy and speed performance of the object tracking. The combination of the morphological transformation and edge enhancement process, such as unsharp masking, may be replaced by other image processing techniques.
[0092] (Decoder side processing) 2 is a flowchart showing a processing procedure 2000A of an image decoding method according to the first embodiment of the present disclosure. In the first step S2001A, an image decoding device 2101A decodes one or more parameters from a bitstream.
[0093] Fig. 12 and Fig. 13 are diagrams showing examples of camera characteristics relating to the installation position of a fixed camera. Fig. 27 and Fig. 28 are diagrams showing examples of camera characteristics relating to a camera mounted on a moving body. Fig. 18 is a diagram showing an example of calculation of the depth and size of an object. Fig. 16 is a flowchart showing an example of a processing procedure S4000 for determining the size of an object. Fig. 17 is a flowchart showing an example of a processing procedure S4100 for determining the depth of an object. The processing corresponding to these figures is similar to the processing on the encoder side, so duplicated explanations will be omitted.
[0094] Next, in step S2002A, the image decoding device 2101A generates pixel samples of the image by decoding the image from the bitstream, where one or more parameters are not used in decoding the image, and the image decoding device 2101A obtains one or more parameters from the bitstream.
[0095] In a final step S2003A, the image decoding device 2101A outputs a signal 2120A including pixel samples of the image and one or more parameters to the second processing device 2102A.
[0096] The second processing unit 2102A performs a predetermined task similar to that performed by the first processing unit 1102A using pixel samples of an image included in the input signal 2120A and one or more parameters. The neural network task may include at least one decision process. An example of a neural network is a convolutional neural network. Examples of neural network tasks include object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, machine-human hybrid vision, or any combination thereof.
[0097] Fig. 14 shows object detection and object segmentation as examples of neural network tasks. Fig. 15 shows object tracking, action recognition, and pose estimation as examples of neural network tasks. The processes corresponding to these figures are similar to those on the encoder side, so redundant explanations will be omitted.
[0098] The second processing device 2102A outputs a signal 2121A indicating the execution result of the neural network task. The signal 2121A may include at least one of the number of detected objects, the confidence level of the detected objects, boundary information or position information of the detected objects, and the classification category of the detected objects. The signal 2121A may be input from the second processing device 2102A to the image decoding device 2101A.
[0099] An example of how the second processing device 2102A utilizes one or more parameters will be described below.
[0100] FIG. 19 is a flowchart showing a processing procedure S5000 for a first example of using one or more parameters. FIG. 20 is a flowchart showing a processing procedure S5100 for a second example of using one or more parameters. FIG. 21 is a flowchart showing a processing procedure S5200 for a third example of using one or more parameters. FIG. 22 is a flowchart showing a processing procedure S5300 for a fourth example of using one or more parameters. FIG. 23 is a flowchart showing a processing procedure S5400 for a fifth example of using one or more parameters. FIG. 24 is a flowchart showing a processing procedure S5500 for a sixth example of using one or more parameters. FIG. 25 is a flowchart showing a processing procedure S5600 for a seventh example of using one or more parameters. FIG. 26 is a flowchart showing a processing procedure S5700 for an eighth example of using one or more parameters. The processing corresponding to these figures is similar to the processing on the encoder side, so duplicated explanations will be omitted.
[0101] 8 is a block diagram showing an example configuration of an image decoding device 2101A according to the first embodiment of the present disclosure. The image decoding device 2101A is configured to decode an input bitstream on a block-by-block basis and output a decoded image. As shown in FIG. 8, the image decoding device 2101A includes an entropy decoding unit 2301, an inverse quantization unit 2302, an inverse transform unit 2303, a block memory 2305, an intra prediction unit 2306, a picture memory 2307, a block memory 2308, an interpolation unit 2309, an inter prediction unit 2310, an analysis unit 2311, and a motion vector prediction unit 2312.
[0102] Next, an exemplary operation flow will be described. The coded bitstream input to the image decoding device 2101A is input to the entropy decoding unit 2301. The entropy decoding unit 2301 decodes the input bitstream and inputs the frequency coefficients, which are decoded values, to the inverse quantization unit 2302. The entropy decoding unit 2301 also obtains the depth and size of the object from the bitstream and inputs this information to the second processing device 2102A. The inverse quantization unit 2302 inversely quantizes the frequency coefficients input from the entropy decoding unit 2301 and inputs the inversely quantized frequency coefficients to the inverse transform unit 2303. The inverse transform unit 2303 inversely transforms the frequency coefficients to generate a difference image and inputs the difference image to an adder. The adder adds the difference image input from the inverse transform unit 2303 to a predicted image input from the intra prediction unit 2306 or the inter prediction unit 2310. The adder inputs an added value 2320 corresponding to the input image to the display device, which then displays the image. The adder also inputs the added value 2320 to the second processing unit 2102A, the block memory 2305, and the picture memory 2307. The added value 2320 is used for further prediction.
[0103] The second processing unit 2102A performs at least one of morphological transformation and edge enhancement processing such as unsharp masking on the sum 2320 based on at least one of the object's depth and size, thereby emphasizing the features of the object included in the input image corresponding to the sum 2320. The second processing unit 2102A performs object tracking, which involves at least a determination process, using the object-enhanced sum 2320 and at least one of the object's depth and size. The object's depth and size improve the accuracy and speed performance of object tracking. Here, the second processing unit 2102A may perform object tracking using position information indicating the position of the object included in the image (e.g., boundary information indicating the boundary surrounding the object) in addition to at least one of the object's depth and size. This further improves the accuracy of object tracking. In this case, the position information is included in the bitstream, and the entropy decoding unit 2301 acquires the position information from the bitstream. The determination result 2321 is input from the second processing unit 2102A to the picture memory 2307 and used for further prediction. For example, by performing processing to emphasize an object based on the determination result 2321 in an input image corresponding to the added value 2320 stored in the picture memory 2307, it is possible to improve the accuracy of subsequent inter prediction. However, input of the determination result 2321 to the picture memory 2307 may be omitted.
[0104] The parser 2311 parses the input bitstream to input some prediction information, such as a block of residual samples, a reference index indicating the reference picture to be used, and a delta motion vector, to the motion vector predictor 2312. The motion vector predictor 2312 predicts the motion vector of the current block based on the prediction information input from the parser 2311. The motion vector predictor 2312 inputs a signal indicating the predicted motion vector to the block memory 2308.
[0105] The intra predictor 2306 and the inter predictor 2310 search for image regions that are most similar to the input image for prediction within the reconstructed image stored in the block memory 2305 or the picture memory 2307. The block memory 2308 fetches blocks of the reconstructed image from the picture memory 2307 using the motion vectors input from the motion vector predictor 2312. The block memory 2308 inputs the blocks of the reconstructed image to the interpolator 2309 for interpolation processing. The interpolated image is input from the interpolator 2309 to the inter predictor 2310 for the inter prediction process.
[0106] 4 is a flowchart showing a processing procedure 2200A of the image decoding method according to the first embodiment of the present disclosure. In the first step S2201A, the entropy decoding unit 2301 decodes the depth and size of the object from the bitstream.
[0107] Next, in step S2202A, the entropy decoding unit 2301 generates pixel samples of the image by entropy decoding the image from the bitstream. The entropy decoding unit 2301 also obtains the depth and size of the object from the bitstream. Here, the depth and size of the object are not used in the entropy decoding of the image. The entropy decoding unit 2301 inputs the obtained depth and size of the object to the second processing device 2102A.
[0108] Next, in step S2203A, the second processing unit 2102A performs a combination of morphological transformations and edge enhancement operations such as unsharp masking on pixel samples of the image based on the depth and size of the object to enhance features of at least one object contained in the image. The object enhancement operation in step S2203A improves the accuracy of the neural network task performed by the second processing unit 2102A in the next step S2204A.
[0109] In the final step S2204A, the second processing unit 2102A performs object tracking, which involves at least a decision process, based on the pixel samples of the image and the depth and size of the object, where the depth and size of the object improve the accuracy and speed performance of the object tracking. The combination of morphological transformation and edge enhancement processes such as unsharp masking may be replaced by other image processing techniques.
[0110] According to this embodiment, the image encoding device 1101A transmits, to the image decoding device 2101A, one or more parameters to be output to the first processing device 1102A for executing predetermined task processing. This allows the image decoding device 2101A to output the one or more parameters received from the image encoding device 1101A to the second processing device 2102A that executes task processing similar to the predetermined task processing. As a result, the second processing device 2102A executes the predetermined task processing based on the one or more parameters input from the image decoding device 2101A, thereby making it possible to improve the accuracy of the task processing in the second processing device 2102A.
[0111] (Second embodiment) In a second embodiment of the present disclosure, a description will be given of how to deal with a case where a camera that outputs images with large distortion, such as a fisheye camera, an ultra-wide-angle camera, or an omnidirectional camera, may be used in the first embodiment.
[0112] (Encoder side processing) 32 is a block diagram showing an example configuration of an encoder 2100B according to the second embodiment of the present disclosure. The encoder 2100B includes a coding unit 2101B and an entropy coding unit 2102B. The entropy coding unit 2102B corresponds to the entropy coding unit 1313 shown in FIG. 7. The coding unit 21021 corresponds to the configuration shown in FIG. 7 excluding the entropy coding unit 1313 and the first processing device 1102A.
[0113] 30 is a flowchart showing a processing procedure 2000B of an image encoding method according to the second embodiment of the present disclosure. In the first step S2001B, an entropy encoding unit 2102B generates a bitstream by entropy encoding an image input from an encoding unit 2101B. The image input to the encoding unit 2101B may be an image output from a camera with large distortion, such as a fisheye camera, an ultra-wide-angle camera, or an omnidirectional camera. The image includes at least one object, such as a person.
[0114] Fig. 33 shows a comparison of the output images from a normal camera and a camera with large distortion. The left side shows the output image from the normal camera, and the right side shows the output image from a camera with large distortion (in this example, an omnidirectional camera).
[0115] Next, in step S2002B, the entropy encoding unit 2102B encodes a parameter set included in the one or more parameters into a bitstream. The parameter set includes boundary information indicating a boundary surrounding an object included in the image and distortion information indicating whether or not the image is distorted.
[0116] The boundary information includes position coordinates of multiple vertices of a bounding box, which is a figure that defines the boundary. Alternatively, the boundary information may include center coordinates, width information, height information, and tilt information of the bounding box. The distortion information includes additional information indicating that the image was captured by a fisheye camera, an ultra-wide-angle camera, or an omnidirectional camera. The boundary information and distortion information may be input to the entropy encoding unit 2102B from the camera or sensor 3101 shown in FIG. 10 or from the pre-processing unit 3202 shown in FIG. 11.
[0117] The parameter sets may be added to the bitstream by being entropy coded or by being stored in the bitstream header.
[0118] The encoder 2100B transmits the bitstream with the parameter set added to the decoder 1100B.
[0119] In the final step S2003B, the entropy encoding unit 2102B outputs the image and the parameter set to the first processing device 1102A. The first processing device 1102A executes a predetermined task process, such as a neural network task, using the input image and parameter set. The neural network task may include at least one determination process. The first processing device 1102A may switch between a machine learning model for distorted images with large distortion and a machine learning model for normal images with small distortion, depending on whether the distortion information in the parameter set includes the additional information.
[0120] 34 to 37 are diagrams showing examples of boundary information. Referring to FIGS. 34 and 35, the boundary information includes the position coordinates of multiple vertices of a bounding box. When the bounding box is defined by a rectangle, the boundary information includes four pixel coordinates (x coordinate and y coordinate) indicating the positions of pixels corresponding to the four vertices a to d. The four pixel coordinates bound the object, and the four pixel coordinates form the shape of a four-sided polygon.
[0121] As shown in Figure 36, multiple bounding boxes may be defined because an image contains multiple objects. Also, because the bounding box is tilted due to image distortion or the like, the side of the bounding box (left or right side) may not be parallel to the side of the screen.
[0122] 37, the shape of the bounding box is not limited to a rectangle, but may be a square, a parallelogram, a trapezoid, a rhombus, etc. Furthermore, if the contour of the object is distorted due to image distortion or the like, the shape of the bounding box may be any trapezoid.
[0123] 34, the boundary information may include center coordinates (x and y coordinates), width information (width), height information (height), and tilt information (angle θ) related to the bounding box. If the bounding box is rectangular, the four pixel coordinates corresponding to the four vertices a to d can be calculated based on the center coordinates, width information, and height information by using the approximation formula shown in FIG.
[0124] (Decoder side processing) 31 is a block diagram showing an example configuration of a decoder 1100B according to the second embodiment of the present disclosure. The decoder 1100B includes an entropy decoding unit 1101B and a decoding unit 1102B. The entropy decoding unit 1101B corresponds to the entropy decoding unit 2301 shown in FIG. 8. The decoding unit 1102B corresponds to the configuration shown in FIG. 8 excluding the entropy decoding unit 2301 and the second processing device 2102A.
[0125] 29 is a flowchart showing a processing procedure 1000B of an image decoding method according to the second embodiment of the present disclosure. In the first step S1001B, an entropy decoding unit 1101B decodes an image from a bitstream received from an encoder 2100B. The image includes at least one object such as a person.
[0126] Next, in step S1002B, the entropy decoding unit 1101B decodes a parameter set from the bitstream received from the encoder 2100B. The parameter set includes boundary information indicating a boundary surrounding an object included in the image and distortion information indicating the presence or absence of distortion in the image.
[0127] In the final step S1003B, the entropy decoding unit 1101B outputs the decoded image and parameter set to the second processing device 2102A. The second processing device 2102A uses the input image and parameter set to execute a predetermined task similar to that executed by the first processing device 1102A. At least one determination process may be executed in the neural network task. The second processing device 2102A may switch between a machine learning model for distorted images with large distortion and a machine learning model for normal images with small distortion, depending on whether the distortion information in the parameter set includes the additional information.
[0128] According to this embodiment, even when a camera that outputs images with large distortion, such as a fisheye camera, an ultra-wide-angle camera, or an omnidirectional camera, is used, a bounding box surrounding an object can be accurately defined. Furthermore, the encoder 2100B transmits a parameter set including boundary information and distortion information to the decoder 1100B. This allows the decoder 1100B to output the parameter set received from the encoder 2100B to the second processing device 2102A. As a result, the second processing device 2102A executes a predetermined task process based on the input parameter set, thereby improving the accuracy of the task process in the second processing device 2102A. [Industrial Applicability]
[0129] The present disclosure is particularly useful when applied to an image processing system that includes an encoder for transmitting images and a decoder for receiving images.
Claims
1. An image encoding device, generating a bitstream by encoding the image; adding one or more parameters not used in encoding the image to the bitstream by encoding the one or more parameters; transmitting the bitstream to which the one or more parameters have been added to an image decoding device; outputting the image and the one or more parameters to a first processing device that executes a predetermined task process; the predetermined task processing includes at least one of object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, and hybrid vision; the image decoding device outputs the image and the one or more parameters to a second processing device that executes the predetermined task processing; Image processing methods.
2. An image decoding device, receiving a bitstream from an image encoding device; Decoding an image from the bitstream; obtaining from the bitstream one or more parameters that have been added to the bitstream by encoding the image and that are not used in decoding the image; the image encoding device outputs the image and the one or more parameters to a first processing device that executes a predetermined task process; the predetermined task processing includes at least one of object detection, object segmentation, object tracking, action recognition, pose estimation, pose tracking, and hybrid vision; the image decoding device outputs the image and the one or more parameters to a second processing device that executes the predetermined task processing; Image processing methods.
Citation Information
Patent Citations
Electronic apparatus, method for controlling thereof, and method for controlling server
US20200053408A1
Object region detection method, object region detection apparatus, and non-transitory computer-readable medium thereof
US20200098132A1
Utilizing a neural network having a two-stream encoder architecture to generate composite digital images
US20210027470A1
Transmitting device, transmitting method, receiving device and receiving method
WO2018043143A1
Encoding device, decoding device, encoding method, and decoding method
WO2019093234A1