Image processing device, image processing method, learning device, generation method, and program
The image processing system addresses the challenge of expressing appropriate texture in each object region by using texture labels and DNNs for texture segmentation and super-resolution, achieving enhanced texture control and image quality adjustment.
Patent Information
- Application Number
- JP2022524368
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-20
- Filing Date
- 2021-05-06
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-05-06
AI Technical Summary
Existing image processing technologies struggle to express appropriate texture in each region of an object due to the difficulty in defining physical parameters for texture control, leading to inadequate texture reproduction or enhancement.
An image processing system that utilizes a control signal generation unit and an image generation unit to process input images using a learned inference model, where texture labels are assigned to each region, enabling direct control of texture through a Deep Neural Network (DNN) for texture segmentation and super-resolution processing.
The system allows for precise control of texture in each image region, aligning with human qualitative senses and improving image quality adjustment by applying texture labels that vary based on local object characteristics, enhancing texture expression and controllability.
Smart Images

Figure 0007729336000001 
Figure 0007729336000002 
Figure 0007729336000003
Abstract
Description
[Technical Field]
[0001] In particular, the present technology relates to an image processing device, an image processing method, a learning device, a generation method, and a program that are capable of generating an image in which an appropriate texture is expressed in each region. [Background technology]
[0002] When adjusting the image quality of display devices such as TVs, there is a need to reproduce or improve texture. Image processing to reproduce or improve texture is usually achieved by combining technologies such as NR (Noise Reduction) processing, super-resolution processing, and contrast / color adjustment processing, or by adjusting the strength of image processing, rather than controlling the texture itself.
[0003] Texture can be considered a qualitative human sensation. Because it is difficult to define physical parameters suitable for expressing texture, it is also difficult to control texture using conventional model-based processing. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-190371 Summary of the Invention [Problem to be solved by the invention]
[0005] Texture can be expressed in various ways, such as fineness, fine grain, shape, gloss, transparency, shadow, texture, and unevenness. The optimal texture varies depending on the characteristics of the object.
[0006] Even if objects in an image are detected by semantic segmentation or the like and processed according to their texture, it is not enough to perform processing to express the same texture for all objects or for all regions of an object. In other words, failure may occur if processing is not performed to express an appropriate texture for each region of an object.
[0007] The present technology has been developed in light of these circumstances, and makes it possible to generate an image in which an appropriate texture is expressed in each region. [Means for solving the problem]
[0008] An image processing device according to one aspect of the present technology includes a control signal generation unit that generates, based on an input image to be processed, a control signal that represents the texture of each region to be realized in an output image of the inference result, and an image generation unit that inputs the input image into an inference model obtained by learning based on a student image generated by applying predetermined image processing to a teacher image and the teacher image in which the texture of each region is expressed by a texture label, and infers the output image in which each region has the texture represented by the control signal.
[0009] A learning device according to another aspect of the present technology includes an acquisition unit that acquires texture labels that represent the texture of each region of a training image, and a learning unit that performs learning using an image generated by performing predetermined image processing on the training image as a student image and the training image as a teacher image in accordance with a control signal that represents the texture of each region of the training image, thereby generating an inference model.
[0010] In one aspect of the present technology, a control signal representing the texture of each region realized in an output image of the inference result is generated based on an input image to be processed, the input image is input to an inference model obtained by learning based on a student image generated by applying a predetermined image processing to a teacher image and the teacher image in which the texture of each region is expressed by a texture label, and an inference is made of the output image in which each region has the texture represented by the control signal.
[0011] In another aspect of the present technology, texture labels representing the texture of each region of a training image are obtained, and an image generated by performing predetermined image processing on the training image is used as a student image. Learning is performed using the training image as a teacher image in accordance with a control signal representing the texture of each region of the training image, and an inference model is generated. [Brief explanation of the drawings]
[0012] [Figure 1] 10A and 10B are diagrams illustrating examples of labels used in image processing of the present technology. [Figure 2] 10A and 10B are diagrams illustrating an example of image processing for controlling texture. [Figure 3] FIG. 10 is a diagram illustrating an example of processing according to an object. [Figure 4] 1 is a diagram illustrating an example of the configuration of an image processing system according to an embodiment of the present technology. [Figure 5] FIG. 2 is a block diagram illustrating an example of the configuration of a learning device. [Figure 6] FIG. 10 is a diagram illustrating an example of setting texture labels. [Figure 7] FIG. 10 is a diagram illustrating an example of training a DNN for texture segmentation detection. [Figure 8] FIG. 1 is a diagram illustrating an example of learning of a DNN for super-resolution processing. [Figure 9] FIG. 10 is a diagram illustrating an example of conversion of texture axis values. [Figure 10] FIG. 1 is a block diagram illustrating an example of the configuration of an image processing device. [Figure 11]FIG. 10 is a diagram illustrating an example of inference using a DNN for texture segmentation detection. [Figure 12] FIG. 10 is a diagram illustrating an example of conversion of texture axis values. [Figure 13] FIG. 10 is a diagram illustrating an example of calculation of a texture axis value. [Figure 14] 10A and 10B are diagrams illustrating examples of adjustment of texture axis values. [Figure 15] FIG. 1 is a diagram illustrating an example of inference using a DNN for super-resolution processing. [Figure 16] 10 is a flowchart illustrating a texture label setting process of the learning device. [Figure 17] 10 is a flowchart illustrating a DNN generation process for detecting texture segmentation performed by a learning device. [Figure 18] 10 is a flowchart illustrating the DNN generation process for super-resolution processing of the learning device. [Figure 19] 10 is a flowchart illustrating an inference process of the image processing device. [Figure 20] FIG. 10 is a diagram illustrating an example of setting an object label and a texture label. [Figure 21] FIG. 10 is a diagram illustrating an example of setting an object label and a texture label. [Figure 22] FIG. 10 is a diagram illustrating an example of setting an object label and a texture label. [Figure 23] FIG. 1 is a diagram illustrating an example of inference using a DNN for super-resolution processing. [Figure 24] FIG. 10 is a diagram illustrating an example of an image quality label. [Figure 25] FIG. 10 is a diagram showing an example of texture labels that incorporate image creation intentions. [Figure 26] FIG. 10 is a diagram showing an image of the image quality of the inference result. [Figure 27] FIG. 1 is a block diagram illustrating an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present technology will be described in the following order. 1. Prerequisites for this technology 2. Image processing system configuration 3. Training the DNN 4. Inference using DNN 5. Operation of the image processing system 6. Label setting example 7. Application Examples 8. Other Examples
[0014] <<Prerequisites for this technology>> FIG. 1 is a diagram showing an example of a label used in the image processing system of the present technology.
[0015] When an input image such as that shown in Figure 1A, which shows a car, is used as the image to be processed, an object label such as that shown in Figure 1B is used for image processing. The object label shown in Figure 1B is information indicating that a car is shown in region #1, which is the region approximately in the center of the input image.
[0016] In the example of Figure 1B, area #1 is simply shown as an oval area, but in reality, an area corresponding to the shape of a car is represented by an object label. Similarly, in other figures described later, the shape of each area will be a shape corresponding to the shape of an object or the like that appears in that area.
[0017] In this technology, not only object labels but also texture labels as shown in FIG. 1C are used.
[0018] The texture label is information that represents the texture of each region. As will be described later, a texture label that is evaluated by a human as being appropriate for expressing the texture of each region is assigned to each region depending on the content of the object or the like depicted in that region.
[0019] In the example of Figure 1C, the texture labels indicate that the texture of area #11, which shows the windshield, is "strong transparency," and the texture of area #12, which shows the headlights, is "strong gloss."
[0020] The texture labels also indicate that the texture of area #13, where the license plate is visible, is "strong clarity of characters," and the texture of area #14, where the side door is visible, is "strong roughness (smoothness)." The texture labels also indicate that the texture of area #15, a part of the floor, is "strong roughness (smoothness)" and "strong gloss."
[0021] In this way, the texture label is information that indicates the type of texture expression that expresses the qualitative texture of each region and the intensity of the texture expressed by that texture expression. The part before the ": (colon)" indicates the type of texture expression, and the part after it indicates the intensity of the texture.
[0022] Texture expressions are defined as fineness, fine grain, shape, gloss, transparency, shadow, texture, matte, unevenness, sizzle, etc. Texture intensity is defined in four levels, for example, weak, medium, strong, and OFF (Unlabeled). Two levels, three levels, or five or more levels may also be defined.
[0023] FIG. 2 is a diagram showing an example of image processing for controlling texture.
[0024] In Figure 2, the left side shows image processing using object labels, and the right side shows image processing using texture labels.
[0025] When image processing for texture control, such as reproducing and improving texture, is performed using object labels, as shown on the left side of A in Figure 2, for example, strong contrast processing is applied to region #21, which contains a car, and strong NR processing is applied to region #22, which contains the sky. Similarly, super-resolution processing (SR (Super Resolution)) with strengths appropriate to the object is applied to other regions.
[0026] In this way, image processing to control texture using object labels is achieved by combining super-resolution processing, contrast and color adjustment processing, enhancement processing, NR processing, etc. for each object area.
[0027] FIG. 3 is a diagram illustrating an example of processing according to an object.
[0028] As shown in Figure 3, when the object is "leaves, trees, grass, flowers, etc. (without shape)," in order to express a sense of detail as texture, processing is performed to make the amplitude of the image signal medium and the bandwidth high. Similarly, for other objects, the type and degree of processing to express texture are set in advance and processing is performed.
[0029] Combining various processes and applying image processing to each region according to preset settings is not realistic in terms of performance, processing volume, scale, adjustment effort, etc. Furthermore, even if image processing is applied according to preset settings, it is not clear whether the desired texture will be achieved.
[0030] On the other hand, if texture control is performed using texture labels, for example, image processing according to "glossiness: strong" and "transparency: strong" is performed on region #31 showing the body of a car, and image processing according to "roughness (smoothness): strong" and "perspective: strong" is performed on region #32 showing the sky, as shown on the right side of A in Figure 2. Similarly, image processing according to each texture is performed on the other regions based on the texture labels.
[0031] Image processing according to texture means image processing to realize that texture. For example, image processing for area #31 is image processing to realize a texture with a strong glossiness and a strong transparency.
[0032] As will be described later, in the image processing system of this technology, image processing is performed using a DNN (Deep Neural Network) to generate an image. Applying image processing according to texture to each region means that an image composed of each region in which that texture is realized is generated.
[0033] In this way, texture labels are introduced into the image processing system of this technology. By introducing texture labels and enabling direct control of texture, image quality control that is in line with human qualitative senses is realized. In other words, it becomes possible to create and adjust images based on human senses.
[0034] Even in areas that contain the same object, the optimal texture varies from area to area depending on the characteristics of the materials that make up the object. By introducing texture labels, it becomes possible to control the texture of each area according to the local characteristics of the object and the image creation policy. In addition, by being able to control the intensity of the texture, it becomes possible to improve the controllability of image quality adjustment.
[0035] In this way, the image processing of this technology provides a new control axis, the texture axis, which is different from the control axis of conventional image quality control. By changing the texture label, it is also possible to provide a control axis specialized for the use case at the image output destination.
[0036] <<Image processing system configuration>> FIG. 4 is a diagram illustrating an example of the configuration of an image processing system according to an embodiment of the present technology.
[0037] The image processing system in Figure 4 is composed of a learning device 1 and an image processing device 2. The learning device 1 and the image processing device 2 may be realized by devices in the same housing, or may be realized by devices in different housings.
[0038] The learning device 1 creates data for learning an inference model such as a DNN. The learning device 1 performs learning using the learning data to generate a DNN.
[0039] As will be described in detail later, through learning in the learning device 1, a DNN that associates a texture label with an area to be texture-controlled is generated. When an image to be processed is input to this DNN, the texture label of each area is output. The DNN that associates the texture label with the area to be texture-controlled is a DNN for texture segmentation detection, which is used for detecting the area to be texture-controlled based on the texture label.
[0040] Also, through learning in the learning device 1, a DNN for super-resolution processing, which is a DNN capable of controlling super-resolution processing using a texture axis value as a control signal, is generated. The texture axis value is a value determined based on the texture label as will be described later. When an image to be processed is input to the DNN for super-resolution processing, a high-resolution image (super-resolution image) subjected to super-resolution processing according to the texture axis value is output.
[0041] The learning device 1 outputs, as a learning DB (Data Base), information on two DNNs, namely, the DNN for texture segmentation detection and the DNN for super-resolution processing, which includes information on the coefficients constituting each layer, to the image processing device 2.
[0042] The image processing device 2 generates a high-resolution image based on the input image by performing inference using the DNN for texture segmentation detection and the DNN for super-resolution processing. For the image processing device 2, for example, images of each frame constituting a moving image captured by a camera are supplied as input images. A moving image of CG (Computer Graphics) may be supplied as the input image, or a still image may be supplied.
[0043] <<Learning of DNN>> <Configuration of Learning Device 1> FIG. 5 is a block diagram showing a configuration example of the learning device 1.
[0044] The learning device 1 is composed of a texture label definition unit 11, a texture label assignment processing unit 12, a degradation processing unit 13, a DNN learning unit 14, a texture axis value conversion unit 15, an object detection unit 16, and a DNN learning unit 17. A Ground Truth image, which is an image for learning, is input to the texture label assignment processing unit 12, the degradation processing unit 13, the object detection unit 16, and the DNN learning unit 17. When the image processing performed in the image processing device 2 is super-resolution processing, the Ground Truth image is a high-resolution image.
[0045] The texture label definition unit 11 outputs information defining the type, intensity, etc. of the texture label to the texture label assignment processing unit 12.
[0046] The texture label assignment processing unit 12 sets texture labels for each region of a GT image (Ground Truth image) in accordance with user operations. When setting texture labels, the user, viewing the GT image, performs an operation to specify the texture label for each region. The texture label assignment processing unit 12 outputs information on the texture label for each region to the DNN learning unit 14 and the texture axis value conversion unit 15. The texture label assignment processing unit 12 functions as an acquisition unit that acquires texture labels that represent the texture of each region of the GT image.
[0047] The degradation processing unit 13 performs degradation processing on the GT image to generate a degraded image. The degradation processing unit 13 outputs the degraded image to the DNN learning unit 14 and the DNN learning unit 17. The degradation processing performed by the degradation processing unit 13 is a down-conversion process for generating an image equivalent to a low-resolution image that is input to the super-resolution processing.
[0048] The DNN learning unit 14 performs learning using the texture labels supplied from the texture label assignment processing unit 12 as teacher data and the degraded images supplied from the degradation processing unit 13 as student data, thereby generating a DNN for texture segmentation detection. The DNN learning unit 14 outputs information such as the coefficients of each layer constituting the DNN for texture segmentation detection as a learning DB 21.
[0049] The texture axis value conversion unit 15 converts the intensity of the texture of each region into a texture axis value based on the texture label supplied from the texture label assignment processing unit 12. The texture axis value conversion unit 15 outputs information on the texture axis value of each region to the DNN learning unit 17.
[0050] The object detection unit 16 performs processing such as semantic segmentation on the GT image to detect objects (objects included in each region) appearing in each region of the GT image. Object detection may be performed by a process other than semantic segmentation. The object detection unit 16 outputs object labels representing objects appearing in each region to the DNN learning unit 17.
[0051] The DNN learning unit 17 performs learning using the GT image as a teacher image and the degraded image supplied from the degradation processing unit 13 as a student image, thereby generating a DNN for super-resolution processing. A DNN having a predetermined network structure such as a GAN (Generative Adversarial Network) is generated as the DNN for super-resolution processing. DNN processing using a GAN and DNN processing such as Style Transfer have a high ability to bring an input image closer to the taste of a group of teacher images that are the correct answer, making it possible to express texture.
[0052] Learning by the DNN learning unit 17 is performed using the texture axis value supplied from the texture axis value conversion unit 15 and the object label supplied from the object detection unit 16 as control signals. For each combination of the texture axis value of each region and the object appearing in each region, coefficients for generating an image with a different texture as an image of each region are learned. The DNN learning unit 17 outputs information such as the coefficients of each layer constituting the DNN for super-resolution processing as a learning DB 22.
[0053] The processing of each part of the learning device 1 will be described in detail below.
[0054] <Texture label settings> FIG. 6 is a diagram showing an example of setting texture labels.
[0055] The texture label is information that indicates the type of texture expression that expresses the texture of each region and the intensity of the texture expressed by that texture expression.
[0056] The texture label assignment for each region of the GT image is performed by a user who views the GT image and evaluates the texture of each region. The user evaluates the texture of each region of the GT image according to the characteristics of each part that constitutes the object and the image creation policy. In response to the user's operation, the texture label assignment processing unit 12 assigns a texture label to each region of the GT image.
[0057] In the example of A in Figure 6, the texture labels "Glossiness: Strong" and "Transparency: Strong" are assigned to the region #71 in the approximate center of the GT image, where the body of the car is shown, and the texture labels "Roughness (Smoothness): Strong" and "Perspective: Strong" are assigned to the region #72, where the sky is shown.
[0058] The texture label "Medium Detail" is assigned to areas #73 and #76, which show distant scenery, and the texture label "Strong Fine Grain" is assigned to areas #74 and #75, which show the road surface.
[0059] In the example of Figure 6B, a texture label of "high detail" is set for the area #81 in the approximate center of the GT image, which contains the flower, and a texture label of "medium fine grain" is set for the area #82, which contains the background.
[0060] A texture label of "Medium Detail" is set for areas #83 and #85, which show the background, and a texture label of "High Gloss" is set for area #84, which shows the pot.
[0061] This type of texture labeling is performed on various GT images.
[0062] By manually labeling ground truth images to evaluate their texture, the qualitative human sense of texture is incorporated into image processing as texture labels.
[0063] The texture label may be set for any region specified by the user, or the results of segmentation using SLIC (Simple Linear Iterative Clustering) or the like may be presented and the texture label may be set for a specified region from among them.
[0064] <Training: DNN for texture segmentation detection> FIG. 7 is a diagram illustrating an example of training a DNN for detecting texture segmentation.
[0065] The DNN for texture segmentation detection is a DNN that links texture labels with areas that are the target of texture control.
[0066] As indicated by the arrow A1 in Fig. 7, learning is performed by the DNN learning unit 14 using texture labels as teacher data and degraded images as student data. In the example of Fig. 7, the texture labels set for the GT image described with reference to Fig. 6B and the degraded image generated based on the GT image are shown as teacher data and student data, respectively. In Fig. 7, the objects in the degraded image are shown in light colors, indicating that the resolution of the degraded image is lower than that of the GT image. This also applies to the subsequent figures.
[0067] By using a DNN for texture segmentation detection generated through such learning, it becomes possible to infer which texture label to assign to each region of the image being processed.
[0068] <Learning: DNN for super-resolution processing> FIG. 8 is a diagram illustrating an example of learning of a DNN for super-resolution processing.
[0069] The DNN for super-resolution processing is a DNN that can control super-resolution processing using the texture axis value as a control signal.
[0070] As indicated by the tip of arrow A11, the process of converting the intensity of the texture of each region represented by the texture label into a texture axis value is performed by the texture axis value conversion unit 15. Learning by the DNN learning unit 17, which uses the GT image as the teacher image and the degraded image as the student image, is performed using the texture axis value indicated by arrow A12 and the object label indicated by arrow A13 as control signals, respectively.
[0071] The object labels used as control signals for training the DNN for super-resolution processing are used to improve the accuracy of the super-resolution processing. By combining the object labels with the texture labels and calculating different coefficients for each combination of object label and texture label, the number of classification patterns can be increased, and the accuracy of inference can be improved.
[0072] Alternatively, only the texture label may be used as the control signal, in which case the object detection unit 16 may not be provided in the learning device 1.
[0073] FIG. 9 is a diagram showing an example of conversion of the texture axis value.
[0074] 9A and 9B show the conversion of the texture axis values of fine grain and fine detail, respectively. The horizontal axis represents the intensity of the texture, and the vertical axis represents the texture axis value. Information representing such a correspondence between the intensity and the texture axis value is provided to the texture axis value conversion unit 15 for each texture label of each texture expression.
[0075] As shown in Fig. 9, reference values for the texture axis value corresponding to the texture intensity of weak / medium / strong / OFF are set. In the example of A in Fig. 9, values V1, V2, V3, and 0 are set as reference values for the texture axis value corresponding to the respective intensities of weak / medium / strong / OFF.
[0076] When the texture label of a certain area is set to "fine graininess: strong", the texture value conversion unit 15 converts the intensity thereof into the texture axis value V3 based on the information in A of FIG. 9. Also, when the texture label of a certain area is set to "refinement: medium", the texture value conversion unit 15 converts the intensity thereof into the texture axis value V12 based on the information in B of FIG. 9.
[0077] By performing learning of the super-resolution processing DNN using such texture axis values as control signals, in the image processing apparatus 2, it becomes possible to control the texture of each area by the texture axis values. At the time of inference in the image processing apparatus 2, when the texture axis value is the intermediate value between two reference values, Volume control is performed to generate an image with the texture of the intermediate intensity.
[0078] Note that the intensity may not be included in the texture label, and only the types of texture expressions may be included. In this case, Volume control according to the ON (labeled) / OFF (Unlabeled) reference values is performed at the time of inference.
[0079] <<Inference Using DNN>> <Configuration of Image Processing Apparatus 2> [[ID=IS=16]]FIG. 10 is a block diagram showing a configuration example of the image processing apparatus 2.
[0080] The image processing apparatus 2 is composed of an object detection unit 31, an inference unit 32, a texture axis value conversion unit 33, an image quality adjustment unit 34, and an inference unit 35. A low-resolution image to be processed is input as an input image to the object detection unit 31, the inference unit 32, and the inference unit 35. The learning DB21 and the learning DB22 output from the learning apparatus 1 are input to the inference unit 32 and the inference unit 35, respectively.
[0081] The object detection unit 31 performs processing such as semantic segmentation on the input image and detects the objects shown in each area of the input image. The object detection unit 31 outputs an object label representing the object shown in each area to the image quality adjustment unit 34 and the inference unit 35.
[0082] The inference unit 32 inputs an input image to the texture segmentation detection DNN and infers texture labels that represent the texture of each region. The inference unit 32 outputs the texture labels of the inference results to the texture axis value conversion unit 33. The texture labels of the inference results also include the likelihood of each texture label.
[0083] The inference unit 32 functions as a texture detection unit that infers a texture label that represents the texture of each region. Since the inference unit 35 and other units perform processing to realize the texture represented by the texture label inferred by the inference unit 32 and generate an output image, the texture label inferred by the inference unit 32 represents the texture of each region realized in the output image.
[0084] The texture axis value conversion unit 33 converts the intensity of the texture of each region into a texture axis value based on the likelihood of the texture label supplied from the inference unit 32. The texture axis value conversion unit 33 outputs information on the texture axis value of each region to the image quality adjustment unit .
[0085] The image quality adjustment unit 34 adjusts the texture axis value of each region determined by the texture axis value conversion unit 33, based on the object label supplied from the object detection unit 31. By adjusting the texture axis value of each region, the image quality of the high-resolution image generated by the inference unit 35 is adjusted.
[0086] The image quality adjustment unit 34 outputs information about the texture axis value of each region after adjustment to the inference unit 35. The information about the texture axis value output from the image quality adjustment unit 34 is used as an inference control signal in the inference unit 35. The image quality adjustment unit 34 functions as a control signal generation unit that generates a control signal that represents the image quality of each region realized in the output image of the inference result.
[0087] The inference unit 35 inputs an input image into the super-resolution processing DNN and performs inference on a high-resolution image. The inference by the inference unit 35 is performed using the texture axis value supplied from the image quality adjustment unit 34 and the object label supplied from the object detection unit 31 as control signals. The inference is performed using coefficients prepared for each combination of the texture axis value of each region and the object appearing in each region.
[0088] The inference unit 35 outputs an image of the inference result as an output image. A configuration for performing processing using the high-resolution image generated by the inference unit 35 is provided downstream of the inference unit 35. In this way, the inference unit 35 functions as an image generation unit that inputs an input image to the DNN for super-resolution processing and performs inference on a high-resolution image in which the texture represented by the texture axis value is realized in each region.
[0089] The processing of each unit of the image processing device 2 will be described in detail below.
[0090] <Inference: DNN for texture segmentation detection> FIG. 11 is a diagram illustrating an example of inference using a DNN for texture segmentation detection.
[0091] As indicated by arrow A21, the inference unit 32 uses an input image, which is a low-resolution image, as input to the DNN for texture segmentation detection, and outputs a texture label as indicated by the tip of arrow A22.
[0092] 11, the texture label "weak detail" is set for the upper left region #91, which shows a distant landscape. The likelihood of the texture label for region #91 is 0.7.
[0093] Similarly, the texture label "Medium Detail" is assigned to the lower left region #92, which shows grass beside the gravel road, and the texture label "Strong Grain" is assigned to the lower center region #93, which shows the gravel road. The texture label "Medium Detail" is assigned to the lower right region #94, which shows grass beside the gravel road, and the texture label "Low Detail" is assigned to the upper right region #95, which shows a distant landscape. The likelihoods of the texture labels for regions #92 to #95, respectively, are 0.8, 0.9, 0.7, and 0.8.
[0094] In this way, the texture label of each region and the likelihood of the texture label represented by a value between 0.0 and 1.0 are output from the texture segmentation detection DNN.
[0095] Inference using the DNN for texture segmentation detection is performed so that the sum of the likelihoods of the texture labels assigned to each region is 1.0.
[0096] For example, the texture label for area #91 is "weak detail," and its likelihood is 0.7. However, texture labels with different intensities, "medium detail," "strong detail," and "off detail," are assigned to area #91, and the likelihood of each is calculated. The sum of the likelihood of the texture label "medium detail," the likelihood of the texture label "strong detail," and the likelihood of the texture label "off detail" is 0.3.
[0097] <Texture axis value conversion> FIG. 12 is a diagram showing an example of conversion of the texture axis value.
[0098] The texture intensity represented by the texture label of each region is converted into a texture axis value in the texture axis value conversion unit 33. Information representing the correspondence between the intensity and the texture axis value, as described with reference to Fig. 9, is provided to the texture axis value conversion unit 33.
[0099] When the texture labels in FIG. 11 are obtained by inference and supplied as indicated by arrow A31 in FIG. 12, the texture intensities of each region are converted into texture axis values as indicated by arrow A32. In the example of FIG. 12, the texture intensities of regions #91 to #95 are converted into texture axis values of 28, 96, 90, 84, and 32, respectively. Note that these texture axis value values are merely examples of conversion, and are obtained by multiplying reference values (reference value 40 for "weak detail," reference value 120 for "medium detail," and reference value 100 for "strong graininess") by likelihood. In practice, the texture axis values are obtained taking into account the reference values of other intensities.
[0100] FIG. 13 is a diagram showing an example of calculation of the texture axis value.
[0101] As shown in Fig. 13, the texture axis value is calculated based on the likelihood of each texture label of the same texture expression and the reference value of the texture axis value. The reference value of the texture axis value is found based on information representing the correspondence between the intensity and the texture axis value.
[0102] For example, the texture axis value for fine grain feeling is calculated by multiplying the reference value corresponding to "fine grain feeling: weak", the reference value corresponding to "fine grain feeling: medium", the reference value corresponding to "fine grain feeling: strong", and the reference value corresponding to "fine grain feeling: OFF" by their respective likelihoods and adding them up.
[0103] <Image quality adjustment> The texture axis value calculated by the texture axis value conversion unit 33 is adjusted by the image quality adjustment unit 34 according to the object label. The texture axis value after adjustment by the image quality adjustment unit 34 becomes a control signal during inference using the super-resolution processing DNN.
[0104] FIG. 14 is a diagram showing an example of adjustment of the texture axis value.
[0105] 14 represents a standard correspondence relationship used for converting the fine-grained texture axis value. The fine-grained texture axis value conversion unit 33 determines the fine-grained texture axis value based on the standard correspondence relationship.
[0106] The dashed line L2 represents the correspondence after adjustment. In the example of FIG. 14, adjustment is made so that the reference values corresponding to each strength of fine graininess are set to values higher than the standard correspondence. Such correspondence between texture strength and texture axis value is set for each object label. The correspondence shown by the dashed line L2 represents the correspondence for rock, stone, and sand.
[0107] The image quality adjustment unit 34 adjusts the fine-grained texture axis values of areas with object labels of rock, stone, and sand to values that correspond to the correspondence relationship of the dashed line L2. This results in an inference that further enhances the fine-grained texture of areas that contain rocks, stones, and sand.
[0108] By making it possible to adjust the texture axis value of each area, i.e., the intensity of the texture, according to the object depicted in each area, it becomes possible to create images for each object, such as changing the sense of detail of the forest or the sense of detail of the animal's fur.
[0109] In addition, by lowering the level of detail for distant trees and forests and increasing the level of detail for nearby trees and forests, it is possible to express textures such as perspective and depth. For example, such textures can be expressed by lowering the texture axis value of detail for areas labeled with distant trees and forests from the reference value used during training, and raising the texture axis value of detail for areas labeled with nearby trees and forests from the reference value used during training. Depth detection, etc., is used to determine the distance of an object.
[0110] When controlling textures such as fine grain and detail using conventional technology, the control is achieved by combining super-resolution processing, enhancement processing, contrast and color adjustment processing, etc., but the expressive power is low and it is not a process that directly controls texture.The above-mentioned processing makes it possible to directly control texture for each object and perform inference.
[0111] Furthermore, even in an area where the same object appears, the texture to be controlled varies from part to part. The above-described processing makes it possible to control the texture for each area of the object. Such texture control using object detection is performed, for example, when the output image of the image processing device 2 is used for display on a display device such as a TV.
[0112] <Inference: DNN for super-resolution processing> FIG. 15 is a diagram illustrating an example of inference using a DNN for super-resolution processing.
[0113] As indicated by arrow A41, an input image that is a low-resolution image is used by the inference unit 35 as input to the super-resolution processing DNN, and a high-resolution image as indicated by the tip of arrow A42 is output. The inference by the inference unit 35 is performed using the texture axis value indicated by arrow A51 and the object label indicated by arrow A52 as control signals.
[0114] <<Image Processing System Operation>> A series of operations of the learning device 1 and image processing device 2 having the above configuration will be described.
[0115] <Operation of learning device 1> The texture label setting process of the learning device 1 will be described with reference to the flowchart of FIG.
[0116] In step S1, the texture label definition unit 11 of the learning device 1 defines the type and intensity of the texture to be controlled in accordance with the image quality adjustment policy and the like.
[0117] In step S2, the object detection unit 16 performs semantic segmentation on the GT image to detect objects appearing in each region of the GT image.
[0118] In step S3, the texture labeling processing unit 12 assigns a texture label to each segmented region in accordance with the user's settings.
[0119] In step S4, the texture label assignment processing unit 12 evaluates / modifies the texture label as appropriate.
[0120] The above process is performed on various GT images, and the amount of texture labels required for DNN training is generated.
[0121] The process of generating a DNN for detecting texture segmentation performed by the learning device 1 will be described with reference to the flowchart in FIG.
[0122] In step S11, the degradation processor 13 performs degradation processing on the GT image.
[0123] In step S12, the DNN learning unit 14 performs learning using the texture labels as training data and the degraded images as student data. Learning by the DNN learning unit 14 is repeated until sufficient accuracy is ensured.
[0124] In step S13, the DNN learning unit 14 generates a DNN for detecting texture segmentation based on the learning result. Information on the coefficients of each layer constituting the DNN for detecting texture segmentation is output to the image processing device 2 as a learning DB 21.
[0125] The process of generating a DNN for super-resolution processing by the learning device 1 will be described with reference to the flowchart in FIG.
[0126] In step S21, the object detection unit 16 performs semantic segmentation on the GT image to detect objects appearing in each region of the GT image.
[0127] In step S22, the texture axis value conversion unit 15 converts the intensity of the texture of each region into a texture axis value based on the texture label.
[0128] In step S23, the DNN learning unit 17 performs learning using the GT image as a teacher image and the degraded image as a student image. Learning by the DNN learning unit 17 is repeated until sufficient accuracy is ensured.
[0129] In step S24, the DNN learning unit 17 generates a DNN for super-resolution processing that can be adjusted using the texture axis value and the object label as control signals based on the learning result. Information on the coefficients of each layer constituting the DNN for super-resolution processing is output to the image processing device 2 as a learning DB 22.
[0130] <Operation of image processing device 2> Next, the inference processing of the image processing device 2 will be described with reference to the flowchart of FIG.
[0131] In step S31, the object detection unit 31 of the image processing device 2 performs semantic segmentation on the input image to detect objects appearing in each region of the input image.
[0132] In step S32, the inference unit 32 inputs the input image to the DNN for texture segmentation detection, and performs inference of a texture label that represents the texture of each region.
[0133] In step S33, the texture axis value conversion unit 33 converts the intensity of the texture of each region into a texture axis value based on the likelihood of the texture label. The texture axis value is calculated based on the likelihood of each texture label as an inference result, as described with reference to FIG. 13 etc.
[0134] In step S34, the image quality adjustment unit 34 adjusts the texture axis value of each region according to the object label.
[0135] In step S35, the image quality adjustment unit 34 adjusts the balance of the overall image quality. The adjustment of the balance of the image quality is performed by appropriately adjusting the texture axis value. The adjustment of the texture axis value for adjusting the balance of the image quality will be described later.
[0136] In step S36, the inference unit 35 inputs the input image to the super-resolution processing DNN and performs inference on a high-resolution image that will become an output image. The inference by the inference unit 35 is performed using the texture axis value supplied from the image quality adjustment unit 34 and the object label supplied from the object detection unit 31 as control signals.
[0137] As described above, the image processing system can realize super-resolution processing that can directly control texture by performing DNN training and inference using DNN based on texture labels that represent human qualitative sensations.
[0138] Super-resolution processing performed in image processing systems is specialized processing that assigns the optimal texture to each region, making it a process with high image restoration and generation capabilities. General-purpose super-resolution processing without specialized processing is prone to falling into an average solution, but this can be prevented. In other words, the image processing system can generate images that express the appropriate texture in each region.
[0139] <<Label setting example>> 20 to 22 are diagrams showing examples of setting object labels and texture labels.
[0140] The images shown on the left side of Figures 20 to 22 are GT images to be labeled. Object detection is performed on the GT image, and objects appearing in each region are detected. For each region containing an object, an object label such as that shown in the center of Figures 20 to 22 is set by the object detection unit 16 during DNN training.
[0141] In the example of Figure 20, the object label "Sky" is set for area #101 of the GT image, which shows the sky, and the object label "Texture (green)" is set for the other areas, areas #102 to #105.
[0142] When such object labels are set, texture labels with different intensities may be set for areas where the same object appears (areas where the same object label is set), as shown in the speech bubble on the right side of Fig. 20. Also, texture labels with different types of texture expression may be set for areas where the same object appears.
[0143] In the example of Figure 20, for area #112 and area #116, which correspond to area #102, which has the same object label "Texture (green)", texture labels of different intensities, "Fineness: Weak" and "Fineness: Strong", are set, respectively.
[0144] In addition, for areas #115 and #116, which correspond to areas that have the same object label "Texture (green)" set, texture labels "Fine / Shape: Strong" and "Fine: Strong", which have different texture expression types, are set, respectively.
[0145] 21, the object label "Car" is set for area #121 of the GT image where a car is shown, and the object label "Sky" is set for area #122 where the sky is shown. Object labels are also set for the other areas, areas #123 to #126.
[0146] When such an object label is set, as shown in the balloon on the right side of FIG. 21, the area to which the object label is set may differ from the area to which the texture label is set.
[0147] In the example of Fig. 21, a texture label of "Glossiness / Transparency: Strong" is set for region #131, which is a portion of region #121 to which the object label "Car" is set. Also, a texture label of "Hardness / Softness: Weak" is set for region #132, which is a portion of region #122 to which the object label "Sky" is set.
[0148] In the example of FIG. 22, an object label of "Animal" is set for the area #141 in the GT image where a dog appears.
[0149] When such object labels are set, multiple types of texture labels may be set for one region, as shown in the balloon on the right side of Fig. 22. Only one type of object label is set for one region.
[0150] In the example of Figure 22, the texture label "Hardness / Softness (Soft): Strong" and the texture label "Fineness / Shape: Strong" are set for the same area as area #141, which has the object label "Animal" set.
[0151] By learning based on such texture labels, it is possible to generate a DNN for texture segmentation detection that can express various textures.
[0152] In addition, the texture labels of the inference results using the DNN for texture segmentation detection will also represent the texture of each region as described above.
[0153] <<Application Examples>> <Application example 1: Image quality adjustment for creators> Although the texture axis value calculated based on the texture label of the inference result of the texture segmentation detection DNN is used as the control signal for the super-resolution processing DNN, the user may be allowed to arbitrarily specify information corresponding to the texture axis value.
[0154] In this case, the user specifies an arbitrary texture for an arbitrary region of the input image, and as shown by arrow A51 in Figure 23, a signal representing the user's specification is used as a control signal for the DNN for super-resolution processing.
[0155] Some users prefer to specify the texture of each area themselves. A function that allows users to freely specify information equivalent to the texture axis value is a function aimed at creators and other users. This allows for a high degree of freedom in image quality adjustment.
[0156] Such image quality adjustment performed in accordance with user operation is performed, for example, as the adjustment of image quality balance in step S35 in Fig. 19. A control signal representing the content after the balance adjustment is used as a control signal for the super-resolution processing DNN.
[0157] The texture labels of the inference results of the texture segmentation detection DNN may be presented as a guide to the user who specifies the texture of each region.
[0158] <Application example 2: Labeling specific to the use case of the output destination> Image quality labels that express image quality different from texture may be used for training the DNN. In this case, instead of a DNN for detecting texture segmentation, a DNN that links image quality labels with regions that are subject to image quality control is generated in the learning device 1.
[0159] For example, an image quality label is set according to the use case at the output destination of the output image of the inference result by the inference unit 35.
[0160] FIG. 24 is a diagram showing an example of an image quality label.
[0161] When the output image of the inference result is used in a game, labels representing areas where people appear and areas where text appears are set as image quality labels.
[0162] When the output image of the inference result is used for electronic zooming for a camera, labels representing the face area, light source area, and reflection area are set as image quality labels.
[0163] To improve the robustness of applications (use cases at the output destination), if the output image is used for Frame Rate Control (FRC), labels representing areas where repetitive patterns appear and areas where captions appear are set as image quality labels. Also, if the output image is used for super-resolution processing, labels representing areas where regularity appears and areas where continuity appears are set as image quality labels.
[0164] Creators may be allowed to set any label relating to image quality as an image quality label.
[0165] In this way, by changing the label, it is possible to create a desired image. Except for the difference in the label, the processing in the image processing system is the same as the processing described above.
[0166] <Application example 3: Use of labels for image creation> By incorporating the user's intentions into the texture labels, the user can train a DNN that can make inferences that take the user's intentions into account. The texture labels that incorporate the user's intentions are set before the DNN is trained.
[0167] FIG. 25 shows examples of texture labels that incorporate image creation intentions.
[0168] The texture labels of areas #151 to #155 shown on the left side of FIG. 25 are normal texture labels set by evaluating the texture according to the actual appearance.
[0169] On the other hand, the texture labels of areas #151 to #155 shown on the right side of Fig. 25 are texture labels that incorporate image creation intentions. The texture labels that incorporate image creation intentions include labels with intensities that differ from normal texture labels.
[0170] FIG. 26 shows an image of the image quality of the inference result using texture labels that incorporate the intention of image creation.
[0171] As shown by the white arrow on the left side of Figure 26, when using a DNN generated based on normal texture labels, the image quality of the output image required as the final output will be targeted at the image quality of the GT image.
[0172] By using a DNN generated based on texture labels that incorporate the intention behind image creation, it is possible to express the image quality of the output image differently from that of a GT image, as shown by the white arrow on the right side of Figure 26.
[0173] <Application Example 4: Image processing other than super-resolution processing> A DNN for image processing other than super-resolution processing, such as contrast / color adjustment processing, SDR-HDR conversion processing, and enhancement processing, may be used in the image processing device 2 instead of the DNN for super-resolution processing.
[0174] Image processing such as contrast and color adjustment and SDR-HDR conversion is a good match for processing to express textures such as gloss, transparency, luster, brightness, and shadows. When enhancement processing is performed, labels can be assigned to objects or areas where emphasis is placed on enhancement adjustments, rather than texture labels.
[0175] The DNN is trained using images that are different from those used to train the DNN for super-resolution processing.
[0176] For example, the learning of the DNN for contrast and color adjustment processing is performed using a GT image as a teacher image and a degraded image obtained by weakening the contrast and lowering the saturation of the GT image as a student image. The image processing performed by the degradation processing unit 13 is processing to weaken the contrast and lower the saturation.
[0177] The DNN for SDR-HDR conversion is trained using HDR images as teacher images and SDR images obtained by tone mapping the HDR images as degradation processing as student images. The image processing performed by the degradation processing unit 13 converts HDR images into SDR images.
[0178] The DNN for the enhancement process is trained using a GT image as a teacher image and a degraded image obtained by removing the high-frequency components of the GT image as a student image. The image processing performed by the degradation processing unit 13 is a process of removing the high-frequency components of the GT image.
[0179] Instead of a DNN for a single process, a DNN for image processing that combines multiple processes, such as super-resolution processing and contrast / color adjustment processing, or SDR-HDR conversion processing and enhancement processing, may be trained and used for inference.
[0180] <Application example 5: Using a DNN for texture segmentation detection as a texture evaluation model> A GT image may be input to a DNN for texture segmentation detection, and texture labels for each region of the GT image may be inferred.
[0181] The texture labels of the inference results are presented to the user and used to evaluate the texture of each region. For example, the user can perform inference based on both the GT image before and after image editing, and check how the texture changes as a result of image editing.
[0182] In this example, the DNN for texture segmentation detection is used as the DNN for texture evaluation. The DNN for texture evaluation is trained using texture labels as training data and GT images as student data.
[0183] <Application Example 6: Semi-supervised learning> The training of the DNN for detecting texture segmentation using a GT image as an input image may be performed by semi-supervised learning. In this case, the texture labels of the inference results obtained by inputting the GT image into the DNN for detecting texture segmentation are used as training data.
[0184] This learning method is effective when there are few texture labels to serve as training data. Rather than using the inference results as training data as is, the accuracy of the inference can be improved by manually evaluating the texture label results and correcting them if necessary.
[0185] <<Other examples>> <Example of computer configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware or a general-purpose personal computer.
[0186] FIG. 27 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program.
[0187] A CPU (Central Processing Unit) 1001 , a ROM (Read Only Memory) 1002 , and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004 .
[0188] An input / output interface 1005 is also connected to the bus 1004. An input unit 1006 including a keyboard, a mouse, etc., and an output unit 1007 including a display, a speaker, etc. are connected to the input / output interface 1005. Also connected to the input / output interface 1005 are a storage unit 1008 including a hard disk, a nonvolatile memory, etc., a communication unit 1009 including a network interface, etc., and a drive 1010 that drives removable media 1011.
[0189] In a computer configured as above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the storage unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing the program.
[0190] The programs executed by the CPU 1001 are installed in the storage unit 1008 by being recorded on, for example, a removable medium 1011 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.
[0191] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0192] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0193] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0194] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0195] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.
[0196] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.
[0197] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0198] <Configuration combination example> The present technology can also be configured as follows.
[0199] (1) a control signal generation unit that generates a control signal representing the texture of each region realized in an output image of the inference result based on an input image to be processed; an image generation unit that inputs the input image into an inference model obtained by performing learning based on a student image generated by performing predetermined image processing on a teacher image and the teacher image in which the texture of each region is expressed by a texture label, and infers the output image in which each region has the texture expressed by the control signal; An image processing device comprising: (2) a texture detection unit that inputs the input image into another inference model obtained by performing learning using, as student data, an image generated by performing the predetermined image processing on a learning image and using, as teacher data, texture labels that represent the texture of each region of the learning image, and infers texture labels that represent the texture of each region realized in the output image; The control signal generation unit generates the control signal based on a texture label of an inference result. The image processing device according to (1) above. (3) Multiple types of texture labels are defined to represent qualitative texture and texture intensity. The image processing device according to (2) above. (4) a conversion unit that converts the intensity of the texture represented by the texture label of the inference result inferred using the other inference model into a numerical value based on likelihood, The control signal generation unit generates the control signal representing the type of texture represented by the texture label of the inference result and the numerical value. The image processing device according to (3) above. (5) The control signal generation unit adjusts the relationship between the intensity of the texture and the numerical value in accordance with the object included in each region. The image processing device according to (4) above. (6) The control signal generating unit generates the control signal according to the texture of each area designated by a user. The image processing device according to (1) above. (7) further comprising an object detection unit that detects an object included in the input image; The learning of the inference model is performed by learning a different coefficient for each object included in the training image, The image generation unit inputs the input image to the inference model in which coefficients corresponding to objects included in the input image are set, and performs inference on the output image. The image processing device according to any one of (1) to (6). (8) The texture of each region is expressed using the texture of the object contained in that region. The image processing device according to any one of (1) to (7). (9) The image processing device generating a control signal representing the texture of each region realized in an output image of the inference result based on the input image to be processed; The input image is input to an inference model obtained by learning based on a student image generated by performing predetermined image processing on a teacher image and the teacher image in which the texture of each region is expressed by a texture label, and an inference is made on the output image in which each region has the texture expressed by the control signal. Image processing methods. (10) On the computer, generating a control signal representing the texture of each region realized in an output image of the inference result based on the input image to be processed; The input image is input to an inference model obtained by learning based on a student image generated by performing predetermined image processing on a teacher image and the teacher image in which the texture of each region is expressed by a texture label, and an inference is made on the output image in which each region has the texture expressed by the control signal. A program for executing a process. (11) an acquisition unit that acquires texture labels that represent the texture of each region of a learning image; a learning unit that performs learning using an image generated by performing predetermined image processing on the learning image as a student image, the learning image as a teacher image, in accordance with a control signal that represents the texture of each region of the learning image, and generates an inference model; A learning device comprising: (12) The image processing apparatus further includes another learning unit that performs learning using an image generated by performing the predetermined image processing on the learning image as student data and a texture label that indicates the texture of each region of the learning image as teacher data, and generates another inference model. The learning device according to (11) above. (13) Multiple types of texture labels are defined to represent qualitative texture and texture intensity. The learning device according to (12) above. (14) a conversion unit that converts the intensity of a texture represented by a texture label that represents the texture of each region of the learning image into a numerical value; The learning unit learns the inference model in accordance with the type of texture represented by a texture label representing the texture of each region of the learning image and the control signal representing the numerical value. The learning device according to (13) above. (15) further comprising an object detection unit that detects an object included in the learning image, The learning unit learns the inference model by calculating a different coefficient for each object included in the learning image. The learning device according to any one of (11) to (14). (16) an image processing unit that performs degradation processing as the predetermined image processing on the learning image; The learning device according to any one of (11) to (15). (17) The acquisition unit acquires texture labels that represent textures of the respective regions of the learning image set in response to an operation by a user. The learning device according to any one of (11) to (16). (18) The learning device Obtain texture labels that represent the texture of each region of the training image, An image generated by performing predetermined image processing on the learning image is used as a student image, and learning is performed using the learning image as a teacher image in accordance with a control signal representing the texture of each region of the learning image, thereby generating an inference model. Generation method. (19) On the computer, Obtain texture labels that represent the texture of each region of the training image, An image generated by performing predetermined image processing on the learning image is used as a student image, and learning is performed using the learning image as a teacher image in accordance with a control signal representing the texture of each region of the learning image, thereby generating an inference model. A program for executing a process. [Explanation of symbols]
[0200] 1 Learning device, 2 Image processing device, 11 Texture label definition unit, 12 Texture label assignment processing unit, 13 Degradation processing unit, 14 DNN learning unit, 15 Texture axis value conversion unit, 16 Object detection unit, 17 DNN learning unit, 31 Object detection unit, 32 Inference unit, 33 Texture axis value conversion unit, 34 Image quality adjustment unit, 35 Inference unit
Claims
1. a control signal generation unit that generates a control signal representing the intensity of the texture of each region realized in an output image of the inference result, based on an input image to be processed; an image generation unit that inputs the input image into an inference model obtained by performing learning based on a degraded image generated by degrading a teacher image and the teacher image in which the texture of each region is expressed by a texture label representing the texture, and that infers that the output image has the texture expressed by the control signal in each region; An image processing device comprising:
2. The image processing apparatus further includes a texture detection unit that inputs the input image to another inference model obtained by performing learning based on an image generated by degrading a training image and texture labels that represent the texture of each region of the training image, and infers texture labels that represent the texture of each region realized in the output image, The control signal generation unit generates the control signal based on a texture label of an inference result. The image processing device according to claim 1 .
3. Multiple types of texture labels are defined that represent the type and intensity of texture. The image processing device according to claim 2 .
4. a conversion unit that converts the intensity of the texture represented by the texture label of the inference result inferred using the other inference model into a numerical value based on likelihood, The control signal generation unit generates the control signal representing the type of texture represented by the texture label of the inference result and the numerical value. The image processing device according to claim 3 .
5. The control signal generation unit adjusts the relationship between the intensity of the texture and the numerical value in accordance with the object included in each region. The image processing device according to claim 4 .
6. The control signal generating unit generates the control signal according to the texture of each area designated by a user. The image processing device according to claim 1 .
7. further comprising an object detection unit that detects an object included in the input image; The learning of the inference model is performed by learning a different coefficient for each object included in the training image, The image generation unit inputs the input image to the inference model in which coefficients corresponding to objects included in the input image are set, and performs inference on the output image. The image processing device according to claim 1 .
8. The texture of each region is expressed using the texture of the object contained in that region. The image processing device according to claim 1 .
9. The image processing device generating a control signal representing the intensity of the texture of each region realized in an output image of the inference result based on the input image to be processed; The input image is input to an inference model obtained by performing learning based on a degraded image generated by degrading a teacher image and the teacher image in which the texture of each region is expressed by a texture label representing the texture, and an inference is performed on the output image in which each region has the texture expressed by the control signal. Image processing methods.
10. On the computer, generating a control signal representing the intensity of the texture of each region realized in an output image of the inference result based on the input image to be processed; The input image is input to an inference model obtained by performing learning based on a degraded image generated by degrading a teacher image and the teacher image in which the texture of each region is expressed by a texture label representing the texture, and an inference is performed on the output image in which each region has the texture expressed by the control signal. A program for executing a process.
11. an acquisition unit that acquires texture labels that represent the texture of each region of a learning image; a learning unit that performs learning based on a degraded image generated by degrading the learning image and the learning image as a teacher image in accordance with a control signal that indicates the intensity of texture in each region of the learning image, and generates an inference model; A learning device comprising:
12. The image processing device further includes a second learning unit that performs learning based on an image generated by degrading the image for learning and a texture label that represents the texture of each region of the image for learning, and generates another inference model. The learning device according to claim 11 .
13. A plurality of types of texture labels are defined that represent the type and intensity of the texture. The learning device according to claim 12.
14. a conversion unit that converts the intensity of a texture represented by a texture label that represents the texture of each region of the learning image into a numerical value; The learning unit learns the inference model in accordance with the type of texture represented by a texture label representing the texture of each region of the learning image and the control signal representing the numerical value. The learning device according to claim 13.
15. further comprising an object detection unit that detects an object included in the learning image, The learning unit learns the inference model by calculating a different coefficient for each object included in the learning image. The learning device according to claim 11 .
16. The image processing device further includes an image processing unit that performs degradation processing on the learning images. The learning device according to claim 11 .
17. The acquisition unit acquires texture labels that represent textures of the respective regions of the learning image set in response to an operation by a user. The learning device according to claim 11 .
18. The learning device Obtain texture labels that represent the texture of each region of the training image, A learning process is performed based on a degraded image generated by degrading the learning image and the learning image as a teacher image, in accordance with a control signal representing the intensity of texture in each region of the learning image, thereby generating an inference model. Generation method.
19. On the computer, Obtain texture labels that represent the texture of each region of the training image, A learning process is performed based on a degraded image generated by degrading the learning image and the learning image as a teacher image, in accordance with a control signal representing the intensity of texture in each region of the learning image, thereby generating an inference model. A program for executing a process.
Citation Information
Patent Citations
Image processing apparatus and program
JP2011171807A
Information processing system, information processing method, and program
JP2018190371A