Image processing apparatus, estimation apparatus, image processing method, and image processing program
By evaluating and selecting high-quality images for processing, the image processing apparatus improves the accuracy of attribute estimation in image processing, addressing the issue of low image quality affecting conventional methods.
Patent Information
- Application Number
- JP2021120343
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-21
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2041-07-21
Smart Images

Figure 0007691877000001 
Figure 0007691877000002 
Figure 0007691877000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an estimation apparatus, an image processing method, and an image processing program.
Background Art
[0002] Conventionally, a technique for estimating attributes such as the age and gender of a person from a face image of the person has been known. For example, Patent Document 1 discloses a technique for estimating the age of a person in a face image by integrating the determination results of each determination machine, which is roughly classified into above or below a predetermined age.
[0003] Further, Patent Document 2 discloses a technique for estimating the attributes of a person in the order in which captured images captured by a camera are stored in a memory and counting the number of persons for each attribute.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] Here, as a method for estimating the attributes of a person from an image of the person, generally, a convolutional neural network (CNN), which is one of the deep learning methods, is used. In the convolutional neural network, an input image is read, convolution and pooling are repeated in the first half to extract important and multiple features in the input image, and in the second half, discrimination (classification) is performed by a fully connected layer and an output layer based on those features.
[0006] In the conventional estimation method, since a learned model learned by machine learning or the like is used to estimate the attribute of the estimation target based on the input image, the estimation result is easily affected by the quality (image quality) of the input image. For example, when the image quality of the input image is low, there arises a problem that the estimation accuracy of the attribute of the estimation target decreases.
[0007] An object of the present invention is to provide an image processing apparatus, an estimation apparatus, an image processing method, and an image processing program capable of improving the estimation accuracy of the attribute of the estimation target.
Means for Solving the Problem
[0008] An image processing apparatus according to one aspect of the present invention is an image processing apparatus that inputs the input image to an estimation apparatus that estimates an attribute of an estimation target based on the input image, and includes a first acquisition processing unit that acquires a captured image of the estimation target, an evaluation processing unit that evaluates the image quality of the captured image acquired by the first acquisition processing unit, and an output processing unit that outputs the captured image as the input image when the evaluation processing unit determines that the image quality of the captured image has a predetermined image quality.
[0009] An estimation apparatus according to another aspect of the present invention is an estimation apparatus that estimates an attribute of an estimation target using the captured image output from the image processing apparatus as the input image, and includes a second acquisition processing unit that acquires the captured image from the image processing apparatus as the input image, and an estimation processing unit that calculates an output value for each of a plurality of classifications of the attribute using a predetermined learned model for the input image acquired by the second acquisition processing unit, calculates a determination value of the attribute based on the calculated output values of the respective classifications, and estimates the attribute of the estimation target based on the calculated determination value.
[0010] An image processing method according to another aspect of the present invention is an image processing method for inputting the input image to an estimation device that estimates an attribute of an estimation target based on the input image, wherein one or more processors perform an acquisition step of acquiring a captured image of the estimation target, an evaluation step of evaluating the image quality of the captured image acquired in the acquisition step, and an output step of outputting the captured image as the input image when it is determined in the evaluation step that the image quality of the captured image has a predetermined image quality.
[0011] An image processing program according to another aspect of the present invention is an image processing program for inputting the input image to an estimation device that estimates an attribute of an estimation target based on the input image, and is a program for causing one or more processors to execute an acquisition step of acquiring a captured image of the estimation target, an evaluation step of evaluating the image quality of the captured image acquired in the acquisition step, and an output step of outputting the captured image as the input image when it is determined in the evaluation step that the image quality of the captured image has a predetermined image quality.
Effects of the Invention
[0012] According to the present invention, it is possible to provide an image processing device, an estimation device, an image processing method, and an image processing program capable of improving the estimation accuracy of an attribute of an estimation target.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Note that the following embodiments are merely examples of embodying the present invention and do not have the character of limiting the technical scope of the present invention.
[0015] The estimation system 100 according to this embodiment is a system that estimates the attributes of an estimation target based on an image of the estimation target. For example, the estimation target may be a person's face, a vehicle, an animal, an item, etc. When the estimation target is a person, the attribute may be, for example, the age of the person. Also, the attribute may be the gender of the person, the presence or absence of glasses, the presence or absence of a mask, the ethnic group, etc. When the estimation target is a vehicle, the attribute may be the vehicle number, the manufacturer, domestic / foreign vehicle, color, etc. The estimation system 100 may estimate one attribute or may estimate a plurality of attributes. By being applied to, for example, digital signage, the estimation system 100 can estimate the ages of passersby on the street, in a store, etc., and can perform crowd flow analysis, marketing, etc. Also, by being applied to, for example, the POS terminal of a store, the estimation system 100 can estimate the age of a customer and can perform marketing, etc.
[0016] In this embodiment, as an example of the estimation target, a person's face will be given, and as an example of the attribute, the age of the person will be given and explained. That is, a person's face is an example of the estimation target of the present invention, and a person's age is an example of the attribute of the present invention.
[0017] [Estimation System 100] FIG. 1 is a diagram showing a schematic configuration of an estimation system 100 according to an embodiment of the present invention. As shown in FIG. 1, the estimation system 100 includes an image processing device 1 and an estimation device 2. The image processing device 1 and the estimation device 2 can communicate via a communication network N1 such as a LAN, the Internet, a WAN, or a public telephone line. Note that the image processing device 1 and the estimation device 2 may be configured as an integrated device (information processing device).
[0018] The estimation device 2 estimates the attributes of the estimation target based on the input image. For example, the estimation device 2 estimates the age of a person based on an image of the person's face. The image processing device 1 inputs the input image to the estimation device 2. For example, the image processing device 1 inputs an image of a person's face captured by the camera 15 as the input image to the estimation device 2.
[0019] [Image Processing Apparatus 1] As shown in FIG. 1, the image processing apparatus 1 includes a control unit 11, a storage unit 12, an operation display unit 13, a communication unit 14, a camera 15, and the like. The image processing apparatus 1 may be an information processing apparatus such as a personal computer, for example.
[0020] The camera 15 is a digital camera that captures an image of a subject and outputs it as digital image data. For example, the camera 15 captures a predetermined area including a person's face at a predetermined frame rate and transmits the captured image data to the control unit 11. Note that the camera 15 may be installed outside the image processing apparatus 1. In this case, the camera 15 transmits the image data to the image processing apparatus 1 via the communication network N1.
[0021] The communication unit 14 is a communication interface for connecting the image processing apparatus 1 to the communication network N1 by wire or wirelessly and performing data communication according to a predetermined communication protocol with other devices (such as the estimation apparatus 2) via the communication network N1.
[0022] The operation display unit 13 is a user interface including a display unit such as a liquid crystal display or an organic EL display that displays various types of information, and an operation unit such as a mouse, a keyboard, or a touch panel that accepts operations.
[0023] The storage unit 12 includes a non-volatile storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive) that stores various types of information. The storage unit 12 stores control programs such as an image processing program 121 for causing the control unit 11 to execute image processing (see FIG. 10) described later. Control programs such as the image processing program 121 are recorded and provided, for example, on a computer-readable non-transitory recording medium, read from the non-transitory recording medium by the reading device of the image processing apparatus 1, and stored in the storage unit 12. The control program may be provided (downloaded) to the image processing apparatus 1 from an external server or the like other than the image processing apparatus 1 via a network and stored in the storage unit 12. Further, the storage unit 12 stores a captured image corresponding to the image data transmitted from the camera 15.
[0024] The control unit 11 includes control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit in which control programs such as a BIOS and an OS for causing the CPU to execute various processes are pre-stored. The RAM is a volatile or non-volatile storage unit that stores various types of information and is used as a temporary storage memory (working area) for various processes executed by the CPU. Then, the control unit 11 controls the image processing apparatus 1 by causing the CPU to execute various control programs pre-stored in the ROM or the storage unit 12.
[0025] Here, as a method for estimating the attributes of a person from an image of the person, a convolutional neural network (CNN), which is one of the deep learning methods, is generally used. In the convolutional neural network, an input image is read, convolution and pooling are repeated in the first half to extract important and multiple features in the input image, and in the second half, discrimination (classification) is performed by a fully connected layer and an output layer based on those features.
[0026] Figure 2 shows the basic structure of a convolutional neural network. The convolutional neural network reads in training data, repeatedly performs convolution and pooling in the first half to extract important and multiple features in the training data, and in the second half, based on those features, performs discrimination using a fully connected layer and an output layer.
[0027] Specifically, in the input layer of the first half processing, the input image, which is the training data, is converted from a multi-dimensional (3D data having the two dimensions of the image and one dimension of the color components RGB) to a one-dimensional vector. In the next convolutional layer, features such as edge extraction are extracted by detecting the shading pattern of the input image. In the next pooling layer, by thinning out only those important feature amounts that have been extracted, the position of the object is allowed to shift while still considering it to be the same object, and the image size (information) is compressed. By connecting multiple stages of convolutional layers and pooling layers, more important features are extracted. Here, as the number of stages of the convolutional layer and the pooling layer increases, the magnitude (gradient) of the feature amount disappears. Therefore, an activation function is inserted after the convolutional layer to emphasize the feature amount and suppress the disappearance of the gradient. In the second half processing, classification is performed based on the feature amounts extracted in the first half processing.
[0028] Figure 3 schematically shows the connection state between the fully connected layer and the output layer. Each circle is an individual node, and for each connection line between the nodes, a connection weight coefficient is multiplied. The number of connections is calculated by the number of nodes (m) in the fully connected layer × the number of outputs (n) in the output layer. The more the number of connections, the better the classification ability, but at the same time, the amount of calculation increases and the performance deteriorates. Also, there is a problem that it takes time to calculate the optimal connection weight coefficient during learning. In the fully connected layer (see Figure 2), the features extracted in the first half are aggregated into each one node, and the connection weight coefficients between the nodes are adjusted to convert them into feature variables for classification. Then, in the last output layer (see Figure 2), the feature variables are output as score values that maximize the probability of being correctly classified. Note that when the output values (score values) of the output layer are added together for the number of outputs, it becomes 1.0 (100%).
[0029] When estimating the attributes (such as age) of a person in a face image, it is common to use an estimation method based on the convolutional neural network. Using learning data (multiple face images) as input images, various feature amounts in the learning data are extracted, and the configuration of the first half process is changed or various parameters (connection weight coefficients of the fully connected layer) are adjusted (learned) so that the output value of the output layer approaches the correct value (age, the range (class) to which the age corresponds) of the learning data of the input image in a comprehensive average manner. And what combines the configuration and estimation parameters optimized by machine learning becomes a "trained model" (see Figure 2).
[0030] For the trained model generated by machine learning by such a learning method, by inputting the face image of the person to be estimated, it becomes possible to estimate the age of the person from the output value (score value).
[0031] Next, an example of an estimation method for estimating attributes from a person's image is given. Figure 4 schematically shows an age estimation method.
[0032] The estimated age is calculated by the sum-of-products operation of each age class ID (node number of the output layer) and the score value corresponding to each age class ID. Note that the age class represents a range obtained by dividing the age at specific ages. In this case, the result of the sum-of-products operation is 3.70, and when the decimal point is truncated, the age class ID becomes "3", so the corresponding age class is estimated to be "13 to 18 years old". Furthermore, by taking into account the decimal point "0.70" and performing linear interpolation within the same age class, it is also possible to estimate the age as 13 + 0.70×(18 - 13) = 16.5 years old.
[0033] Note that since these score values vary, it is desirable to estimate the age based on the average value of multiple input images (face images).
[0034] Here, an example of the training data used in machine learning for generating a learned model is shown in FIG. 5. Each of the face images 1 to 10 represents a face image of a person in each age range. For each training data folder, for example, "serial number # age lower limit - age upper limit" is given as the folder name. A plurality of face image data corresponding to that age range are stored in each training data folder, and these are collectively used as training data (training data set). The "number of age classes" is the number of divisions of the classified ages. In this example, the number of age classes is "10 classes".
[0035] Note that the training data itself is not stored in the learned model. Instead, the learned model is obtained by extracting their features using a convolutional neural network and finding the optimal configuration and parameters such that an estimated value matching the training data is output.
[0036] By the way, in the conventional estimation method, in order to estimate the attribute of the object to be estimated based on the input image using a learned model learned by machine learning or the like, the estimation result is easily affected by the quality (image quality) of the input image. For example, when the image quality of the input image is low, there arises a problem that the estimation accuracy of the attribute of the object to be estimated decreases. Hereinafter, a specific example of this problem will be described.
[0037] FIG. 6 shows the transition of the score values of the edge, contrast, and nose-to-eye aspect ratio of the extracted face images and the total score value of these indicators in the image data (moving image data) captured in a predetermined period (about 10 seconds). The face images P1 to P4 show the face images of the same person with different orientations, image qualities, etc. As shown in FIG. 6, the face image P1 indicates that the total score value is high and the image quality (image quality) is high. In contrast, the face image P2 is, for example, an image with the face facing down, and the score value of the contrast is low. Also, the face image P3 is, for example, an image with the face facing sideways, and the score value of the nose-to-eye aspect ratio is low. Also, the face image P4 is an overall blurred image, and the score value of the edge is low. That is, the face images P2 to P4 indicate low image quality.
[0038] Here, when estimating the age of a person using the face image P1 as the input image, it is possible to estimate the accurate age. On the other hand, when estimating the age of a person using the face images P2 to P4 as the input images, it becomes impossible to estimate the accurate age. Thus, when the image quality of the input image is low, there arises a problem that the estimation accuracy of the attributes to be estimated decreases. In contrast, according to the image processing apparatus 1 according to the present embodiment, as described below, it is possible to improve the estimation accuracy of the attributes to be estimated.
[0039] Specifically, as shown in FIG. 1, the control unit 11 of the image processing apparatus 1 includes various processing units such as an acquisition processing unit 111, an evaluation processing unit 112, and an output processing unit 113. Note that the control unit 11 functions as the various processing units by executing various processes according to the image processing program 121 by the CPU. Also, some or all of the processing units included in the control unit 11 may be configured by electronic circuits. Note that the image processing program 121 may be a program for causing a plurality of processors to function as the various processing units.
[0040] The acquisition processing unit 111 acquires a captured image of the object to be estimated from the camera 15. Specifically, the acquisition processing unit 111 sequentially acquires captured images of a person captured by the camera 15 at a predetermined frame rate. Also, the acquisition processing unit 111 acquires a plurality (predetermined number) of captured images of the person who is the object to be estimated. The acquisition processing unit 111 stores the acquired captured images in the storage unit 12. The acquisition processing unit 111 is an example of the first acquisition processing unit of the present invention.
[0041] The evaluation processing unit 112 evaluates the image quality of the captured image acquired by the acquisition processing unit 111. Specifically, the evaluation processing unit 112 extracts a face image from the captured image, calculates feature amounts related to the extracted face image, and calculates an evaluation value of the face image based on the calculated feature amounts. For example, the evaluation processing unit 112 calculates an evaluation value of the face image based on the respective feature amounts of edges, contrast, and image size in the extracted face image.
[0042] In addition, the evaluation processing unit 112 calculates an evaluation value for each of a plurality of captured images (face images) related to the person to be estimated. The evaluation processing unit 112 is an example of the evaluation processing unit of the present invention.
[0043] Hereinafter, a specific example will be described. FIG. 7 shows a face image extracted from a captured image of the person to be estimated. Note that a well-known technique such as MTCNN (multi-task cascaded convolutional neural networks) can be applied to the technique of detecting a face from a captured image and cutting out a face image. For example, the evaluation processing unit 112 detects the position of the face in the captured image and the positions of main organs such as eyes, nose, and mouth, and cuts out an image within a predetermined range to generate a face image. In addition, the evaluation processing unit 112 calculates the size in the original image and the feature amount of the periphery of the eyes (the rectangular frame area in FIG. 7) for the generated face image.
[0044] As the feature amount of the periphery of the eyes, it is conceivable to apply a noise filter, a high-pass filter, and a low-pass filter and use the difference between their maximum values and minimum values. For example, as shown in FIG. 8, the evaluation processing unit 112 performs filter processing on the face image using a noise filter Fd, a high-pass filter Fh, and a low-pass filter Fl. The evaluation processing unit 112 obtains the difference (edge evaluation value dd) between the maximum value and the minimum value of the result of the filter processing by the high-pass filter Fh, and obtains the difference (contrast evaluation value dr) between the maximum value and the minimum value of the result of the filter processing by the low-pass filter Fl. When the image size is "dw" and the adjustment value is "mag", the evaluation value F of the face image is obtained by the following formula. F=(dd×dr×dw×mag / 128.0);
[0045] When the acquisition processing unit 111 acquires a plurality of captured images, the evaluation processing unit 112 generates a plurality of face images and calculates the evaluation value F for each face image. For example, when the acquisition processing unit 111 acquires five captured images, the evaluation processing unit 112 generates five face images and calculates five evaluation values F for each face image.
[0046] When the evaluation processing unit 112 determines that the image quality of the captured image has a predetermined image quality, the output processing unit 113 outputs the captured image to the estimation device 2 as an input image. Specifically, the output processing unit 113 outputs the face image as an input image based on the evaluation value F of the face image calculated by the evaluation processing unit 112.
[0047] For example, the output processing unit 113 outputs the top predetermined number of face images with high evaluation value F among the plurality of face images as input images. For example, when the evaluation processing unit 112 generates 5 face images of the person to be estimated and calculates the evaluation value F for each of the 5 face images, the output processing unit 113 extracts the top 3 face images with high evaluation value F among the 5 face images and outputs the extracted 5 face images as input images. The output processing unit 113 is an example of the output processing unit of the present invention.
[0048] As another embodiment, the evaluation processing unit 112 may determine whether the evaluation value F of the face image exceeds a predetermined evaluation value (threshold). Further, the evaluation processing unit 112 may determine whether each evaluation value F corresponding to a plurality of face images exceeds a predetermined evaluation value. In this case, when the evaluation processing unit 112 determines that the evaluation value F of the face image exceeds the predetermined evaluation value, the output processing unit 113 outputs the face image as an input image. Further, the output processing unit 113 outputs one or more face images among the plurality of face images whose evaluation value F exceeds the predetermined evaluation value as input images.
[0049] As still another embodiment, until the number of face images whose evaluation value F of the face image exceeds the predetermined evaluation value reaches a predetermined number (for example, 3), the camera 15 may execute a process of capturing a person, or the acquisition processing unit 111 may execute a process of acquiring a captured image.
[0050] According to the above configuration, for the face image of the person to be estimated, a face image with good image quality can be extracted as the input image of the estimation device 2. For example, one or more face images having a predetermined image quality (high image quality) among a plurality of face images of the person to be estimated can be extracted as the input image of the estimation device 2.
[0051] [Estimation Device 2] As shown in FIG. 1, the estimation device 2 includes a control unit 21, a storage unit 22, an operation display unit 23, a communication unit 24, etc. The image processing device 1 may be an information processing device such as a personal computer, for example.
[0052] The communication unit 24 is a communication interface for connecting the estimation device 2 to the communication network N1 by wire or wirelessly and performing data communication according to a predetermined communication protocol with other devices (such as the image processing device 1) via the communication network N1.
[0053] The operation display unit 23 is a user interface including a display unit such as a liquid crystal display or an organic EL display for displaying various information, and an operation unit such as a mouse, a keyboard, or a touch panel for receiving operations.
[0054] The storage unit 22 includes a non-volatile storage device such as an HDD or an SSD for storing various information. The storage unit 22 stores (stores) control programs such as an estimation program 221 for causing the control unit 21 to execute the estimation process (see FIG. 11) described later. Control programs such as the estimation program 221 are recorded and provided, for example, on a computer-readable non-transitory recording medium, read from the non-transitory recording medium by a reading device of the estimation device 2, and stored in the storage unit 22. The control program may be provided (downloaded) to the estimation device 2 from an external server or the like other than the estimation device 2 via a network and stored in the storage unit 22. Further, the storage unit 22 also stores information such as a learned model 222 machine-learned by a predetermined learning method and an estimation result estimated by the estimation device 2.
[0055] The control unit 21 includes control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit in which control programs such as BIOS and OS for causing the CPU to execute various processes are pre-stored. The RAM is a volatile or non-volatile storage unit that stores various information and is used as a temporary storage memory (working area) for various processes executed by the CPU. Then, the control unit 21 controls the estimation device 2 by causing the CPU to execute various control programs pre-stored in the ROM or the storage unit 22.
[0056] Specifically, the control unit 21 of the estimation device 2 according to the present embodiment includes various processing units such as an acquisition processing unit 211, an estimation processing unit 212, a correction processing unit 213, and an output processing unit 214. Note that the control unit 21 functions as the various processing units by causing the CPU to execute various processes according to the estimation program 221. Also, some or all of the processing units included in the control unit 21 may be configured by electronic circuits. Note that the estimation program 221 may be a program for causing a plurality of processors to function as the various processing units.
[0057] Note that the control unit 21 functions as the various processing units by causing the CPU to execute various processes according to the estimation program 221 using the learned model 222. Also, some or all of the processing units included in the control unit 21 may be configured by electronic circuits. Note that the estimation program 221 may be a program for causing a plurality of processors to function as the various processing units.
[0058] Here, the learned model 222 is a single learned model generated based on learning data in which an image (for example, a face image) of an estimation target (for example, a person) and an attribute (for example, age) of the estimation target are associated with each other.
[0059] The acquisition processing unit 211 acquires a captured image of the estimation target from the image processing apparatus 1. Here, the acquisition processing unit 211 acquires a plurality of face images of a person who is the estimation target from the image processing apparatus 1. The acquisition processing unit 211 is an example of the second acquisition processing unit of the present invention.
[0060] The estimation processing unit 212 uses the learned model 222 to estimate an attribute (for example, age) from the output value of the output layer corresponding to the attribute using the face image acquired by the acquisition processing unit 211 as an input image. Specifically, the estimation processing unit 212 calculates output values for each of a plurality of classifications (age classes) of an attribute using the learned model 222 for the input image acquired by the acquisition processing unit 211, calculates a determination value of the attribute (age) based on the calculated output values of each classification, and estimates the attribute of the face image based on the calculated determination value. The estimation processing unit 212 is an example of the estimation processing unit of the present invention.
[0061] For example, as shown in FIG. 4, the estimation processing unit 212 estimates the age of the face image by performing a sum-of-products operation of the ID (i) = 0 to (N - 1) (N is the number of age classes) of the output layer for age and the respective score values (output values). Here, the number of age classes is "10", and as a result of performing a sum-of-products operation of the ID (i) = 0 to 9 of the output layer and the respective score values, the determination value is "3.70". Therefore, the estimation processing unit 212 estimates that the age of the face image belongs to the age class of "13 to 18 years old". Note that the estimation processing unit 212 may calculate the age corresponding to the output value by linearly interpolating using the minimum age and the maximum age among the plurality of ages included in the estimated age class. In the example shown in FIG. 4, in the estimated age class of "13 to 18 years old", the minimum age is "13 years old" and the maximum age is "18 years old". Therefore, the estimation processing unit 212 calculates the estimated age to be 16.5 years old (= 13 + 0.70 × (18 - 13)).
[0062] In addition, the estimation processing unit 212 estimates the age for each of the plurality of face images, and determines the average value of each estimation result as the estimated age of the person.
[0063] The output processing unit 214 outputs the estimation result. For example, the output processing unit 214 causes the operation display unit 23 to display the estimation result (such as "13 to 18 years old" or "16.5 years old") estimated for the input face image. Further, the output processing unit 214 may transmit the estimation result to another device (such as the image processing device 1) via the communication network N1.
[0064] Here, when the image quality of the face image input from the image processing device 1 is low, the age determination accuracy may decrease. For example, when the focus of the face image extracted by the image processing device 1 is not sufficiently in focus and is blurred, the estimated age may be lower (younger) than the actual age. This is because in the case of a face image with high image quality, it is possible to appropriately estimate the age considering elements such as fine texture due to wrinkles, etc., while in the case of a face image with low image quality, it becomes difficult to consider such elements, resulting in a tendency for the age to be estimated lower.
[0065] Therefore, in order to further improve the age estimation accuracy, the correction processing unit 213 may correct the age estimation result estimated by the estimation processing unit 212. Specifically, the correction processing unit 213 corrects the estimation result of the attribute estimated by the estimation processing unit 212 when the feature amount related to the estimation target included in the captured image is less than or equal to a predetermined value. For example, the correction processing unit 213 corrects the age class estimated by the estimation processing unit 212 to the next higher class when the evaluation value dd of the edge is less than or equal to a predetermined evaluation value (threshold value). For example, when the evaluation value dd of the edge of the face image is less than or equal to the predetermined evaluation value and the estimation processing unit 212 estimates that the age of the face image belongs to the age class of "13 to 18 years old", the correction processing unit 213 estimates that the age of the face image belongs to the age class of "19 to 29 years old". Note that the image processing device 1 inputs the information of the evaluation value dd together with the extracted face image to the estimation device 2, and the correction processing unit 213 executes the correction processing based on the evaluation value dd acquired from the image processing device 1. The correction processing unit 213 is an example of the correction processing unit of the present invention.
[0066] As another embodiment, when the estimated age result for the input image (face image) is not appropriate, the estimation processing unit 212 may determine it as an estimation error. Specifically, for the face image, when the score value (output value) of each calculated age class is equal to or less than a predetermined score value, or when the maximum score value (the maximum output value of the present invention) among the calculated score values of each age class is equal to or less than a predetermined score value (the predetermined output value of the present invention), the estimation processing unit 212 determines that the age estimation for the face image is an error. Also, when the variance value of the score values of each age class is greater than a certain value (threshold value), there will be no peak. Therefore, as another embodiment, the estimation processing unit 212 may calculate the variance value of the score values of each age class and determine it as an estimation error when the variance value is greater than the threshold value.
[0067] Note that when an estimation error occurs, the score values of each age class become uniformly low, and the determination value tends to be the age class in the center. For example, when the image quality is good, as shown by the score value A in FIG. 9A, the score values of each age class are normally distributed. On the other hand, when the image quality is not good, as shown by the score value B in FIG. 9A, the score values of each age class become uniformly (evenly) low, and the determination value tends to be the class in the center (here, "4.75").
[0068] Therefore, in order to determine that the estimation of the image corresponding to the score value B is an error, the estimation processing unit 212 subtracts a correction value (for example, "0.3") from the score values of each age class, for example. As a result, as shown in FIG. 9B, since the score value B becomes "0.00" in each age class, the determination value cannot be calculated. For this reason, the estimation processing unit 212 determines that the age estimation of the face image corresponding to the score value B is an error. Also, as shown in FIG. 9B, the estimation processing unit 212 can appropriately estimate the age of the face image corresponding to the score value A.
[0069] As another embodiment, when the score value for each age class is less than a predetermined value (e.g., "0.5"), the estimation processing unit 212 may determine that the score value is an error and convert it to "0.00". Even in this method, while appropriately estimating the age of the face image corresponding to the score value A, it is possible to determine that the estimation of the age of the face image corresponding to the score value B is an error.
[0070] According to the estimation system 100 according to this embodiment, since an image with good image quality can be extracted from the images of the estimation target captured by the camera 15 and input to the estimation device 2, it is possible to accurately estimate the attributes of the estimation target.
[0071] [Image Processing] Hereinafter, with reference to FIG. 10, an example of the procedure of image processing executed by the control unit 11 of the image processing apparatus 1 will be described.
[0072] Note that the present invention can be regarded as an invention of an image processing method for executing one or more steps included in the image processing. Also, one or more steps included in the image processing described here may be appropriately omitted. Also, the execution order of each step in the image processing may be different within a range that produces the same operational effects. Furthermore, here, the case where the control unit 11 executes each step in the image processing is described as an example, but in other embodiments, one or more processors may execute each step in the image processing in a distributed manner.
[0073] The control unit 11 of the image processing apparatus 1 executes a process of extracting an image to be the target of the estimation process by the estimation device 2 and outputting it to the estimation device 2.
[0074] First, in step S1, the control unit 11 determines whether a predetermined number of captured images have been acquired. The predetermined number is set to, for example, "5". For example, when the control unit 11 acquires five captured images of a specific person from the camera 15 (S1: Yes), the process proceeds to step S2. Step S1 is an example of the acquisition step of the present invention.
[0075] In step S2, the control unit 11 generates a face image from each acquired captured image. Specifically, the control unit 11 detects the positions of the face and major organs such as eyes, nose, and mouth in each captured image, cuts out an image within a predetermined range, and generates a face image (see FIG. 7).
[0076] Next, in step S3, the control unit 11 evaluates each generated face image. Specifically, the control unit 11 calculates the size in the original image and the feature amount of the periphery of the eyes (the rectangular frame area in FIG. 7) for each face image.
[0077] For example, as shown in FIG. 8, the control unit 11 performs filter processing on each face image using a noise filter Fd, a high-pass filter Fh, and a low-pass filter Fl, and calculates an evaluation value F of each face image according to the following formula. Step S3 is an example of the evaluation step of the present invention. F=(dd×dr×dw×mag / 128.0);
[0078] Next, in step S4, the control unit 11 extracts a plurality of top face images with high calculated evaluation values F. For example, the control unit 11 extracts the top 3 face images in descending order of the evaluation value F among 5 face images.
[0079] Next, in step S5, the control unit 11 outputs the extracted face images to the estimation device 2. Here, the control unit 11 outputs the extracted 3 face images to the estimation device 2. Note that the control unit 11 may output information regarding the evaluation value F (for example, the evaluation value dd of the edge, the evaluation value dr of the contrast, the image size dw, etc.) to the estimation device 2 together with the face images. The face images output by the image processing device 1 are used as input images for estimation targets in the estimation device 2. Step S5 is an example of the output step of the present invention.
[0080] [Estimation process] Hereinafter, an example of the procedure of the estimation process executed by the control unit 21 of the estimation device 2 will be described with reference to FIG. 11.
[0081] Note that the present invention can be regarded as an invention of an estimation method for executing one or more steps included in the above-described estimation process. Further, one or more steps included in the above-described estimation process may be appropriately omitted. Also, the execution order of each step in the above-described estimation process may be different as long as the same operational effects are produced. Furthermore, here, the case where the control unit 21 executes each step in the above-described estimation process is described as an example, but in other embodiments, one or more processors may execute each step in the above-described estimation process in a distributed manner.
[0082] The estimation device 2 executes an estimation process according to an estimation program 221 using a learned model 222 that estimates the age of a person based on a face image of the person input from the image processing device 1.
[0083] First, in step S11, the control unit 21 determines whether or not a face image of an estimation target has been acquired. When the control unit 11 acquires the face image from the image processing device 1 (S11: Yes), the process proceeds to step S12. Here, the control unit 21 acquires three face images of a specific person from the image processing device 1.
[0084] In step S12, for each of the three acquired face images, the control unit 21 calculates a score value (output value) for each of a plurality of age classes using the learned model 222.
[0085] Next, in step S13, the control unit 21 corrects the calculated score value for each age class. For example, the control unit 21 subtracts a correction value "0.3" from the calculated score value for each age class (see FIG. 9B). Note that in the estimation process of the present invention, the process of step S13 may be omitted.
[0086] Next, in step S14, the control unit 21 calculates an age determination value based on the corrected score values. The control unit 21 calculates an age determination value for each of the three face images.
[0087] Next, in step S15, the control unit 21 determines whether an error has occurred in the calculation process of the determination value. If no calculation error occurs (refer to the score value A in FIG. 9B), the process proceeds to step S16. On the other hand, if a calculation error occurs (refer to the score value B in FIG. 9B), the process proceeds to step S151. Note that the control unit 21 may determine whether a calculation error has occurred for each face image.
[0088] In step S151, the control unit 21 outputs an estimation error. For example, when a calculation error occurs for three face images, the control unit 21 outputs an estimation error.
[0089] In step S16, the control unit 21 estimates the age based on the determination value. For example, the control unit 21 estimates the age for each of the three face images, and determines the average value of each estimation result as the estimated age of the person to be estimated.
[0090] Finally, in step S17, the control unit 21 outputs the estimation result. For example, the control unit 21 causes the operation display unit 23 to display the estimation result (such as "13 - 18 years old" or "16.5 years old") estimated for the input face image. Further, the control unit 21 may transmit the estimation result to other devices (such as the image processing device 1) via the communication network N1.
[0091] Note that as another embodiment, when the evaluation value dd of the edge of the face image acquired from the image processing device 1 is less than or equal to a predetermined evaluation value (threshold value), the control unit 21 may correct the age class estimated in step S16 to the next higher class. For example, when the evaluation value dd of the edge of the face image is less than or equal to the predetermined evaluation value and the age of the face image is estimated to belong to the age class of "13 - 18 years old" in step S16, the control unit 21 corrects the age class to "19 - 29 years old".
[0092] As described above, the image processing apparatus 1 according to the present embodiment is an apparatus that inputs the input image to an estimation apparatus 2 that estimates an attribute of an estimation target based on the input image. Further, the image processing apparatus 1 acquires a captured image of the estimation target, evaluates the image quality of the acquired captured image, and when it is determined that the image quality of the captured image has a predetermined image quality, outputs the captured image as the input image.
[0093] Further, the estimation apparatus 2 according to the present embodiment acquires the captured image as the input image, calculates an output value for each of a plurality of classifications of the attribute using a predetermined learned model for the acquired input image, calculates a determination value of the attribute based on the calculated output value for each classification, and estimates the attribute of the estimation target based on the calculated determination value.
[0094] Thereby, the attribute of the estimation target can be estimated from an image with good image quality among the images of the estimation target. Therefore, it becomes possible to improve the estimation accuracy of the attribute of the estimation target.
[0095] Note that the image processing apparatus 1 and the estimation apparatus 2 may be configured as an integrated device. Further, the device may be a digital signage or a store terminal (POS terminal). Further, the camera 15 may be installed in the digital signage, may be installed in the store terminal, or may be installed in the store.
[0096] In the above-described embodiment, a configuration for estimating one attribute (age) from a face image is given as an example, but the present invention may be a configuration for estimating a plurality of attributes from a captured image.
Explanation of Reference Numerals
[0097] 1: Image processing apparatus 2: Estimation apparatus 100: Estimation system 111: Acquisition processing unit 112: Evaluation processing unit 113: Output processing unit 211: Acquisition processing unit 212: Estimation processing unit 213: Correction Processing Unit 214: Output Processing Unit
Claims
1. An image processing apparatus that inputs the input image to an estimation apparatus that estimates an attribute of an estimation target based on the input image, a first acquisition processing unit that acquires a captured image of the estimation target; an evaluation processing unit that performs a filter process on the captured image acquired by the first acquisition processing unit to calculate a feature amount of a specific part of the estimation target, and calculates an evaluation value of the captured image based on the feature amount; an output processing unit that outputs the captured image as the input image based on the evaluation value of the captured image calculated by the evaluation processing unit; comprising: The evaluation processing unit calculates, as a first evaluation value of an edge in the captured image, a difference between a maximum value and a minimum value of a result of a filter process by a high-pass filter, and calculates, as a second evaluation value of contrast in the captured image, a difference between a maximum value and a minimum value of a result of a filter process by a low-pass filter, and calculates an evaluation value of the captured image based on the first evaluation value and the second evaluation value. An image processing apparatus.
2. The first acquisition processing unit acquires a plurality of the captured images of the estimation target, the evaluation processing unit calculates the evaluation value for each of the plurality of the captured images, the output processing unit outputs, as the input image, a predetermined number of the captured images having a high evaluation value among the plurality of the captured images, The image processing apparatus according to claim 1.
3. The first acquisition processing unit acquires a plurality of the captured images of the estimation target, the evaluation processing unit calculates the evaluation value for each of the plurality of the captured images, the output processing unit outputs, as the input image, one or more of the captured images whose evaluation value exceeds a predetermined evaluation value, The image processing apparatus according to claim 1.
4. An estimation apparatus that estimates an attribute of the estimation target using the captured image output from the image processing apparatus according to any one of claims 1 to 3 as the input image, a second acquisition processing unit that acquires the captured image as the input image from the image processing apparatus; for the input image acquired by the second acquisition processing unit, using a predetermined learned model, calculates an output value for each of a plurality of classifications of the attribute, calculates a determination value of the attribute based on the calculated output value of each classification, and estimates the attribute of the estimation target based on the calculated determination value. An estimation processing unit; When the evaluation value of the edge in the captured image, which is the difference between the maximum value and the minimum value of the result of the filter process by the high-pass filter, is less than or equal to a threshold value, a correction processing unit that corrects the estimation result of the attribute estimated by the estimation processing unit; An estimation device comprising the same.
5. The correction processing unit corrects the estimation result of the attribute estimated by the estimation processing unit when the feature amount regarding the estimation target included in the captured image is less than or equal to a predetermined value. The estimation device according to claim 4.
6. When the output value of each classification calculated for the input image is less than or equal to a predetermined output value, or when the maximum output value among the output values of each classification calculated is less than or equal to a predetermined output value, the estimation processing unit determines that the estimation of the attribute for the input image is an error. The estimation device according to claim 4.
7. An image processing method for inputting the input image to an estimation device that estimates the attribute of an estimation target based on the input image, wherein one or more processors perform an acquisition step of acquiring a captured image of the estimation target; an evaluation step of performing a filter process on the captured image acquired in the acquisition step to calculate a feature amount of a specific part of the estimation target, and calculating an evaluation value of the captured image based on the feature amount; an output step of outputting the captured image as the input image based on the evaluation value of the captured image calculated in the evaluation step; and perform In the evaluation step, the difference between the maximum value and the minimum value of the result of the filter process by the high-pass filter is calculated as a first evaluation value of the edge in the captured image, the difference between the maximum value and the minimum value of the result of the filter process by the low-pass filter is calculated as a second evaluation value of the contrast in the captured image, and the evaluation value of the captured image is calculated based on the first evaluation value and the second evaluation value. An image processing method.
8. An image processing program for inputting the input image to an estimation device that estimates the attribute of an estimation target based on the input image, the program comprising: an acquisition step of acquiring a captured image of the estimation target; an evaluation step of performing a filter process on the captured image acquired in the acquisition step to calculate a feature amount of a specific part of the estimation target, and calculating an evaluation value of the captured image based on the feature amount; an output step of outputting the captured image as the input image based on the evaluation value of the captured image calculated in the evaluation step; cause to be executed by one or more processors, In the evaluation step, calculate the difference between the maximum value and the minimum value of the result of the filtering process by the high-pass filter as the first evaluation value of the edge in the captured image, calculate the difference between the maximum value and the minimum value of the result of the filtering process by the low-pass filter as the second evaluation value of the contrast in the captured image, and calculate the evaluation value of the captured image based on the first evaluation value and the second evaluation value. An image processing program.
Citation Information
Patent Citations
Fingerprint registration method and related product
CN107330374A
Image recognition method and device based on convolutional neural network
CN107944458A
Posture recognition method, device and system and computer readable storage medium
CN112733722A
Age estimation method and age estimation device
JP2009271885A
Attribute-based head-count totaling device, attribute-based head-count totaling method and attribute-based head-count totaling system
JP2010033474A