Estimation system, estimation method, and estimation program
Patent Information
- Application Number
- JP2021120342
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-21
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-07-21
AI Technical Summary
Conventional methods for estimating multiple attributes, such as gender and age from a face image, face increased processing load and compromised estimation accuracy due to high computational demands and errors propagating between attribute estimations.
A single trained model is used to estimate multiple attributes simultaneously by separating output layers for each attribute, reducing the total number of connections and enabling parallel processing, thereby improving estimation accuracy while minimizing processing load.
The method achieves improved estimation accuracy for multiple attributes with reduced processing load and memory requirements, allowing for faster and more accurate simultaneous estimation of attributes like gender and age.
Smart Images

Figure 00000012_0000 
Figure 00000013_0000 
Figure 00000013_0001
Abstract
Description
Technical Field
[0001] The present invention relates to an estimation system, an estimation method, and an estimation program for estimating attributes of an estimation target.
Background Art
[0002] Conventionally, a technique for estimating attributes such as the age and gender of a person from a face image of the person has been known. For example, in Patent Document 1, for each of a plurality of predetermined age groups, a score representing the probability that the face of a person in an image corresponds to that age group is calculated, a portion that adversely affects the attribute estimation in the image is identified based on the state of the face, the score is corrected so that the influence of the portion is reduced, and the age group corresponding to the corrected score representing the highest probability among the corrected scores for each age group is regarded as the attribute of the person.
[0003] Also, Patent Document 2 discloses a technique for estimating the age group of a user by comparing the feature information of a face image with the feature information stored in a learning result storage means.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] Here, a common method for estimating a person's attributes from an image of that person is the convolutional neural network (CNN), a deep learning technique. In this convolutional neural network, an input image is read, and in the first half, important and multiple features of the input image are extracted by repeatedly performing convolution and pooling. In the second half, identification (classification) is performed by fully connected layers and an output layer based on these features.
[0006] Conventionally, methods that estimate a person's gender and age from their face using, for example, the aforementioned convolutional neural network, have the following problems. For example, when estimating gender and age as a pair using a single trained model, the number of outputs (number of fully connected layers) becomes large. Specifically, the number of fully connected layers is the product of the number of genders "2", the number of age classes "N", and the number of nodes in the fully connected layer "m" (2 × N × m). As a result, the computational load of the estimation process increases, leading to a problem of increased processing load.
[0007] Another possible method for estimating a person's gender and age is to use two pre-trained models: one for estimating gender and another for estimating age using the gender estimation result. However, with this method, if there is an error in gender estimation, the age estimation will also be incorrect. This leads to the problem of not being able to obtain sufficient estimation accuracy.
[0008] Thus, with conventional techniques, it is difficult to improve estimation accuracy while reducing the processing load when estimating multiple attributes (e.g., gender and age) of an object (e.g., a person).
[0009] The object of the present invention is to provide an estimation system, estimation method, and estimation program that can improve estimation accuracy while reducing the processing load when estimating multiple attributes of an estimated target. [Means for solving the problem]
[0010] An estimation system according to one aspect of the present invention is a system comprising: an acquisition processing unit that acquires an image of the target to be estimated; and an estimation processing unit that uses a single trained model generated based on training data in which the image of the target to be estimated and each of the plurality of attributes of the target to be estimated are associated with each other, and uses the image acquired by the acquisition processing unit as an input image to estimate a first attribute from the first output value of a first output layer corresponding to a first attribute included in the plurality of attributes, and estimates a second attribute from the second output value of a second output layer corresponding to a second attribute included in the plurality of attributes.
[0011] Another aspect of the present invention relates to an estimation method in which one or more processors perform an acquisition step of acquiring an image to be estimated, and an estimation step of using a single trained model generated based on training data in which the image to be estimated and each of the multiple attributes to be estimated are associated with each other, using the image acquired in the acquisition step as an input image, estimating the first attribute from the first output value of a first output layer corresponding to the first attribute included in the multiple attributes, and estimating the second attribute from the second output value of a second output layer corresponding to the second attribute included in the multiple attributes.
[0012] Another aspect of the present invention relates to an estimation program which causes one or more processors to execute an acquisition step of acquiring an image to be estimated, and an estimation step of using a single trained model generated based on training data in which the image to be estimated and each of the multiple attributes of the estimated object are associated with each other, using the image acquired in the acquisition step as an input image, estimating the first attribute from the first output value of a first output layer corresponding to the first attribute included in the multiple attributes, and estimating the second attribute from the second output value of a second output layer corresponding to the second attribute included in the multiple attributes. [Effects of the Invention]
[0013] According to the present invention, it is possible to provide an estimation system, estimation method, and estimation program that can improve estimation accuracy while reducing the processing load when estimating multiple attributes of an estimated target. [Brief explanation of the drawing]
[0014] [Figure 1] Figure 1 is a block diagram showing the configuration of an estimation system according to an embodiment of the present invention. [Figure 2] Figure 2 shows an example of the basic structure of a convolutional neural network according to an embodiment of the present invention. [Figure 3] Figure 3 is a schematic diagram showing the coupling state between the fully connected layer and the output layer included in the convolutional neural network according to an embodiment of the present invention. [Figure 4A] Figure 4A is a schematic diagram illustrating an example of a gender estimation method according to an embodiment of the present invention. [Figure 4B] Figure 4B is a schematic diagram illustrating an example of an age estimation method according to an embodiment of the present invention. [Figure 5] Figure 5 shows an example of training data according to an embodiment of the present invention. [Figure 6] Figure 6 shows an example of an estimation method using a conventional pre-trained model. [Figure 7] Figure 7 shows an example of an estimation method using a conventional pre-trained model. [Figure 8] Figure 8 shows an example of an estimation method using a trained model according to an embodiment of the present invention. [Figure 9] Figure 9 is a flowchart illustrating an example of the procedure for the estimation process performed in the estimation system according to an embodiment of the present invention. [Modes for carrying out the invention]
[0015] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Note that the following embodiments are merely examples of embodying the present invention and do not have the character of limiting the technical scope of the present invention.
[0016] The estimation system 1 according to this embodiment is a system that estimates each of a plurality of attributes in an image based on the image to be estimated. For example, the object to be estimated is a human face, a vehicle, an animal, an article, or the like. When the object to be estimated is a human, the plurality of attributes include the gender and age of the human. Further, the plurality of attributes may include the presence or absence of glasses, the presence or absence of a mask, the ethnic type, etc. of the human. When the object to be estimated is a vehicle, the plurality of attributes include the vehicle number, the manufacturer, domestic / imported vehicle, color, etc. The estimation system 1 can be applied to, for example, digital signage to estimate the gender and age of passersby on the street, in front of a store, etc., and perform crowd flow analysis, marketing, etc. Also, the estimation system 1 can be applied to, for example, the POS terminal of a store to estimate the gender and age of customers and perform marketing, etc.
[0017] In this embodiment, as an example of the object to be estimated, a human face will be cited, and as an example of the plurality of attributes, the gender and age of the human will be cited for explanation. That is, the human face is an example of the object to be estimated in the present invention, the gender of the human is an example of the first attribute in the present invention, and the age of the human is an example of the second attribute in the present invention. The attributes estimated by the estimation system 1 are not limited to two, and may be three or more.
[0018] FIG. 1 is a diagram showing a schematic configuration of the estimation system 1 according to an embodiment of the present invention. As shown in FIG. 1, the estimation system 1 includes a control unit 11, a storage unit 12, an operation display unit 13, a communication unit 14, etc. The estimation system 1 may be an information processing device such as a personal computer, for example.
[0019] The communication unit 14 is a communication interface for connecting the estimation system 1 to a network by wire or wireless connection and for performing data communication with other devices via the network in accordance with a predetermined communication protocol.
[0020] The operation display unit 13 is a user interface comprising a display unit such as a liquid crystal display or an organic EL display that displays various types of information, and an operation unit such as a mouse, keyboard, or touch panel that accepts input.
[0021] The storage unit 12 includes a non-volatile storage device such as an HDD (Hard Disk Drive) or SSD (Solid State Drive) that stores various types of information. The storage unit 12 stores control programs such as an estimation program 121 that cause the control unit 11 to execute the estimation process described later (see Figure 9). The control programs such as the estimation program 121 are provided, for example, by being recorded on a computer-readable non-temporary recording medium, and are read from the non-temporary recording medium by the reading device of the estimation system 1 and stored in the storage unit 12. The control programs may also be provided (downloaded) to the estimation system 1 via a network from an external server other than the estimation system 1 and stored in the storage unit 12. The storage unit 12 also stores information such as a trained model 122 that has been machine-learned using the learning method described later, and estimation results estimated by the estimation system 1.
[0022] The control unit 11 includes control devices such as a CPU, ROM, and RAM. The CPU is a processor that performs various arithmetic operations. The ROM is a non-volatile memory unit that stores control programs such as a BIOS and OS in advance to allow the CPU to perform various operations. The RAM is a volatile or non-volatile memory unit that stores various information and is used as a temporary memory (work area) for the various operations performed by the CPU. The control unit 11 controls the estimation system 1 by executing various control programs that are pre-stored in the ROM or memory unit 12 using the CPU.
[0023] Here, a convolutional neural network (CNN), a type of deep learning technique, is generally used as a method to estimate the attributes of a person from an image of that person. In this convolutional neural network, an input image is read, and in the first half, important and multiple features of the input image are extracted by repeatedly performing convolution and pooling, and in the second half, identification (classification) is performed by fully connected layers and an output layer based on those features.
[0024] Figure 2 shows the basic structure of a convolutional neural network. The convolutional neural network loads training data, extracts important and multiple features from the training data by repeatedly performing convolution and pooling in the first half, and then performs classification based on these features using fully connected layers and an output layer in the second half.
[0025] Specifically, in the input layer of the first half of the processing, the input image, which is the training data, is converted from multidimensional data (3D data combining the 2D image and the 1D color component RGB) to a 1D vector. In the next convolutional layer, features such as edge detection are extracted by detecting the grayscale pattern of the input image. In the following pooling layer, the extracted important features are thinned out to allow for positional shifts so that the object is considered the same object even if its position changes, and the image size (information) is compressed. By connecting multiple convolutional and pooling layers, more important features are extracted. Here, if the number of convolutional and pooling layers increases, the magnitude (gradient) of the features will disappear, so an activation function is inserted after the convolutional layer to enhance the features and suppress the disappearance of the gradient. In the second half of the processing, classification is performed based on the features extracted in the first half of the processing.
[0026] Figure 3 schematically shows the connection state between the fully connected layer and the output layer. Each circle represents an individual node, and a connection weight coefficient is multiplied for each connection line between nodes. The number of connections is calculated by the number of nodes in the fully connected layer (m) × the number of outputs in the output layer (n). As the number of connections increases, the classification ability improves, but the computational load increases, and performance decreases. In addition, there is the challenge that it takes time to calculate the optimal connection weight coefficient during training. In the fully connected layer (see Figure 2), the features extracted in the first half are aggregated into one node each, and the connection weight coefficients between each node are adjusted for classification and converted into feature variables. Finally, in the output layer (see Figure 2), the feature variables are output as a score value that maximizes the probability of correct classification. Note that if the output values (score values) of the output layer are added up for each output, it equals 1.0 (100%).
[0027] When estimating the attributes (gender and age) of a person in a facial image, it is common to use an estimation method based on the aforementioned convolutional neural network. The input image is training data (multiple facial images), various features are extracted from this training data, and the configuration of the first processing is changed, or various parameters (connection weight coefficients of the fully connected layer) are adjusted (learned) so that the output value of the output layer approaches the ground truth (gender, age, and the age range (class)) for the training data of the input image on an overall average basis. The resulting "trained model" (see Figure 2) is a combination of the optimized configuration and estimation parameters achieved through machine learning.
[0028] By inputting a facial image of the person to be estimated into a trained model generated through machine learning using this learning method, it becomes possible to estimate the gender and age of that person from the output value (score value).
[0029] Next, we will give an example of an estimation method for determining gender and age from a person's image. Figure 4A schematically shows the method for estimating gender, and Figure 4B schematically shows the method for estimating age.
[0030] For example, the gender of the person with the higher score in the output layer is used as the estimated gender. In this example, the score corresponding to male is 0.22, and the score corresponding to female is 0.78 (= 1.0 - male score), so the person being estimated is estimated to be female.
[0031] The estimated age is calculated by performing a sum-of-products operation between each age class ID (node number in the output layer) and the score value corresponding to each age class ID. An age class represents a range of ages divided into specific age groups. In this case, the result of the sum-of-products operation is 3.70, and truncating the decimal point gives the age class ID as "3," so the corresponding age class is estimated to be "13-18 years old." Furthermore, by taking the decimal point "0.70" into account and performing linear interpolation within the same age class, the age can also be estimated as 13 + 0.70 × (18-13) = 16.5 years old.
[0032] Since these score values fluctuate, it is desirable to estimate gender and age using the average value of multiple input images.
[0033] Here, Figure 5 shows an example of training data used in machine learning when generating a trained model. Face images 1 to 10 each represent face images of men in each age range, and face images 11 to 20 each represent face images of women in each age range. Each training data folder is given a folder name such as "sequence number#gender#lower age limit-upper age". Each training data folder stores multiple face image data corresponding to that age range, and these are collectively used as training data (training dataset). The "number of age classes" is the number of age divisions, and in this example, there are "10 classes" for each gender, so the total number of age classes for both male and female genders is "20 classes".
[0034] It should be noted that the training data itself is not stored in the trained model. Instead, its features are extracted using a convolutional neural network, and the optimal configuration and parameters are determined to output estimated values that match the training data, resulting in the trained model.
[0035] By the way, conventional methods that estimate a person's gender and age from their face using, for example, the aforementioned convolutional neural network, have the following problems. For example, when estimating gender and age as a pair using a single trained model A, the number of outputs (number of fully connected layers) becomes large. Specifically, the number of fully connected layers is the product of the number of genders "2", the number of age classes "N", and the number of nodes in the fully connected layer "m" (2 × N × m). As a result, the computational load of the estimation process increases, leading to a problem of increased processing load. The following explains a specific example of this problem.
[0036] Figure 6 schematically illustrates the estimation method using the conventional pre-trained model A. Pre-trained model A estimates a person's gender and age as a pair and consists of a single pre-trained model.
[0037] In this estimation method, the number of outputs is the total number of combinations of gender class and age class; therefore, in the example in Figure 6, the total number of combinations is "20". A score value is then calculated for each output layer. The sum of all score values is "1".
[0038] First, the sum of the score values corresponding to ID(i)=0 to 9 in the output layer (degree of maleness) is compared with the sum of the score values corresponding to ID(i)=10 to 19 in the output layer (degree of femaleness), and the larger value is estimated to be the gender. In this case, the degree of maleness is "0.89" and the degree of femaleness is "0.11", so the gender of the input face image is estimated to be male.
[0039] Next, the age of the face image is estimated by performing a sum-of-products operation on the ID of the output layer corresponding to the estimated gender and its score value. In this case, since the person has already been determined to be male, the sum-of-products operation on the corresponding output layer ID(i)=0~9 and its respective score value is "1.20", so the age of the face image is estimated to be in the "4~6 year old" age class. In the case of a female, the output layer ID(i)=10~19, so when performing the sum-of-products operation, the number of male age classes "10" is subtracted from the ID value before proceeding.
[0040] In the estimation method using the pre-trained model A, all gender and age classes are combined. If we denote the number of genders as "2", the number of age classes as "N", and the number of nodes in the fully connected layer as "m", then the number of fully connected layers becomes "2 × N × m". This results in a large amount of computation for the estimation process, leading to increased processing load.
[0041] Another estimation method could be to use two pre-trained models: a pre-trained model B1 for estimating gender and a pre-trained model B2 for estimating age using the gender estimation result. However, with this method, if there is an error in gender estimation, the age estimation will also be incorrect. This leads to the problem of not being able to obtain sufficient estimation accuracy. A specific example of this problem will be explained below.
[0042] Figure 7 schematically illustrates the estimation method using two conventional pre-trained models B1 and B2. In the estimation method shown in Figure 7, the gender of the face image is first estimated using the pre-trained model B1 in the first step, and then, depending on the result, the age of the face image is estimated in the second step using either the pre-trained model B2-M (for males) or the pre-trained model B2-F (for females), which is either male or female.
[0043] In this estimation method, the score for males is "0.78" and the score for females is "0.22," so the gender of the input face image is estimated to be male.
[0044] Next, the age is estimated using the pre-trained male age model B2-M. The sum of products calculated by the output layer's ID(i)=0~9 and their respective score values is "1.20," thus estimating the age of the face image to be in the "4~6 year old" age class. If the gender is determined to be female, the pre-trained female age model B2-F is used, and the age is estimated in the same way as for males.
[0045] As shown in Figure 7, the estimation method estimates gender and age separately and stepwise. The total number of connections is "m × (2 + N)", which is less than that of the trained model A (see Figure 6). However, a total of three trained models are required: one trained model B1 for gender estimation, one trained model B2-M for male age estimation, and one trained model B2-F for female age estimation. Furthermore, applying three trained models increases memory usage, and performance decreases because the age is estimated for the estimated gender after the gender has been estimated. In addition, if the gender estimation is incorrect, the subsequent age estimation will also be incorrect, resulting in a decrease in estimation accuracy.
[0046] Thus, with conventional estimation methods (see Figures 6 and 7), it is difficult to improve estimation accuracy while reducing the processing load when estimating multiple attributes (e.g., gender and age) of an object (e.g., a person). In contrast, according to the estimation system 1 of this embodiment, as shown below, it is possible to improve estimation accuracy while reducing the processing load when estimating multiple attributes of an object.
[0047] Specifically, the control unit 11 of the estimation system 1 according to this embodiment includes various processing units such as an acquisition processing unit 111, an estimation processing unit 112, and an output processing unit 113, as shown in Figure 1. The control unit 11 functions as these various processing units by executing various processes according to the estimation program 121 using the CPU. Some or all of the processing units included in the control unit 11 may be composed of electronic circuits. The estimation program 121 may be a program that causes multiple processors to function as these various processing units.
[0048] The control unit 11 functions as one of the various processing units by executing various processes according to the estimation program 121 using the trained model 122. Furthermore, some or all of the processing units included in the control unit 11 may be composed of electronic circuits. The estimation program 121 may also be a program that causes multiple processors to function as the various processing units.
[0049] Here, the trained model 122 is a single trained model generated based on training data in which an image of the target to be estimated (e.g., a person) (e.g., a facial image), a first attribute of the target to be estimated (e.g., gender), and a second attribute of the target to be estimated (e.g., age) are associated with each other. In addition, the trained model of the present invention may be generated based on training data in which an image of the target to be estimated and each of the three or more attributes of the target to be estimated are associated with each other.
[0050] Figure 8 schematically shows the estimation method using the pre-trained model 122 according to this embodiment. In the pre-trained model 122, separate output layers are provided for gender and age, and the output layers are configured so that the sum of the output values (score values) of each output layer is "1".
[0051] The acquisition processing unit 111 acquires the captured image of the target for estimation. Here, the acquisition processing unit 111 acquires the face image of the person who is the target for estimation. For example, the acquisition processing unit 111 acquires the face image of a person captured by a camera (not shown) connected to the estimation system 1 via the communication unit 14. The acquisition processing unit 111 sequentially acquires the face images captured by the camera at a predetermined frame rate. The camera may be included in the estimation system 1. The acquisition processing unit 111 is an example of the acquisition processing unit of the present invention.
[0052] The estimation processing unit 112 uses a single pre-trained model 122 to take the face image acquired by the acquisition processing unit 111 as an input image, to estimate gender from the first output value of the first output layer corresponding to gender, and to estimate age from the second output value of the second output layer corresponding to age. Furthermore, the estimation processing unit 112 uses a single pre-trained model 122 to simultaneously estimate gender and age from the input face image. The estimation processing unit 112 is an example of the estimation processing unit of the present invention.
[0053] Specifically, the estimation processing unit 112 calculates a first output value for each of the multiple gender classifications (gender classes), and estimates the gender of the face image based on the multiple first output values calculated. In addition, the estimation processing unit 112 calculates a second output value for each of the multiple age classifications (age classes), and estimates the age of the face image based on the multiple second output values calculated.
[0054] Furthermore, the estimation processing unit 112 outputs an output value of the total number obtained by adding the number of gender classifications (number of genders) and the number of age classifications (number of age classes). For example, the estimation processing unit 112 outputs an output value (second output value) of the total number "12", which is the sum of the number of gender classifications "2" and the number of age classifications "10".
[0055] The estimation processing unit 112 then estimates the age of the person in the face image based on the result of a sum-of-products operation between the second output value calculated for each of the multiple age classes and the corresponding age class.
[0056] In the example shown in Figure 8, the estimation processing unit 112 estimates the gender of the face image to be male because the first output value (score value) for males is "0.78" and the second output value (score value) for females is "0.22".
[0057] Furthermore, the estimation processing unit 112 estimates the age of the face image by performing a sum-of-products operation on the ID(i)=0~(N-1)(N is the number of age classes) of the output layer for age and the respective score values (second output values). In this case, the number of age classes is "10", and the sum-of-products operation on the ID(i)=0~9 of the output layer and the respective score values results in "1.20", so the estimation processing unit 112 estimates that the age of the face image belongs to the "4~6 year old" age class. Note that the estimation processing unit 112 estimates the age in the same way whether the gender is estimated to be male or female.
[0058] Here, the estimation processing unit 112 calculates the age corresponding to the second output value by linear interpolation using the minimum and maximum ages among the multiple ages included in the estimated age class. In the example shown in Figure 8, in the estimated age class "4-6 years old", the minimum age is "4 years old" and the maximum age is "6 years old". Therefore, the estimation processing unit 112 calculates the estimated age as 4.4 years old (= 4 + 0.20 × (6 - 4)).
[0059] The output processing unit 113 outputs the estimation result. For example, the output processing unit 113 displays the estimation result ("male", "4-6 years old", or "4.4 years old") estimated for the input face image on the operation display unit 13. The output processing unit 113 may also transmit the estimation result to other devices via the communication unit 14.
[0060] According to the estimation method of this embodiment, the trained model 122 can simultaneously estimate gender and age with a single trained model, similar to the conventional trained model A (see Figure 6), but the total number of connections is "m × (2 + N)", which is less than the total number of connections of trained model A, "2 × N × m". Therefore, the training time is shorter than that of trained model A. In addition, since gender and age can be estimated simultaneously with a single trained model 122, memory can be reduced and performance can be improved. Furthermore, since age can be estimated independently of the gender estimation result, as with the conventional trained model B, estimation accuracy can also be improved.
[0061] [Estimation Processing] The following describes an example of the procedure for the estimation process performed by the control unit 11 of the estimation system 1, with reference to Figure 9.
[0062] Furthermore, the present invention can be understood as an estimation method for performing one or more steps included in the estimation process. The one or more steps included in the estimation process described herein may be omitted as appropriate. Also, the execution order of each step in the estimation process may differ to the extent that similar effects are produced. Moreover, while the case where the control unit 11 executes each step in the estimation process is described here as an example, in other embodiments, one or more processors may distribute and execute each step in the estimation process.
[0063] The estimation system 1 is fitted with, for example, a trained model 122 (see Figure 8) that estimates a person's gender and age. The control unit 11 uses the trained model 122 to perform estimation processing according to the estimation program 121.
[0064] First, in step S1, the control unit 11 determines whether or not it has acquired the face image to be estimated. If the control unit 11 has acquired the face image (S1: Yes), the process proceeds to steps S21 and S22. Step S1 is an example of the acquisition steps of the present invention.
[0065] In step S21, the control unit 11 calculates a first output value (score value) corresponding to gender. Specifically, the control unit 11 calculates a first output value of the first output layer corresponding to male and female, respectively. For example, as shown in Figure 8, the control unit 11 calculates "0.78" as the first output value (score value) for male and "0.22" as the first output value (score value) for female.
[0066] In step S31, following step S21, the control unit 11 estimates the gender of the face image. For example, the control unit 11 estimates the gender of the person with the larger score value as the gender to be estimated. In this case, the control unit 11 estimates the gender of the face image to be male, with a score value of "0.78". After step S31, the process moves on to step S4.
[0067] Meanwhile, in step S22, the control unit 11 calculates a second output value (score value) corresponding to age. Specifically, the control unit 11 calculates a score value (second output value) for each ID(i)=0 to (N-1) of the output layer for age. For example, as shown in Figure 8, the control unit 11 calculates "0.20" as the score value for age class "0", "0.33" as the score value for age class "1", and "0.00" as the score value for age class "9". In this way, the control unit 11 calculates 10 score values (second output values) corresponding to each age class.
[0068] In step S32, following step S22, the control unit 11 estimates the age of the face image. For example, the control unit 11 estimates the age by performing a sum-of-products operation on the ID(i)=0~(N-1) (where N is the number of age classes) of the output layer for age and the respective score values (second output values). Here, the number of age classes is "10", and the sum-of-products operation on the ID(i)=0~9 of the output layer and the respective score values results in "1.20", so the control unit 11 estimates that the face belongs to the "4~6 year old" age class. Alternatively, the control unit 11 may estimate the age as 4.4 years old (=4+0.20×(6-4)) by performing linear interpolation. After step S32, the process moves to step S4.
[0069] Thus, the control unit 11 executes the gender estimation process in steps S21 and S31 and the age estimation process in steps S22 and S32 in parallel and individually using a single trained model 122. Alternatively, the control unit 11 may execute the gender estimation process and the age estimation process together (or simultaneously). Steps S21 and S31, and steps S22 and S32 are examples of estimation steps of the present invention.
[0070] In step S4, the control unit 11 outputs the estimation result. For example, the control unit 11 displays the estimation result ("male", "4-6 years old") estimated for the input face image on the operation display unit 13. The control unit 11 may also transmit the estimation result to other devices via the communication unit 14.
[0071] As described above, the estimation system 1 according to this embodiment acquires an image of the target to be estimated, and uses a single trained model 122 generated based on training data in which the image of the target to be estimated and each of the multiple attributes of the target to be estimated are associated with each other, to estimate the first attribute from the first output value of the first output layer corresponding to the first attribute included in the multiple attributes, and estimate the second attribute from the second output value of the second output layer corresponding to the second attribute included in the multiple attributes.
[0072] In other words, estimation system 1 estimates multiple attributes using a single pre-trained model 122. Estimation system 1 generates a single pre-trained model 122 that holds a configuration and parameters optimized to classify into the attribute information of the training data, based on features extracted from training data consisting of multiple face images and their associated attribute information (gender, age, etc.). By inputting face image data of the same type (face) and resolution as the training data into this pre-trained model 122, the desired attribute information is simultaneously output and estimated. Furthermore, in estimation system 1, the number of outputs of the pre-trained model 122 is the sum of the number of categories for each attribute information. In the case of gender and the number of age classes (N), the number of outputs will be "2 + N".
[0073] Furthermore, the estimation system 1 obtains an output value by performing a sum-of-products operation on the output values belonging to the relevant attribute information. The integer part identifies the classification index of the output belonging to the relevant attribute information, and the decimal part represents the proportion of the attribute in the classification index of that attribute. By multiplying the difference between the upper and lower limits of the classification index of the relevant attribute by the decimal part and adding this result to the lower limit, a more detailed estimate of the classification of that attribute can be obtained.
[0074] Furthermore, this estimate can also be applied to the estimation of gender. In the example in Figure 8, the result of the sum-of-products operation is "0 × 0.78 + 1 × 0.22 = 0.22". Since the lower limit of the corresponding attribute class is "0" and the upper limit is "1", the result is "0 + 0.22 × (1 - 0) = 0.22". Since gender is binary, rounding it to the nearest integer gives "0", which is estimated to be male.
[0075] Furthermore, the trained model 122 is generated using machine learning, which involves taking the training data (see Figure 5) as input images, extracting various features from the training data, and modifying the configuration of the preprocessing and adjusting various parameters (connection weight coefficients of the fully connected layer) so that the first output value of the first output layer for gender and the second output value of the second output layer for age approach the ground truth values (gender, age, and the range (class) to which age corresponds) for the training data of the input image. This is achieved through optimized configuration and estimated parameters.
[0076] Thus, by using a single pre-trained model 122, multiple attributes of the target to be estimated can be estimated together (or simultaneously). Compared to conventional estimation methods (see Figures 6 and 7), it is possible to improve estimation accuracy while reducing the processing load of the estimation. In addition, the training time of the pre-trained model 122 can be shortened.
[0077] The estimation system 1 may include a generation processing unit that generates the trained model 122. Alternatively, the estimation system 1 may obtain the trained model 122 from an information processing unit that generates the trained model 122. For example, the estimation system 1 may download the trained model 122 via a network and store it in the storage unit 12.
[0078] Furthermore, the estimation system 1 may consist of a single information processing device (estimation device) and may be configured to be installed (connected) to other devices. For example, the estimation system 1 may be built into a digital signage display. Alternatively, the estimation system 1 may be built into a store terminal (POS terminal) in a store. [Explanation of Symbols]
[0079] 1: Estimation System 11: Control Unit 12: Storage section 13: Operation display section 14: Communications Department 111: Acquisition Processing Unit 112: Estimation Processing Unit 113: Output Processing Unit 121: Estimation Program 122: Pre-trained model
Claims
1. an acquisition processing unit that acquires a captured image of the estimation target; an estimation processing unit that uses a single trained model generated based on learning data in which the image of the estimation target and each of a plurality of attributes of the estimation target are associated with each other, and estimates, with the captured image acquired by the acquisition processing unit as an input image, a first attribute from a first output value of a first output layer corresponding to a first attribute included in the plurality of attributes, and estimates the second attribute from a second output value of a second output layer corresponding to a second attribute included in the plurality of attributes; An estimation system comprising:
2. the estimation processing unit simultaneously estimates the first attribute and the second attribute. The estimation system of claim 1 .
3. The estimation processing unit calculating the first output value for each of a plurality of classifications of the first attribute, and estimating the first attribute for the input image based on the calculated plurality of first output values; calculating the second output value for each of a plurality of categories of the second attribute, and estimating the second attribute for the input image based on the calculated plurality of second output values; The estimation system according to claim 1 or 2.
4. the estimation processing unit outputs an output value of a total number obtained by adding together the number of categories of the first attribute and the number of categories of the second attribute. The estimation system according to any one of claims 1 to 3.
5. the trained model is generated based on training data in which a face image of a person, a gender of the person, and an age of the person are associated with each other; the acquisition processing unit acquires a face image of a person, the estimation processing unit uses the trained model to estimate the gender of the person to be estimated from the first output value of the first output layer corresponding to gender, with the face image acquired by the acquisition processing unit as an input image, and estimates the age of the person to be estimated from the second output value of the second output layer corresponding to age; The estimation system according to any one of claims 1 to 4.
6. The estimation processing unit outputs an output value of a total number obtained by adding together the number of genders into which the genders are classified and the number of multiple age classes into which the ages are classified. The estimation system according to claim 5 .
7. the estimation processing unit estimates the age of the person to be estimated based on a result of a product-sum operation between the second output value calculated for each of the plurality of age classes and the corresponding age class; The estimation system according to claim 6 .
8. the estimation processing unit calculates the age corresponding to the second output value by linear interpolation using a minimum age and a maximum age among a plurality of ages included in the estimated age class. The estimation system according to claim 7 .
9. one or more processors, an acquisition step of acquiring a captured image of the estimation target; an estimation step of using a single trained model generated based on learning data in which the image of the estimation target and each of the multiple attributes of the estimation target are associated with each other, with the captured image acquired in the acquisition step as an input image, estimating the first attribute from a first output value of a first output layer corresponding to a first attribute included in the multiple attributes, and estimating the second attribute from a second output value of a second output layer corresponding to a second attribute included in the multiple attributes; Estimation method to perform.
10. an acquisition step of acquiring a captured image of the estimation target; an estimation step of using a single trained model generated based on learning data in which the image of the estimation target and each of the multiple attributes of the estimation target are associated with each other, with the captured image acquired in the acquisition step as an input image, estimating the first attribute from a first output value of a first output layer corresponding to a first attribute included in the multiple attributes, and estimating the second attribute from a second output value of a second output layer corresponding to a second attribute included in the multiple attributes; An estimation program for causing one or more processors to execute the above.