Line-of-sight estimation device, line-of-sight estimation method, model generation device, and model generation method

By acquiring correction and ground truth information, and using machine learning-generated models to infer the direction of gaze, the problem of poor accuracy of gaze direction caused by individual differences in the central fovea position is solved, achieving high-precision and high-speed gaze inference.

CN114787861BActive Publication Date: 2026-01-16OMRON CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080085841.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-10
Publication Date
2026-01-16
Estimated Expiration
2040-01-10

AI Technical Summary

Technical Problem

In existing technologies, due to individual differences in the position of the central concave area, the accuracy of predicting the direction of sight is poor, making it difficult to maintain high accuracy across different individuals.

Method used

By acquiring correction information, including feature information and ground truth information, and using a learned inference model generated by machine learning, the direction of gaze is inferred by combining the object image, thus correcting for individual differences.

Benefits of technology

It improves the accuracy of line-of-sight prediction, reduces information processing costs, and enables high-speed line-of-sight prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787861B_ABST
    Figure CN114787861B_ABST
Patent Text Reader

Abstract

An aspect of the present application relates to a line-of-sight estimation device that estimates a line-of-sight direction of a subject using not only a subject image that reflects the eyes of the subject but also correction information that includes feature information related to the line-of-sight of the eyes of the subject observing a prescribed direction and true value information that indicates a true value of the prescribed direction. Thus, in the line-of-sight estimation device, the line-of-sight direction of the subject can be estimated on the basis of consideration of individual differences. Therefore, improvement in estimation accuracy of the line-of-sight direction of the subject can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a line-of-sight estimation device, a line-of-sight estimation method, a model generation device, and a model generation method. BACKGROUND

[0002] In recent years, various technologies are being developed to estimate the line-of-sight direction of a subject. As an example of a method of estimating the line-of-sight direction, the corneal reflection method is known. In the corneal reflection method, a bright spot (Purkinje's spot) is generated on the cornea by light emitted from a light source, and the line-of-sight is estimated from the positional relationship between the generated bright spot and the pupil. According to this method, the line-of-sight direction can be estimated with high accuracy without depending on the orientation of the face or the like. However, in this method, if the bright spot cannot be generated on the cornea, it is difficult to estimate the line-of-sight direction. Therefore, the range in which the line-of-sight direction can be estimated is limited. In addition to this, it is possible to be affected by a change in the position of the head, resulting in a decrease in the estimation accuracy of the line-of-sight direction.

[0003] As another example of a method of estimating the line-of-sight direction, a method using the shape of the pupil is known. In this method, the shape of the eyeball is regarded as a sphere, and the outline of the pupil is regarded as a circle, and the fact that the appearance shape of the pupil becomes elliptical with the movement of the eyeball is utilized. That is, in this method, the pupil shape of the subject reflected in the captured image is fitted, and the line-of-sight direction is estimated from the inclination and the ratio of the major axis to the minor axis of the obtained pupil shape (ellipse). According to this method, since the calculation method is simple, it is possible to reduce the processing cost taken for the estimation of the line-of-sight direction, and to speed up the processing. However, if the pupil shape cannot be accurately acquired, it is possible to cause a decrease in the estimation accuracy of the line-of-sight direction. Therefore, in a case where the resolution of the pupil image in the captured image obtained due to the reason that the head is far from the capturing device, the performance of the capturing device is low, or the like, the fitting of the pupil shape is difficult, and thus it is possible to make the estimation of the line-of-sight direction difficult.

[0004] On the other hand, in Patent Literature 1, a method of estimating the line-of-sight direction using a learned model such as a neural network is proposed. In the method proposed in Patent Literature 1, a partial image reflecting the eye is extracted from a captured image obtained by capturing the face of the subject, and the line-of-sight direction of the subject is estimated from the extracted partial image using the learned model. According to the method proposed in Patent Literature 1, it is expected to realize a system capable of estimating the line-of-sight direction robustly and with high accuracy with respect to a change in the position of the head of the subject or the like.

[0005] PRIOR ART DOCUMENTS

[0006] PATENT LITERATURE

[0007] Patent Literature 1: Japanese Patent Application Publication No. 2019-028843 SUMMARY

[0008] Technical problem to be solved by the invention

[0009] The present inventors have found the following problems in the existing method. That is, it is known that the human retina has a fovea in the center, and the fovea contributes to vision in a high-fineness central visual field. Therefore, the direction of the line of sight of a person can be defined by a line connecting the fovea and the center of the pupil. The position of the fovea differs from person to person. That is, the fovea is not necessarily located at the exact center of the retina, and the position thereof can differ depending on individual differences. It is difficult to determine the position of the fovea of each person from a captured image obtained by a capturing device.

[0010] In the existing method, a model for estimating the direction of the line of sight is constructed based on data obtained from a subject. However, there are individual differences in the position of the fovea between a subject who is the object of estimation of the direction of the line of sight in a use scenario and the subject, and even if the pupil is similarly reflected in a captured image, the direction of the line of sight can differ. Therefore, in the existing method, there is a problem that the estimation accuracy of the direction of the line of sight can be degraded due to the individual differences in the position of the fovea.

[0011] An aspect of the present invention is made in view of the above-described actual circumstances, and aims to provide a technology capable of estimating the direction of the line of sight of a subject with high accuracy.

[0012] Technical solution for solving the problem

[0013] The present invention adopts the following configuration in order to solve the above-described technical problem.

[0014] That is, an aspect of the present invention relates to a line-of-sight estimation device including: an information acquisition unit that acquires correction information including feature information related to a line of sight of an eye of a subject observing a predetermined direction, and true value information indicating a true value of the predetermined direction observed by the eye of the subject; an image acquisition unit that acquires a subject image in which an eye of an existing subject is reflected; an estimation unit that estimates a direction of the line of sight of the subject reflected in the subject image using a completed-estimation model generated by machine learning, the completed-estimation model being trained by the machine learning to output an output value suitable for correct answer information indicating a true value of a direction of the line of sight of a subject reflected in a learning subject image with respect to input of learning correction information and a learning subject image obtained from a subject, the estimation of the direction of the line of sight being constituted by inputting the acquired subject image and the correction information to the completed-estimation model, performing an operation process of the completed-estimation model, and thereby acquiring, from the completed-estimation model, an output value corresponding to an estimation result of the direction of the line of sight of the subject reflected in the subject image; and an output unit that outputs information related to the estimation result of the direction of the line of sight of the subject.

[0015] In this configuration, in order to estimate the line-of-sight direction of the subject, not only the subject image that reflects the subject's eyes but also the correction information that includes the characteristic information and the true value information is used. The characteristic information is related to the line-of-sight of the subject's eyes that observes a prescribed direction. The true value information indicates the true value of the prescribed direction. According to the characteristic information and the true value information, it is possible to grasp the characteristic of the eyes that form the line-of-sight with respect to the known direction (i.e., the individuality of the subject's line-of-sight) from the true value. Therefore, according to this configuration, by further using the correction information in the estimation of the line-of-sight direction, it is possible to correct the difference in the line-of-sight direction caused by the individual difference between the subject and the subject. That is, it is possible to estimate the line-of-sight direction of the subject on the basis of taking into account the individual difference. Therefore, it is possible to achieve the improvement of the estimation accuracy of the line-of-sight direction of the subject.

[0016] In the line-of-sight estimation device according to the above aspect, the correction information can include the characteristic information and the true value information corresponding to a plurality of different prescribed directions, respectively. According to this configuration, since it is possible to more accurately grasp the individuality of the subject's line-of-sight from the correction information with respect to a plurality of different directions, it is possible to achieve further improvement of the estimation accuracy of the line-of-sight direction of the subject.

[0017] In the line-of-sight estimation device according to the above aspect, the characteristic information and the true value information can be constituted by including a correction characteristic amount related to the correction, which is derived by coupling the characteristic information and the true value information. The learned estimation model can have a first extractor and an estimator. The execution of the operation processing of the learned estimation model can be constituted by: inputting the acquired subject image to the first extractor, and executing the operation processing of the first extractor, thereby acquiring an output value corresponding to the first characteristic amount related to the subject image from the first extractor; and inputting the correction characteristic amount and the acquired first characteristic amount to the estimator, and executing the operation processing of the estimator. According to this configuration, it is possible to provide a learned estimation model that can appropriately estimate the line-of-sight direction of the subject from the subject image and the correction information. In addition, according to this configuration, by reducing the amount of information of the correction information, it is possible to reduce the cost of the information processing for estimating the line-of-sight direction of the subject, and thus it is possible to achieve the high speed of the information processing.

[0018] In the line-of-sight estimation device according to the above aspect, the feature information can be constituted by a second feature amount related to a reference image that represents an eye of the subject who is looking in the prescribed direction. The information acquisition unit can also have a coupler. The acquisition of the correction information can be constituted by: acquiring the second feature amount; acquiring the true value information; and inputting the acquired second feature amount and the true value information to the coupler and executing an operation process of the coupler, thereby acquiring an output value corresponding to the correction feature amount from the coupler. According to this constitution, the operation process for deriving the correction feature amount is not executed in the line-of-sight direction estimation process, but is executed in the acquisition process of the correction information. Therefore, it is possible to suppress the processing cost of the estimation process. In particular, in a mode in which the acquisition process of the subject image and the line-of-sight direction estimation process are repeatedly executed, if the derivation of the correction feature amount is completed, in the repeated operation, the already derived correction feature amount can be repeatedly used, and the execution of the acquisition process of the correction information can be omitted. Therefore, it is possible to reduce the cost of the series of operation processes, and thus, it is possible to achieve the high speed of the series of operation processes.

[0019] In the line-of-sight estimation device according to the above aspect, the information acquisition unit can also have a second extractor. The acquisition of the second feature amount can be constituted by: acquiring the reference image; and inputting the acquired reference image to the second extractor and executing an operation process of the second extractor, thereby acquiring an output value corresponding to the second feature amount from the second extractor. According to this constitution, it is possible to appropriately acquire the feature information (second feature amount) that represents the feature of the line of sight of the eye of the subject who is looking in the prescribed direction.

[0020] In the line-of-sight estimation device according to the above aspect, the learned estimation model can have a first extractor and an estimator. The execution of the operation process of the learned estimation model can be constituted by: inputting the acquired subject image to the first extractor and executing an operation process of the first extractor, thereby acquiring an output value corresponding to a first feature amount related to the subject image from the first extractor; and inputting the feature information, the true value information, and the acquired first feature amount to the estimator and executing an operation process of the estimator. According to this constitution, it is possible to provide a learned estimation model that can appropriately estimate the line-of-sight direction of the subject from the subject image and the correction information.

[0021] In the line-of-sight estimation device according to the above aspect, the feature information can be constituted by a second feature quantity related to a reference image that represents an eye of the subject who is looking in the prescribed direction. The information acquisition section can have a second extractor. The acquisition of the correction information can be constituted by: acquiring the reference image; inputting the acquired reference image to the second extractor and executing an operation process of the second extractor, thereby acquiring an output value corresponding to the second feature quantity from the second extractor; and acquiring the true value information. According to this constitution, the feature information (second feature quantity) that represents the feature of the line-of-sight of the eye of the subject who is looking in the prescribed direction can be appropriately acquired. In addition, in a mode in which the acquisition process of the subject image and the estimation process of the line-of-sight direction are repeatedly executed, if the derivation of the second feature quantity is completed, in the repeated operation, the already derived second feature quantity can be repeatedly used, and the execution of the acquisition process of the correction information can be omitted. Therefore, the cost of the series of operation processes for specifying the line-of-sight direction of the subject can be reduced, and thus the speedup of the series of operation processes can be achieved.

[0022] In the line-of-sight estimation device according to the above aspect, the feature information can be constituted by a reference image that represents an eye of the subject who is looking in the prescribed direction. The learned estimation model can have a first extractor, a second extractor, and an estimator. The execution of the operation process of the learned estimation model can be constituted by: inputting the acquired subject image to the first extractor and executing an operation process of the first extractor, thereby acquiring an output value corresponding to a first feature quantity related to the subject image from the first extractor; inputting the reference image to the second extractor and executing an operation process of the second extractor, thereby acquiring an output value corresponding to a second feature quantity related to the reference image from the second extractor; and inputting the acquired first feature quantity, the acquired second feature quantity, and the true value information to the estimator and executing an operation process of the estimator. According to this constitution, the learned estimation model that can appropriately estimate the line-of-sight direction of the subject from the subject image and the correction information can be provided.

[0023] In the line-of-sight estimation device according to the above aspect, the learned estimation model can include a first converter and an estimator. The learned estimation model can be configured to execute the operation processing by inputting the acquired object image to the first converter, executing the operation processing of the first converter, and thereby acquiring an output value corresponding to a first heat map related to the line-of-sight direction of the object person from the first converter, and inputting the acquired first heat map, the feature information, and the true value information to the estimator, and executing the operation processing of the estimator. According to this configuration, it is possible to provide a learned estimation model capable of appropriately estimating the line-of-sight direction of the object person from the object image and the correction information.

[0024] In the line-of-sight estimation device according to the above aspect, the feature information can be constituted by a second heat map related to the line-of-sight direction of the eye observing the predetermined direction, the second heat map being derived from a reference image of the eye of the object person observing the predetermined direction. The information acquisition unit can include a second converter. The acquisition of the correction information can be configured by acquiring the reference image, inputting the acquired reference image to the second converter, executing the operation processing of the second converter, and thereby acquiring an output value corresponding to the second heat map from the second converter, acquiring the true value information, and converting a third heat map related to the true value of the predetermined direction into the true value information. The input of the first heat map, the feature information, and the true value information to the estimator can be configured by inputting the first heat map, the second heat map, and the third heat map to the estimator. According to this configuration, by adopting a common heat map form as the data form on the input side, it is possible to make the configuration of the estimator relatively simple, and by easily integrating each information (the feature information, the true value information, and the object image) in the estimator, it is possible to improve the estimation accuracy of the estimator.

[0025] In the line-of-sight estimation device according to the above aspect, the acquisition of the object image and the estimation of the line-of-sight direction of the object person by the estimation unit can be repeatedly executed by the image acquisition unit. According to this configuration, it is possible to continuously perform the estimation of the line-of-sight direction of the object person.

[0026] In the line-of-sight estimation device according to the above aspect, the information acquisition unit can acquire the correction information by observing the line-of-sight of the object person using a sensor after outputting an instruction to the object person to observe a predetermined direction. According to this configuration, it is possible to appropriately and simply acquire the correction information representing the individuality of the line-of-sight of the object person.

[0027] An aspect of the present application can also be a device that generates a learned prediction model usable in the line-of-sight prediction device of each of the above-described modes. For example, an aspect of the present application relates to a model generation device including: a first acquisition unit that acquires learning correction information including learning feature information related to a line of sight of a subject's eye observing a predetermined direction and learning true value information indicating a true value of the predetermined direction observed by the subject's eye; a second acquisition unit that acquires a plurality of learning data sets each composed of a combination of a learning object image in which the subject's eye is reflected and correct answer information indicating a true value of a line-of-sight direction of the subject reflected in the learning object image; and a machine learning unit that performs machine learning of a prediction model using the plurality of learning data sets acquired, the performance of the machine learning being constituted by training the prediction model so as to output an output value suitable for the corresponding correct answer information with respect to an input of the learning object image and the learning correction information for each of the learning data sets.

[0028] As other modes of the line-of-sight prediction device and the model generation device of each of the above-described modes, an aspect of the present application can be an information processing method, a program, or a computer-readable storage medium storing such a program. Here, the computer-readable storage medium refers to a medium that stores information such as a program by electrical, magnetic, optical, mechanical, or chemical action. In addition, a line-of-sight prediction system according to an aspect of the present application can be constituted by the line-of-sight prediction device and the model generation device of any of the above-described modes.

[0029] For example, an aspect of the present application relates to a line-of-sight estimation method executed by a computer, the method including: acquiring correction information including feature information related to a line of sight of an eye of a subject observing a predetermined direction and true value information indicating a true value of the predetermined direction observed by the eye of the subject; acquiring a subject image in which the eye of the existing subject is reflected; estimating a line-of-sight direction of the subject reflected in the subject image using a learned estimation model generated by machine learning, the learned estimation model being trained by the machine learning to output an output value suitable for correct answer information with respect to input of learning correction information and a learning subject image, the learning correction information being of the same kind as the correction information, the learning subject image being of the same kind as the subject image, the correct answer information indicating a true value of a line-of-sight direction of a subject reflected in the learning subject image, the estimation of the line-of-sight direction being constituted by inputting the acquired subject image and the correction information into the learned estimation model, performing an operation process of the learned estimation model, and acquiring an output value corresponding to an estimation result of the line-of-sight direction of the subject reflected in the subject image from the learned estimation model; and outputting information related to the estimation result of the line-of-sight direction of the subject.

[0030] Further, for example, an aspect of the present application relates to a model generation method executed by a computer, the method including: acquiring learning correction information including learning feature information related to a line of sight of an eye of a subject observing a predetermined direction and learning true value information indicating a true value of the predetermined direction observed by the eye of the subject; acquiring a plurality of learning data sets each constituted by a combination of a learning subject image in which the eye of the existing subject is reflected and correct answer information indicating a true value of a line-of-sight direction of the subject reflected in the learning subject image; and implementing machine learning of an estimation model using the acquired plurality of learning data sets, the implementation of the machine learning being constituted by training the estimation model with respect to each of the learning data sets in a manner that outputs an output value suitable for corresponding correct answer information with respect to input of the learning subject image and the learning correction information.

[0031] (EFFECT OF INVENTION)

[0032] According to the present application, a line-of-sight direction of a subject can be estimated with high precision. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 An example of a scene in which the present application is applied is schematically illustrated.

[0034] Figure 2 An example of a hardware configuration of a model generation device to which the embodiment relates is schematically illustrated.

[0035] Figure 3 An example of a hardware configuration of the line-of-sight estimation device according to the embodiment is schematically illustrated.

[0036] Figure 4A An example of a software configuration of the model generation device according to the embodiment is schematically illustrated.

[0037] Figure 4B An example of a software configuration of the model generation device according to the embodiment is schematically illustrated.

[0038] Figure 5A An example of a software configuration of the line-of-sight estimation device according to the embodiment is schematically illustrated.

[0039] Figure 5B An example of a software configuration of the line-of-sight estimation device according to the embodiment is schematically illustrated.

[0040] Figure 6 An example of a processing procedure of the model generation device according to the embodiment is shown.

[0041] Figure 7 An example of a processing procedure of the line-of-sight estimation device according to the embodiment is shown.

[0042] Figure 8 An example of a scenario of acquiring correction information according to the embodiment is schematically illustrated.

[0043] Figure 9 An example of a software configuration of the model generation device according to the modified example is schematically illustrated.

[0044] Figure 10 An example of a software configuration of the line-of-sight estimation device according to the modified example is schematically illustrated.

[0045] Figure 11 An example of a software configuration of the model generation device according to the modified example is schematically illustrated.

[0046] Figure 12 An example of a software configuration of the line-of-sight estimation device according to the modified example is schematically illustrated.

[0047] Figure 13A An example of a software configuration of the model generation device according to the modified example is schematically illustrated.

[0048] Figure 13B An example of a software configuration of the model generation device according to the modified example is schematically illustrated.

[0049] Figure 14 An example of a software configuration of the line-of-sight estimation device according to the modified example is schematically illustrated. Detailed Implementation

[0050] Hereinafter, an embodiment of the present invention (hereinafter also referred to as "this embodiment") will be described with reference to the accompanying drawings. However, the embodiments described below are merely illustrative of the present invention in all respects. Of course, various modifications and variations can be made without departing from the scope of the present invention. That is, when implementing the present invention, specific configurations conforming to the embodiments may be appropriately adopted. It should be noted that natural language is used to describe the data in this embodiment; however, more specifically, computer-recognizable analog languages, instructions, parameters, machine language, etc., are used for specification.

[0051] §1 Applicable Examples

[0052] Figure 1 An example of a scenario in which the present invention is applicable is illustrated schematically. For example... Figure 1 As shown, the line-of-sight estimation system 100 according to this embodiment includes a model generation device 1 and a line-of-sight estimation device 2.

[0053] The model generation apparatus 1 according to this embodiment is a computer configured to generate a learned prediction model 3 capable of predicting the gaze direction of a subject. Specifically, the model generation apparatus 1 according to this embodiment acquires learning correction information 50, which includes learning feature information and learning truth information. The learning feature information is related to the gaze of a subject observing a predetermined direction. The learning truth information represents the truth value of the predetermined direction observed by the subject's eyes. The predetermined direction is the gaze direction known based on the truth value. The specific value of the predetermined direction is not particularly limited and can be appropriately selected according to the embodiment. As an example of the predetermined direction, it is preferable to select a direction that is likely to appear in the scenario of predicting the gaze direction of a subject.

[0054] Furthermore, the model generation apparatus 1 according to this embodiment acquires multiple learning datasets 51, each of which is composed of a learning object image 53 reflecting the eyes of a subject and corrective information 55. The corrective information 55 represents the ground truth of the subject's gaze direction reflected in the learning object image 53. The multiple learning datasets 51 may also include learning datasets obtained from subjects observing in a specified direction, similar to the learning correction information 50. Furthermore, "for learning" refers to something used for machine learning. The description of "for learning" can be omitted.

[0055] Further, the present embodiment involves a model generation device 1 that implements machine learning of a prediction model 3 using a plurality of learning data sets 51 acquired. The implementation of the machine learning is constituted by training the prediction model 3 in a manner that outputs an output value suitable for corresponding correct answer information 55 with respect to an input of a learning target image 53 and learning correction information 50 for each learning data set 51. Thereby, it is possible to generate a learned prediction model 3 that acquires an ability to predict a line-of-sight direction of a subject in a target image from correction information and the target image. Further, "learned" can also be referred to as "trained".

[0056] On the other hand, the line-of-sight prediction device 2 is a computer configured to predict a line-of-sight direction of a target person R using the learned prediction model 3 generated. Specifically, the line-of-sight prediction device 2 according to the present embodiment acquires correction information 60 including feature information and true value information with respect to the target person R. The correction information 60 is data of the same kind as the above-mentioned learning correction information 50 obtained from the subject. The target person R can be the same person as the subject, or can be a different person.

[0057] The feature information is related to the line-of-sight of the eyes of the target person R observing a prescribed direction. The feature information can include components related to the features of the eyes forming the line-of-sight of the prescribed direction, and the form thereof can not be particularly limited and can be appropriately determined according to the embodiment. For example, the feature information can be constituted by a reference image that reflects the eyes of the target person observing the prescribed direction. Alternatively, the feature information can be constituted by a feature amount of the line-of-sight extracted from the reference image. The feature information is data of the same kind as the above-mentioned learning feature information.

[0058] The true value information indicates the true value of the prescribed direction observed by the eyes of the target person R. The data form of the true value, that is, the form of the representation of the line-of-sight direction, can be any form indicating information related to the line-of-sight direction, and can not be particularly limited and can be appropriately selected according to the embodiment. The line-of-sight direction can be represented by, for example, an angle such as a pitch angle and an azimuth angle. Alternatively, the line-of-sight direction can be represented by a position of a fixation (hereinafter, also referred to as "fixation position") within a field of view. The angle or the fixation position can be directly represented by a numerical value, or can be represented by a degree or a probability using a heat map. The true value information is data of the same kind as the above-mentioned learning true value information.

[0059] The correction information 60 including the feature information and the true value information can be constituted by directly including the feature information and the true value information as different data (for example, in a form capable of being separated), or can be constituted by including information (for example, a correction feature amount described later) derived by coupling the feature information and the true value information. An example of the constitution of the correction information 60 will be described later.

[0060] Further, the gaze estimation device 2 according to the present embodiment acquires an object image 63 that represents the eyes of the existing subject R. In the present embodiment, the gaze estimation device 2 is connected to the camera S, and can acquire the object image 63 from the camera S. The object image 63 can be any image that can include an image of the eyes of the subject R. For example, the object image 63 can be an image as obtained by the camera S, or can be a partial image extracted from the obtained image. The partial image can be obtained by extracting a range that represents at least one of the eyes from the image obtained by the camera S, for example. The extraction of the partial image can use known image processing.

[0061] Next, the gaze estimation device 2 according to the present embodiment estimates the gaze direction of the eyes of the subject R represented in the object image 63 using the learned estimation model 3 generated by the above-described machine learning. The estimation of the gaze direction is constituted by inputting the acquired object image 63 and the correction information 60 to the learned estimation model 3, and executing the operation processing of the learned estimation model 3, thereby acquiring an output value corresponding to the estimation result of the gaze direction of the eyes of the subject R represented in the object image 63 from the learned estimation model 3. Further, the gaze estimation device 2 according to the present embodiment outputs information related to the result of estimating the gaze direction of the subject R.

[0062] As described above, in the present embodiment, in order to estimate the gaze direction of the subject R, not only the object image 63 that represents the eyes of the subject R, but also the correction information 60 that includes the feature information and the true value information are used. According to the feature information and the true value information, it is possible to grasp the feature of the eyes that form the gaze direction for the known direction (i.e., the individuality of the gaze direction of the subject R) from the true value. Therefore, according to the present embodiment, by further using the correction information 60 in the estimation of the gaze direction, it is possible to correct the difference in the gaze direction caused by the individual difference between the subject and the subject R. That is, it is possible to estimate the gaze direction of the subject R on the basis of taking into account the individual difference. Therefore, according to the present embodiment, in the gaze estimation device 2, it is expected to improve the accuracy of estimating the gaze direction of the subject R. Further, according to the model generation device 1 according to the present embodiment, it is possible to generate the learned estimation model 3 that can estimate the gaze direction of the subject R with such high accuracy.

[0063] The present embodiment can be applied to all scenes of estimating the gaze direction of the subject R. As an example of the scene of estimating the gaze direction, for example, there can be mentioned a scene of estimating the gaze direction of a driver of a vehicle, a scene of estimating the gaze direction of a user who communicates with a robot device, a scene of estimating the gaze direction of a user in a user interface and using the obtained estimation result for input, and the like. The driver and the user are examples of the subject R. The estimation result of the gaze direction can be appropriately used according to each scene.

[0064] Further, in Figure 1 example, the model generation device 1 and the line-of-sight estimation device 2 are connected to each other via a network. The kind of the network can be appropriately selected from, for example, the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, and the like. However, the method of exchanging data between the model generation device 1 and the line-of-sight estimation device 2 can not be limited to such an example, and can be appropriately selected according to the embodiment. For example, between the model generation device 1 and the line-of-sight estimation device 2, data can be exchanged by using a storage medium.

[0065] Further, in Figure 1 example, the model generation device 1 and the line-of-sight estimation device 2 are each constituted by a different computer. However, the constitution of the line-of-sight estimation system 100 according to the embodiment can not be limited to such an example, and can be appropriately determined according to the embodiment. For example, the model generation device 1 and the line-of-sight estimation device 2 can be an integrated computer. Further, for example, at least one of the model generation device 1 and the line-of-sight estimation device 2 can be constituted by a plurality of computers.

[0066] §2 Constitution Example

[0067] [Hardware Constitution]

[0068] < Model Generation Device >

[0069] Figure 2 An example of the hardware constitution of the model generation device 1 according to the embodiment is schematically illustrated. As Figure 2 indicated, the model generation device 1 according to the embodiment is a computer electrically connected to a control section 11, a storage section 12, a communication interface 13, an external interface 14, an input device 15, an output device 16, and a driver 17. Further, in Figure 2 , the communication interface and the external interface are described as "communication I / F" and "external I / F".

[0070] The control section 11 includes a CPU (Central Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), and the like as a hardware processor, and is constituted to perform information processing according to a program and various data. The storage section 12 is an example of a storage, and is constituted by, for example, a hard disk drive, a solid state drive, and the like. In the embodiment, the storage section 12 stores various information such as a model generation program 81, a plurality of data sets 120, learning result data 125, and the like.

[0071] The model generation program 81 is a program for causing the model generation device 1 to perform the following information processing for generating the learned estimation model 3 by implementing machine learning.Figure 6 The model generation program 81 includes a series of commands of the information processing. Each of the data sets 120 is constituted by a combination of the learning image 121 and the correct answer information 123. The learning result data 125 indicates information related to the completed learning estimation model 3 generated by machine learning. In the present embodiment, the learning result data 125 is generated as a result of executing the model generation program 81. Details will be described later.

[0072] The communication interface 13 is, for example, a wired LAN (Local Area Network) module, a wireless LAN module, or the like, and is an interface for performing wired or wireless communication via a network. The model generation device 1 can also perform data communication via a network between the communication interface 13 and another information processing device. The external interface 14 is, for example, a USB (Universal Serial Bus) port, a dedicated port, or the like, and is an interface for connecting with an external device. The kind and the number of the external interface 14 can be arbitrarily selected. The model generation device 1 can be connected with a camera for obtaining the learning image 121 via at least one of the communication interface 13 and the external interface 14.

[0073] The input device 15 is, for example, a mouse, a keyboard, or the like, and is a device for performing input. In addition, the output device 16 is, for example, a display, a speaker, or the like, and is a device for performing output. An operator such as a user can operate the model generation device 1 using the input device 15 and the output device 16.

[0074] The driver 17 is, for example, a CD driver, a DVD driver, or the like, and is a drive device for reading various information such as a program stored in the storage medium 91. The storage medium 91 is a medium that stores various information such as a program by electric, magnetic, optical, mechanical, or chemical action in a manner that a computer and other devices, machines, or the like can read the stored program. At least any one of the above-described model generation program 81 and the plurality of data sets 120 can also be stored in the storage medium 91. The model generation device 1 can also acquire at least any one of the above-described model generation program 81 and the plurality of data sets 120 from the storage medium 91. Furthermore, in the present embodiment, the model generation program 81 and the plurality of data sets 120 are stored in the storage medium 91. However, the model generation program 81 and the plurality of data sets 120 can also be stored in the storage medium 91 and the storage device 18. Figure 2 In the present embodiment, a disk-type storage medium such as a CD, a DVD, or the like is exemplified as an example of the storage medium 91. However, the kind of the storage medium 91 is not limited to the disk-type, and can be a type other than the disk-type. As the storage medium other than the disk-type, for example, a semiconductor memory such as a flash memory can be exemplified. The kind of the driver 17 can be arbitrarily selected according to the kind of the storage medium 91.

[0075] Further, regarding the specific hardware configuration of the model generation device 1, the constituent elements can be omitted, replaced, and added as appropriate in accordance with the embodiment. For example, the control section 11 can also include a plurality of hardware processors. The hardware processors can be constituted by a microprocessor, an FPGA (field-programmable gate array), a DSP (digital signal processor), or the like. The storage section 12 can also be constituted by a RAM and a ROM included in the control section 11. At least any one of the communication interface 13, the external interface 14, the input device 15, the output device 16, and the drive 17 can also be omitted. The model generation device 1 can also be constituted by a plurality of computers. In this case, the hardware configurations of the respective computers can be identical or different. In addition, the model generation device 1 can be a general-purpose server device, a PC (Personal Computer), or the like, in addition to being designed as an information processing device dedicated to the service provided.

[0076] <Line-of-sight Estimation Device>

[0077] Figure 3 An example of the hardware configuration of the line-of-sight estimation device 2 related to the present embodiment is schematically illustrated. As shown in FIG. 2, the line-of-sight estimation device 2 related to the present embodiment is a computer electrically connected with a control section 21, a storage section 22, a communication interface 23, an external interface 24, an input device 25, an output device 26, and a drive 27. Figure 3

[0078] The control section 21 to the drive 27 and the storage medium 92 of the line-of-sight estimation device 2 can each be configured similarly to the respective control section 11 to the drive 17 and the storage medium 91 of the model generation device 1 described above. The control section 21 is configured to include a CPU, a RAM, a ROM, and the like as hardware processors, and to perform various information processing in accordance with programs and data. The storage section 22 is constituted by, for example, a hard disk drive, a solid state drive, or the like. In the present embodiment, the storage section 22 stores various information such as the line-of-sight estimation program 82, the correction information 60, and the learning result data 125.

[0079] The line-of-sight estimation program 82 is a program for causing the line-of-sight estimation device 2 to perform the information processing (described later) for estimating the direction of the line of sight of the subject person R visualized in the subject image 63 using the learned estimation model 3. Figure 7 The line-of-sight estimation program 82 contains a series of commands for the information processing. At least any one of the line-of-sight estimation program 82, the correction information 60, and the learning result data 125 can also be stored in the storage medium 92. In addition, the line-of-sight estimation device 2 can acquire at least any one of the line-of-sight estimation program 82, the correction information 60, and the learning result data 125 from the storage medium 92.​

[0080] In addition, in Figure 3 In the example, the line-of-sight estimation device 2 is connected with the camera S (imaging device) via the external interface 24. Thereby, the line-of-sight estimation device 2 can acquire the subject image 63 from the camera S. However, the connection method with the camera S can not be limited to such an example, and can be appropriately selected according to the embodiment. In a case where the camera S is provided with a communication interface, the line-of-sight estimation device 2 can also be connected with the camera S via the communication interface 23. The kind of the camera S can be appropriately selected according to the embodiment. The camera S can be, for example, a general RGB camera, a depth camera, an infrared camera, or the like. The camera S can be appropriately configured to image the eye of the subject R.

[0081] Further, regarding the detailed hardware configuration of the line-of-sight estimation device 2, the constituent elements can be appropriately omitted, replaced, and added according to the embodiment. For example, the control section 21 can also include a plurality of hardware processors. The hardware processor can be constituted by a microprocessor, an FPGA, a DSP, or the like. The storage section 22 can also be constituted by a RAM and a ROM included in the control section 21. At least any one of the communication interface 23, the external interface 24, the input device 25, the output device 26, and the driver 27 can also be omitted. The line-of-sight estimation device 2 can also be constituted by a plurality of computers. In this case, the hardware configurations of the respective computers can be consistent or inconsistent. In addition, the line-of-sight estimation device 2 can be a general server device, a general PC, a PLC (programmable logic controller), or the like, in addition to being designed as an information processing device dedicated to the provided service.

[0082] [Software Configuration]

[0083] < Model Generation Device >

[0084] Figure 4A and Figure 4B An example of the software configuration of the model generation device 1 according to the embodiment is schematically illustrated. The control section 11 of the model generation device 1 loads the model generation program 81 stored in the storage section 12 into a RAM. Then, the control section 11 controls the respective constituent elements by interpreting and executing the commands included in the model generation program 81 loaded into the RAM by the CPU, thereby causing the model generation device 1 to function as a computer provided with a collection section 111, a first acquisition section 112, a second acquisition section 113, a machine learning section 114, and a saving processing section 115 as software modules. That is, in the embodiment, the respective software modules of the model generation device 1 are realized by the control section 11 (CPU). Figure 4A Figure 4B and

[0085] ​The collection section 111 acquires a plurality of data sets 120. Each data set 120 is constituted by a combination of a learning image 121 that visualizes the eyes of a subject and correct answer information 123. The correct answer information 123 indicates a true value of the line of sight direction of the subject visualized in the corresponding learning image 121. The first acquisition section 112 acquires learning correction information 50 that includes learning feature information 502 and learning true value information 503. The learning feature information 502 is related to the line of sight of the eyes of the subject who observes a prescribed direction. The learning true value information 503 indicates, for the corresponding learning feature information 502, a true value of the prescribed direction observed by the eyes of the subject. In the present embodiment, the acquisition of the learning correction information 50 can utilize a data set 120 obtained for a subject who observes a prescribed direction. The learning correction information 50 can also include learning feature information 502 and learning true value information 503 that respectively correspond to a plurality of different prescribed directions. That is, a plurality of prescribed directions for grasping the individuality of the line of sight of a person can be set, and the correction information can include feature information and true value information regarding each of the set prescribed directions.

[0086] The second acquisition section 113 acquires a plurality of learning data sets 51 that are respectively constituted by a combination of a learning target image 53 that visualizes the eyes of a subject and correct answer information 55 that indicates a true value of the line of sight direction of the subject visualized in the learning target image 53. In the present embodiment, each of the above-described data sets 120 can be used as a learning data set 51. That is, the above-described learning image 121 can be used as a learning target image 53, and the above-described correct answer information 123 can be used as correct answer information 55. The machine learning section 114 performs machine learning of the estimation model 3 using the acquired plurality of learning data sets 51. The performance of the machine learning is constituted by training the estimation model 3 in such a manner that, for each learning data set 51, an output value that is appropriate for the corresponding correct answer information 55 is output in relation to the input of the learning target image 53 and the learning correction information 50.

[0087] The constitution of the estimation model 3 can be appropriately decided according to the embodiment as long as it can perform an operation for estimating the line of sight direction of a person from correction information and a target image, and can not be particularly limited. In addition, the data form of the correction information can be appropriately decided according to the embodiment as long as it includes a component related to feature information and true value information (i.e., a component related to the features of the eyes that form a known direction of the line of sight), and can not be particularly limited. The steps of the machine learning can be appropriately decided according to the constitution of the estimation model 3 and the correction information.

[0088] As Figure 4BAs shown, in the present embodiment, the estimation model 3 is provided with an extractor 31 and an estimator 32. The extractor 31 is an example of the first extractor. In the present embodiment, the feature information and the true value information are constituted by including a correction feature quantity related to correction that is derived by coupling the feature information and the true value information. That is, the correction information is constituted by the correction feature quantity. The coupling can be simply forming the information into one, or can include forming the information into one and compressing the information. In the present embodiment, in order to acquire the correction feature quantity, an extractor 35 and a coupler 36 are used. The extractor 35 is an example of the second extractor.

[0089] The extractor 31 is configured to accept input of an image (a target image) that represents an eye of an existing person, and output an output value corresponding to a feature quantity related to the input image. In other words, the extractor 31 is configured to extract a feature quantity from an image that represents an eye of an existing person. The estimator 32 is configured to accept input of a feature quantity calculated by the extractor 31 and a correction feature quantity, and output an output value corresponding to a result of estimation of a line-of-sight direction of a person represented in a corresponding image (i.e., an image input to the extractor 31 in order to obtain the input feature quantity). In other words, the estimator 32 is configured to estimate a line-of-sight direction of a person from a feature quantity of an image and a correction feature quantity. The output of the extractor 31 is connected to the input of the estimator 32.

[0090] The extractor 35 is configured to accept input of an image that represents an eye of an existing person, and output an output value corresponding to a feature quantity related to the input image, similarly to the extractor 31. The extractor 35 can use an extractor common to the extractor 31 (i.e., the extractor 35 is the same as the extractor 31), or can use an extractor different from the extractor 31 (i.e., the extractor 35 is different from the extractor 31). The coupler 36 is configured to accept input of feature information and true value information, and output an output value corresponding to a correction feature quantity related to correction that is derived by coupling the input feature information and the true value information. In the present embodiment, the feature information is constituted by a feature quantity related to a reference image that represents an eye of a person (a target person) who observes a predetermined direction. By providing the reference image to the extractor 35 and performing operation processing of the extractor 35, an output value corresponding to a feature quantity of the reference image can be obtained from the extractor 35. Therefore, the output of the extractor 35 is connected to the input of the coupler 36. The data form of each feature quantity can not be particularly limited, and can be appropriately selected according to the embodiment.

[0091] As Figure 4AAs shown, the machine learning section 114 first prepares the learning model 4 having the extractor 41 and the estimator 43 in order to generate a trained extractor that can be used as each of the extractors (31, 35). The extractor 41 corresponds to each of the extractors (31, 35). The output of the extractor 41 is connected to the input of the estimator 43. The estimator 43 is configured to accept the input of the feature quantity calculated by the extractor 41 and output an output value corresponding to the estimation result of the gaze direction of the person appearing in the corresponding image (i.e., the image input to the extractor 41 in order to obtain the input feature quantity).

[0092] The machine learning section 114 performs machine learning of the learning model 4 using the acquired plurality of data sets 120. The machine learning section 114 inputs the learning image 121 included in each data set 120 to the extractor 41 and executes the operation processing of the extractor 41 and the estimator 43. Through this operation processing, the machine learning section 114 acquires, from the estimator 43, an output value corresponding to the estimation result of the gaze direction of the subject appearing in the learning image 121. In the machine learning of the learning model 4, the machine learning section 114 trains the learning model 4 in such a manner that the output value obtained from the estimator 43 through the operation processing is adapted to the correct answer information 123 for each data set 120. As a result of this machine learning, a component related to the eyes of the subject included in the learning image 121 is included in the output (i.e., the feature quantity) of the trained extractor 41 so that the gaze direction of the subject can be estimated in the estimator 43.

[0093] In the case where the same extractor is used for each of the extractors (31, 35), the trained extractor 41 generated through machine learning can be commonly used for each of the extractors (31, 35). In this case, the amount of information of each of the extractors (31, 35) can be reduced, and the cost of machine learning can be suppressed. On the other hand, in the case where different extractors are used for each of the extractors (31, 35), the machine learning section 114 can also prepare different learning models 4 at least for part of the extractors 41, perform respective machine learning. Furthermore, the trained extractors 41 generated through the respective machine learning can be used as each of the extractors (31, 35). Each of the extractors (31, 35) can directly use the trained extractors 41, or can use a copy of the trained extractors 41. Similarly, in the case where a plurality of predetermined directions are set, the extractor 35 can be prepared for each of the different directions set respectively, or can be prepared commonly for the plurality of different directions set. In the case where the extractor 35 is prepared commonly for the plurality of different directions, the amount of information of the extractor 35 can be reduced, and the cost of machine learning can be suppressed.

[0094] Next, as shown in FIG. 4, the machine learning section 114 performs machine learning of the learning model 4 using the plurality of data sets 120 acquired in advance. The machine learning section 114 trains the learning model 4 in such a manner that the output value obtained from the estimator 43 through the operation processing is adapted to the correct answer information 123 for each data set 120. Figure 4BAs shown, the machine learning unit 114 prepares the learning model 30 including the extractor 35, the coupler 36, and the inference model 3. In the present embodiment, the machine learning unit 114 implements machine learning of the learning model 30 in a manner that eventually acquires the ability of the inferrer 32 of the inference model 3 to infer the gaze direction of a person. During the machine learning of the learning model 30, the output of the coupler 36 is connected to the input of the inferrer 32. Thus, in the machine learning of the learning model 30, both the inferrer 32 and the coupler 36 are trained.

[0095] During the machine learning, the first acquisition unit 112 acquires the learning correction information 50 using the extractor 35 and the coupler 36. Specifically, the first acquisition unit 112 acquires a learning reference image 501 that represents the eye of a subject who is looking in a prescribed direction and learning true value information 503 that represents the true value of the prescribed direction. The first acquisition unit 112 can also acquire a learning image 121 included in the data set 120 obtained for the subject who is looking in the prescribed direction as the learning reference image 501 and acquire the correct answer information 123 as the learning true value information 503.

[0096] The first acquisition unit 112 inputs the acquired learning reference image 501 to the extractor 35 and executes the operation processing of the extractor 35. Thus, the first acquisition unit 112 acquires, from the extractor 35, an output value corresponding to the feature amount 5021 related to the learning reference image 501. In the present embodiment, the learning feature information 502 is constituted by this feature amount 5021.

[0097] Next, the first acquisition unit 112 inputs the acquired feature amount 5021 and the learning true value information 503 to the coupler 36 and executes the operation processing of the coupler 36. Thus, the first acquisition unit 112 acquires, from the coupler 36, an output value corresponding to the correction-related feature amount 504 derived by the coupling of the learning feature information 502 and the learning true value information 503. The feature amount 504 is an example of a learning correction feature amount. In the present embodiment, the learning correction information 50 is constituted by this feature amount 504. The first acquisition unit 112 can acquire the learning correction information 50 by these operation processes and using the extractor 35 and the coupler 36.

[0098] Further, in a case where a plurality of prescribed directions are set, the first acquisition unit 112 can also acquire the learning reference image 501 and the learning true value information 503 for each of the plurality of different prescribed directions. The first acquisition unit 112 can also input each learning reference image 501 to the extractor 35 and execute the operation processing of the extractor 35. Thereby, the first acquisition unit 112 can also acquire each feature quantity 5021 from the extractor 35. Next, the first acquisition unit 112 can also input each acquired feature quantity 5021 and the learning true value information 503 for each prescribed direction to the coupler 36 and execute the operation processing of the coupler 36. Through these operation processes, the first acquisition unit 112 can also acquire a feature quantity 504 derived by coupling the learning feature information 502 and the learning true value information 503 for each of the plurality of different prescribed directions. In this case, the feature quantity 504 can include information that aggregates the learning feature information 502 and the learning true value information 503 for each of the plurality of different prescribed directions. However, the method of acquiring the feature quantity 504 can not be limited to this example. As another example, the feature quantity 504 can also be calculated for each of the different prescribed directions. In this case, a common coupler 36 can be used in the calculation of the feature quantity 504, or different couplers 36 can be used for each of the different prescribed directions.

[0099] Further, the second acquisition unit 113 acquires a plurality of learning data sets 51 each composed of a combination of the learning target image 53 and the correct answer information 55. In the present embodiment, the second acquisition unit 113 can also use at least any one of the plurality of data sets 120 collected as the learning data set 51. That is, the second acquisition unit 113 can acquire the learning image 121 of the data set 120 as the learning target image 53 of the learning data set 51 and acquire the correct answer information 123 of the data set 120 as the correct answer information 55 of the learning data set 51.

[0100] The machine learning unit 114 inputs the learning target image 53 included in each of the acquired learning data sets 51 to the extractor 31, and executes the operation processing of the extractor 31. Through this operation processing, the machine learning unit 114 acquires the feature amount 54 related to the learning target image 53 from the extractor 31. Next, the machine learning unit 114 inputs the acquired feature amount 54 and the feature amount 504 (learning correction information 50) acquired from the coupler 36 to the estimator 32, and executes the operation processing of the estimator 32. Through this operation processing, the machine learning unit 114 acquires the output value corresponding to the estimation result of the line-of-sight direction of the subject appearing in the learning target image 53 from the estimator 32. In the machine learning of the learning model 30, the machine learning unit 114 trains the learning model 30 for each of the learning data sets 51 in such a manner that the output value obtained from the estimator 32 is fitted to the corresponding correct answer information 55, along with the calculation of the above-mentioned feature amount 504 and the operation processing of the above-mentioned estimation model 3.

[0101] The training of the learning model 30 can also include the training of each of the extractors (31, 35). Alternatively, through the machine learning of the above-mentioned learning model 4, each of the extractors (31, 35) is trained to acquire the ability to extract a feature amount including a component from which the line-of-sight direction of a person can be estimated from an image. Therefore, in the training of the learning model 30, the training of each of the extractors (31, 35) can also be omitted. Through the machine learning of the learning model 30, the coupler 36 can acquire the ability to derive a correction feature amount that is useful for estimating the line-of-sight direction of a person by coupling the feature information and the true value information. In addition, the estimator 32 can acquire the ability to appropriately estimate the line-of-sight direction of a person appearing in a corresponding image from the feature amount of the image obtained by the extractor 31 and the correction feature amount obtained by the coupler 36.

[0102] Further, in the machine learning of the learning model 30, the learning reference image 501 and the learning true value information 503 used in the calculation of the feature quantity 504 are preferably derived from the same subject as the learning data set 51 used in the training thereof. That is, it is assumed that the learning reference image 501, the learning true value information 503, and the plurality of learning data sets 51 are acquired from a plurality of different subjects, respectively. In this case, the respective origins are preferably identified so that the learning reference image 501, the learning true value information 503, and the plurality of learning data sets 51 obtained from the same subject are used for the machine learning of the learning model 30. The respective origins (i.e., the subjects) can be identified by additional information such as an identifier, for example. In the case where the learning reference image 501, the learning true value information 503, and the plurality of learning data sets 51 are acquired from a plurality of data sets 120, each data set 120 can also include the additional information. In this case, the subjects as the respective origins can be identified from the additional information, whereby the learning reference image 501, the learning true value information 503, and the plurality of learning data sets 51 obtained from the same subject can be used for the machine learning of the learning model 30.

[0103] The saving processing section 115 generates information related to the learning model 30 (i.e., the learning-completed extractor 31, the learning-completed coupler 36, and the learning-completed estimation model 3) as the learning result data 125. Further, the saving processing section 115 saves the generated learning result data 125 to a prescribed storage area.

[0104] (Example of Configuration of Each Model)

[0105] Each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36 are constituted by a model capable of machine learning with an operation parameter. The machine learning model used, respectively, can be of any type as long as it can perform the respective operation processing, and can be appropriately selected according to the embodiment. In the present embodiment, each extractor (31, 35, 41) uses a convolutional neural network. In addition, each estimator (32, 43) and the coupler 36 use a fully connected neural network.

[0106] As Figure 4A and Figure 4BAs shown, each extractor (31, 35, 41) has a convolution layer (311, 351, 411) and a pooling layer (312, 352, 412). The convolution layer (311, 351, 411) is configured to perform a convolution operation on the data provided. The convolution operation corresponds to a process of calculating the correlation of the data provided with a prescribed filter. For example, by performing a convolution of an image, a shade pattern similar to the shade pattern of the filter can be detected from the input image. The convolution layer (311, 351, 411) has a neuron (node) that is a neuron corresponding to the convolution operation and is coupled with a part of the region of the input or the output of the layer disposed further forward (input side) than the layer itself. The pooling layer (312, 352, 412) is configured to perform a pooling process. The pooling process discards part of the information of the position of the strong response of the data provided with respect to the filter, and achieves invariance of the response with respect to the slight positional change of the feature occurring in the data. For example, in the pooling process, the maximum value in the filter can be extracted, and the values other than this can be deleted.

[0107] The number of the convolution layer (311, 351, 411) and the pooling layer (312, 352, 412) included in each extractor (31, 35, 41) can not be particularly limited, and can be appropriately decided according to the embodiment. In the example of FIG. 11, the number of the convolution layer (311, 351, 411) and the pooling layer (312, 352, 412) included in each extractor (31, 35, 41) is 2, but the number can be appropriately decided according to the embodiment. Figure 4A and Figure 4B In the example of FIG. 11, the convolution layer (311, 351, 411) disposed further on the input side (left side of the figure) constitutes the input layer. In addition, the pooling layer (312, 352, 412) disposed further on the output side (right side of the figure) constitutes the output layer. However, the structure of each extractor (31, 35, 41) can not be limited to such an example. The configuration of the convolution layer (311, 351, 411) and the pooling layer (312, 352, 412) can be appropriately decided according to the embodiment. For example, the convolution layer (311, 351, 411) and the pooling layer (312, 352, 412) can be alternately disposed. Or, one or a plurality of the pooling layer (312, 352, 412) can be disposed after a plurality of the convolution layer (311, 351, 411) is disposed in series. In addition, the type of the layer included in each extractor (31, 35, 41) can not be limited to the convolution layer and the pooling layer. In each extractor (31, 35, 41), for example, other types of layers such as a normalization layer, an exit layer, a fully connected layer, and the like can be included.

[0108] In the present embodiment, the structure of each extractor (31, 35) is derived from the structure of the extractor 41 used respectively. In the case where each extractor (31, 35) is prepared separately, the structure can be identical or different between the extractor 31 and the extractor 35. Likewise, in the case where a plurality of prescribed directions are set and different extractors 35 are prepared for each of the different prescribed directions set, the structure of the extractors 35 prepared for each prescribed direction can be identical or at least a part of the extractors 35 can be different from the other extractors 35.

[0109] On the other hand, each speculator (32, 43) and the coupler 36 have one or more fully connected layers (321, 431, 361). The number of the fully connected layers (321, 431, 361) possessed by each speculator (32, 43) and the coupler 36 can not be particularly limited and can be appropriately decided according to the embodiment. In the case where a plurality of fully connected layers are possessed, the fully connected layer disposed at the most input side constitutes an input layer and the fully connected layer disposed at the most output side constitutes an output layer. The fully connected layers disposed between the input layer and the output layer constitute intermediate (hidden) layers. In the case where one fully connected layer is possessed, the one fully connected layer functions as the input layer and the output layer.

[0110] Each fully connected layer (321, 431, 361) has one or more neurons (nodes). The number of the neurons (nodes) included in each fully connected layer (321, 431, 361) can not be particularly limited and can be appropriately selected according to the embodiment. The number of the neurons included in the input layer can be decided, for example, according to the data such as the feature quantity, the true value information, and the form thereof. In addition, the number of the neurons included in the output layer can be decided, for example, according to the data such as the feature quantity, the speculating result, and the form thereof. Each neuron included in each fully connected layer (321, 431, 361) is connected to all the neurons of the adjacent layer. However, the connection relationship of each neuron can not be limited to such an example and can be appropriately set according to the embodiment.

[0111] The weights (connection weights) are set for each connection of the convolution layers (311, 351, 411) and the fully connected layers (321, 431, 361). A threshold value is set in each neuron, and the output of each neuron is basically determined depending on whether the sum of the products of each input and each weight exceeds the threshold value. The threshold value can also be expressed by an activation function. In this case, the output of each neuron is determined by inputting the sum of the products of each input and each weight to the activation function and performing the operation of the activation function. The type of the activation function can be arbitrarily selected. The weights of the connections between the neurons included in the convolution layers (311, 351, 411) and the fully connected layers (321, 431, 361) and the threshold values of the neurons are examples of the operation parameters used in the operation processing of each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36.

[0112] Further, the data form of the input and output of each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36 can not be particularly limited and can be appropriately determined depending on the embodiment. For example, the output layer of each estimator (32, 43) can also be configured to directly output the estimation result (for example, regression). Alternatively, the output layer of each estimator (32, 43) can also be configured to, for example, have one or more neurons for each class of the recognition target and indirectly output the estimation result in a manner that outputs, from each neuron, a probability or the like that corresponds to the class. In addition, the input layer of each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36 can also be configured to accept the input of other data in addition to the input data of the above-mentioned reference image, the target image, the feature amount, the true value information, and the like. Arbitrary pre-processing can be applied to the input data before being input to the input layer.

[0113] In the machine learning of the above-described learning model 4, the machine learning unit 114 repeatedly adjusts the values of the operation parameters of each extractor 41 and each estimator 43 with respect to each data set 120 so that the error between the output value obtained from the estimator 43 by the above-described operation processing and the correct answer information 123 becomes small. Thereby, the trained extractor 41 can be generated. In addition, in the machine learning of the above-described learning model 30, the machine learning unit 114 repeatedly adjusts the values of the operation parameters of each extractor (31, 35), the coupler 36, and the estimator 32 with respect to the learning reference image 501, the learning true value information 503, and each learning data set 51 so that the error between the output value obtained from the estimator 32 by the above-described operation processing and the correct answer information 55 becomes small. In the machine learning of the learning model 30, the adjustment of the values of the operation parameters of each extractor (31, 35) can also be omitted. Thereby, the trained learning model 30 can be generated.

[0114] The saving processing section 115 generates learning result data 125 for reproducing the learned estimation model 3 (extractor 31 and estimator 32), learned extractor 35, and learned coupler 36 generated by machine learning. The configuration of the learning result data 125 can be arbitrary as long as each can be reproduced. For example, the saving processing section 115 generates information indicating the values of the operation parameters of the generated learned estimation model 3, learned extractor 35, and learned coupler 36 as the learning result data 125. Depending on the case, information indicating the configuration of each can also be included in the learning result data 125. The configuration can be determined, for example, according to the number of layers from the input layer to the output layer in the neural network, the type of each layer, the number of neurons included in each layer, the connection relationship between the neurons of adjacent layers, and the like. The saving processing section 115 saves the generated learning result data 125 to a prescribed storage area.

[0115] Further, in the present embodiment, for the convenience of explanation, an example in which the results of machine learning of each extractor (31, 35), estimator 32, and coupler 36 are saved as one learning result data 125 is described. However, the form of saving the learning result data 125 can not be limited to such an example. The results of machine learning of each extractor (31, 35), estimator 32, and coupler 36 can also be saved as different data.

[0116] <LINE OF SIGHT ESTIMATION DEVICE>

[0117] Figure 5A and Figure 5B An example of the software configuration of the line of sight estimation device 2 related to the present embodiment is schematically illustrated. The control section 21 of the line of sight estimation device 2 loads the line of sight estimation program 82 stored in the storage section 22 to the RAM. Further, the control section 21 controls each constituent element by interpreting and executing the commands included in the line of sight estimation program 82 loaded to the RAM by the CPU. Thus, as shown in Figure 5A and Figure 5B As shown in FIG. 8, the line of sight estimation device 2 related to the present embodiment functions as a computer provided with a data acquisition section 211, an image acquisition section 212, an estimation section 213, and an output section 214 as software modules. That is, in the present embodiment, each software module of the line of sight estimation device 2 is realized by the control section 21 (CPU) as with the above-described model generation device 1.

[0118] The information acquisition section 211 acquires the correction information 60 including the feature information 602 related to the line of sight of the eye of the subject R observing a prescribed direction and the true value information 603 indicating the true value of the prescribed direction observed by the eye of the subject R. As shown in Figure 5AAs illustrated, in the present embodiment, the information acquisition unit 211 has the completed extractor 35 and the completed coupler 36 by holding the learning result data 125. The information acquisition unit 211 acquires the reference image 601 that represents the eye of the subject R who observes the prescribed direction. The information acquisition unit 211 inputs the acquired reference image to the completed extractor 35 and executes the operation processing of the extractor 35. Thereby, the information acquisition unit 211 acquires the output value corresponding to the feature quantity 6021 related to the reference image 601 from the extractor 35. The feature quantity 6021 is an example of the second feature quantity. In the present embodiment, the feature information 602 is constituted by this feature quantity 6021. In addition, the information acquisition unit 211 acquires the true value information 603. Further, the information acquisition unit 211 inputs the acquired feature quantity 6021 and the true value information 603 to the completed coupler 36 and executes the operation processing of the coupler 36. Thereby, the information acquisition unit 211 acquires the output value corresponding to the correction-related feature quantity 604 derived by coupling the feature information 602 and the true value information 603 from the coupler 36. The feature quantity 604 is an example of the correction feature quantity. In the present embodiment, the correction information 60 is constituted by this feature quantity 604. The information acquisition unit 211 can acquire the correction information 60 (the feature quantity 604) by these operation processes and using the completed extractor 35 and the completed coupler 36.

[0119] Further, in correspondence with the generation process of the completed estimation model 3 described above, the correction information 60 can also include the feature information 602 and the true value information 603 corresponding to a plurality of different prescribed directions respectively. In this case, as with the generation process described above, the information acquisition unit 211 can also acquire the reference image 601 and the true value information 603 for a plurality of different prescribed directions respectively. The information acquisition unit 211 can also input each reference image 601 to the completed extractor 35 and execute the operation processing of the extractor 35, thereby acquiring each feature quantity 6021 from the extractor 35. Next, the information acquisition unit 211 can also input each acquired feature quantity 6021 and the true value information 603 for each prescribed direction to the completed coupler 36 and execute the operation processing of the coupler 36. Thereby, the information acquisition unit 211 can also acquire the correction-related feature quantity 604 from the coupler 36. In this case, the feature quantity 604 can include information that aggregates the feature information 602 and the true value information 603 of each of a plurality of different prescribed directions. However, the method of acquiring the feature quantity 604 can not be limited to this example. As another example, the feature quantity 604 can also be calculated for each of different prescribed directions in correspondence with the generation process described above. In this case, a common coupler 36 can be used in the calculation of the feature quantity 604, or different couplers 36 can be used for each of different prescribed directions respectively.

[0120] As Figure 5BAs shown, the image acquisition unit 212 acquires the subject image 63 of the subject R's eyes. The estimation unit 213 has a learning-completed estimation model 3 generated by machine learning by retaining the learning result data 125. The estimation unit 213 estimates the line-of-sight direction of the subject R's eyes reflected in the subject image 63 using this learning-completed estimation model 3. As this estimation processing, the estimation unit 213 inputs the acquired subject image 63 and the correction information 60 into the learning-completed estimation model 3 and executes the operation processing of the learning-completed estimation model 3. Thereby, the estimation unit 213 acquires the output value corresponding to the estimation result of the line-of-sight direction of the subject R's eyes reflected in the subject image 63 from the learning-completed estimation model 3.

[0121] The operation processing of the learning-completed estimation model 3 can be appropriately decided according to the configuration of the learning-completed estimation model 3. In the present embodiment, the learning-completed estimation model 3 has a learning-completed extractor 31 and an estimator 32. First, the estimation unit 213 inputs the acquired subject image 63 into the learning-completed extractor 31 and executes the operation processing of the extractor 31. By this operation processing, the estimation unit 213 acquires the output value corresponding to the feature quantity 64 related to the subject image 63 from the extractor 31. The feature quantity 64 is an example of the first feature quantity. The above-mentioned feature quantity 6021 and the feature quantity 64 can also be respectively replaced with an image feature quantity. Next, the estimation unit 213 inputs the feature quantity 604 acquired by the information acquisition unit 211 and the feature quantity 64 acquired from the extractor 31 into the estimator 32 and executes the operation processing of the estimator 32. In the present embodiment, the operation processing of the learning-completed estimation model 3 is configured by executing the operation processing of these extractor 31 and estimator 32. As a result of these operation processing, the estimation unit 213 can obtain the output value corresponding to the estimation result of the line-of-sight direction of the subject R's eyes reflected in the subject image 63 from the estimator 32. The output unit 214 outputs information related to the estimation result of the subject R's line-of-sight direction.

[0122] <Other>

[0123] As for each software module of the model generation device 1 and the line-of-sight estimation device 2, detailed description will be given in the action example described later. Further, in the present embodiment, an example in which each software module of the model generation device 1 and the line-of-sight estimation device 2 is implemented by a general-purpose CPU is described. However, part or all of the above software modules can also be implemented by one or a plurality of dedicated processors. That is, each of the above modules can also be implemented as a hardware module. In addition, as for the software configuration of each of the model generation device 1 and the line-of-sight estimation device 2, omission, replacement, and addition of software modules can also be appropriately made according to the embodiment.

[0124] §3 Action Example

[0125] [Model Generation Device]

[0126] Figure 6 is a flowchart showing an example of a processing procedure of the model generation device 1 according to the present embodiment. The processing procedure described below is an example of the model generation method. However, the processing procedure described below is merely an example, and each step can be changed within a possible range. Furthermore, each processing procedure described below can omit, replace, and add steps as appropriate according to the embodiment.

[0127] (step S101).

[0128] In step S101, the control section 11 operates as the collection section 111 to collect a plurality of data sets 120 for learning from the subjects. Each data set 120 is constituted by a combination of a learning image 121 that shows the subject's eye and correct answer information 123 that shows the true value of the subject's line-of-sight direction shown in the learning image 121.

[0129] Each data set 120 can be generated as appropriate. For example, a camera S or a camera of the same kind as the camera S and subjects are prepared. The number of subjects can be determined as appropriate. The subjects are instructed to look in various directions, and the face of the subject looking in the instructed direction is photographed using the camera. Thereby, the learning image 121 can be acquired. The learning image 121 can also be an image as it is obtained by the camera. Alternatively, the learning image 121 can be generated by applying some kind of image processing to the image obtained by the camera. Information showing the true value of the line-of-sight direction instructed to the subject is associated with the acquired learning image 121 as the correct answer information 123. In the case where a plurality of subjects are present, additional information such as an identifier of the subject can be further associated so as to identify the source of the data set 120. Through these processes, each data set 120 can be generated. Furthermore, the method of acquiring the learning image 121 and the correct answer information 123 can employ a method of the same kind as the method of acquiring the reference image 601 and the true value information 603 described later. Figure 8

[0130] ​Each of the data sets 120 can be automatically generated by an action of the computer or can be manually generated by at least partly containing an operation of an operator. In addition, the generation of each of the data sets 120 can be performed by the model generation apparatus 1 or can be performed by another computer other than the model generation apparatus 1. In a case where the data sets 120 are generated by the model generation apparatus 1, the control section 11 automatically or manually performs the above-described generation processing by an operation of the operator via the input device 15, thereby acquiring a plurality of the data sets 120. On the other hand, in a case where each of the data sets 120 is generated by another computer, the control section 11 acquires a plurality of the data sets 120 generated by another computer, for example, via a network, the storage medium 91, or the like. It is also possible that a part of the data sets 120 are generated by the model generation apparatus 1 and the other data sets 120 are generated by one or a plurality of other computers.

[0131] The number of the acquired data sets 120 can not be particularly limited and can be appropriately selected depending on the embodiment. When a plurality of the data sets 120 are acquired, the control section 11 causes the processing to proceed to the next step S102.

[0132] (Step S102)

[0133] In the step S102, the control section 11 functions as a machine learning section 114 and performs machine learning of the learning model 4 using the collected plurality of the data sets 120. In this machine learning, the control section 11 trains the extractor 41 and the estimator 43 in such a manner that the output value (estimation result of the line-of-sight direction) obtained from the estimator 43 by inputting the learning image 121 to the extractor 41 is adapted to the corresponding correct answer information 123 for each of the data sets 120. Further, it is not necessary that all of the collected data sets 120 are used for the machine learning of the learning model 4. The data sets 120 used in the machine learning of the learning model 4 can be appropriately selected.

[0134] As an example, first, the control section 11 prepares neural networks respectively constituting the extractor 41 and the estimator 43 which are processing targets of the machine learning. The structure of each of the neural networks (for example, the number of layers, the kind of each layer, the number of neurons included in each layer, the connection relationship of the neurons of the adjacent layers to each other, and the like), the initial value of the weight of the connection between each of the neurons, and the initial value of the threshold value of each of the neurons can be provided by a template or can be provided by an input of an operator. In addition, in a case where relearning is performed, the control section 11 can prepare the extractor 41 and the estimator 43 in accordance with the learning result data obtained by the machine learning in the past.

[0135] Next, the control section 11 executes the training processing of the extractor 41 and the estimator 43 using the learning image 121 of each data set 120 as the training data (input data) and the correct answer information 123 as the teacher data (teacher signal, label). In this training processing, a stochastic gradient descent method, a mini-batch gradient descent method, or the like can be used.

[0136] For example, the control section 11 inputs the learning image 121 into the extractor 41 and executes the operation processing of the extractor 41. That is, the control section 11 inputs the learning image 121 into the input layer (in the example, the convolution layer 411 disposed at the most input side) of the extractor 41 and executes the operation processing of the forward propagation of each layer (411, 412) such as the firing determination of the neuron in order from the input side. Through this operation processing, the control section 11 acquires the output value corresponding to the feature quantity extracted from the learning image 121 from the output layer (in the example, the pooling layer 412 disposed at the most output side) of the extractor 41. Figure 4A Figure 4A

[0137] Next, the control section 11 inputs the output value (feature quantity) obtained into the input layer (the fully connected layer 431 disposed at the most input side) of the estimator 43 in the same manner as the operation processing of the extractor 41 and executes the operation processing of the forward propagation of the estimator 43. Through this operation processing, the control section 11 acquires the output value corresponding to the estimation result of the line-of-sight direction of the subject reflected in the learning image 121 from the output layer (the fully connected layer 431 disposed at the most output side) of the estimator 43.

[0138] Next, the control section 11 calculates the error between the output value obtained from the output layer of the estimator 43 and the correct answer information 123. In the calculation of the error (loss), a loss function can be used. The loss function is a function that evaluates the difference (that is, the degree of difference) between the output of the machine learning model and the correct answer, and the larger the difference value between the output obtained from the output layer and the correct answer, the larger the value of the error calculated by the loss function. The kind of the loss function used to calculate the error can not be particularly limited and can be appropriately selected according to the embodiment.

[0139] The control section 11 calculates the error of the value of each operation parameter (the weight of the connection between each neuron, the threshold value of each neuron, and the like) of the extractor 41 and the estimator 43 in order from the output side by the back propagation method and using the gradient of the error of the output value calculated. The control section 11 updates the value of each operation parameter of the extractor 41 and the estimator 43 according to each error calculated. The degree of update of the value of each operation parameter can be adjusted according to the learning rate. The learning rate can be provided by the designation of the operator or provided as a set value in the program.

[0140] ​​The control section 11 adjusts the values of the respective operation parameters of the extractor 41 and the estimator 43 for each data set 120 through the above series of update processes so that the sum of the errors of the calculated output values becomes smaller. For example, the control section 11 can repeatedly adjust the values of the respective operation parameters of the extractor 41 and the estimator 43 through the above series of update processes until a prescribed condition is satisfied, such as a prescribed number of times of execution, the sum of the errors of the calculated values being below a threshold value, and the like.

[0141] As a result of the machine learning, the control section 11 can generate a trained learning model 4 that has acquired the ability to appropriately estimate the gaze direction of the subject represented in the learning image 121 for each data set 120. In addition, the output (i.e., the feature quantity) of the trained extractor 41 contains components related to the eyes of the subject included in the learning image 121 so that the gaze direction of the subject can be appropriately estimated in the estimator 43. When the machine learning of the learning model 4 is completed, the control section 11 causes the process to proceed to the next step S103.

[0142] (Step S103)

[0143] In step S103, the control section 11 prepares a learning model 30 including the estimation model 3 using the learning result of the extractor 41.

[0144] In the present embodiment, the control section 11 prepares each extractor (31, 35) from the learning result of the extractor 41. That is, the control section 11 uses the trained extractor 41 generated through step S102 or a copy thereof as each extractor (31, 35). In the case where each extractor (31, 35) is prepared separately, or in the case where a plurality of prescribed directions are set and the extractor 35 is prepared separately for each of the different prescribed directions set, the control section 11 can prepare a separate learning model 4 and perform respective machine learning in the above step S102. Also, the control section 11 can use the trained extractors 41 generated through the respective machine learning or copies thereof as each extractor (31, 35).

[0145] In addition, the control section 11 prepares the neural networks that respectively constitute the estimator 32 and the coupler 36. As with the above extractor 41, the structure of the neural networks that respectively constitute the estimator 32 and the coupler 36, the initial values of the weights of the connections between the respective neurons, and the initial values of the thresholds of the respective neurons can be provided through a template or through input by an operator. In addition, in the case where relearning is performed, the control section 11 can prepare the estimator 32 and the coupler 36 from learning result data obtained through machine learning in the past. When the learning model 30 constituted by each extractor (31, 35), the estimator 32, and the coupler 36 is prepared, the control section 11 causes the process to proceed to the next step S104.

[0146] (Step S104)

[0147] In step S104, the control unit 11 operates as the first acquisition unit 112 to acquire learning correction information 50, which includes learning feature information 502 and learning truth information 503.

[0148] In this embodiment, the control unit 11 acquires learning correction information 50 using the extractor 35 and the coupler 36. Specifically, the control unit 11 first acquires a learning reference image 501 reflecting the eyes of a subject observing a specified direction, and learning truth information 503 representing the truth value of the specified direction (viewing direction) observed by the subject as reflected in the learning reference image 501. The control unit 11 may also acquire the learning image 121 contained in the dataset 120 obtained for subjects observing a specified direction as the learning reference image 501, and acquire the correct solution information 123 as the learning truth information 503. Alternatively, the control unit 11 may acquire the learning reference image 501 and the learning truth information 503 separately from the dataset 120. The method for acquiring the learning reference image 501 and the learning truth information 503 can be the same as the method for generating the dataset 120.

[0149] Next, the control unit 11 inputs the acquired learning reference image 501 into the input layer of the extractor 35. Figure 4B In this example, the convolutional layer 351, located closest to the input side, performs the forward propagation operation of the extractor 35. Through this operation, the control unit 11 obtains data from the output layer of the extractor 35 (…). Figure 4B In the example, the pooling layer 352, located on the output side, acquires the output value corresponding to the feature quantity 5021 (learning feature information 502) related to the learning reference image 501. Next, the control unit 11 inputs the acquired feature quantity 5021 and the learning truth information 503 into the input layer of the coupler 36 (the fully connected layer 361 located on the input side) and performs forward propagation processing of the coupler 36. Through this processing, the control unit 11 acquires the output value corresponding to the correction-related feature quantity 504 from the output layer of the coupler 36 (the fully connected layer 361 located on the output side).

[0150] In the present embodiment, the control section 11 can acquire the learning correction information 50 composed of the feature quantity 504 by these operation processes, and using the extractor 35 and the coupler 36. Further, as described above, in a case where a plurality of prescribed directions are set, the control section 11 can also acquire the learning reference image 501 and the learning true value information 503 for each of the plurality of different prescribed directions. Moreover, the control section 11 can also acquire the learning correction information 50 including the learning feature information 502 and the learning true value information 503 for each of the plurality of different prescribed directions by performing the operation processes of the extractor 35 and the coupler 36 for each. When the learning correction information 50 is acquired, the control section 11 causes the process to proceed to the next step S105.

[0151] (Step S105)

[0152] In step S105, the control section 11 functions as a second acquisition section 113, and acquires a plurality of learning data sets 51 each composed of a combination of the learning target image 53 and the correct answer information 55.

[0153] In the present embodiment, the control section 11 can also use at least any one of the plurality of data sets 120 collected as the learning data set 51. That is, the control section 11 can acquire the learning image 121 of the data set 120 as the learning target image 53 of the learning data set 51, and acquire the correct answer information 123 of the data set 120 as the correct answer information 55 of the learning data set 51. Alternatively, the control section 11 can acquire each learning data set 51 separately from the above-described data set 120. The method of acquiring each learning data set 51 can be the same as the method of generating the data set 120.

[0154] The number of learning data sets 51 acquired can not be particularly limited, and can be appropriately selected according to the embodiment. When a plurality of learning data sets 51 are acquired, the control section 11 causes the process to proceed to the next step S106. Further, the timing at which the process of step S105 is executed can not be limited to this example. As long as the process of step S105 is executed before the process of the later-described step S106 is executed, the process of step S105 can be executed at any timing.

[0155] (Step S106)

[0156] In step S106, the control section 11 functions as a machine learning section 114, and performs machine learning of the estimation model 3 using the plurality of learning data sets 51 acquired. In this machine learning, the control section 11 trains the estimation model 3 so as to output an output value suitable for the corresponding correct answer information 55 with respect to the input of the learning target image 53 and the learning correction information 50 for each learning data set 51.

[0157] In this embodiment, the control unit 11 uses the learning object image 53, the learning reference image 501, and the learning ground truth information 503 of each learning dataset 51 as training data, and uses the positive solution information 55 of each learning dataset 51 as teacher data, and performs training processing of the learning model 30 including the inference model 3. In this training processing, stochastic gradient descent, mini-batch gradient descent, etc., can be used.

[0158] For example, the control unit 11 inputs the learning object images 53 contained in each learning dataset 51 into the input layer of the extractor 31. Figure 4B In this example, the convolutional layer 311, located closest to the input side, performs the forward propagation operation of the extractor 31. Through this operation, the control unit 11 obtains data from the output layer of the extractor 31 (…). Figure 4B In the example, the pooling layer 312, which is configured on the output side, obtains the output value corresponding to the feature quantity 54 extracted from the learning object image 53.

[0159] Next, the control unit 11 inputs the feature quantity 504 obtained from the coupler 36 and the feature quantity 54 obtained from the extractor 31 into the input layer (the fully connected layer 321 located on the input side) of the inferr 32, and performs forward propagation operation processing of the inferr 32. Through this operation processing, the control unit 11 obtains the output value corresponding to the inference result of the subject's gaze direction reflected in the learning object image 53 from the output layer (the fully connected layer 321 located on the output side) of the inferr 32.

[0160] Next, the control unit 11 calculates the error between the output value obtained from the output layer of the inferr 32 and the corresponding positive solution information 55. Similar to the machine learning in learning model 4 described above, any loss function can be used in calculating the error. The control unit 11 calculates the errors of the values ​​of each operational parameter of each extractor (31, 35), coupler 36, and inferr 32 sequentially from the output side using the error backpropagation method and the gradient of the calculated output value error. The control unit 11 updates the values ​​of each operational parameter of each extractor (31, 35), coupler 36, and inferr 32 based on the calculated errors. Similar to the machine learning in learning model 4 described above, the degree to which the values ​​of each operational parameter are updated can be adjusted according to the learning rate.

[0161] The control section 11 performs the above series of update processes in conjunction with the calculation of the feature quantity 504 of step S104 and the operation processing of the above estimation model 3. Thereby, the control section 11 adjusts the values of the operation parameters of each extractor (31, 35), the coupler 36, and the estimator 32 with respect to the learning reference image 501, the learning true value information 503, and each learning data set 51 so that the sum of the errors of the calculated output values becomes smaller. As with the machine learning of the above learning model 4, the control section 11 can repeatedly adjust the values of the operation parameters of each extractor (31, 35), the coupler 36, and the estimator 32 through the above series of update processes until a prescribed condition is satisfied.

[0162] Further, as described above, the subject as each respective source can also be identified so that the learning reference image 501, the learning true value information 503, and the plurality of learning data sets 51 obtained from the same subject are used for the machine learning of the learning model 30. In addition, each extractor (31, 35) is trained to acquire the ability to extract a feature quantity containing a component capable of estimating the gaze direction of a person from an image through the machine learning of the above learning model 4. Therefore, in the above update processes, the process of adjusting the values of the operation parameters of each extractor (31, 35) can be omitted. In addition, the process of step S104 can be executed at any timing before the operation processing of the estimator 32 is performed. For example, the process of step S104 can also be executed after the operation processing of the extractor 31 is performed.

[0163] As a result of this machine learning, the control section 11 can generate a trained learning model 30 that acquires the ability to appropriately estimate the gaze direction of a person from the learning reference image 501, the learning true value information 503, and the learning target image 53 with respect to each learning data set 51. That is, the control section 11 can generate a learned coupler 36 that acquires the ability to derive a correction feature quantity that is beneficial for the estimation of the gaze direction of a person with respect to each learning data set 51. In addition, the control section 11 can generate a learned estimator 32 that acquires the ability to appropriately estimate the gaze direction of a person appearing in the corresponding image from the feature quantity of the image obtained by the extractor 31 and the correction feature quantity obtained by the coupler 36 with respect to each learning data set 51. When the machine learning of the learning model 30 is completed, the control section 11 causes the processing to proceed to the next step S107.

[0164] (Step S107)

[0165] In step S107, the control section 11 functions as a saving processing section 115, generates information related to the learned learning model 30 (the estimation model 3, the extractor 35, and the coupler 36) generated by machine learning as learning result data 125, and saves the generated learning result data 125 to a predetermined storage area.

[0166] The predetermined storage area can be, for example, a RAM in the control section 11, the storage section 12, an external storage device, a storage medium, or a combination thereof. The storage medium can be, for example, a CD, a DVD, or the like, and the control section 11 can store the learning result data 125 in the storage medium via the drive 17. The external storage device can be, for example, a data server such as a NAS (Network Attached Storage). In this case, the control section 11 can store the learning result data 125 in the data server via a network using the communication interface 13. Alternatively, the external storage device can be, for example, an external storage device connected to the model generation device 1 via the external interface 14.

[0167] When the saving of the learning result data 125 is completed, the control section 11 ends the processing related to the present action example.

[0168] Further, the generated learning result data 125 can be applied to the control device 2 at any timing. For example, the control section 11 can transmit the learning result data 125 to the gaze estimation device 2 as the processing of step S107 or separately from the processing of step S107. The gaze estimation device 2 can acquire the learning result data 125 by receiving the transmission. Alternatively, for example, the gaze estimation device 2 can acquire the learning result data 125 by accessing the model generation device 1 or the data server via a network using the communication interface 23. Alternatively, for example, the gaze estimation device 2 can acquire the learning result data 125 via the storage medium 92. Alternatively, for example, the learning result data 125 can be embedded in the gaze estimation device 2 in advance.

[0169] Further, the control section 11 can update or newly create the learning result data 125 by repeatedly performing the processing of steps S101 to S107 (or steps S104 to S107) periodically or aperiodically. At the time of the repetition, at least a part of the data used in the machine learning can be appropriately changed, corrected, added, deleted, or the like. Further, the control section 11 can update the learning result data 125 held by the gaze estimation device 2 by providing the gaze estimation device 2 with the updated or newly generated learning result data 125 using any method.

[0170] [Gaze Estimation Device]

[0171] Figure 7 is a flowchart showing an example of a processing procedure of the line-of-sight estimation device 2 according to the present embodiment. The processing procedure described below is an example of a line-of-sight estimation method. However, the processing procedure described below is merely an example, and each step can be changed within a possible range. Furthermore, each processing procedure described below can omit, replace, and add steps as appropriate according to the embodiment.

[0172] (Step S201)

[0173] In step S201, the control section 21 functions as the information acquisition section 211, and acquires the correction information 60 including the feature information 602 and the true value information 603.

[0174] Figure 8 An example of a method of acquiring the correction information 60 is schematically illustrated. In the present embodiment, first, the control section 21 outputs an instruction to cause the subject R to observe a prescribed direction. In this case, the control section 21 can output the instruction to cause the subject R to observe the prescribed direction by using the output device 26. The output device 26 is a device that outputs an instruction to the subject R. The output device 26 can be appropriately selected according to the embodiment. The output device 26 can include a display device such as a display, a speaker, and the like. Figure 8 In the example of FIG. 6, the output device 26 includes the display 261. The control section 21 displays a mark M on the display 261 at a position corresponding to the prescribed direction. Furthermore, the control section 21 outputs an instruction to the subject R to observe the mark M displayed on the display 261. The output form of the instruction can be appropriately selected according to the embodiment. In the case where the output device 26 includes a speaker, the output of the instruction can also be performed by sound via the speaker. In addition, in the case where the output device 26 includes a display device such as the display 261, the output of the instruction can also be performed by image display via the display device. After the instruction is output, the control section 21 captures the face of the subject R observing the mark M by the camera S. The camera S is an example of a sensor that can observe the line-of-sight of the subject R. Thus, the control section 21 can acquire the reference image 601 that reflects the eyes of the subject observing the prescribed direction. In addition, the control section 21 can of course acquire the true value information 603 according to the output instruction.

[0175] Moreover, the prescribed direction can not be limited to the mark M displayed on the display 261, and can be appropriately determined according to the embodiment. For example, in the case of the scene in which the driver's line-of-sight direction is estimated, when the position of the camera S is determined, the positional relationship of the installed object such as the inside rearview mirror and the camera S is prescribed. In this way, in the case where the object whose position is prescribed with respect to the sensor of the line-of-sight of the observation target R exists, the control section 21 can also output the instruction to observe the object. In this way, in the case where the object whose positional relationship with the sensor of the line-of-sight of the observation target R is prescribed exists, the control section 21 can also output the instruction to the observation target R to observe the object. According to this method, the reference image 601 and the corresponding true value information 603 that exhibit the individuality of the line-of-sight of the observation target R can be appropriately and simply acquired. In addition, in the above-described model generation scene and the scene in which the line-of-sight is estimated (the scene in which the model is used), the prescribed direction can not be completely identical. In order to cope with this situation, a plurality of different prescribed directions can also be set, and in the scene in which the model is used, data of at least an arbitrary one of the prescribed directions (the reference image 601 and the true value information 603 in the present embodiment) can be randomly selected.

[0176] Next, the control section 21 refers to the learning result data 125 to set the extractor 35 and the coupler 36 in which the learning is completed. The control section 21 inputs the acquired reference image 601 to the input layer of the extractor 35 in which the learning is completed, and executes the operation processing of the forward propagation of the extractor 35. Through this operation processing, the control section 21 acquires the output value corresponding to the feature amount 6021 (the feature information 602) related to the reference image 601 from the output layer of the extractor 35 in which the learning is completed. Next, the control section 21 inputs the acquired feature amount 6021 and the true value information 603 to the input layer of the coupler 36 in which the learning is completed, and executes the operation processing of the forward propagation of the coupler 36. Through this operation processing, the control section 21 acquires the output value corresponding to the correction-related feature amount 604 from the output layer of the coupler 36 in which the learning is completed. In the present embodiment, the control section 21 can acquire the correction information 60 composed of the feature amount 604 by these operation processes and using the extractor 35 and the coupler 36 in which the learning is completed.

[0177] Moreover, as described above, the correction information 60 can also include the feature information 602 and the true value information 603 corresponding to a plurality of different prescribed directions, respectively, in correspondence with the generation process of the learning-completed estimation model 3. In the present embodiment, the control section 21 can also acquire the correction information 60 composed of the feature amount 604 by executing the above-described acquisition process for each of the different prescribed directions Figure 8), thereby acquiring the plurality of different prescribed directions each of the reference images 601 and the true value information 603. Also, the control section 21 can acquire the correction information 60 (feature amount 604) including the plurality of different prescribed directions each of the feature information 602 and the true value information 603 by performing the operation processing for each of the extractors 35 and the couplers 36 for which the learning is completed. When the observation information 60 is acquired, the control section 21 causes the process to proceed to the next step S202

[0178] (Step S202)

[0179] In step S202, the control section 21 functions as an image acquisition section 212, and acquires the object image 63 in which the eyes of the subject R are reflected. In the present embodiment, the control section 21 controls the operation of the camera S via the external interface 24 to capture the subject R. Thereby, the control section 21 can directly acquire the object image 63 which is the object of the line-of-sight direction estimation processing from the camera S. The object image 63 can be either a dynamic image or a static image. However, the path to acquire the object image 63 can not be limited to such an example. For example, the camera S can be controlled by another computer. In this case, the control section 21 can indirectly acquire the object image 63 from the camera S via the other computer. When the object image 63 is acquired, the control section 21 causes the process to proceed to the next step S203.

[0180] (Step S203)

[0181] In step S203, the control section 21 functions as an estimation section 213, and estimates the line-of-sight direction of the eyes of the subject R reflected in the object image 63 using the learned estimation model 3. In this estimation processing, the control section 21 inputs the acquired object image 63 and the correction information 60 to the learned estimation model 3, and performs the operation processing of the learned estimation model 3. Thereby, the control section 21 acquires the output value corresponding to the estimation result of the line-of-sight direction of the eyes of the subject R reflected in the object image 63 from the learned estimation model 3.

[0182] In the present embodiment, first, the control section 21 refers to the learning result data 125 to set the completed extractor 31 and the completed estimator 32. Next, the control section 21 inputs the acquired object image 63 to the input layer of the completed extractor 31, and executes the operation processing of the forward propagation of the completed extractor 31. Through this operation processing, the control section 21 acquires the output value corresponding to the feature quantity 64 related to the object image 63 from the output layer of the completed extractor 31. Next, the control section 21 inputs the feature quantity 604 acquired through the step S201 and the feature quantity 64 acquired from the completed extractor 31 to the input layer of the completed estimator 32, and executes the operation processing of the forward propagation of the completed estimator 32. Through this operation processing, the control section 21 can acquire the output value corresponding to the estimation result of the line-of-sight direction of the subject R reflected in the object image 63 from the output layer of the completed estimator 32. That is, in the present embodiment, the estimation of the line-of-sight direction of the subject R reflected in the object image 63 is achieved by providing the object image 63 and the correction information 60 to the completed estimation model 3, and executing the operation processing of the forward propagation of the completed estimation model 3. Further, the processing of the above step S201 can be executed at any timing before the operation processing of the completed estimator 32 is executed. For example, the processing of the above step S201 can be executed after the operation processing of the completed extractor 31 is executed. When the estimation processing of the line-of-sight direction is completed, the control section 21 advances the processing to the next step S204.

[0183] (Step S204)

[0184] In the step S204, the control section 21 functions as the output section 214, and outputs the information related to the estimation result of the line-of-sight direction of the subject R.

[0185] The output destination and the content of the output information can be appropriately decided according to the embodiment. For example, the control section 21 can directly output the estimation result of the line-of-sight direction to a memory such as a RAM, the storage section 22, or the output device 26. The control section 21 can also make a history of the line-of-sight direction of the subject R by outputting the estimation result of the line-of-sight direction to the memory.

[0186] Further, for example, the control section 21 can execute certain information processing using the result of the estimation of the line-of-sight direction. Also, the control section 21 can output the result of the execution of the information processing as information related to the result of the estimation. As an example, assume a scenario in which the line-of-sight direction of the driver is estimated in order to monitor the state of the driver of the vehicle. In this scenario, the control section 21 can determine whether the driver is squinting or not based on the estimated line-of-sight direction. Also, in the case where it is determined that squinting is occurring, the control section 21 can execute processing that instructs the driver to look in a direction suitable for driving or to reduce the speed of the vehicle as the output processing of step S204. As another example, assume a scenario in which the line-of-sight direction of the estimation target person R is estimated in the user interface. In this scenario, the control section 21 can execute processing that executes an application corresponding to an icon existing in the estimated line-of-sight direction or changes the display range in such a manner that a display existing in the estimated line-of-sight direction comes to the center of the display device as the output processing of step S204. When the information related to the result of the estimation of the line-of-sight direction is output, the control section 21 advances the processing to the next step S205.

[0187] (Step S205)

[0188] In step S205, it is determined whether the estimation processing of the line-of-sight direction is repeated. The criterion for determining whether the estimation processing is repeated can be appropriately decided according to the embodiment.

[0189] As the criterion for determination, for example, the period or the number of times during which the processing is repeated can be set. In this case, the control section 21 can determine whether the estimation processing of the line-of-sight direction is repeated or not based on whether the period or the number of times during which the estimation processing of the line-of-sight direction is executed reaches a prescribed value. That is, in the case where the period or the number of times during which the estimation processing is executed does not reach the prescribed value, the control section 21 can determine that the estimation processing of the line-of-sight direction is repeated. On the other hand, in the case where the period or the number of times during which the estimation processing is executed reaches the prescribed value, the control section 21 can determine that the estimation processing of the line-of-sight direction is not repeated.

[0190] Further, for example, the control section 21 can repeat the estimation processing of the line-of-sight direction until an instruction to end is given via the input device 25. In this case, during the period in which the instruction to end is not given, the control section 21 can determine that the estimation processing of the line-of-sight direction is repeated. On the other hand, after the instruction to end is given, the control section 21 can determine that the estimation processing of the line-of-sight direction is not repeated.

[0191] In a case where it is determined that the estimation process of the line-of-sight direction is repeated, the control section 21 returns the process to step S202, and repeatedly executes the acquisition process of the object image 63 (step S202) and the estimation process of the line-of-sight direction of the object person R (step S203). Thereby, the estimation of the line-of-sight direction of the object person R can be continuously performed. On the other hand, in a case where it is determined that the estimation process of the line-of-sight direction is not repeated, the control section 21 stops the repeated execution of the estimation process of the line-of-sight direction, and ends the processing procedure involved in the present action example.

[0192] If the derivation of the correction information 60 (feature quantity 604) of step S201 is completed, the already derived correction information 60 can be repeatedly used in each cycle in which the estimation process of the line-of-sight direction is executed, as long as the correction information 60 is not updated. Therefore, the process of step S201 can be omitted in each cycle in which the estimation process of the line-of-sight direction is executed, as in the present embodiment. However, it is not necessary that the process of step S201 is omitted in all cycles in which the estimation process of the line-of-sight direction is executed. In a case where the correction information 60 is updated, step S201 can be executed again at an arbitrary timing. In addition, the process of step S204 can be omitted in at least a part of the cycles.

[0193] [Features]

[0194] As described above, in the present embodiment, in step S203, in order to estimate the line-of-sight direction of the object person R, not only the object image 63 which reflects the eyes of the object person R but also the correction information 60 including the feature information 602 and the true value information 603 is used. According to the feature information 602 and the true value information 603, the individuality of the line-of-sight of the object person R in the known direction can be grasped according to the true value. Therefore, according to the present embodiment, the line-of-sight direction of the object person R reflected in the object image 63 can be estimated on the basis of the personal difference between the subject and the object person R which can be grasped from the correction information 60. Therefore, the improvement of the estimation accuracy of the line-of-sight direction of the object person R in step S203 can be achieved. By using the correction information 60, for the object person R whose eyes are not directed to the line-of-sight direction due to strabismus or the like, it is also expected that the estimation accuracy of the line-of-sight direction is improved. In the present embodiment, the correction information 60 can also include the feature information 602 and the true value information 603 corresponding to a plurality of different prescribed directions, respectively. Thereby, the individuality of the line-of-sight of the object person R can be more accurately grasped according to the correction information 60 with respect to a plurality of different prescribed directions. Therefore, the further improvement of the estimation accuracy of the line-of-sight direction of the object person R can be achieved. According to the model generation apparatus 1 involved in the present embodiment, by the processes of steps S101 to S107, the learning completed estimation model 3 which can estimate the line-of-sight direction of the object person R with such high accuracy can be generated.

[0195] In addition, in the present embodiment, instead of directly using the reference image 601 and the true value information 603 as the correction information 60, a feature quantity 604 that is derived by extracting the feature quantity 6021 from the reference image 601 and coupling the resulting feature information 602 and the true value information 603 is used as the correction information 60. Thereby, the amount of information of the correction information 60 can be reduced. In addition, in the present embodiment, the derivation of the feature quantity 604 is performed within the processing of step S201. In a case where the estimation processing of the line-of-sight direction of the subject R is repeatedly performed, the derived feature quantity 604 can be repeatedly used in each cycle. Thereby, the processing cost of step S203 can be suppressed. Thus, according to the present embodiment, the speedup of the estimation processing of the line-of-sight direction of the subject R in step S203 can be achieved.

[0196] Further, by the extractor 35 that is completed with learning, the feature quantity 6021 (feature information 602) that contains a component related to the feature of the line-of-sight of the subject R who observes a prescribed direction can be appropriately extracted from the reference image 601. In addition, by the coupler 36 that is completed with learning, the feature quantity 604 that contains a component that is a collection of the feature of the line-of-sight of the subject R who observes a prescribed direction and the true value of the prescribed direction can be appropriately derived from the feature quantity 6021 and the true value information 603. Thus, in the estimation model 3 that is completed with learning, the line-of-sight direction of the subject R can be appropriately estimated from the feature quantity 604 and the object image 63.

[0197] §4 Modification

[0198] The above-described embodiments of the present application have been described in detail, but the above-described description is merely an example of the present application in all aspects. Needless to say, various modifications or changes can be made without departing from the scope of the present application. For example, the following modifications can be made. In addition, the same reference numerals are used for the same constituent elements as those of the above-described embodiments, and the description is appropriately omitted for the same points as those of the above-described embodiments. The following modifications can be appropriately combined.

[0199] <4.1>

[0200] In the above-described embodiment, the camera S is used in the acquisition of the correction information 60. However, the sensor for observing the line of sight of the subject R is not limited to such an example. The sensor can be of any type as long as it can observe the feature of the line of sight of the subject R, and can be appropriately selected according to the embodiment. The sensor can use, for example, a scleral contact lens with a coil, an electro-oculogram (EOG) sensor, or the like. In this case, the line-of-sight estimation device 2 can also, as in the above-described embodiment, observe the line of sight of the subject R by the sensor after outputting an instruction to the subject R to look in a prescribed direction. The feature information 602 can be acquired from the sensing data obtained by this observation. The acquisition of the feature information 602 can use, for example, a search coil method, an EOG method, or the like.

[0201] <4.2>

[0202] In the above-described embodiment, the estimation model 3 is constituted by the extractor 31 and the estimator 32. The correction information 60 is constituted by the feature quantity 604 derived from the reference image 601 and the true value information 603 by the extractor 35 and the coupler 36. The estimator 32 is configured to accept the input of the feature quantity 604 derived by the coupler 36 and the feature quantity 64 related to the subject image 63. However, the constitution of the estimation model 3 and the correction information 60 can not be limited to such an example.

[0203] For example, the estimation model 3 can further include the coupler 36. In this case, the correction information 60 can be constituted by the feature information 602 and the true value information 603. The process of step S201 can be constituted by acquiring the reference image 601, acquiring the feature quantity 6021 (feature information 602) related to the reference image 601 by inputting the reference image 601 to the extractor 35 and executing the operation process of the extractor 35, and acquiring the true value information 603. The process of step S203 can further include the process of deriving the feature quantity 604 from the feature quantity 6021 and the true value information 603 by the coupler 36.

[0204] In addition, for example, the estimation model 3 can further include the extractor 35 and the coupler 36. In this case, the feature information 602 can be constituted by the reference image 601. The process of step S201 can be constituted by acquiring the reference image 601 and the true value information 603. The correction information 60 can be constituted by the reference image 601 and the true value information 603. The process of step S203 can further include the process of deriving the feature quantity 604 from the reference image 601 and the true value information 603 by the extractor 35 and the coupler 36.

[0205] In addition, for example, in the line-of-sight estimation device 2, the extractor 35 can be omitted. In this case, the control section 21 can directly acquire the feature information 602. As an example, in a case where the feature information 602 is constituted by the feature quantity 6021, the process of extracting the feature quantity 6021 from the reference image 601 can be performed by another computer. The control section 21 can acquire the feature quantity 6021 from the other computer. As another example, the feature information 602 can be constituted by the reference image 601. Thereby, the coupler 36 can be constituted so as to accept the input of the reference image 601 and the true value information 603.

[0206] Figure 9 An example of a software configuration of the model generation device 1 that generates the estimation model 3A involved in the first modification example is schematically illustrated. Figure 10 An example of a software configuration of the line-of-sight estimation device 2 that utilizes the estimation model 3A involved in the first modification example is schematically illustrated. In the first modification example, the coupler 36 is omitted. Thereby, in the processing order of the model generation device 1 and the line-of-sight estimation device 2, the process of deriving the correction feature quantity from the feature information and the true value information is omitted. The estimator 32A is constituted so as to accept the input of the feature information, the true value information, and the feature quantity related to the object image. That is, the feature information and the true value information are directly input to the estimator 32A instead of the correction feature quantity. Except for these points, the first modification example is constituted similarly to the above-described embodiment. The estimator 32A has one or a plurality of fully connected layers 321A similarly to the above-described embodiment. The estimation model 3A is constituted by the extractor 31 and the estimator 32A.

[0207] As Figure 9 illustrated, in the first modification example, the model generation device 1 can generate the learned estimation model 3A (the extractor 31 and the estimator 32A) and the extractor 35 by the same processing order as the above-described embodiment except for the point that the training process of the above-described coupler 36 is omitted. In the above-described step S107, the control section 11 generates information related to the learned estimation model 3A and the extractor 35 generated by machine learning as the learning result data 125A. Further, the control section 11 saves the generated learning result data 125A to a prescribed storage area. The learning result data 125A can be provided to the line-of-sight estimation device 2 at any timing.

[0208] Similarly, as Figure 10As shown, the gaze estimation device 2 can estimate the gaze direction of the subject R by the same processing sequence as the above-described embodiment, except that the operation processing of the coupler 36 is omitted. In the above-described step S201, the control section 21 acquires the reference image 601 and the true value information 603. The control section 21 inputs the acquired reference image 601 to the extractor 35, and executes the operation processing of the extractor 35. Thereby, the control section 21 acquires the output value corresponding to the feature quantity 6021 (feature information 602) related to the reference image 601 from the extractor 35. In the first modification example, the correction information 60 is constituted by the feature quantity 6021 (feature information 602) and the true value information 603.

[0209] In the above-described step S203, the control section 21 estimates the gaze direction of the subject R appearing in the subject image 63 using the learned estimation model 3A. Specifically, the control section 21 inputs the acquired subject image 63 to the extractor 31, and executes the operation processing of the extractor 31. Through this operation processing, the control section 21 acquires the feature quantity 64 related to the subject image 63 from the extractor 31. Next, the control section 21 inputs the feature quantity 6021 (feature information 602), the true value information 603, and the feature quantity 64 to the estimator 32A, and executes the operation processing of the estimator 32A. Through this operation processing, the control section 21 can acquire the output value corresponding to the estimation result of the gaze direction of the subject R appearing in the subject image 63 from the estimator 32A.

[0210] According to the first modification example, as in the above-described embodiment, in the learned estimation model 3A, the gaze direction of the subject R can be appropriately estimated from the feature information 602 (feature quantity 6021), the true value information 603, and the subject image 63. By using the feature information 602 and the true value information 603, it is possible to achieve improvement of the estimation accuracy of the gaze direction of the subject R in step S203. In addition, in the case where the estimation processing of the gaze direction of the subject R is repeatedly performed, the feature quantity 6021 (feature information 602) derived through step S201 can be repeatedly used in each cycle. Correspondingly, it is possible to achieve speeding up of the estimation processing of the gaze direction of the subject R in step S203.

[0211] Further, in this first modification example, the extractor 35 can also be omitted in the gaze estimation device 2. In this case, the control section 21 can also acquire the feature information 602 directly as in the above-described embodiment. In the case where the feature information 602 is constituted by the reference image 601, the estimator 32A can also be constituted so as to accept input of the reference image 601, the true value information 603, and the feature quantity 64.

[0212] Figure 11 An example of a software configuration of a model generation device 1 that generates an estimation model 3B involved in the second modification example is schematically illustrated.Figure 12 An example of a software configuration of the line-of-sight estimation device 2 using the second modified example of the estimation model 3A is schematically illustrated. In the second modified example, the estimation model 3B further includes the extractor 35. That is, the estimation model 3B includes the extractors (31, 35) and the estimator 32B. Accordingly, the feature information is constituted by the reference image. Except for these points, the second modified example is configured similarly to the first modified example. The estimator 32B is configured similarly to the estimator 32A. The estimator 32B includes one or more fully connected layers 321B similarly to the first modified example.

[0213] As Figure 11 indicated, in the second modified example, the model generation device 1 can generate the completed estimation model 3B by the same processing sequence as the first modified example. In the above step S107, the control section 11 generates information related to the completed estimation model 3B generated by machine learning as the learning result data 125B. Further, the control section 11 saves the generated learning result data 125B to a prescribed storage area. The learning result data 125B can be provided to the line-of-sight estimation device 2 at any time.

[0214] Similarly, as Figure 12 indicated, the line-of-sight estimation device 2 can estimate the line-of-sight direction of the subject R by the same processing sequence as the first modified example. In the above step S201, the control section 21 acquires the reference image 601 and the true value information 603. In the above step S203, the control section 21 estimates the line-of-sight direction of the subject R appearing in the object image 63 using the completed estimation model 3B. Specifically, the control section 21 inputs the acquired object image 63 to the extractor 31 and executes the operation processing of the extractor 31. Through this operation processing, the control section 21 acquires the feature quantity 64 related to the object image 63 from the extractor 31. In addition, the control section 21 inputs the acquired reference image 601 to the extractor 35 and executes the operation processing of the extractor 35. Thereby, the control section 21 acquires the output value corresponding to the feature quantity 6021 related to the reference image 601 from the extractor 35. The processing sequence of each extractor (31, 35) can be arbitrary. Next, the control section 21 inputs the feature quantity 6021, the true value information 603, and the feature quantity 64 to the estimator 32B and executes the operation processing of the estimator 32B. Through this operation processing, the control section 21 can acquire the output value corresponding to the estimation result of the line-of-sight direction of the subject R appearing in the object image 63 from the estimator 32B.

[0215] According to the second variation, similarly to the above embodiment, after learning the inference model 3B, the gaze direction of the object R can be appropriately inferred based on the reference image 601 (feature information), the truth information 603, and the object image 63. By utilizing the feature information and the truth information 603, the accuracy of inferring the gaze direction of the object R in step S203 can be improved.

[0216] Figure 13A and Figure 13B An example of the software configuration of the model generation device 1 for generating the speculative model 3C involved in the third variation is illustrated schematically. Figure 14 An example of the software configuration of the gaze estimation device 2 using the inference model 3C involved in the third variation is illustrated. In this third variation, the gaze direction among features such as the gaze direction is represented using a heatmap. The heatmap represents the direction of a person's gaze through an image. The value of each pixel in the heatmap corresponds, for example, to the degree to which the person is gazing at that location. If the sum of the values ​​of the pixels is normalized to 1, the value of each pixel can represent the probability that the person is gazing at that location.

[0217] Accordingly, Figure 13A and Figure 13B As shown, each extractor (31, 35, 41) is replaced by a converter (31C, 35C, 41C). Converter 31C is an example of a first converter, and converter 35C is an example of a second converter. Each converter (31C, 35C, 41C) is configured to accept an image of a person's eyes as input and output a heatmap derived from the input image related to the person's gaze direction. That is, each converter (31C, 35C, 41C) is configured to convert an image of a person's eyes into a heatmap related to the gaze direction.

[0218] like Figure 13A As shown, in learning model 4, extractor 41 is replaced by transformer 41C, and inferrer 43 is omitted. Transformer 41C includes convolutional layer 415, pooling layer 416, unpooling layer 417, and deconvolutional layer 418. Unpooling layer 417 is configured to perform the inverse operation of pooling processing of pooling layer 416. Deconvolutional layer 418 is configured to perform the inverse operation of convolutional operation of convolutional layer 415.

[0219] The number of layers 415-418 can be appropriately determined according to the implementation method. The anti-pooling layer 417 and the deconvolution layer 418 are positioned closer to the output side than the convolutional layer 415 and the pooling layer 416. Figure 13AIn the example, the convolution layer 415 disposed at the most input side constitutes the input layer, and the deconvolution layer 418 disposed at the most output side constitutes the output layer. However, the structure of the converter 41C can not be limited to such an example, and can be appropriately decided according to the embodiment. The converter 41C can include other kinds of layers such as a normalization layer, an exit layer, and the like. In the present modification example, as with the above embodiment, the machine learning of the converter 41C is performed first, and the trained converter 41C is used as each converter (31C, 35C). Thus, the structure of each converter (31C, 35C) is derived from the converter 41C. Further, as with the above embodiment, each converter (31C, 35C) can use a common converter or different converters.

[0220] In addition, as shown in Figure 13B and Figure 14 In the present modification example, the estimation model 3C includes the converter 31C and the estimator 32C. The feature information is constituted by the heat map related to the line-of-sight direction of the eye observing the prescribed direction, which is derived from the reference image in which the person (the subject R) observing the prescribed direction is reflected. The estimator 32C is configured to accept the input of the heat map derived from the subject image, the feature information, and the true value information, and output an output value corresponding to the estimation result of the line-of-sight direction of the person reflected in the subject image. In the present modification example, the feature information is constituted by the heat map related to the line-of-sight direction of the eye observing the prescribed direction, which is derived from the reference image. The true value information is converted into a heat map related to the true value of the prescribed direction. Accordingly, the input of the heat map derived from the subject image, the feature information, and the true value information is constituted by the heat map derived from the subject image, the heat map (feature information) derived from the reference image, and the heat map derived from the true value information.

[0221] In the example of Figure 13B and Figure 14 The estimator 32C includes, in order from the input side, the concatenation layer 325, the convolution layer 326, and the conversion layer 327. The concatenation layer 325 is configured to concatenate each of the input heat maps. The conversion layer 327 is configured to convert the output from the convolution layer 326 into the estimation result of the line-of-sight direction. The concatenation layer 325 and the conversion layer 327 can be appropriately constituted by a plurality of neurons (nodes). Further, the structure of the estimator 32 can not be limited to such an example, and can be appropriately decided according to the embodiment. The estimator 32C can include other kinds of layers such as a pooling layer, a fully connected layer, and the like. Except for these points, the third modification example is constituted as with the above embodiment. The model generation device 1 generates the trained estimation model 3C through the same processing sequence as the above embodiment. In addition, the line-of-sight estimation device 2 estimates the line-of-sight direction of the subject R using the trained estimation model 3C through the same processing sequence as the above embodiment.

[0222] (Processing sequence of the model generation device)

[0223] like Figure 13A As shown, in step S102 above, the control unit 11 performs machine learning on the converter 41C using multiple datasets 120. For example, firstly, the control unit 11 inputs the learning images 121 from each dataset 120 into the converter 41C and performs computational processing on the converter 41C. Thus, the control unit 11 obtains the output value from the converter 41C corresponding to the heatmap converted from the learning images 121.

[0224] Furthermore, the control unit 11 converts the corresponding forward resolution information 123 into a heatmap 129. The method for converting the forward resolution information 123 into a heatmap 129 can be appropriately selected according to the implementation. For example, the control unit 11 prepares an image of the same size as the heatmap output by the converter 41C. Then, in the prepared image, the control unit 11 configures a predetermined distribution, such as a Gaussian distribution, centered on the position corresponding to the true value of the viewing direction represented by the forward resolution information 123. The maximum value of the distribution can be appropriately determined. Thus, the forward resolution information 123 can be converted into a heatmap 129.

[0225] Next, the control unit 11 calculates the error between the output value obtained from the converter 41C and the heatmap 129. The subsequent machine learning processing can be the same as in the embodiment described above. The control unit 11 uses the gradient of the calculated error in the output value to sequentially calculate the error of each operational parameter of the converter 41C from the output side using the error backpropagation method, and updates the value of each operational parameter based on the calculated error.

[0226] Through the aforementioned series of update processes, the control unit 11 adjusts the values ​​of each operational parameter of the converter 41C for each dataset 120 to reduce the sum of errors in the calculated output values. The control unit 11 can also repeatedly adjust the values ​​of each operational parameter of the converter 41C until the specified conditions are met. The result of this machine learning is that the control unit 11 can generate a trained converter 41C for each dataset 120, which acquires the ability to appropriately convert an image reflecting a person's eyes into a heatmap related to the gaze direction of that eye.

[0227] Next, as Figure 13B As shown, in step S103 above, the control unit 11 converts the converter 41C into individual converters (31C, 35C). Thus, the control unit 11 prepares a learning model consisting of the prediction model 3C and the converter 35C.

[0228] In the step S104 described above, the control section 11 acquires the learning feature information 502C using the converter 35C. Specifically, the control section 11 acquires the learning reference image 501 and the learning true value information 503 as in the embodiment described above. Next, the control section 11 inputs the acquired learning reference image 501 to the converter 35C and executes the operation processing of the converter 35C. Through this operation processing, the control section 11 acquires, from the converter 35C, an output value corresponding to the learning heat map 5021C related to the gaze direction of the eye observing the prescribed direction derived from the learning reference image 501. In the present modification, the learning feature information 502C is constituted by this heat map 5021C. In addition, the control section 11 converts the learning true value information 503 into a heat map 5031. In this conversion, the same method as the method of converting the correct answer information 123 into the heat map 129 described above can be used. Through these operation processings, the control section 11 can acquire the learning correction information constituted by two heat maps (5021C, 5031). As in the embodiment described above, the control section 11 can also acquire the learning reference image 501 and the learning true value information 503 for a plurality of different prescribed directions respectively. Also, the control section 11 can acquire the heat maps (5021C, 5031) for the plurality of different prescribed directions respectively by executing the respective operation processings. In the step S105 described above, the control section 11 acquires a plurality of learning data sets 51 as in the embodiment described above.

[0229] In the step S106 described above, the control section 11 implements the machine learning of the estimation model 3 using the acquired plurality of learning data sets 51. In the present modification, the control section 11 inputs the learning target image 53 of each learning data set 51 to the converter 31C and executes the operation processing of the converter 31C. Through this operation processing, the control section 11 acquires, from the converter 31C, an output value corresponding to the heat map 54C converted from the learning target image 53. The control section 11 inputs each heat map (5021C, 5031, 54C) to the estimator 32C and executes the operation processing of the estimator 32C. Through this operation processing, the control section 11 acquires, from the estimator 32C, an output value corresponding to the estimation result of the gaze direction of the subject reflected in the learning target image 53.

[0230] Next, the control section 11 calculates the error between the output value obtained from the estimator 32C and the corresponding correct answer information 55. The processing of the machine learning thereafter can be the same as in the embodiment described above. The control section 11 calculates the error of the value of each operation parameter of the learning model from the output side in order using the gradient of the error of the calculated output value by the error backpropagation method and updates the value of each operation parameter based on the calculated error.

[0231] The control section 11 performs the above series of update processes in conjunction with the operation processing of the converter 35C and the operation processing of the estimation model 3C, thereby adjusting the values of the respective operation parameters of the learning model with respect to the learning reference image 501, the learning true value information 503, and each of the learning data sets 51 so as to make the sum of the errors of the calculated output values smaller. The control section 11 can also repeatedly adjust the values of the respective operation parameters of the learning model until a prescribed condition is satisfied. As a result of this machine learning, the control section 11 can generate a trained learning model with respect to each of the learning data sets 51, which has acquired the ability to appropriately estimate the gaze direction of a person from the learning reference image 501, the learning true value information 503, and the learning target image 53.

[0232] Further, the subjects as respective origins can also be identified in the same manner as in the above-described embodiments, to use the learning reference image 501, the learning true value information 503, and the plurality of learning data sets 51 obtained from the same subject for the machine learning of the learning model. In addition, during the period in which the values of the respective operation parameters are repeatedly adjusted, the heat map 5031 obtained from the learning true value information 503 can be repeatedly used, whereby the process of converting the learning true value information 503 into the heat map 5031 can be omitted. The learning true value information 503 can also be converted into the heat map 5031 in advance. In addition, each of the converters (31C, 35C) is trained by the machine learning of the above-described converter 41C to acquire the ability to convert an image of an existing person's eye into a heat map related to the gaze direction of that eye. Therefore, in the above-described update process, the process of adjusting the values of the respective operation parameters of each of the converters 31C, 35C can be omitted. In this case, during the period in which the values of the respective operation parameters are repeatedly adjusted, the operation results of each of the converters (31C, 35C) can be repeatedly used. That is, the operations to derive each of the heat maps (5021C, 5031) can not be repeatedly performed.

[0233] In the above-described step S107, the control section 11 generates information related to the trained estimation model 3C and the trained converter 35C generated by the machine learning as learning result data 125C. The control section 11 saves the generated learning result data 125 to a prescribed storage area. The learning result data 125C can be provided to the gaze estimation device 2 at any time.

[0234] (Processing sequence of the gaze estimation device)

[0235] As Figure 14 indicated in the present modified example, the information acquisition section 211 has the trained converter 35C by retaining the learning result data 125C, and the estimation section 213 has the trained estimation model 3C. The trained estimation model 3C is provided with the trained converter 31C and the estimator 32C.

[0236] In the step S201 described above, the control section 21 acquires the reference image 601 and the true value information 603. The control section 21 inputs the acquired reference image 601 to the learned converter 35C, and executes the operation processing of the converter 35C. Thereby, the control section 21 acquires, from the learned converter 35C, the output value corresponding to the heat map 6021C derived from the reference image 601, which is related to the gaze direction of the eye observing the prescribed direction. The heat map 6021C is an example of the second heat map. In this modified example, the feature information 602C is constituted by this heat map 6021C. In addition, the control section 21 converts the true value information 603 into the heat map 6031 related to the true value of the prescribed direction. In this conversion, the same method as the method of converting the above-mentioned correct answer information 123 into the heat map 129 can be used. The heat map 6031 is an example of the third heat map. Thereby, the control section 21 can acquire the correction information constituted by each heat map (6021C, 6031). Further, the control section 21 can acquire the reference image 601 and the true value information 603 of each of a plurality of different prescribed directions, as in the above-mentioned embodiment. Moreover, the control section 21 can acquire the heat map (6021C, 6031) of each of a plurality of different prescribed directions by executing the respective operation processing.

[0237] In the step S203 described above, the control section 21 estimates the gaze direction of the object person R appearing in the object image 63, using the learned estimation model 3C. Specifically, the control section 21 inputs the acquired object image 63 to the learned converter 31C, and executes the operation processing of the converter 31C. By this operation processing, the control section 21 acquires, from the learned converter 31C, the output value corresponding to the heat map 64C derived from the object image 63, which is related to the gaze direction of the object person R. The heat map 64C is an example of the first heat map. Next, the control section 21 inputs each heat map (6021C, 6031, 64) to the learned estimator 32C, and executes the operation processing of the estimator 32C. By this operation processing, the control section 21 can acquire, from the learned estimator 32C, the output value corresponding to the estimation result of the gaze direction of the object person R appearing in the object image 63.

[0238] According to the third modification, as with the above-described embodiment, in the learned model 3C, the line-of-sight direction of the subject R can be appropriately estimated from the feature information 602C, the true value information 603, and the subject image 63. By using the feature information 602C and the true value information 603, it is possible to achieve an improvement in the estimation accuracy of the line-of-sight direction of the subject R in step S203. In addition, the number of parameters tends to increase and the operation speed tends to decrease in the fully connected layer as compared with the convolution layer. In contrast, according to the third modification, even if the fully connected layer is not used, it is possible to configure the respective converters (31C, 35C) and the estimator 32C. Therefore, it is possible to make the amount of information of the learned model 3C small, and it is possible to achieve an improvement in the processing speed of the learned model 3C. Furthermore, by adopting the common heat map form as the data form on the input side, it is possible to make the structure of the estimator 32C relatively simple, and by easily integrating the respective information (the feature information, the true value information, and the subject image) in the estimator 32C, it is expected to improve the estimation accuracy of the estimator 32C.

[0239] In addition, in the third modification, the configuration of the learned model 3C can not be limited to such an example. The true value information 603 can be directly input to the estimator 32C without being converted into the heat map 6031. The feature information 602C can be input to the estimator 32C in a form different from the heat map 6021C. For example, the feature information 602C can be input to the estimator 32C in the form of the feature amount as with the above-described embodiment. Before being input to the estimator 32C, the feature information 602C and the true value information 603 can be coupled.

[0240] In addition, the estimator 32C can output the estimation result of the line-of-sight direction in the form of the heat map. In this case, the conversion layer 327 can be omitted in the estimator 32C. The control section 21 can determine the line-of-sight direction of the subject R on the basis of the center of gravity of the heat map, the position of the pixel of the maximum value, or the like. It is easier to estimate the true value from the heat map for training as compared with estimating the numerical value from the heat map for training, and it is easy to generate a learned model having a high estimation accuracy. Therefore, by adopting the heat map as the data form on the input side and the output side, it is expected to improve the estimation accuracy of the line-of-sight direction of the learned model 3C. In addition to this, a scenario in which an organ point of the face of the subject R is detected together with the line-of-sight direction is assumed. In this case, in the detection method in recent years, the detection result of the organ point of the face is sometimes expressed in the form of the heat map. According to this method, it is possible to merge the heat map indicating the estimation result of the line-of-sight direction into the heat map indicating the detection result of the organ point of the face, and it is possible to output each result in a single display. Furthermore, it is possible to configure each learned model in a single form, and thus it is possible to improve the real-time performance. In addition, in this method, at least any one of the true value information 603 and the feature information 602C can be input to the estimator 32C in a form different from the heat map.

[0241] Further, for example, in the gaze estimation device 2, the converter 35C can be omitted. In this case, the control section 21 can directly acquire the feature information 602C. As an example, in a case where the feature information 602C is constituted by the heat map 6021C, the process of converting the reference image 601 into the heat map 6021C can be performed by another computer. The control section 21 can acquire the heat map 6021C from the other computer. As another example, the feature information 602C can be constituted by the reference image 601. Thereby, the estimator 32C can be configured to accept input of the reference image 601.

[0242] <4.3>

[0243] In the above-described embodiments, each extractor (31, 35, 41) uses a convolutional neural network. Each estimator (32, 43) and the coupler 36 use a fully connected neural network. However, the type of neural network that can be used for each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36 can not be limited to such examples. Each extractor (31, 35, 41) can use a fully connected neural network, a recurrent neural network, or the like. Each estimator (32, 43) and the coupler 36 can use a convolutional neural network, a recurrent neural network.

[0244] Further, in the learning model 30, each component can not necessarily be separated. A combination of two or more components can be constituted by one neural network. For example, the estimation model 3 (the extractor 31 and the estimator 32) can be constituted by one neural network.

[0245] Further, the type of machine learning model that constitutes each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36 can not be limited to a neural network. Each extractor (31, 35, 41), each estimator (32, 43), and the coupler 36 can use, for example, a support vector machine, a regression model, a decision tree model, or the like.

[0246] Further, in the above-described embodiments, the estimation model 3, the extractor 35, and the coupler 36 that have completed learning can be generated by another computer other than the model generation device 1. In a case where the machine learning of the learning model 4 is performed by the other computer, the process of step S102 can be omitted from the processing sequence of the model generation device 1. In a case where the machine learning of the learning model 30 is performed by the other computer, the processes of steps S103 to S107 can be omitted from the processing sequence of the model generation device 1. The first acquisition section 112 and the second acquisition section 113 can be omitted from the software configuration of the model generation device 1. In a case where the gaze estimation system 100 does not use the results of the machine learning of the model generation device 1, the model generation device 1 can be omitted from the gaze estimation system 100.

[0247] <4.4>

[0248] In the above-described embodiments, for example, the correction information 60 can also be provided in advance by performing the process of step S201 and the like within the initial setting process. In this case, the process of step S201 can be omitted from the processing sequence of the line-of-sight estimation device 2. In addition, in a case where the correction information 60 is not changed after being acquired, the extractor 35 and the coupler 36 that have completed learning can be omitted or deleted in the line-of-sight estimation device 2. At least a part of the process of acquiring the correction information 60 can be performed by another computer. In this case, the line-of-sight estimation device 2 can also acquire the correction information 60 by acquiring the operation result of the other computer.

[0249] In addition, in the above-described embodiments, the line-of-sight estimation device 2 can also not repeatedly perform the line-of-sight direction estimation process. In this case, the process of step S205 can be omitted from the processing sequence of the line-of-sight estimation device 2.

[0250] In addition, in the above-described embodiments, the data set 120 can also not be used in the acquisition of each learning data set 51 and the learning correction information 50. In addition to this, in a case where machine learning of the learning model 4 is performed by another computer, the process of step S101 can be omitted from the processing sequence of the model generation device 1. The collection unit 111 can be omitted from the software configuration of the model generation device 1.

[0251] Explanation of Reference Signs

[0252] 1 … model generation device, 11 … control unit, 12 … storage unit, 13 … communication interface, 14 … external interface, 15 … input device, 16 … output device, 17 … driver, 111 … collection unit, 112 … first acquisition unit, 113 … second acquisition unit, 114 … machine learning unit, 115 … saving processing unit, 120 … data set, 121 … learning image, 123 … correct answer information, 125 … learning result data, 81 … model generation program, 91 … storage medium, 2 … line-of-sight estimation device, 21 … control unit, 22 … storage unit, 23 … communication interface, 24 … external interface, 25 … input device, 26 … output device, 27 … driver, 211 … information acquisition unit, 212 … image acquisition unit, 213 … estimation unit, 214 … output unit, 261 … display, M … marker, 82 … line-of-sight estimation program, 92 … storage medium, 30 … learning model, 3 … estimation model, 31 … extractor (first extractor), 311 … convolution layer, 312 … pooling layer, 32 … estimator, 321 … fully connected layer, 35 … extractor (second extractor), 351 … convolution layer, 352 … pooling layer, 36 … coupling unit, 361 … fully connected layer, 4 … learning model, 41 … extractor, 411 … convolution layer, 412 … pooling layer, 43 … estimator, 431 … fully connected layer, 50 … learning correction information, 501 … learning reference image, 502 … learning feature information, 5021 … feature amount, 503 … learning true value information, 504 … feature amount, 51 … learning data set, 53 … learning object image, 54 … feature amount, 55 … correct answer information, 60 … correction information, 601 … reference image, 602 … feature information, 6021 … feature amount (second feature amount), 603 … true value information, 604 … feature amount (correction feature amount), 63 … object image, 64 … feature amount (first feature amount), R … object person, S … camera.

Claims

1. A gaze estimation device comprising: an information acquisition unit configured to acquire correction information including feature information related to a gaze of an eye of a subject observing a predetermined direction and true value information indicating a true value of the predetermined direction observed by the eye of the subject; an image acquisition unit configured to acquire a subject image representing the eye of the subject; an estimation unit configured to perform machine learning of an estimation model using a plurality of learning data sets, and estimate a gaze direction of the subject represented in the subject image using the learned estimation model, the learned estimation model being trained to output an output value suitable for correct answer information indicating a true value of a gaze direction of a subject represented in a learning subject image with respect to input of learning correction information and the learning subject image obtained from the subject by performing the machine learning for each of the learning data sets, the plurality of learning data sets each being constituted by a combination of the learning subject image and the correct answer information, the estimation of the gaze direction being constituted by inputting the acquired subject image and the correction information into the learned estimation model, and performing an operation process of the learned estimation model to acquire an output value corresponding to an estimation result of the gaze direction of the subject represented in the subject image from the learned estimation model; and an output unit configured to output information related to the estimation result of the gaze direction of the subject. 2.The gaze estimation device according to claim 1, wherein the correction information includes the feature information and the true value information corresponding to a plurality of different predetermined directions, respectively. 3.The gaze estimation device according to claim 1 or 2, wherein the correction information including the feature information and the true value information is constituted by including a correction feature quantity related to correction derived by coupling the feature information and the true value information, the learned estimation model includes a first extractor and an estimator, the operation process of the learned estimation model is constituted by: inputting the acquired subject image into the first extractor, and performing an operation process of the first extractor to acquire an output value corresponding to a first feature quantity related to the subject image from the first extractor; and inputting the correction feature quantity and the acquired first feature quantity into the estimator, and performing an operation process of the estimator. 4.The gaze estimation device according to claim 3, wherein the feature information is constituted by a second feature quantity related to a reference image representing the eye of the subject observing the predetermined direction, the information acquisition unit includes a coupler, the acquisition of the correction information is constituted by: acquiring the second feature quantity; acquiring the true value information; and inputting the acquired second feature quantity and the true value information into the coupler, and performing an operation process of the coupler to acquire an output value corresponding to the correction feature quantity from the coupler. 5.The gaze estimation device according to claim 4, wherein the information acquisition unit further includes a second extractor, the acquisition of the second feature quantity is constituted by: inputting a reference image representing the eye of the subject observing the predetermined direction into the second extractor, and performing an operation process of the second extractor to acquire an output value corresponding to the second feature quantity from the second extractor, and the acquisition of the true value information is constituted by: inputting a true value image representing the predetermined direction into the second extractor, and performing an operation process of the second extractor to acquire an output value corresponding to the true value information from the second extractor. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The acquisition of the second feature quantity is configured by: acquiring the reference image; and inputting the acquired reference image into the second extractor and executing the operation processing of the second extractor, thereby acquiring an output value corresponding to the second feature quantity from the second extractor.

6. The line-of-sight estimation device according to claim 1 or 2, wherein the learned estimation model includes a first extractor and an estimator, the execution of the operation processing of the learned estimation model is configured by: inputting the acquired object image into the first extractor and executing the operation processing of the first extractor, thereby acquiring an output value corresponding to a first feature quantity related to the object image from the first extractor; and inputting the feature information, the true value information, and the acquired first feature quantity into the estimator and executing the operation processing of the estimator.

7. The line-of-sight estimation device according to claim 6, wherein the feature information is configured by a second feature quantity related to a reference image of an eye of the object who is looking in the prescribed direction, the information acquisition unit includes a second extractor, the acquisition of the correction information is configured by: acquiring the reference image; inputting the acquired reference image into the second extractor and executing the operation processing of the second extractor, thereby acquiring an output value corresponding to the second feature quantity from the second extractor; and acquiring the true value information.

8. The line-of-sight estimation device according to claim 1 or 2, wherein the feature information is configured by a reference image of an eye of the object who is looking in the prescribed direction, the learned estimation model includes a first extractor, a second extractor, and an estimator, the execution of the operation processing of the learned estimation model is configured by: inputting the acquired object image into the first extractor and executing the operation processing of the first extractor, thereby acquiring an output value corresponding to a first feature quantity related to the object image from the first extractor; inputting the reference image into the second extractor and executing the operation processing of the second extractor, thereby acquiring an output value corresponding to a second feature quantity related to the reference image from the second extractor; and inputting the acquired first feature quantity, the acquired second feature quantity, and the true value information into the estimator and executing the operation processing of the estimator.

9. The line-of-sight estimation device according to claim 1 or 2, wherein the learned estimation model includes a first converter and an estimator, the execution of the operation processing of the learned estimation model is configured by: inputting the acquired object image into the first converter and executing the operation processing of the first converter, thereby acquiring an output value corresponding to a first heat map related to the line-of-sight direction of the object from the first converter; and inputting the acquired first heat map, the feature information, and the true value information into the estimator and executing the operation processing of the estimator.

10. The line-of-sight estimation device according to claim 9, wherein ​ ​ ​ ​ The feature information is constituted by a second heat map related to a line-of-sight direction of an eye observing the prescribed direction, the second heat map being derived from a reference image of an eye of the subject observing the prescribed direction, The information acquisition section has a second converter, The acquisition of the correction information is constituted by: acquiring the reference image; inputting the acquired reference image into the second converter and executing an operation process of the second converter, thereby acquiring an output value corresponding to the second heat map from the second converter; acquiring the true value information; and converting a third heat map related to a true value of the prescribed direction into the true value information, The first heat map, the feature information, and the true value information are inputted into the estimator by inputting the first heat map, the second heat map, and the third heat map into the estimator.

11. The line-of-sight estimation device according to claim 1, wherein The acquisition of the object image and the estimation of the line-of-sight direction of the subject by the estimation section are repeatedly executed by the image acquisition section.

12. The line-of-sight estimation device according to claim 1, wherein The information acquisition section acquires the correction information by observing a line-of-sight of the subject with a sensor after outputting an instruction to the subject to observe a prescribed direction.

13. A line-of-sight estimation method, comprising the steps of: acquiring correction information including feature information related to a line-of-sight of an eye of a subject observing a prescribed direction and true value information indicating a true value of the prescribed direction observed by the eye of the subject; acquiring an object image of an eye of a subject; implementing machine learning of an estimation model using a plurality of learning data sets, and estimating a line-of-sight direction of the subject imaged in the object image using a learned estimation model generated by the machine learning, the learned estimation model is trained to output an output value suitable for correct answer information by executing the machine learning for each of the learning data sets, the correct answer information indicating a true value of a line-of-sight direction of a subject imaged in a learning object image, the learning data sets being constituted by combinations of the learning object image and the correct answer information, the plurality of learning data sets are constituted by combinations of the learning object image and the correct answer information, the estimation of the line-of-sight direction is constituted by inputting the acquired object image and the correction information into the learned estimation model and executing an operation process of the learned estimation model, thereby acquiring an output value corresponding to an estimation result of the line-of-sight direction of the subject imaged in the object image from the learned estimation model; and outputting information related to the estimation result of the line-of-sight direction of the subject.

14. A model generation device, comprising: a first acquisition section that acquires learning correction information including learning feature information related to a line-of-sight of an eye of a subject observing a prescribed direction and learning true value information indicating a true value of the prescribed direction observed by the eye of the subject; ​ a second acquisition unit that acquires a plurality of learning data sets each composed of a combination of a learning object image that represents an eye of a subject and correct answer information that represents a true value of a line of sight direction of the subject represented in the learning object image; and a machine learning unit that performs machine learning of a prediction model using the acquired plurality of learning data sets, the performance of the machine learning being constituted by training the prediction model so as to output an output value appropriate for the correct answer information corresponding thereto with respect to input of the learning object image and the learning correction information.

15. A model generation method comprising the following steps performed by a computer: acquiring learning correction information including learning feature information related to a line of sight of an eye of a subject observing a prescribed direction and learning true value information representing a true value of the prescribed direction observed by the eye of the subject; acquiring a plurality of learning data sets each composed of a combination of a learning object image that represents an eye of a subject and correct answer information that represents a true value of a line of sight direction of the subject represented in the learning object image; and performing machine learning of a prediction model using the acquired plurality of learning data sets, the performance of the machine learning being constituted by training the prediction model so as to output an output value appropriate for the correct answer information corresponding thereto with respect to input of the learning object image and the learning correction information.

Citation Information

Patent Citations

  • Information processing apparatus for estimating person's line of sight and estimation method, and learning device and learning method

    JP2019028843A

  • Neural network training for three dimensional (3D) gaze prediction with calibration parameters

    CN110321773A