Information processing method and information processing device

By using three-dimensional computer graphics technology to generate a three-dimensional model of virtual avatars in line of sight estimation, data deviation and error problems are solved, and the accuracy of the line of sight estimation model and data set quality are improved.

CN119948531APending Publication Date: 2025-05-06SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068135.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-11
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to collect large-scale and unbiased data in line of sight estimation, resulting in error and data deviation problems, affecting the accuracy of line of sight estimation.

Method used

By using three-dimensional computer graphics technology to generate a three-dimensional model of virtual avatars, control the state and rendering conditions of the model, and generate high-quality learning data sets for training the line-of-sight estimation model.

Benefits of technology

It improves the accuracy and processing accuracy of the line-of-sight estimation model, reduces data bias, and enhances the quality of the learning data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948531A_ABST
    Figure CN119948531A_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing method and an information processing device that can improve the quality of a training data set using a three-dimensional model based on a three-dimensional CG. The information processing apparatus controls a state of a three-dimensional model of a person in a three-dimensional virtual space and a rendering condition for rendering the three-dimensional model, and generates input data including both a three-dimensional model image and training data based on the state and the rendering condition of the three-dimensional model. The three-dimensional model image is an image obtained by rendering the three-dimensional model, and the training data includes correct data for the three-dimensional model. The present technology can be applied, for example, to a process for generating a training data set for a gaze estimation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an information processing method and an information processing apparatus, and more particularly, to an information processing method and an information processing apparatus for machine learning. Background Art

[0002] In recent years, with the development of deep learning technology, the accuracy of identifying people or actions in images has been significantly improved, and technologies for estimating non-verbal information about people in images have been actively developed. Among them, the technology for estimating the gaze of a person, such as the object of attention or the object of attention, has attracted great attention.

[0003] In gaze estimation techniques, machine learning is typically performed using a learning dataset consisting of a pair of input data and correct data. The input data pair consists of images of a sample person, similar to other non-verbal information estimation techniques. Here, the correct data includes gaze information indicating the correct answer for the sample's gaze direction.

[0004] In addition, in machine learning, it is important to collect a learning dataset. In other words, it is important to collect a large amount of high-quality learning data.

[0005] Here, there are roughly two types of methods for collecting learning data for visual line estimation.

[0006] The first method is a method for collecting learning data using samples including actual people. In this method, for example, learning data including a pair of facial images of a person with real-time actions and gaze directions is collected (for example, see Patent Document 1).

[0007] The second method is a method for collecting learning data using a sample including a three-dimensional model of a person (hereinafter referred to as an avatar) using three-dimensional computer graphics (CG). In this method, for example, learning data is collected, the learning data including a pair of monocular images obtained by rendering the eye region of one eye of the avatar and line of sight information based on the inclination of the avatar's eyeball object in the monocular image (for example, see Non-Patent Document 1).

[0008] Reference List

[0009] Patent Literature

[0010] Patent Document 1: Japanese Patent Application 2021-190041

[0011] Non-patent literature

[0012] Non-patent literature 1: Summary of the Invention

[0013] Problems to be solved by the present invention

[0014] The first method is currently unpopular because errors (noise) are included in the gaze information included in the correct data, and because it is difficult to collect large-scale, unbiased data. Data bias here refers to deviations in the position of the face in the facial image, unnecessary correlation between the position and tilt of the face when the gaze is directed in a specific direction, and so on. Specifically, for example, if there are many facial images with the face positioned downward in the image, or if the gaze direction is rightward, it is assumed that there are many facial images with the face pointing to the right, etc.

[0015] On the other hand, the second method has advantages such as the ability to accurately annotate gaze information, the ability to automatically generate large amounts of data, and the ability to generate a variety of eye region images by varying the contours and texture of the face. Consequently, the learning datasets collected using the second method have recently become commonly used to train gaze estimation models.

[0016] From the above, it is expected that the use of virtual images can improve the quality of learning datasets.

[0017] This technology was developed in light of this situation and aims to improve the quality of learning datasets using 3D models using 3D CG. As a result, the accuracy of machine learning is improved. Furthermore, the precision of processing using the learning models obtained through machine learning is enhanced.

[0018] Solution to the problem

[0019] In an information processing method according to the first aspect of the present technology, an information processing device controls the state of a three-dimensional model of a person in a three-dimensional virtual space and the rendering conditions for rendering the three-dimensional model, and based on the state of the three-dimensional model and the rendering conditions, generates input data including a three-dimensional model image and learning data, wherein the three-dimensional model image is an image obtained by rendering the three-dimensional model, and the learning data includes correct data about the three-dimensional model.

[0020] An information processing device according to a second aspect of the present technology includes: an estimation unit that performs estimation processing on a person by using a learning model, the learning model being generated by learning using a learning data set, the learning data set being a collection of learning data generated based on a state of a three-dimensional model of a person in a three-dimensional virtual space and a rendering condition for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering condition, the learning data including input data, the input data including a three-dimensional model image and correction data about the three-dimensional model, the three-dimensional model image being an image obtained by rendering the three-dimensional model.

[0021] In an information processing method according to the second aspect of the present technology, an information processing device performs estimation processing about a person by using a learning model, which is generated by learning using a learning data set, which is a set of learning data sets generated based on the state of a three-dimensional model of a person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering conditions, the learning data includes input data, the input data includes a three-dimensional model image and correction data about the three-dimensional model, and the three-dimensional model image is an image obtained by rendering the three-dimensional model.

[0022] In a first aspect of the present technology, the state of a three-dimensional model of a person in a three-dimensional virtual space and the rendering conditions for rendering the three-dimensional model are controlled, and based on the state of the three-dimensional model and the rendering conditions, input data including a three-dimensional model image and learning data are generated, where the three-dimensional model image is an image obtained by rendering the three-dimensional model, and the learning data includes correct data about the three-dimensional model.

[0023] In a second aspect of the present technology, estimation processing about a person is performed by using a learning model, the learning model is generated by learning using a learning data set, the learning data set is a collection of learning data generated based on the state of a three-dimensional model of the person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering conditions, the learning data includes input data, the input data includes a three-dimensional model image and correction data about the three-dimensional model, the three-dimensional model image is an image obtained by rendering the three-dimensional model. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 are diagrams illustrating examples of monocular images for learning and monocular images for fine adjustment.

[0025] Figure 2 This is a block diagram showing a first embodiment of an information processing system to which the present technology is applied.

[0026] Figure 3 It shows Figure 2 Block diagram of a configuration example of the learning dataset generation unit in .

[0027] Figure 4 Is used to illustrate the Figure 2 Flowchart of information processing performed by the information processing system in .

[0028] Figure 5 This is a flowchart for explaining the details of the learning data set generation process.

[0029] Figure 6 is a diagram illustrating an example of an object generated in a CG space.

[0030] Figure 7 : are diagrams illustrating examples of directions in which the orientation of an avatar's face changes.

[0031] Figure 8 : are diagrams illustrating examples of directions in which the orientation of an avatar's face changes.

[0032] Figure 9 : are diagrams illustrating examples of directions in which the orientation of an avatar's face changes.

[0033] Figure 10 : are diagrams illustrating examples of directions in which the orientation of an avatar's face changes.

[0034] Figure 11 : are diagrams illustrating examples of directions in which the orientation of an avatar's face changes.

[0035] Figure 12 is a diagram showing an example of a moving direction of a camera object.

[0036] Figure 13 is a diagram illustrating an example of a method for changing the position of the face of an avatar.

[0037] Figure 14 is a diagram illustrating an example of the position of a gaze point object.

[0038] Figure 15 is a diagram illustrating an example of the position of a gaze point object.

[0039] Figure 16 is a diagram illustrating an example of the relationship between the position of each object and the avatar image.

[0040] Figure 17 is a diagram illustrating an example of an avatar image.

[0041] Figure 18 is a diagram illustrating an example of an avatar image.

[0042] Figure 19 : is a diagram showing an example of an avatar in a state similar to the state of the white of the eye.

[0043] Figure 20 This is a block diagram showing a second embodiment of an information processing system to which the present technology is applied.

[0044] Figure 21 It shows Figure 20 Block diagram of a configuration example of the learning dataset generation unit in .

[0045] Figure 22 Is used to illustrate the Figure 21 Flowchart of details of the learning dataset generation process performed by the information processing system in.

[0046] Figure 23 This is a block diagram showing a third embodiment of an information processing system to which the present technology is applied.

[0047] Figure 24 is a block diagram showing a configuration example of a learning dataset supplementing unit.

[0048] Figure 25 Is used to illustrate the Figure 23 Flowchart of information processing performed by the information processing system in .

[0049] Figure 26 This is a flowchart for explaining the details of the learning data set generation process.

[0050] Figure 27 is a block diagram showing a configuration example of a computer. DETAILED DESCRIPTION

[0051] Hereinafter, a mode for carrying out the present technology will be described. The description will be given in the following order. 1. Background Technology

[0053] 2. First Implementation

[0054] 3. Second embodiment

[0055] 4. Third Implementation

[0056] 5. Modifications

[0057] 6. Others

[0058] <<1. Background Technology>>

[0059] First, the background of the present technology will be described.

[0060] As described in Non-Patent Document 1, conventional machine learning using an avatar's gaze estimation model uses a learning dataset consisting of a pair of monocular images and gaze information based on the monocular images. In this case, for example, when estimating the gaze direction based on an image obtained by capturing a real person using a gaze estimation model obtained through machine learning, the following four problems arise.

[0061] 1. Because gaze estimation is performed for each eye using a monocular image cut out from the captured image, there are cases where the estimated gaze directions of each eye do not intersect due to estimation errors. In this case, because the gaze estimation model cannot cope with this situation, another algorithm is needed to estimate the final gaze direction.

[0062] 2. When fine-tuning the gaze estimation model using a captured image obtained by capturing the user during operation, there are cases where the characteristics of the monocular image used for fine-tuning obtained from the captured image are greatly different from the characteristics of the monocular image used for learning included in the learning data. For example, Figure 1 A schematically shows an example of a monocular image used for learning, and Figure 1 B schematically illustrates an example of a monocular image used for fine-tuning. For example, the monocular image used for learning and the monocular image used for fine-tuning may differ significantly in terms of eye tilt, eye size, number of pixels, etc. Therefore, it may be difficult to process the monocular image used for learning and the monocular image used for fine-tuning in the same manner.

[0063] 3. In conventional machine learning models using avatar gaze estimation models, the resolution of the rendered monocular image is also essentially constant because the positional relationship between the avatar's eyes and the camera is essentially constant. However, during operation, the relative position between the camera and the person dynamically changes, so monocular images are not always obtained at the same resolution. Furthermore, due to perspective, distortion, and other factors in the captured image, the appearance of the eye varies depending on the position of the face in the captured image. Therefore, when performing gaze estimation using a monocular image cut out of the captured image, estimation accuracy can sometimes decrease.

[0064] 4. Non-Patent Document 1 does not consider the tilt of the face in the roll direction (the direction in which the head is tilted). Therefore, when cutting out an eye area image from a captured image in which the face is tilted in the roll direction, a mechanism for separately detecting the tilt of the face is required.

[0065] On the other hand, the present technology solves these problems and improves the accuracy of a learning model (hereinafter, referred to as an estimation model) that performs estimation processing on a person, such as a line of sight estimation model.

[0066] <<2. First embodiment>>

[0067] Next, we will refer to Figures 2 to 18 A first embodiment of the present technology is described.

[0068] <Configuration Example of Information Processing System 101>

[0069] First, refer to Figure 2 , describes a configuration example of an information processing system 101 to which the present technology is applied.

[0070] The information processing system 101 is a system that performs machine learning using an avatar that is a three-dimensional model of a person of CG and performs estimation processing on the person based on the result of the machine learning.

[0071] The information processing system 101 includes a learning dataset generating unit 111 , a learning dataset accumulating unit 112 , a learning unit 113 , and an estimating unit 114 .

[0072] The learning dataset generation unit 111 generates an avatar in a CG space and generates a learning dataset using the avatar. The learning dataset generation unit 111 accumulates the generated learning dataset in the learning dataset accumulation unit 112.

[0073] The learning unit 113 performs machine learning using the learning dataset accumulated in the learning dataset accumulation unit 112 and generates an estimation model for performing estimation processing on a person. The learning unit 113 supplies the estimation model to the estimation unit 114.

[0074] Estimation unit 114 includes a system, device, or program that performs a person estimation process based on a captured image obtained by capturing an actual person using an estimation model. For example, estimation unit 114 estimates non-verbal information (e.g., at least one of a state or characteristics) about the person based on the captured image. Furthermore, estimation unit 114 further performs various types of processing based on the results of the estimation process, as needed.

[0075] <Configuration Example of Learning Dataset Generating Unit 111>

[0076] Figure 3 A configuration example of the learning dataset generating unit 111 is shown.

[0077] The learning data set generating unit 111 includes an object generating unit 151 , a state control unit 152 , and a learning data generating unit 153 .

[0078] The object generation unit 151 generates various objects in the CG space. For example, the object generation unit 151 generates an avatar in the CG space, a camera object for controlling the rendering conditions of the avatar (virtual imaging of the avatar), and a gaze point object that indicates the position at which the avatar is looking. The object generation unit 151 provides information about each generated object to the state control unit 152.

[0079] The state control unit 152 controls the conditions for generating learning data by controlling, for example, the state of the CG space and the state of each object in the CG space. The state of the CG space includes, for example, states related to drawing conditions (virtual imaging conditions) at the time of drawing. Specifically, for example, the state of the CG space includes the light beam or lighting state of the CG space, the background of the CG space, and the like. The state of each object in the CG space includes, for example, the state of the above-mentioned virtual image, camera object, and gaze point object. The state control unit 152 provides information about the CG space (hereinafter, referred to as CG space information) to the learning data generating unit 153. The CG space information includes, for example, information about the state of the CG space and the state of each object in the CG space.

[0080] Furthermore, the state control unit 152 instructs the object generation unit 151 to generate an object in the CG space as necessary.

[0081] The learning data generating unit 153 generates learning data based on the state of the CG space and the state of each object in the CG space. The learning data includes input data and correct data for the input data.

[0082] The input data includes an image (hereinafter, referred to as an avatar image) obtained by rendering an avatar (virtually imaged by a camera object) in a CG space.

[0083] The correct data includes information indicating a correct answer for at least one of the states or characteristics of the avatar in the avatar image. For example, when learning an estimation model for estimating a person's line of sight (line of sight estimation model), the correct data includes line of sight information, i.e., information indicating a correct answer for the direction of the avatar's line of sight.

[0084] The learning data generating unit 153 adds the generated learning data to the learning data set accumulated in the learning data set accumulating unit 112 .

[0085] Note that an example will be described below in which the information processing system 101 learns a visual line estimation model for estimating a visual line direction of a person and estimates the visual line direction of the person using the visual line estimation model.

[0086] <Information Processing by Information Processing System 101>

[0087] Next, we will refer to Figure 4 The flowchart of FIG. 1 describes information processing performed by the information processing system 101 .

[0088] In step S1 , the learning dataset generating unit 111 performs a learning dataset generating process.

[0089] Here, we will refer to Figure 5 The flowchart of FIG. 4 describes the details of the learning dataset generation process.

[0090] In step S51, the object generation unit 151 generates each object in the CG space. Figure 6 As shown, the object generation unit 151 generates an avatar 201 , a camera object 202 , and a gaze point object 203 in a CG space.

[0091] In the above conventional technology, a monocular image can be generated by simply specifying the rotation angle of the eyeball of one eye of the avatar. On the other hand, in this technology, a mechanism is required to guide the sight lines of both eyes of the avatar to the same coordinates.

[0092] On the other hand, the avatar 201 includes an eyeball object in each eye and can control the direction of the eyeball of each eye separately. Therefore, for example, the avatar 201 can guide the sight of both eyes to the gaze point object 203 (the same coordinate in the CG space).

[0093] Note that, in the following, the type of the avatar 201 may differ depending on the person, but the mark is basically not distinguished.

[0094] The object generation unit 151 supplies the generated information about each object to the state control unit 152 .

[0095] In step S52 , the state control unit 152 updates the state of each object.

[0096] When operating a gaze estimation model, the conditions used to capture images of people serving as gaze estimation targets vary. For example, the relative position and pose between the person's face and the camera, as well as the direction of the person's eyes, can change in various ways. As a result, the position, size, and orientation of the person's face, as well as the direction of their eyes, change in the captured image.

[0097] On the other hand, for example, the state control unit 152 changes the state of each object in the CG space so that the position, size, and direction of the face of the avatar 201 and the direction of the eyeballs in the avatar image vary more widely for one avatar 201. Specifically, for example, the state control unit 152 changes the relative position and posture of the face of the avatar 201 and the camera object 202, and the position of the gaze point object 203 relative to the eyes of the avatar 201.

[0098] For example, Figure 7 As shown, the state control unit 152 changes the orientation of the face of the avatar 201 around the three axes of roll, pitch, and yaw. As a result, the orientation of the face of the avatar 201 changes in the direction of arrow A11 (roll direction), the direction of arrow A12 (pitch direction), and the direction of arrow A13 (yaw direction).

[0099] More specifically, for example, Figure 8 As shown in FIG. 1 , the roll angle of the face of the avatar 201 changes with a predetermined step size within a predetermined range. Figure 8 As shown in FIG. 1B , the roll angle is based on the direction in which the face of the avatar 201 faces forward (0°). Then, as Figure 8 As shown in A of FIG. 1 , the direction in which the face of the avatar 201 is tilted to the right is set to the negative direction, and as shown in FIG. Figure 8 As shown in C, the direction in which the face of the avatar 201 is tilted to the left is set as the positive direction.

[0100] For example, when the range of the roll angle change of the face of the avatar 201 is -25° to +25° and the step size is 5°, the roll angle of the face of the avatar 201 changes in 11 steps. As a result, the above-mentioned Problem 4 of the prior art is solved.

[0101] For example, Figure 9 As shown in FIG. 1 , the pitch angle of the face of the avatar 201 changes within a predetermined range by a predetermined step length. Figure 9 As shown in FIG. 1B , the pitch angle is based on the direction in which the face of the avatar 201 faces forward (0°). Then, as Figure 9 As shown in FIG. 1A , the direction in which the face of the avatar 201 is tilted upward is set to the negative direction, and as shown in FIG. Figure 9 As shown in C, the direction in which the face of the avatar 201 is tilted downward is set as the positive direction.

[0102] For example, in a case where the range of the change in the pitch angle of the face of the avatar 201 is -25° to +25°, and the step size is 5°, the pitch angle of the face of the avatar 201 changes in 11 steps.

[0103] For example, Figure 10 As shown in FIG. 1 , the deflection angle of the face of the avatar 201 changes within a predetermined range with a predetermined step length. Figure 10 As shown in FIG. 1B , the yaw angle is based on the direction in which the face of the avatar 201 faces forward (0°). Then, as Figure 10 As shown in A of FIG. 1 , the rightward direction of the face of the avatar 201 is set to the negative direction, and as Figure 10As shown in C, the leftward direction of the face of the avatar 201 is set as the positive direction.

[0104] For example, in a case where the range of the change in the yaw angle of the face of the avatar 201 is −25° to +25° and the step size is 5°, the yaw angle of the face of the avatar 201 changes in 11 steps.

[0105] Note that, for example, Figure 11 As shown in A to C of FIG, the direction of the face of the avatar 201 can be changed by combining two or more of the roll angle, the pitch angle, and the yaw angle.

[0106] In addition, for example, Figure 12 As shown, the state control unit 152 translates the camera target 202 in the left-right, up-down, and front-back directions relative to the face of the avatar 201. As a result, the relative position between the face of the avatar 201 and the camera target 202 changes. Furthermore, in conjunction with the change in the orientation of the avatar 201's face, the relative posture between the face of the avatar 201 and the camera target 202 changes.

[0107] Note that the relative position between the face of the avatar 201 and the camera object 202 may be changed by moving only the avatar 201 or moving both the avatar 201 and the camera object 202, for example.

[0108] For example, Figure 13 An example is shown in which a camera mounted on a laptop personal computer (PC) 204 is assumed to be the camera subject 202 .

[0109] For example, the state control unit 152 moves the position of the face of the avatar 201 within the virtual imaging range A21 of the camera object 202 .

[0110] For example, the distance of the avatar 201 to the camera object 202 changes in five stages from distance D1 to distance D5. In addition, at each distance, the position of the face of the avatar 201 in the up-down direction (height direction) and the left-right direction (lateral direction) changes. For example, at Figure 13 In the example shown in FIG. 2 , when the distance between the avatar 201 and the camera object 202 is distance D5, the face position of the avatar 201 is set to a total of 20 positions within the imaging range A21, including four positions at predetermined intervals in the vertical direction and five positions at predetermined intervals in the horizontal direction. Then, for example, by setting the face position of the avatar 201 to 20 positions at each distance, the relative position between the face of the avatar 201 and the camera object 202 changes in 100 different ways.

[0111] Note that, hereinafter, an example will be described in which the relative position between the avatar 201 and the camera object 202 is changed by moving the position of the camera object 202 .

[0112] Furthermore, for example, the relative posture between the face of the avatar 201 and the camera object 202 can be changed by changing only the posture of the camera object 202 or by changing the postures of both the avatar 201 and the camera object 202. Furthermore, not only the direction of the face of the avatar 201 but also the direction of the entire body of the avatar 201 can be changed.

[0113] Note that, hereinafter, an example will be described in which the direction of the face of the avatar 201 is changed to change the relative posture between the face of the avatar 201 and the camera object 202 .

[0114] In addition, for example, Figure 14 As shown, the state control unit 152 moves the gaze point object 203 in the left-right and up-down directions on the camera plane. The camera plane is, for example, a plane perpendicular to the optical axis of the camera object 202 at the front end of the camera object 202. Rectangular frames arranged in a grid pattern on the camera plane indicate candidates for the position of the gaze point object 203. Among them, the rectangular frame indicated by the diagonal line pattern indicates the current position of the gaze point object 203.

[0115] Figure 15 Another example of a candidate for the position of the gaze point object 203 is shown. Figure 12 For example, Figure 15 An example is shown in which a camera mounted on a laptop PC 204 is assumed to be the camera object 202 .

[0116] For example, on the camera plane P1 of the camera object 202, a total of 81 points (i.e., 9 points in the left-right direction × 9 points in the up-down direction, with the optical axis of the camera object 202 as the center) are set as candidates for the position of the gaze point object 203. For example, the interval between the candidate positions is set to 5 cm in the up-down direction and the left-right direction.

[0117] By moving the position of the gaze point object 203 in this manner, the relative position between the face of the avatar 201 and the gaze point object 203 changes.

[0118] Note that each foveation object 203 does not have to be arranged on the camera plane. In addition, the foveation objects 203 do not have to be arranged on the same plane.

[0119] As described above, the state control unit 152 changes the combination of the states of the objects by, for example, a predetermined algorithm to expand the variation of the avatar image. Then, in step S52, the state control unit 152 updates the state of each target to one of the combinations of the states of each target for which learning data has not yet been generated.

[0120] return Figure 5 In step S53, the state control unit 152 directs the line of sight of the avatar 201 toward the gazing point object 203. For example, the state control unit 152 calculates the relative position between the eyeball object of each eye of the avatar 201 and the gazing point object 203. Based on the relative position between the eyeball object of each eye of the avatar 201 and the gazing point object 203, the state control unit 152 calculates the direction (rotation angle) of the eyeball object of each eye when the line of sight of each eye of the avatar 201 faces the direction of the gazing point object 203.

[0121] State control unit 152 rotates the eyeball object for each eye of avatar 201 based on the calculated rotation angle. For example, the eyeball object is rotated so that a line connecting the center of the eyeball object and the center point of the iris of each eye of avatar 201 (the pupil centerline) points toward gaze point object 203. In this case, for example, the deviation between the pupil centerline and the actual fovea, as seen in real people, can be reflected. Furthermore, considering the dynamic movement of the gaze, eye movements such as saccades and drift can be reflected in the movement of the eyeball object.

[0122] As a result, the line of sight of the avatar 201 is directed toward the gaze point object 203. That is, both eyes of the avatar 201 are gazing in the direction of the gaze point object 203.

[0123] The state control unit 152 supplies the learning data generating unit 153 with CG space information including information indicating the state of the CG space and the state of each object in the CG space set by the processing of steps S52 and S53 .

[0124] In step S54, the learning data generating unit 153 generates the line of sight information of the avatar 201. Specifically, the learning data generating unit 153 generates the line of sight information of the avatar 201 based on one or more of the relative position and posture between the face of the avatar 201 and the camera object 202 in the CG space, the rotation angle of the eyeball object of each eye of the avatar 201, and the position of the gaze point object.

[0125] The gaze information of the avatar 201 includes information regarding the gaze direction of the avatar 201. For example, when viewed from the camera object 202, the gaze direction of the avatar 201 is represented by the rotation angle of the eyeball object for each eye of the avatar 201. Alternatively, for example, the gaze direction of the avatar is represented by the relative position of the point where the gaze of each eye of the avatar 201 intersects with the camera object 202. Alternatively, the gaze direction of the avatar is represented by, for example, the relative position of the gaze point object 203 with respect to the camera object 202. Alternatively, for example, the gaze direction of the avatar 201 can be represented by coordinates in the avatar image rather than by the relative position with respect to the camera object 202.

[0126] Note that it is desirable to narrow the line of sight direction of the avatar 201 in the line of sight information to a value such that even when the lines of sight of the person's respective eyes do not intersect due to estimation errors, etc. during operation, the line of sight direction of the person can be set to one. In this case, for example, the intersection of the lines of sight of the avatar's two eyes, the coordinates of the gaze point object 203, etc. are used as the line of sight direction of the avatar 201.

[0127] It should be noted that, for example, when the lines of sight of the two eyes of the virtual image 201 do not intersect, the learning data generating unit 153 can select a line of sight direction that is closer to the gaze point object 203 or can calculate a line of sight direction based on the line of sight directions of the two eyes.

[0128] This makes it possible to accurately estimate the direction of a person's gaze without performing calculations outside the gaze estimation model, even if the gazes of the person's two eyes do not intersect when operating. That is, the above-mentioned problem 1 of the prior art is solved.

[0129] In step S55, the learning data generation unit 153 generates an avatar image. For example, the learning data generation unit 153 renders an image of the avatar 201 captured by the camera object 202 in a CG space based on the relative position and posture of the avatar 201 and the camera object 202. As a result, an avatar image including the face of the avatar 201 is generated.

[0130] Please note that Figure 16 As shown, even if the positional relationship between the avatar 201 and the gaze point object 203 is the same, if the position of the camera object 202 is different, the generated avatar image is changed.

[0131] Figure 16Figure A shows an example where the positions of the camera object 202 and the gaze point object 203 are substantially the same. In this case, the line of sight of the avatar 201 is directed toward the camera object 202. Therefore, in the avatar image IM11 generated in this state, for example, the avatar 201 faces forward at approximately the center.

[0132] Figure 16 FIG. B shows an example in which the positions of the camera object 202 and the gaze point object 203 are separated from each other. In this case, the line of sight of the avatar 201 is directed in a direction different from the camera object 202. Therefore, in the avatar image IM12 generated in this state, for example, the avatar 201 faces diagonally to the lower left at the lower right corner of the image.

[0133] Therefore, even if the avatar 201 and the gaze point object 203 do not move, the gaze information must be recalculated if at least one of the relative position and posture of the avatar 201 or the camera object 202 changes. This is a process that is not performed in the prior art, where the eyeball is always located at the center of the monocular image.

[0134] Figure 17 and Figure 18 Shows an example of an avatar image.

[0135] Figure 17 An example of an avatar image in a case where the avatar 201 is not looking in the direction of the camera object 202 is shown.

[0136] Figure 18 A to C show examples of avatar images in a case where the direction of the face of the avatar 201 and the relative position between the face of the avatar 201 and the camera object 202 are different.

[0137] Figure 18 A in FIG. 1 shows an example of an avatar image in a case where the direction of the line of sight does not coincide with the direction of the face of the avatar 201 .

[0138] Figure 18 B and C show examples of avatar images in which the face and sight of the avatar 201 are oriented diagonally to the upper right. However, since the relative position of the camera object 202 to the avatar 201 is Figure 18 The difference between B and C in FIG. 1 is that the position and size of the avatar 201 in the avatar image are different.

[0139] Back to Figure 5, in step S56, the learning data generating unit 153 generates learning data based on the avatar image and the line of sight information, and adds the learning data to the learning data set. Specifically, the learning data generating unit 153 generates input data including (data of) the avatar image generated in the process of step S55. The learning data generating unit 153 generates correct data including the line of sight information of the avatar 201 generated in the process of step S54. The learning data generating unit 153 generates learning data including the input data and the correct data. As a result, the avatar image included in the input data and the line of sight information included in the correct data are paired. That is, the line of sight information represents the correct answer of the line of sight direction of the avatar 201 in the avatar image.

[0140] The learning data generating unit 153 adds the generated learning data to the learning data set accumulated in the learning data set accumulating unit 112 .

[0141] In step S57 , the state control unit 152 determines whether the quantity and quality of the learning data set are sufficient.

[0142] For example, the state control unit 152 determines whether the data amount of the learning dataset accumulated in the learning dataset accumulation unit 112 is sufficient.

[0143] In addition, for example, the state control unit 152 determines whether the variation of the learning dataset accumulated in the learning dataset accumulation unit 112 is sufficient. For example, the variation of the learning dataset is determined based on at least one of a variation of the avatar image included in the input data of each piece of learning data included in the learning dataset or a variation of the line of sight information included in the correct data.

[0144] For example, the variant of the avatar image is determined based on at least one of the position of the face, the size of the face, the direction of the face, the direction of the eyeballs, or other characteristics of the avatar 201 in the avatar image. As the characteristics of the avatar 201, for example, at least one of the race, gender, age, facial makeup, face size, skin color, etc. of the avatar is assumed.

[0145] The variation of the line-of-sight information is determined based on, for example, the variation of the line-of-sight direction represented by the line-of-sight information.

[0146] When the data amount of the learning dataset is still insufficient, or when the variation of the learning dataset is still insufficient, the state control unit 152 determines that at least one of the quantity and quality of the learning dataset is still insufficient, and the processing proceeds to step S58.

[0147] In step S58 , the learning data set generating unit 111 changes the avatar 201 as needed.

[0148] For example, if the variation of the learning data generated using the current avatar 201 is sufficient, the state control unit 152 instructs the object generation unit 151 to change the characteristics of the avatar 201 so as to expand the variation of the characteristics of the avatar 201 in the learning data set. For example, the state control unit 152 instructs the object generation unit 151 to change at least one of the race, gender, age, facial makeup, face size, or skin color of the avatar.

[0149] On the other hand, the object generation unit 151 generates a new avatar 201 in the CG space according to an instruction of the state control unit 152 and deletes the old avatar 201. The object generation unit 151 supplies information about the generated avatar 201 to the state control unit 152.

[0150] On the other hand, in a case where the variations of the learning data generated using the current avatar 201 are still insufficient, the learning data generating unit 153 determines to continue generating the learning data while maintaining the current avatar 201 .

[0151] Thereafter, the process returns to step S52 , and the processes of steps S52 to S58 are repeatedly performed until it is determined in step S57 that the quantity and quality of the learning data set are sufficient.

[0152] On the other hand, in step S57 , when the data amount of the learning dataset is sufficient and the variation of the learning dataset is sufficient, the state control unit 152 determines that the quantity and quality of the learning dataset are sufficient, and the learning dataset generation process ends.

[0153] Back to Figure 4 In step S2, the learning unit 113 performs machine learning using the generated learning dataset and generates a line of sight estimation model. Specifically, the learning unit 113 performs machine learning using the learning dataset accumulated in the learning dataset accumulation unit 112 (for example, using a neural network-based learning method). The learning unit 113 provides the line of sight estimation model obtained through machine learning to the estimation unit 114.

[0154] For example, a gaze estimation model is a neural network-based model that estimates the gaze direction of a person in a captured image based on pixel information of a captured image obtained by capturing a target person and feature point information of the person's face extracted from an input image.

[0155] It should be noted that, for example, a line of sight estimation model may also be generated by statistically analyzing the position and orientation of the face, the amount of movement of the iris, and the like in the avatar image, as well as the line of sight information of the correct data.

[0156] In step S3, estimation unit 114 performs gaze estimation processing using the generated gaze estimation model. For example, estimation unit 114 estimates the gaze direction of the target person based on a captured image obtained by capturing the person using the gaze estimation model. Specifically, for example, estimation unit 114 inputs the captured image into the gaze estimation model and obtains gaze information indicating the gaze direction of the person, which is output from the gaze estimation model.

[0157] It should be noted that, for example, the estimation unit 114 can use the gaze estimation model to perform various applications. For example, the estimation unit 114 can use the gaze estimation model to perform an application to estimate which part of the shared material the reviewer is viewing during an online meeting. As a result, the presenter can add a detailed description of the content that the reviewer was interested in during the presentation, or can supplement and explain content that the student may have overlooked.

[0158] After this, the information processing ends.

[0159] As described above, the quality of learning datasets using avatars can be improved. Specifically, for example, large learning datasets can be automatically generated with minimal deviation. This results in improved learning accuracy for gaze estimation models while minimizing the burden of generating learning datasets. Furthermore, the generated gaze estimation model can be used to estimate a person's gaze direction with high accuracy. This resolves Problem 3 of the prior art.

[0160] In addition, the traceability of learning data is improved.

[0161] For example, the learning data generation unit 153 stores the script used to generate each piece of learning data. Each script includes, for example, the type of the avatar 201, the state of the CG space, and the state of each object in the CG space when the learning data was generated. The state of each object in the CG space includes the position and posture of the avatar 201 in the CG space, the position and posture of the camera object 202 in the CG space, the position of the gaze point object 203 in the CG space, and the like.

[0162] As a result, the type of the avatar 201 used to generate each piece of learning data, the state of the CG space, and the state of each object in the CG space can be confirmed later.

[0163] In addition, even if the learning data is not stored, the learning data can be reproduced by storing the script. As a result, for example, when the learning data set is made public, it becomes unnecessary to consider the privacy of individuals included in the learning data set.

[0164] Furthermore, for example, even if the learning dataset is not held, in the case of requesting disclosure of the learning dataset as long as the avatar 201 has authority, the request can be easily processed using a script.

[0165] <<3. Second embodiment>>

[0166] Next, refer to Figures 19 to 22 , describing a second embodiment of the present technology.

[0167] For example, in the case of generating a learning dataset while changing the state of each object in the CG space as described above, it is assumed that the state of the avatar 201 becomes abnormal depending on the combination of the states of the objects.

[0168] Here, the abnormal state of the avatar 201 is, for example, a state in which the avatar image of the avatar 201 in this state becomes an unrealistic image, etc., and in the case where the avatar image is used for machine learning, the learning efficiency may be reduced.

[0169] For example, Figure 19 This shows an example in which the state of the avatar 201 becomes abnormal.

[0170] In this example, the face of avatar 201 is tilted significantly downward to the left. Meanwhile, gaze point object 203 is positioned at a position tilted significantly downward to the right from the face of avatar 201. Consequently, the line of sight of avatar 201's eyes is tilted significantly downward to the right. However, since the direction of the line of sight differs significantly from the direction of avatar 201's face, the iris of avatar 201 is barely visible, and both eyes appear almost white.

[0171] As a result, in the avatar image IM21 rendered by the camera object 202, both eyes of the avatar 201 become almost white, and the gaze direction of the avatar 201 becomes unclear. Therefore, the gaze direction of the avatar 201 of the avatar image IM21 does not match the gaze direction indicated by the gaze information of the correct data, and learning efficiency may be reduced.

[0172] On the other hand, the second embodiment of the present technology improves the quality of the learning dataset by excluding avatar images corresponding to the avatar 201 in an abnormal state, such as the avatar image IM21 , from the learning dataset.

[0173] <Configuration Example of Information Processing System 301>

[0174] Figure 20 An example of the configuration of an information processing system 301 as a second embodiment of an information processing system to which the present technology is applied is shown. Figure 2 Portions corresponding to those of the information processing system 101 in FIG. 1 are denoted by the same reference numerals, and descriptions thereof are appropriately omitted.

[0175] The information processing system 301 is different from the information processing system 101 in that a learning dataset generating unit 311 is provided instead of the learning dataset generating unit 111 .

[0176] <Configuration Example of Learning Dataset Generating Unit 311>

[0177] Next, refer to Figure 21 , describes an example of the configuration of the learning data set generating unit 311. Note that in the figure, Figure 3 Portions corresponding to those of the learning dataset generating unit 111 in FIG. 1 are denoted by the same reference numerals, and descriptions thereof will be appropriately omitted.

[0178] The learning dataset generating unit 311 is different from the learning dataset generating unit 111 in that a state determining unit 351 is added.

[0179] The state determination unit 351 acquires CG space information from the state control unit 152. Based on the state of the CG space and the state of each object in the CG space included in the CG space information, the state determination unit 351 determines whether the state of the avatar 201 is abnormal. If the state of the avatar 201 is determined to be abnormal, the state determination unit 351 notifies the state control unit 152 of the abnormal state of the avatar 201. On the other hand, if the state of the avatar 201 is determined to be normal, the state determination unit 351 provides the CG space information acquired from the state control unit 152 to the learning data generation unit 153.

[0180] <Processing in Information Processing System 301>

[0181] Next, the processing of the information processing system 301 will be described.

[0182] It should be noted that the information processing system 301 processes information according to the above Figure 4 The flowchart is executed similarly to the information processing of the information processing system 101. However, the difference between this processing and the processing of the information processing system 101 is that Figure 22 The flowchart of executes the learning data set generation process in step S1.

[0183] <Learning Dataset Generation Process>

[0184] Here, we will refer to Figure 22 The flowchart of describes the details of the learning dataset generation process of the information processing system 301.

[0185] In steps S101 and S102, the Figure 5 The processing in steps S51 and S52 is similar to the processing in FIG.

[0186] In step S103, similar to Figure 5 In the process of step S53 in , the state control unit 152 directs the line of sight of the avatar 201 toward the gaze point object 203. The state control unit 152 provides the state determination unit 351 with CG space information including information indicating the state of the CG space and the state of each object in the CG space.

[0187] In step S104 , the state determination unit 351 determines whether the state of the avatar 201 is abnormal.

[0188] Specifically, the state determination unit 351 determines whether the virtual image 201 is in an abnormal state based on one or more of, for example, the relative position between the face of the virtual image 201 and the camera object 202, the relative posture between the face of the virtual image 201 and the camera object 202, and the direction (rotation angle) of the eye object of the virtual image 201.

[0189] For example, as the abnormal state of the avatar 201 , a state in which at least one eye of the avatar 201 is not visible in the avatar image is assumed.

[0190] As an example of a state in which at least one eye of the avatar 201 is not visible in the avatar image, for example, a state in which at least one eye of the avatar 201 protrudes from the view angle of the camera object 202 (hereinafter referred to as an out-of-view state) is assumed. In other words, at least one eye of the avatar 201 protrudes from the avatar image.

[0191] For example, the out-of-view state can be detected based on the relative position between the eyeball object of each eye of the avatar 201 and the camera object 202 .

[0192] As another example of a state in which at least one eye of the avatar 201 is not visible in the avatar image, for example, a state (hereinafter referred to as a hidden state) in which the face of the avatar 201 faces a direction different from the camera object 202 and at least one eye of the avatar 201 is hidden when viewed from the camera object 202 is assumed. That is, at least one eye of the avatar 201 is hidden in the avatar image.

[0193] For example, a hidden state occurs when the face of the avatar 201 is located on a line segment connecting an eyeball object of at least one eye of the avatar 201 and the camera object 202. Therefore, the hidden state can be detected based on, for example, the relative position between the eyeball object of each eye of the avatar 201 and the camera object 202, and the relative position between the face of the avatar 201 and the camera object 202.

[0194] Furthermore, as an abnormal state of the avatar 201, for example, a state in which at least one eye of the avatar 201 is in a state similar to the white of the eye as viewed from the camera subject 202 (hereinafter referred to as the white of the eye state) is assumed. That is, in the avatar image, at least one eye of the avatar 201 is in a state similar to the white of the eye.

[0195] Note that the state similar to the white of the eye is, for example, a state in which the ratio or area of ​​the region of the iris in the eye of the avatar 201 is smaller than a predetermined threshold value.

[0196] For example, in the above reference Figure 19 The white-of-the-eye state occurs when the difference between the face direction and the sight direction of the avatar 201 is too large. Therefore, the white-of-the-eye state can be detected based on, for example, the face direction of the avatar 201 and the direction (rotation angle) of the eyeball of each eye.

[0197] For example, when the out-of-view state, hidden state, and white-of-the-eye state are not detected, the state determination unit 351 determines that the state of the avatar 201 is normal, and supplies the CG space information acquired from the state control unit 152 to the learning data generation unit 153. Then, the process proceeds to step S105.

[0198] In steps S105 to S107, the Figure 5 The process is similar to the process in steps S54 to S56 in . Then, the process proceeds to step S108.

[0199] On the other hand, in step S104, if at least one of the out-of-view state, hidden state, or white-of-the-eye state is detected, state determination unit 351 determines that the state of avatar 201 is abnormal and notifies state control unit 152 of the abnormal state of avatar 201. Thereafter, steps S105 through S107 are skipped, and processing proceeds to step S108. In other words, if the state of avatar 201 is abnormal, no learning data is generated, and learning data based on avatar 201 determined to be abnormal is not added to the learning dataset.

[0200] In step S108, similar to Figure 7 In the process of step S57, it is determined whether the quantity and quality of the learning data set are sufficient. In the case where it is determined that at least one of the quantity and quality of the learning data set is still insufficient, the process proceeds to step S109.

[0201] In step S109, Figure 7 The processing in step S58 is similar to that in step S58, and the virtual image 201 is changed as needed.

[0202] Afterward, the process returns to step S102, and steps S102 through S109 are repeatedly executed until step S108 determines that the quantity and quality of the learning data set are sufficient. As a result, learning data is generated while the state of each object and avatar 201 in the CG space is changing. However, if the state of avatar 201 is abnormal, no learning data is generated.

[0203] On the other hand, in step S108 , in the case where it is determined that the quantity and quality of the learning dataset are sufficient, the learning dataset generation process ends.

[0204] As described above, learning data generated when the avatar 201 is abnormal is excluded, and the quality of the learning data set is improved. As a result, the accuracy of the gaze estimation model is improved.

[0205] <<4. Third embodiment>>

[0206] Next, refer to Figures 23 to 26 , describing a third embodiment of the present technology.

[0207] When a learning dataset is generated and used using images obtained by actually photographing a person, as in Patent Document 1, deviations in the learning dataset may occur. Details of deviations in the learning dataset will be described later, but for example, deviations in the face position or gaze direction of a person in the captured image are assumed.

[0208] Gaze estimation models learned using bias learning datasets tend to be non-robust models for bias components. For example, there are cases where the gaze direction is estimated to be right only when the position of a person's face is on the right side in the captured image.

[0209] On the other hand, in the third embodiment of the present technology, by using the virtual image 201 to supplement the deficiencies of the acquired learning dataset, the deviation of the learning dataset is reduced, and the quantity and quality of the learning dataset are improved.

[0210] <Configuration Example of Information Processing System 401>

[0211] Figure 23 An example of the configuration of an information processing system 401 as a third embodiment of an information processing system to which the present technology is applied is shown. Figure 2 Portions corresponding to those of the information processing system 101 in FIG. 1 are denoted by the same reference numerals, and descriptions thereof are appropriately omitted.

[0212] The information processing system 401 is different from the information processing system 101 in that a learning dataset acquisition unit 411 and a learning dataset supplementation unit 412 are added and the learning dataset generation unit 111 is deleted.

[0213] The learning dataset acquisition unit 411 acquires a learning dataset and accumulates the acquired learning dataset in the learning dataset accumulation unit 112 .

[0214] The learning dataset supplementing unit 412 supplements the deficiency of the acquired learning dataset accumulated in the learning dataset accumulating unit 112 .

[0215] <Configuration Example of Learning Dataset Supplementation Unit 412>

[0216] Figure 24 The configuration example of the learning data set supplement unit 412 is shown. Figure 21 The parts corresponding to those of the learning dataset generating unit 311 in are denoted by the same reference numerals, and descriptions thereof are appropriately omitted.

[0217] The learning dataset supplementation unit 412 includes a data analysis unit 451, a supplementation planning unit 452, and a learning dataset generation unit 453. The learning dataset generation unit 453 differs from the learning dataset generation unit 311 in that an object generation unit 461 and a state control unit 462 are provided instead of the object generation unit 151 and the state control unit 152.

[0218] The data analyzing unit 451 analyzes the acquired learning dataset accumulated in the learning dataset accumulating unit 112. The data analyzing unit 451 supplies information indicating the analysis result of the learning dataset to the supplementation planning unit 452.

[0219] The supplementary planning unit 452 creates a plan for supplementing the deficiencies of the learning dataset based on the analysis result of the learning dataset. The supplementary planning unit 452 provides information indicating the supplementary plan of the learning dataset to the object generation unit 461 and the state control unit 462.

[0220] The object generation unit 461 generates various objects in the CG space based on the supplementary planning of the learning dataset. For example, the object generation unit 461 generates the avatar 201, the camera object 202, the gaze point object 203, and the like in the CG space. The object generation unit 461 provides information about each generated object to the state control unit 462.

[0221] The state control unit 462 controls the conditions for generating learning data by controlling the state of the CG space and the state of each object in the CG space based on the supplementary planning of the learning data set. The state control unit 462 provides the state determination unit 351 with CG space information including information indicating the state of each object in the CG space and the state of the CG space.

[0222] Furthermore, the state control unit 462 instructs the object generation unit 461 to generate an object in the CG space as necessary.

[0223] <Information Processing by Information Processing System 401>

[0224] Next, we will refer to Figure 25 The flowchart of FIG. 40 describes information processing performed by the information processing system 401 .

[0225] In step S201 , the learning dataset acquisition unit 411 acquires a learning dataset.

[0226] Note that the method for acquiring a learning dataset is not particularly limited. For example, the learning dataset acquisition unit 411 automatically collects learning datasets publicly available on the Internet, etc. For example, the learning dataset acquisition unit 411 acquires a learning dataset input by a user.

[0227] In addition, an image of a sample of a person used as input data of a learning dataset (hereinafter, referred to as a sample image) may be an image obtained by actually capturing the person (captured image) or an image generated by CG (avatar image).

[0228] Furthermore, for example, the learning dataset acquisition unit 411 may collect sample images disclosed on the Internet or the like, and generate a learning dataset using each sample image and correct data given to the collected sample images.

[0229] The learning dataset acquisition unit 411 accumulates the acquired learning dataset in the learning dataset accumulation unit 112 .

[0230] In step S202 , the learning dataset supplementing unit 412 performs learning dataset supplementing processing.

[0231] Here, we will refer to Figure 26 The flowchart describes the details of the learning dataset supplementation process.

[0232] In step S251 , the data analyzing unit 451 analyzes the acquired learning data set. That is, the data analyzing unit 451 analyzes the learning data set acquired in the process of step S201 and accumulated in the learning data set accumulating unit 112 .

[0233] The data analysis unit 451 detects insufficiency of the learning data based on the analysis result of the learning data set.

[0234] Here, the insufficiency of the learning dataset is, for example, an insufficiency of at least one of the quantity or quality of the learning dataset. The insufficiency of the quantity of the learning dataset is, for example, an insufficiency of the data volume of the learning dataset. The insufficiency of the quantity of the learning dataset includes, for example, a variant of insufficiency of the learning dataset.

[0235] The insufficient variation of the learning dataset is represented by, for example, the deviation of the learning dataset. The deviation of the learning dataset includes, for example, the deviation of the input data and the deviation of the correct data.

[0236] Deviations in the input data include, for example, sample deviations and sample image deviations. Sample deviations include, for example, sample feature deviations. Sample feature deviations include, for example, at least one deviation in race, gender, age, facial makeup, facial size, and facial color. Sample image deviations include, for example, deviations in information obtained from the face in the sample image. Information obtained from the face in the sample image includes, for example, at least one deviation in facial position, facial size, facial orientation, eye orientation, and the correlation between gaze direction and facial orientation.

[0237] For example, the deviation of the correct data includes the deviation of the sight line information. The deviation of the sight line information includes, for example, the deviation of the sight line direction indicated by the sight line information.

[0238] The data analysis unit 451 provides information indicating a detection result of insufficient learning data to the supplementary planning unit 452 .

[0239] In step S252 , the supplementation planning unit 452 plans the supplementation of the learning data based on the analysis result.

[0240] Specifically, for example, the supplement planning unit 452 plans the supplementary learning data set so that the learning data set has a sufficient amount of data and corrects the deviation of the learning data set. For example, the supplement planning unit 452 plans the characteristics of the virtual image 201 when generating the learning data for supplementation and the state of each object in the CG space.

[0241] The characteristics of the avatar include, for example, at least one of race, gender, age, facial makeup, face size, face color, and the like.

[0242] The state of each object in the CG space includes, for example, the relative position and posture of the face of the avatar 201 and the camera object 202 , the relative position of the gaze point object 203 with respect to the eyes of the avatar 201 , and the like.

[0243] For example, when the distance between the face of the sample in the acquired learning dataset and the camera is substantially constant, the state of each object in the CG space is planned so that avatar images having different distances between the face of the avatar 201 and the camera object 202 are generated. For example, when the distance between the face of the sample in the sample image and the camera is mostly approximately 50 cm, the state of each object in the CG space is planned so that avatar images are generated in a state where the distance between the face of the avatar 201 and the camera object 202 is 40 cm, 60 cm, or 70 cm.

[0244] For example, in the acquired learning dataset, in the case where there are most sample images in which the sample looks downward at the camera, the state of each object in the CG space is planned so that an avatar image is generated in which the avatar 201 looks upward at the camera object 202. For example, the state of each object in the CG space is planned so that avatar images are generated in the state in which the gaze point object is arranged at positions 5 cm, 10 cm, and 15 cm above the camera object 202.

[0245] For example, when the acquired learning dataset includes only sample images of a specific race, the generation of the avatar 201 is planned so that the types of races of the avatar 201 increase. For example, the generation of the avatar 201 is planned so that the ratio of the races of the avatar 201 is constant.

[0246] The supplement planning unit 452 generates a supplement parameter list that parameterizes the characteristics of the avatar 201 and the state of each object in the CG space when generating the learning data for supplementation based on the supplement planning of the learning data set. The supplement planning unit 452 provides the supplement parameter list to the object generation unit 461 and the state control unit 462.

[0247] In step S253, the object generation unit 461 generates various objects in the CG space based on the supplementary plan. For example, the object generation unit 461 sets the characteristics of the avatar 201 based on the supplementary parameter list and generates the avatar 201 in the CG space with the set characteristics. The object generation unit 461 generates the camera object 202 and the gaze point object 203 in the CG space. The object generation unit 461 provides information about each generated object to the state control unit 462.

[0248] In step S254, the state control unit 462 updates the state in the CG space based on the supplementary plan. Specifically, similar to Figure 5 In the processing in step S52, the state control unit 462 updates the state of each object to one of the combinations of states of each object for which learning data has not yet been generated based on the supplementary parameter list.

[0249] In steps S255 to S259, the Figure 22 The processing is similar to the processing in steps S103 to S107 in .

[0250] In this case, for example, an avatar image is generated in a format similar to the sample images included in the acquired learning dataset. As a result, during learning, the learning data included in the acquired learning dataset and the supplementary learning data can be processed similarly without being distinguished from each other. This resolves the aforementioned problem 2 of the prior art.

[0251] In step S260, the state control unit 462 determines whether the supplementation of the learning data set is completed. If parameters for which no learning data has been generated remain in the supplement parameter list, the state control unit 462 determines that the supplementation of the learning data set is not completed, and the process proceeds to step S261.

[0252] In step S261 , the learning data set generating unit 453 changes the avatar 201 as needed based on the supplementary plan.

[0253] For example, state control unit 462 determines, based on the supplementary parameter list, whether learning data for all parameters has been generated using current avatar 201. If it is determined that learning data for all parameters has been generated using current avatar 201, state control unit 462 selects one of the characteristics of avatar 201 for which learning data has not yet been generated, based on the supplementary parameter list. State control unit 462 instructs object generation unit 461 to change the characteristic of avatar 201 to the selected characteristic.

[0254] On the other hand, the object generation unit 461 generates a new avatar 201 in the CG space according to an instruction of the state control unit 462 and deletes the old avatar 201. The object generation unit 461 provides information on the generated avatar 201 to the state control unit 462.

[0255] On the other hand, in the event that it is determined that learning data for all parameters has not yet been generated using the current avatar 201 , the state control unit 462 determines to continue generating learning data while maintaining the current avatar 201 .

[0256] Thereafter, the process returns to step S254 , and the processes of steps S254 to S261 are repeatedly performed until it is determined in step S260 that the supplementation of the learning data set has ended.

[0257] On the other hand, in step S260 , when no parameter for which learning data has not been generated remains in the supplemented parameter list, the state control unit 462 determines that the supplementation of the learning dataset has ended, and the learning dataset supplementation process ends.

[0258] Back to Figure 25 In step S203, the learning unit 113 performs machine learning using the supplemented learning dataset and generates a line of sight estimation model. Specifically, the learning unit 113 performs machine learning using a predetermined learning method using the supplemented learning dataset accumulated in the learning dataset accumulation unit 112. The learning unit 113 provides the line of sight estimation model obtained through machine learning to the estimation unit 114.

[0259] In step S204, similar to Figure 4 In the process of step S3 in , the visual line estimation process is performed using the generated visual line estimation model.

[0260] As described above, by supplementing the acquired learning dataset, the quantity and quality of the learning dataset are improved. The accuracy of the line of sight estimation model is improved by performing machine learning using the supplemented learning dataset.

[0261] Furthermore, the third embodiment of the present technology can be applied to a case where fine-tuning of a visual line estimation model of a specific user (hereinafter, referred to as a target user) is performed and the visual line estimation model of the target user is optimized.

[0262] For example, in the case of performing fine-tuning of a visual line estimation model for a target user, additional learning is performed using a learning dataset obtained by using a captured image obtained by capturing the target user.

[0263] In this case, for example, when the amount of learning dataset acquired using captured images is insufficient or the deviation of the learning dataset is large, the robustness of the visual line estimation model is reduced by optimizing the visual line estimation model for the learning dataset.

[0264] On the other hand, for example, as described above, the acquired learning dataset can be supplemented with an avatar 201 similar to the target user to improve the quantity and quality of the target user's learning dataset. In addition, the amount of data in the learning dataset acquired using the captured image obtained by pre-capturing the target user can be reduced, and the load required for generating the learning dataset can be reduced.

[0265] Note that whether the avatar 201 is similar to the target user can be determined by, for example, comparing the feature quantities of facial makeup. For example, the avatar 201 having the closest ratio between the distance between the inner corners of the target user's eyes and the length of the entire face is used as the avatar 201 similar to the target user.

[0266] Then, by performing additional learning using a supplementary learning dataset, the accuracy of fine-tuning is improved, and the accuracy of estimating the target user's gaze by the gaze estimation model is improved. For example, the robustness of the gaze estimation model is improved.

[0267] 5. Modifications

[0268] Hereinafter, modifications of the above-described embodiment of the present technology will be described.

[0269] For example, two or more avatars 201 may be generated in the CG space at once.

[0270] For example, in a case where at least one eye of the avatar 201 is in a state similar to the white of the eye, the rotation angle of the eyeball object of the eye may be corrected within a range not to change to the white of the eye.

[0271] The above description describes an example in which learning data is generated while changing the relative position and posture between the face of avatar 201 and camera object 202, as well as the relative position between the face of avatar 201 and gaze point object 203. However, learning data can also be generated by changing another state of each object. For example, the relative position and posture between the entire body or a portion other than the face of avatar 201 and camera object 202, and the relative position between the entire body or a portion other than the face of avatar 201 and gaze point object 203 are assumed as such states of each object. Furthermore, for example, the facial expressions and hand gestures of avatar 201 are also assumed.

[0272] For example, learning data can be generated when the state of the CG space is changed. As such a state of the CG space, for example, a light beam or the lighting state of the CG space, the background of the CG space, etc. are assumed.

[0273] The gazing point object 203 only needs to be able to specify coordinates of the direction (rotation angle) of the eyeball object for controlling the avatar 201, and the appearance is not limited to the above example. In addition, the gazing point object 203 is not necessarily visible.

[0274] In addition to or instead of the abnormality determination of the state of the avatar 201 , abnormality determination for detecting an avatar image causing a degradation in the quality of the learning dataset may be performed.

[0275] The present technology can also be applied to, for example, a case where a learning data set is generated for an estimation model that performs estimation processing of non-verbal information about a person other than the line of sight (for example, at least one of the state or characteristics of the person). For example, the present technology can be applied to a case where a learning data set is generated for an estimation model that estimates the posture or emotion of a person. For example, the present technology can be applied to a case where a learning data set is generated for an estimation model that performs lip reading to estimate the content of a person's speech through the movement of the person's lips. For example, the present technology can be applied to a case where a learning data set is generated for an estimation model that estimates the characteristics of a person such as race, gender, and age.

[0276] 6. Others

[0277] <Computer Configuration Example>

[0278] The above series of processes can be performed by hardware or software. In the case where the series of processes are performed by software, a program configuring the software is installed in the computer. Here, examples of computers include computers incorporated into dedicated hardware, general-purpose personal computers that can perform various functions by installing various programs, etc.

[0279] Figure 27 : is a block diagram showing a configuration example of hardware of a computer that executes the above-described series of processes by a program.

[0280] In a computer 1000 , a central processing unit (CPU) 1001 , a read-only memory (ROM) 1002 , and a random access memory (RAM) 1003 are interconnected via a bus 1004 .

[0281] An input / output interface 1005 is further connected to the bus 1004 . An input unit 1006 , an output unit 1007 , a storage unit 1008 , a communication unit 1009 , and a drive 1010 are connected to the input / output interface 1005 .

[0282] The input unit 1006 includes input switches, buttons, a microphone, an imaging element, and the like. The output unit 1007 is formed using a display, a speaker, and the like. The storage unit 1008 is formed using a hard disk, a nonvolatile memory, and the like. The communication unit 1009 is formed using a network interface and the like. The drive 1010 drives a removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0283] In the computer 1000 configured as described above, for example, a program stored in the storage unit 1008 is loaded into the RAM 1003 via the input / output interface 1005 and the bus 1004 by the CPU 1001 , and the above-described series of processes is executed.

[0284] For example, the program executed by the computer 1000 (CPU 1001) can be provided by being recorded in the removable medium 1011 as a package medium or the like. In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0285] In the computer 1000, by mounting the removable medium 1011 on the drive 1010, the program can be installed in the storage unit 1008 via the input / output interface 1005. Also, the program can be received by the communication unit 1009 via a wired or wireless transmission medium to be installed on the storage unit 1008. In addition, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.

[0286] It should be noted that the program executed by the computer may be a program that sequentially executes a plurality of processes in the time series described in this specification, or may be a program that executes processes in parallel or at necessary timing such as when a call is made.

[0287] In addition, in this specification, a system means an assembly of multiple components (devices, modules (components), etc.), and it does not matter whether all the components are in the same housing. Therefore, multiple devices housed in separate housings and connected to each other via a network and a single device including multiple modules housed in a single housing are both systems.

[0288] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the gist of the present technology.

[0289] For example, the present technology may be embodied in cloud computing, where functions are performed in a shared manner via a network.

[0290] Furthermore, each step described in the above flowcharts may be performed by one device or may be performed by a plurality of devices in a shared manner.

[0291] Furthermore, in the case where a plurality of processes are included in one step, the plurality of processes included in one step may be executed by one device or may be executed by a plurality of devices in a shared manner.

[0292] <Configuration combination example>

[0293] The present technology may also have the following configurations. (1)

[0295] An information processing method, comprising:

[0296] controlling the state of a three-dimensional model of a person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model; and

[0297] Based on the state of the 3D model and the rendering conditions, input data including a 3D model image and learning data are generated. The 3D model image is an image obtained by rendering the 3D model, and the learning data includes correct data about the 3D model. (2)

[0299] According to the information processing method of (1),

[0300] The state of the three-dimensional model includes at least one of a position or a posture of the three-dimensional model in the three-dimensional space. (3)

[0302] According to the information processing method of (2),

[0303] The state of the three-dimensional model further includes the direction of the eyeball of the three-dimensional model in the three-dimensional space. (4)

[0305] According to the information processing method of (3),

[0306] The correct data includes sight line information indicating the sight line direction of the three-dimensional model. (5)

[0308] According to the information processing method of (4),

[0309] The three-dimensional model image includes both eyes of the three-dimensional model, and

[0310] The sight line information indicates a sight line direction of both eyes with respect to the three-dimensional model. (6)

[0312] According to the information processing method of any one of (3) to (5),

[0313] wherein the information processing device further controls the position of a gaze point indicating the position at which the three-dimensional model is gazing in the virtual space, and

[0314] The direction of the eye of the 3D model is controlled based on the relative position between the eye of the 3D model and the gaze point. (7)

[0316] According to the information processing method of any one of (2) to (6),

[0317] The information processing device controls the relative position and posture between the three-dimensional model and the camera object for controlling rendering conditions. (8)

[0319] The information processing method according to any one of (2) to (7), wherein the posture of the three-dimensional model includes a direction of a face of the three-dimensional model in the three-dimensional space. (9)

[0321] According to the information processing method of any one of (1) to (8),

[0322] The information processing device controls the state and rendering conditions of the three-dimensional model and expands the state of the three-dimensional model in the three-dimensional model image or the change of at least one of the correct data in the learning data set as a set of learning data. (10)

[0324] According to the information processing method of (9),

[0325] The learning dataset is used to learn a learning model for estimating at least one of a state or a characteristic of a person. (11)

[0327] According to the information processing method of any one of (1) to (10),

[0328] wherein the information processing device determines the state of the three-dimensional model, and

[0329] In a case where the information processing apparatus determines that the state of the three-dimensional model is abnormal, the information processing apparatus does not add learning data based on the three-dimensional model determined to be abnormal to the learning data set that is a collection of learning data. (12)

[0331] According to the information processing method of (11),

[0332] Among them, the state in which the three-dimensional model is abnormal is a state in which learning efficiency may be reduced in a case where learning data including a three-dimensional model image based on the three-dimensional model is used for learning. (13)

[0334] According to the information processing method of (12),

[0335] The state in which the three-dimensional model is abnormal is at least one of a state in which at least one eye of the three-dimensional model is not included in the three-dimensional model image or a state in which at least one eye of the three-dimensional model resembles the white of an eye. (14)

[0337] The information processing method according to any one of (1) to (13),

[0338] The information processing device analyzes the acquired learning data set, and

[0339] The analysis result based on the learning dataset supplements at least one deficiency in the quantity or quality of the learning dataset. (15)

[0341] According to the information processing method of (14),

[0342] wherein the information processing device detects deviations in the learning data set,

[0343] Learning data for correcting the bias of the learning dataset is generated, and the learning data is added to the learning dataset. (16)

[0345] The information processing method according to any one of (1) to (15),

[0346] Here, the information processing device generates a three-dimensional model and changes the characteristics of the three-dimensional model. (17)

[0348] The information processing method according to any one of (1) to (16),

[0349] The information processing device further controls the state of the virtual space. (18)

[0351] The information processing method according to any one of (1) to (17),

[0352] The state of the three-dimensional model includes at least one of an expression or a gesture of the three-dimensional model. (19)

[0354] An information processing device, comprising:

[0355] An estimation unit that performs estimation processing on a person by using a learning model, the learning model is generated by learning using a learning data set, the learning data set is a collection of learning data generated based on the state of a three-dimensional model of a person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering conditions, the learning data includes input data, the input data includes a three-dimensional model image and correction data about the three-dimensional model, the three-dimensional model image is an image obtained by rendering the three-dimensional model. (20)

[0357] An information processing method, comprising:

[0358] An information processing device performs estimation processing about a person by using a learning model that utilizes a learning data set, the learning data set being a collection of learning data sets generated based on the state of a three-dimensional model of the person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering conditions, the learning data including input data, the input data including a three-dimensional model image and correction data about the three-dimensional model, the three-dimensional model image being an image obtained by rendering the three-dimensional model. (twenty one)

[0360] An information processing method includes: an information processing device performs learning of a learning model, the learning model uses a learning data set to perform estimation processing on a person, the learning data set is a collection of learning data generated based on the state of a three-dimensional model of the person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering conditions, the learning data includes input data, the input data includes a three-dimensional model image and correction data about the three-dimensional model, and the three-dimensional model image is an image obtained by rendering the three-dimensional model.

[0361] It should be noted that the effects described in this specification are merely examples and are not limiting, and other effects can be achieved.

[0362] Reference Symbols List

[0363] 101 Information Processing Systems

[0364] 111 Learning Dataset Generation Unit

[0365] 112 Learning Dataset Accumulation Unit

[0366] 113 learning units

[0367] 114 Estimation Unit

[0368] 151 Object Generation Unit

[0369] 152 Status Control Unit

[0370] 153 Learning Data Generation Unit

[0371] 301 Information Processing Systems

[0372] 311 Learning Dataset Generation Unit

[0373] 351 Status determination unit

[0374] 401 Information Processing System

[0375] 411 Learning Dataset Acquisition Unit

[0376] 412 Learning Dataset Supplementary Unit

[0377] 451 Data Analysis Unit

[0378] 452 Supplementary Planning Unit

[0379] 453 Learning Dataset Generation Unit

[0380] 461 Object Generation Unit

[0381] 462 Status Control Unit

Claims

1. An information processing method, comprising: controlling the state of a three-dimensional model of a person in a three-dimensional virtual space and rendering conditions for rendering the three-dimensional model; and Based on the state of the three-dimensional model and the rendering condition, input data including a three-dimensional model image and learning data are generated, wherein the three-dimensional model image is an image obtained by rendering the three-dimensional model, and the learning data includes correct data about the three-dimensional model.

2. The information processing method according to claim 1, in, The state of the three-dimensional model includes at least one of a position or a posture of the three-dimensional model in the virtual space.

3. The information processing method according to claim 2, in, The state of the three-dimensional model also includes the direction of the eyeball of the three-dimensional model in the virtual space.

4. The information processing method according to claim 3, in, The correct data includes sight line information indicating a sight line direction of the three-dimensional model.

5. The information processing method according to claim 4, wherein: The three-dimensional model image includes both eyes of the three-dimensional model, and The sight line information indicates a sight line direction for both eyes of the three-dimensional model.

6. The information processing method according to claim 3, in, The information processing device further controls the position of a gaze point indicating a position where the three-dimensional model is looking in the virtual space, and The direction of the eye of the three-dimensional model is controlled based on the relative position between the eye of the three-dimensional model and the gaze point.

7. The information processing method according to claim 2, in, The information processing device controls the relative position and posture between the three-dimensional model and a camera object for controlling the rendering conditions.

8. The information processing method according to claim 2, wherein: The pose of the three-dimensional model includes the direction of the face of the three-dimensional model in the virtual space.

9. The information processing method according to claim 1, wherein: The information processing device controls the state of the three-dimensional model and the rendering condition, and expands a variation of at least one of the state of the three-dimensional model or the correct data in the three-dimensional model image in a learning data set which is a set of the learning data.

10. The information processing method according to claim 9, in, The learning data set is used to learn a learning model for estimating at least one of a state or a characteristic of a person.

11. The information processing method according to claim 1, in, The information processing device determines the state of the three-dimensional model, and In a case where the information processing device determines that the state of the three-dimensional model is abnormal, the information processing device does not add the learning data based on the three-dimensional model determined to be abnormal to the learning data set which is the set of the learning data.

12. The information processing method according to claim 11, in, The state in which the three-dimensional model is abnormal is a state in which learning efficiency may be reduced if the learning data including the three-dimensional model image based on the three-dimensional model is used for learning.

13. The information processing method according to claim 12, in, The state in which the three-dimensional model is abnormal is at least one of a state in which at least one eye of the three-dimensional model is not included in the three-dimensional model image or a state in which at least one eye of the three-dimensional model resembles the white of an eye.

14. The information processing method according to claim 1, in, The information processing device analyzes the acquired learning data set, and At least one deficiency of the quantity or quality of the learning data set is supplemented based on the analysis result of the learning data set.

15. The information processing method according to claim 14, in, The information processing device detects deviations in the learning data set, The learning data for correcting the deviation of the learning data set is generated, and the learning data is added to the learning data set.

16. The information processing method according to claim 1, in, The information processing device generates the three-dimensional model and changes the characteristics of the three-dimensional model.

17. The information processing method according to claim 1, in, The information processing device further controls the state of the virtual space.

18. The information processing method according to claim 1, in, The state of the three-dimensional model includes at least one of an expression or a gesture of the three-dimensional model.

19. An information processing device, comprising: An estimation unit performs estimation processing on a person by using a learning model, wherein the learning model is generated by learning using a learning data set, wherein the learning data set is a collection of learning data generated based on a state of a three-dimensional model of a person in a three-dimensional virtual space and a rendering condition for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering condition, wherein the learning data includes input data, wherein the input data includes a three-dimensional model image and correction data on the three-dimensional model, wherein the three-dimensional model image is an image obtained by rendering the three-dimensional model.

20. An information processing method, comprising: An information processing device performs estimation processing on a person by using a learning model, wherein the learning model is generated by learning using a learning data set, wherein the learning data set is a collection of learning data generated based on a state of a three-dimensional model of a person in a three-dimensional virtual space and a rendering condition for rendering the three-dimensional model while changing the state of the three-dimensional model and the rendering condition, wherein the learning data includes input data, wherein the input data includes a three-dimensional model image and correction data about the three-dimensional model, wherein the three-dimensional model image is an image obtained by rendering the three-dimensional model.