Information processing device, information processing method, and program
The information processing system addresses the issue of unnatural avatar expressions by using a combination of sensors and animation curves to transition to a stable recovery expression when facial recognition fails, ensuring a natural appearance.
Patent Information
- Application Number
- PCT/JP2025/004868
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-02-14
- Publication Date
- 2025-09-04
AI Technical Summary
Existing technologies fail to maintain a natural facial expression on an avatar when the user's facial expression cannot be recognized due to factors like movement or the face being out of the camera's range, leading to unnatural freezing or sudden changes in the avatar's expression.
An information processing system that utilizes a combination of camera images, sensors, and motion capture systems to determine important feature points, and when recognition fails, smoothly transitions the avatar's expression to a stable recovery expression using animation curves and context information to maintain a natural appearance.
Ensures the avatar displays a natural and stable facial expression even during lost states by smoothly transitioning to a recovery expression, reducing the perception of unnatural freezing or sudden changes.
Smart Images

Figure JP2025004868_04092025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present technology relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program that enable an avatar to display a natural facial expression even in a lost state.
[0002] There is a technology that recognizes a user's facial expression based on a camera image captured by a camera and reflects the user's facial expression in an avatar. For example, Patent Document 1 describes a technology that recognizes a user's emotion based on a voice waveform and reflects an expression corresponding to the user's emotion in an avatar.
[0003] International Publication No. 2023 / 068067
[0004] Hereinafter, a state in which the user's facial expression cannot be recognized based on the camera image, for example, when the user moves or looks down and the user's face is not captured by the camera, will be referred to as a lost state.
[0005] In the lost state, the user's facial expression is not properly reflected on the avatar. For example, the avatar's face may freeze and retain the unnatural expression it had just had before the lost state. Furthermore, when the lost state is resolved, the avatar's facial expression may suddenly change, which may make other users looking at the avatar feel unnatural. It is not desirable for the avatar's face to freeze or its facial expression to suddenly change frequently.
[0006] This technology was developed in consideration of such circumstances, and makes it possible for an avatar to display natural facial expressions even in a lost state.
[0007] An information processing device according to one aspect of the present technology includes an avatar control unit that, when acquiring the intensity of a feature point of a user is successful, controls a parameter value for the feature point of an avatar to be a value corresponding to the intensity of the feature point of the user, and, when acquiring the intensity of the feature point of the user fails, controls the parameter value to be a value for when the acquisition of the intensity of the feature point of the user fails.
[0008] An information processing method according to one aspect of the present technology includes, when acquisition of intensities of feature points of a user is successful, controlling a value of a parameter for the feature point of an avatar to be a value corresponding to the intensity of the feature point of the user, and, when acquisition of intensities of the feature point of the user fails, controlling the value of the parameter to be a value for when failure occurs.
[0009] A program according to one aspect of the present technology causes a computer to execute a process of controlling, when acquisition of intensities of a user's feature points is successful, a parameter value for the feature points of an avatar to become a value corresponding to the intensity of the feature points of the user, and, when acquisition of intensities of the user's feature points is unsuccessful, controlling the parameter value to become a value for the case of failure.
[0010] In one aspect of the present technology, when acquisition of the intensity of a feature point of a user is successful, the value of a parameter for the feature point of an avatar is controlled to be a value corresponding to the intensity of the feature point of the user, and when acquisition of the intensity of the feature point of the user fails, the value of the parameter is controlled to be a value for when failure occurs.
[0011] 1 is a block diagram showing an example of a configuration of an information processing system according to an embodiment of the present technology. FIG. 1 is a diagram showing an example of a configuration of an information processing system on the user's side. FIG. 2 is a diagram showing an example of an image displayed on an output device and a display. FIG. 3 is a diagram showing an example of a facial expression of the user's own user and an expression expressed by an avatar. FIG. 4 is a diagram showing an example of a facial expression of the user's own user and an expression expressed by an avatar when using the information processing system of the present technology. FIG. 5 is a first diagram explaining a processing flow for making an avatar express a recovery expression. FIG. 6 is a second diagram explaining a processing flow for making an avatar express a recovery expression. FIG. 7 is a third diagram explaining a processing flow for making an avatar express a recovery expression. FIG. 8 is a fourth diagram explaining a processing flow for making an avatar express a recovery expression. FIG. 9 is a diagram showing an example of an animation curve. FIG. 10 is a diagram showing a flow for changing an avatar's expression from a recovery expression to a realistic expression. FIG. 11 is a diagram showing examples of modal information and context information. FIG. 12 is a diagram showing an example of a time series of loss possibility values. FIG. 13 is a diagram showing examples of the strength of the user's own user's feature points and parameter values for the avatar's feature points when the loss possibility value is equal to or greater than a threshold. FIG. 14 is a diagram showing an example of an avatar's expression when the loss possibility value is equal to or greater than a threshold. FIG. 15 is a diagram showing an example of a screen for defining preset expressions. 23. FIG. 24 is a diagram showing an example of an avatar without a mouth. FIG. 25 is a diagram showing an example of an avatar having feature points different from those of the user. FIG. 26 is a diagram showing an example of a screen for inputting feature points to be reflected in an avatar. FIG. 27 is a diagram showing examples of the facial expressions of an avatar before the user removes the HMD and after the user puts the HMD back on. FIG. 28 is a diagram showing an example of a UI that prompts the user to adjust the wearing position of the HMD. FIG. 29 is a block diagram showing an example of the functional configuration of an information processing device. FIG. 29 is a flowchart illustrating processing by an information processing device to reflect a user's facial expressions in an avatar. FIG. 29 is a flowchart illustrating processing at the time of loss performed in step S5 of FIG. 23. FIG. 29 is a flowchart illustrating processing at the time of re-recognition performed in step S7 of FIG. 23. FIG. 29 is a flowchart illustrating processing by an information processing device 11 to perform an interaction according to a loss possibility value. FIG. 29 is a block diagram showing an example of the hardware configuration of a computer.
[0012] Hereinafter, embodiments of the present technology will be described in the following order: 1. Embodiment of the present technology 2. Configuration and operation of an information processing device
[0013] 1. Embodiments of the Present Technology Overview of Information Processing System FIG. 1 is a block diagram showing an example configuration of an information processing system according to an embodiment of the present technology.
[0014] The information processing system in Fig. 1 recognizes a user's facial expression based on a camera image captured by a camera and reflects the user's facial expression in an avatar. Furthermore, if the information processing system in Fig. 1 cannot recognize the user's facial expression based on the camera image, it controls the parameter values of the avatar or recognizes the user's facial expression based on information other than the camera image, thereby making the avatar display a natural facial expression.
[0015] In the following, a state in which the user's facial expression cannot be recognized based on the camera image (a state in which the user's facial expression cannot be acquired), for example, when the user moves or looks down and the user's face is not captured by the camera, is referred to as a lost state.
[0016] As shown in FIG. 1 , the information processing system of the present technology includes a camera 1 , a sensor 2 , a motion capture system 3 , a microphone 4 , an information processing device 11 , and an output device 21 .
[0017] The camera 1 may be a webcam, a single-lens reflex camera, a mirrorless camera, a camera provided in a smart device, a camera provided in an HMD (Head Mounted Display), an infrared camera, a depth sensor, etc. The camera 1 supplies a camera image obtained by capturing, for example, an upper body including at least the face of the user to the information processing device 11.
[0018] The sensor 2 is configured by an IMU (Inertial Measurement Unit), an electromyographic sensor, etc. The sensor 2 detects the posture and movement of the user and supplies sensor information indicating the detection results to the information processing device 11.
[0019] The motion capture system 3 detects the user's movements using markers and motion sensors worn by the user, and supplies the information processing device 11 with movement information indicating the user's movements.
[0020] The microphone 4 collects the user's voice and supplies the obtained sound information to the information processing device 11 .
[0021] The information processing device 11 is composed of a PC, a smart device, an HMD, a TV (television receiver), etc. The information processing device 11 recognizes the facial expression of the user based on information supplied from the camera 1, the sensor 2, the motion capture system 3, and the microphone 4, and reflects the user's facial expression in an avatar.
[0022] 2, for example, a web camera serving as camera 1 captures a face of user U1 to obtain a camera image. The information processing device 11 recognizes the facial expression of user U1 based on the camera image captured by camera 1 and causes avatar A1 to display a facial expression corresponding to the facial expression of user U1. The avatar A1 displaying a facial expression corresponding to the user's facial expression is displayed on, for example, a display 11A connected to the information processing device 11.
[0023] In the information processing system of the present technology, facial expressions are converted into data as a combination of AUs (Action Units) and the positions (movements) of landmarks.
[0024] An AU is a movement that combines one or more facial muscles. The information processing device 11 recognizes the user's facial expression, for example, as the intensity of each AU, using a FACS (Facial Action Coding System) that converts facial expressions appearing on a person's face into machine-recognizable data based on a combination of AUs. The information processing device 11 controls the values of parameters (parameter values) set for the avatar based on the recognized intensity of each AU, thereby reflecting the user's facial expression in the avatar. The parameters set for the avatar correspond to, for example, each AU.
[0025] Furthermore, the information processing device 11 recognizes the user's facial expression as, for example, the movement of each landmark on the user's face, and controls the values of parameters set for the avatar based on the recognized movement of each landmark, thereby reflecting the user's facial expression in the avatar. The parameters set for the avatar correspond to, for example, each landmark.
[0026] Furthermore, the information processing device 11 can recognize the AU intensity and landmark movement of a part of the user's face, and estimate the expression of the user's entire face based on the recognition results, thereby estimating the AU intensity and landmark movement of other parts of the user's face.
[0027] Hereinafter, AUs and landmarks are also referred to as feature points, and the strength of an AU or the degree of movement of a landmark is also referred to as the strength of a feature point. In summary, the information processing device 11 recognizes (acquires) the strength of a user's feature points based on the camera image captured by the camera 1, and controls the parameter values for the avatar's feature points based on the strength of the user's feature points.
[0028] The information processing device 11 supplies an image including an avatar showing an expression corresponding to the user's expression and sound information supplied from the microphone 4 to the output device 21 in FIG.
[0029] The output device 21 is, for example, a device used by a user at a remote location from the user of the camera 1, and is configured by a PC, a smart device, an HMD, a TV, a camera, etc. The output device 21 includes a display 31 and a speaker 32. Hereinafter, the user of the camera 1, etc. will also be referred to as the "own user," and the user of the output device 21 will also be referred to as the "other user."
[0030] The display 31 displays an image supplied from the information processing device 11. The speaker 32 outputs a sound indicated by the sound information supplied from the information processing device 11.
[0031] FIG. 3 is a diagram showing an example of an image displayed on the output device 21 and the display 31. As shown in FIG.
[0032] Another user U2, who is in a remote location relative to the user U1 (FIG. 2), wears an HMD as an output device 21 on his / her head, as shown in A of FIG. 3. On a display 31 of the output device 21, an image including an avatar A1 showing an expression corresponding to the expression of the user U1 is displayed, as shown in B of FIG. 3, and the voice of the user U1 is output from a speaker 32.
[0033] The information processing device 11 may be a PC used by the user, or a server connected to the camera 1 via a network.
[0034] The functions of the information processing device 11 may be realized by a PC and a server sharing or cooperating in processing. For example, the PC recognizes the intensity of feature points based on camera images, and the server controls the parameter values of an avatar to generate an image including the avatar, and transmits the image to the output device 21.
[0035] Furthermore, the functions of the information processing device 11 may be realized by sharing or jointly performing processing among a PC, a server, and the output device 21. For example, the information processing device 11 supplies feature point information indicating the intensities of the user's feature points to the output device 21, and the output device 21 controls the parameter values of the avatar based on the feature point information to generate an image including the avatar.
[0036] FIG. 4 is a diagram showing examples of facial expressions of the user and facial expressions displayed on an avatar.
[0037] As shown on the left side of FIG. 4, when the facial expression of the user U1 is in a neutral state, the avatar A1 displays a neutral facial expression.
[0038] As shown in the center left of FIG. 4, when the user U1 starts to laugh, the avatar A1 shows an expression corresponding to the expression of the user U1 when he or she starts to laugh.
[0039] 4, if the user U1 looks down while smiling, the information processing device 11 will be unable to recognize the facial expression of the user U1 based on the camera image (lost state). Therefore, during the lost state, the avatar A1 will remain motionless and will have the unnatural facial expression it had just before entering the lost state.
[0040] Here, an example will be described in which the information processing device 11 is unable to recognize the entire facial expression of the user U1, but there are also cases in which the information processing device 11 is unable to recognize feature points of only a portion of the face of the user U1.
[0041] As shown on the right side of Figure 4, when the user U1 turns to face forward, the lost state is resolved and the information processing device 11 is able to recognize the facial expression of the user U1 again based on the camera image, the avatar A1 suddenly shows an expression corresponding to the facial expression of the user U1.
[0042] It is not desirable for the face of the avatar A1 to stop moving or for the facial expression of the avatar A1 to suddenly change frequently.
[0043] Possible factors on the user's side that can cause a lost state include: - The user's face appears out of the shooting range of camera 1, such as when the user leaves their seat or falls asleep; - The user's feature points disappear from the camera image, such as when the user laughs out loud and leans back, or wipes away tears; - The user's feature points disappear from the camera image unintentionally, such as when the user sneezes or their posture deteriorates over time; - The user's feature points become hidden by parts of the body other than the face, such as when the user rests their chin on their hand, their eyes are hidden by their bangs, or their mouth is hidden by their beard.
[0044] Possible factors on the device side that can cause the device to go into a lost state include: - The camera position changing - The HMD slipping off - The HMD camera becoming foggy due to the user's sweat or moisture.
[0045] Possible environmental factors that can cause the user to become lost include: - Recognizing the facial expressions of other people around the user - Recognizing the facial expressions of people on TV or on posters - Changes in the way the light hits the user.
[0046] Possible software-related factors that can cause a lost state include: - A decrease in frame rate - An error due to a conflict with other processing.
[0047] Therefore, in the information processing system of the present technology, if the information processing device succeeds in recognizing (acquiring) the intensity of the feature points of its own user, it controls the parameter values for the feature points of the avatar to be values corresponding to the intensity of the feature points of its own user, and if it fails to recognize the intensity of the feature points of its own user, it controls the parameter values to be values for when the recognition fails.
[0048] FIG. 5 is a diagram showing examples of facial expressions of the user and facial expressions displayed on an avatar when the information processing system of the present technology is used.
[0049] As shown on the left side of Figure 5, when the information processing device 11 is able to recognize the facial expression of the user U1 (when the information processing device 11 is able to successfully acquire the intensities of the feature points of the user), the avatar A1 displays a real facial expression that corresponds to the facial expression of the user U1.
[0050] As shown in the center of Figure 5, when the user U1 looks down and enters a lost state (when the intensity of the user's feature points fails to be acquired), the avatar A1 displays a lost expression, which is the expression immediately before entering the lost state.
[0051] 5, if the lost state continues (if the intensities of the feature points of the user fail to be acquired), the facial expression of the avatar A1 changes from the lost expression to a recovered expression. The recovered expression is the facial expression that is displayed on the avatar A1 when it is in the lost state.
[0052] When the lost state continues, the facial expression of the avatar A1 is changed from the lost expression to a recovery expression instead of remaining unchanged as the lost expression, thereby making it possible for the avatar A1 to continue to show a natural facial expression even if the lost state occurs.
[0053] <Flow of Making Avatar Show Recovery Facial Expression> Next, with reference to FIGS. 6 to 9, a flow of processing of making an avatar show a recovery facial expression will be described.
[0054] (A) Determination of Main Feature Points While the information processing device 11 is able to recognize the facial expression of the user, the information processing device 11 determines, from among the facial feature points of the user, feature points that are considered to be important for reflecting the facial expression of the user in the avatar as main feature points. Here, the information processing device 11 determines whether the main feature points are in a lost state.
[0055] For example, a feature point that satisfies the condition of the following formula (1) is determined as a main feature point: (1−k)×X t-1 + k × X t >Th...(1)
[0056] In equation (1), k is a coefficient, X is the intensity of the feature point, t is the time (frame number of the camera image), and Th is a threshold. In equation (1), a feature point whose intensity in several frames from the past frame to the current frame exceeds the threshold Th is determined as a main feature point.
[0057] The time series of the intensity of each feature point on the user's face is shown in the lower part of Fig. 6. In the lower part of Fig. 6, the horizontal axis represents time, and the vertical axis represents the intensity of the feature point.
[0058] From time t-2 to t, the intensities of feature points x, y, and z are equal to or greater than the threshold indicated by the dashed dotted line, and therefore feature points x, y, and z are determined as the main feature points.
[0059] As shown by the oval in the upper left of Fig. 6, feature point x indicates whether the outer corner of the eye narrows, and the narrower the outer corner of the eye of user U1, the higher the intensity of feature point x. Feature point y indicates whether the cheek lifts, and the more the cheek of user U1 lifts, the higher the intensity of feature point y. Feature point z indicates whether the corners of the mouth lift, and the more the corners of the mouth of user U1 lift, the higher the intensity of feature point z.
[0060] While the information processing device 11 can recognize the facial expression of the user U1, the parameter values for the feature points x, y, and z of the avatar A1 are controlled to be the same as the intensities of the feature points x, y, and z of the user U1, for example, as shown in the upper right side of Fig. 6. While the information processing device 11 can recognize the facial expression of the user U1, the parameter values of the feature points other than the main feature points of the avatar A1 are also controlled to be the same as the intensities of the feature points of the user U1.
[0061] Note that the main feature points may be determined based on the amount of change in intensity relative to the intensity of the feature points when the user U1's facial expression is neutral. Furthermore, the main feature points may be determined based on context information obtained by analyzing the speech of the user U1. For example, when the user U1 is uttering cheerful words, feature points whose intensity increases when the user U1 has a cheerful facial expression are determined as the main feature points. Other methods may also be used to determine the main feature points.
[0062] The main feature points are updated according to the passage of time, fluctuations in the amount of change in the feature intensity, the content of the conversation, the level of excitement in the conversation, and so on.
[0063] (B) Loss of Feature Points At time t+1, if the user U1 looks down as shown in the upper left of FIG. 7 and the information processing device 11 is no longer able to recognize the feature points x and z, the intensities of the feature points x and z recognized by the information processing device 11 will drop sharply, as shown by the dashed line in the lower part of FIG. 7.
[0064] In this case, the information processing device 11 determines that feature point x and feature point y have been lost. In order to suppress the influence of noise, it is preferable to determine whether or not a main feature point has been lost based on the average intensity over a period of several frames, or to determine that a main feature point has been lost when a series of frames has low intensity.
[0065] At time t+1 when feature point x and feature point y are lost, avatar A1 displays the facial expression that was displayed at time t, immediately before time t+1, as a lost facial expression, as shown in the upper right side of FIG. 7 .
[0066] (C) Determination of Return Point When the feature point x and the feature point y are lost, the information processing device 11 changes the facial expression of the avatar A1 from a lost facial expression to a returned facial expression. Hereinafter, the parameter value for expressing the returned facial expression is also referred to as a return point.
[0067] The return point is selected from among candidate return points. The candidate return points may be calculated when a major feature point is lost or may be set in advance. For example, as shown in FIG. 8 , the parameter value for feature point x of avatar A1 is controlled within a range from 0.0 to 1.0, and the parameter value for feature point x for expressing a lost facial expression (at time t+1) is 0.2. In this case, three parameter values, 0.0, 0.5, and 1.0, are set as candidate return points for feature point x, as shown by the white dots at the bottom of FIG. 8 .
[0068] The candidate return points may be set, for example, to the minimum or maximum parameter values for the feature points of avatar A1, or may be set at equal intervals within the range of possible parameter values. The intervals (resolution) at which return points are set may differ for each feature point. It is also possible to set candidate return points or vary the range of possible return points (maximum values) by taking into account context information and the impression of the facial expression (smiling, sad, etc.) expressed by avatar A1 before losing the main feature point.
[0069] The information processing device 11 selects, for example, a parameter value of 0.0 from the three parameter candidates as the return point, and changes the parameter value for the feature point x so that the parameter value becomes 0.0 at time t+2, as shown by the thick arrow at the bottom of Fig. 9. As a result, even when the user U1 is looking down as shown in the upper left of Fig. 9, the avatar A1 will show a return expression with the corners of its eyes wide open as shown in the upper right of Fig. 9.
[0070] In this way, the candidate closest to the parameter value for expressing the lost facial expression is selected as the return point. By selecting the candidate closest to the parameter value for expressing the lost facial expression as the return point, the amount of change in the facial expression of avatar A1 when changing from the lost facial expression to the returned facial expression can be reduced, and the change in the facial expression of avatar A1 can be made natural.
[0071] Note that the return point may be selected from the candidates in consideration of the balance with other major feature points. For example, the return points of feature points y and z may be selected based on the return point of feature point x, or the candidates for the return points of feature points y and z may be limited. Also, for example, the return points of feature points x and z may be selected based on the return point of feature point y that is not considered lost, or the candidates for the return points of feature points x and z may be limited.
[0072] The stable facial expression before the avatar entered the lost state may be displayed as the recovery facial expression. A stable facial expression is an expression that the user judges not to be unnatural among the facial expressions displayed by the avatar before the avatar entered the lost state. The user can specify such a stable facial expression (a not unnatural facial expression) during the period when the real facial expression is displayed by the avatar.
[0073] Alternatively, a neutral facial expression may be displayed as the return facial expression. A neutral facial expression is, for example, a facial expression in which the parameter values for all feature points are 0. The neutral facial expression may be defined by the user during advance calibration, or a facial expression displayed by the avatar for a predetermined period of time or longer may be defined as the neutral facial expression. When a neutral facial expression is displayed as the return facial expression, the return point is a fixed value set (defined) in advance.
[0074] The above describes an example in which only the parameter values for the lost major feature points are changed to the return points. Since the impression of an avatar's facial expression is often determined by the balance of each feature point, if the number of major feature points determined to be lost is greater than a predetermined number, the information processing device 11 may determine that all feature points are lost.
[0075] (D) Method of Changing from Lost Facial Expression to Returned Facial Expression If an avatar's facial expression suddenly changes from a lost facial expression to a returned facial expression, other users may feel that it is unnatural. Therefore, the information processing device 11 sets an animation curve that gradually changes parameter values as the avatar's facial expression changes from a lost facial expression to a returned facial expression.
[0076] FIG. 10 is a diagram showing an example of an animation curve.
[0077] An animation curve set for expressing a realistic facial expression is shown in A of Fig. 10. When the information processing device 11 can recognize the user's facial expression, the parameter value gradually changes from the minimum value to the maximum value in a quadratic manner over seven frames from time t10 to time t17, as shown in A of Fig. 10, thereby expressing a realistic facial expression on the avatar.
[0078] As shown in FIG. 10B, when a feature point is lost at time t14 and the return point is set to the maximum parameter value for that feature point, the parameter value gradually changes from the value at the time of loss to the return point in accordance with the original animation curve, causing the avatar to show a return expression.
[0079] The animation curve may be updated according to the difference between the parameter value at the time of loss and the return point. For example, as shown in C of Fig. 10, the parameter value gradually changes from the value at the time of loss to the return point in accordance with an animation curve in which the parameter value changes quadratically from the minimum value to the maximum value over nine frames from time t30 to time t39, causing the avatar to express a return expression.
[0080] An animation curve based on the predicted time required for the lost state to be resolved may be set for the expression of the recovered facial expression.
[0081] For example, if it is predicted that the lost state will be resolved soon, the information processing device 11 changes the parameter value in accordance with an animation curve in which the parameter value changes linearly. If it is predicted that the lost state will be resolved soon, the information processing device 11 may cause the avatar to continue to express a lost expression without expressing a recovery expression.
[0082] For example, if the lost state is predicted to continue for a predetermined period or longer, the information processing device 11 changes the parameter value in accordance with the longest animation curve that makes the change in facial expression look natural. If the lost state is predicted to continue for a predetermined period or longer, the information processing device 11 may turn on an automatic animation mode. In the automatic animation mode, the avatar performs, for example, a preset movement.
[0083] When a lost state occurs due to dropped frames in camera images or feature point information, the information processing device 11 can predict the time required for the lost state to be resolved. Also, when a lost state occurs due to a user leaning backward, the information processing device 11 can predict the time required for the lost state to be resolved by learning past recognition results of the user's facial expressions.
[0084] (E) Method of Changing from a Returned Expression to a Realistic Expression FIG. 11 is a diagram showing a flow of changing an avatar's expression from a returned expression to a realistic expression.
[0085] If the facial expression of avatar A1 changes suddenly when the lost state is resolved, other users may feel that it is unnatural. Therefore, even if the facial expression loss state is resolved, the restored facial expression is displayed for several frames, as shown in the center of Figure 11.
[0086] Thereafter, the information processing device 11 gradually changes the parameter value of each feature point, causing the avatar A1 to express a realistic facial expression, as shown on the right side of Fig. 11. In other words, when the lost state is resolved, the information processing device 11 gradually changes the value of each feature point after a predetermined delay. The information processing device 11 may change the parameter value of each feature point to a parameter value for expressing a realistic facial expression in accordance with an animation curve in which the parameter value changes quadratically. The animation curve may be updated according to the difference between the return point and the parameter value for expressing the realistic facial expression.
[0087] When the lost state is resolved, if the user's position has changed compared to before the lost state, the position and size of the avatar may be gradually changed, similar to the facial expression.
[0088] (F) Feedback for Lost Facial Expressions Images of an avatar expressing a lost facial expression may be recorded by the information processing device 11. In this case, the user can later view the image of the avatar expressing the lost facial expression and provide feedback on their evaluation of the lost facial expression (whether they want the avatar to express it again). When a lost facial expression that has been fed back by the user as an expression that the user does not want to express again is expressed, the information processing device 11, for example, reduces the number of frames required to determine whether a feature point has been lost, and causes the avatar to immediately express a restored facial expression.
[0089] Other users may provide feedback on their evaluations of the lost facial expressions of their own users' avatars. For example, if a user's avatar frequently displays a lost facial expression, other users may turn off the display of that avatar, set the avatar to a fixed facial expression, or turn on the auto-animation mode for that avatar.
[0090] <Example of using modal information or context information other than camera images> When a lost state occurs, a return point may be determined based on a parameter value for expressing a lost expression and modal information or context information other than camera images.
[0091] FIG. 12 is a diagram showing an example of modal information and context information.
[0092] The modal information is information (sensor information) about the user detected by a sensor other than the camera 1. In the example of FIG. 12, the modal information used to determine the return point includes sound information and gesture information. Sound information may record laughter or crying. Gesture information may record user movements such as clapping hands or wiping the area around the eyes with hands.
[0093] The modal information may be detected by the sensor 2, the motion capture system 3, and the microphone 4 after the lost state has occurred, or may be detected before the lost state has occurred. When a loss possibility value (described later) is equal to or greater than a predetermined threshold, detection of the modal information may be started.
[0094] Furthermore, when the information processing system allows multiple users to communicate with each other via their avatars, the information processing device 11 can use context information related to the communication via the avatars to determine the return point of the user's avatar. The context information related to the communication via the avatars includes the main feature points of the other users, the result of facial expression estimation based on the main feature points of the user, the result of estimation of the context of the conversation between users, etc.
[0095] For example, if it is estimated that the user has an angry expression based on the intensity of the user's main feature points before the user entered the lost state, a preset expression corresponding to the angry expression is displayed as the recovery expression. The duration of the preset expression is also determined based on the context information. If the lost state is not resolved even after the preset expression has been displayed for the duration determined based on the context information, for example, a neutral expression is displayed. Note that a fixed time may be set in advance for each expression as the duration of the preset expression.
[0096] <Calculation of Loss Probability Value> The information processing device 11 calculates a loss probability value (loss probability information) indicating the possibility of a lost state (failure to acquire the intensity of the feature points of the user) based on information supplied from the camera 1, the sensor 2, the motion capture system 3, and the microphone 4. Specifically, the information processing device 11 calculates the loss probability value based on the optical flow between frames of the camera image, sensor information from the IMU sensor, the context of the conversation, the characteristics of the user, whether the user is holding a tool, etc.
[0097] For example, it is thought that the user is more likely to enter a lost state if they perform actions such as laughing out loud, thinking deeply, looking down, looking up at the sky, or resting their chin on their hand, so if the user is likely to perform these actions, the lost probability value will be high.
[0098] 13 is a diagram showing an example of a time series of loss probability values, in which the horizontal axis represents time and the vertical axis represents loss probability values.
[0099] 13 , when the loss possibility value becomes equal to or greater than a threshold value at time t51, the information processing device 11 performs an interaction according to the loss possibility value in preparation for a lost state. For example, the information processing device 11 sets a gentle animation curve for the facial movement of the avatar.
[0100] FIG. 14 is a diagram showing an example of the strength of the feature points of the user and the parameter values of the feature points of the avatar when the loss possibility value is equal to or greater than a threshold value.
[0101] Assume that a feature point is lost at time t62, as shown by the dashed line in the upper part of Fig. 14. If the strength of the feature point of the user is directly reflected in the parameter value of the avatar, the avatar will show a lost expression after time t62, as shown by the dashed line in the middle part of Fig. 14.
[0102] Suppose that the loss possibility value exceeds a threshold value at time t61, which is before time t62. In this case, the information processing device 11 does not simply reflect the strength of the feature points of the user in the parameter values of the avatar, but rather extends the strength of the feature points of the user in the time direction and then reflects the strength in the parameter values of the avatar's feature points. In other words, when the loss possibility value is equal to or greater than the threshold value, the information processing device 11 performs an interaction that controls the avatar to slow down its movements.
[0103] As a result, even after time t62, the avatar's facial expression does not become immobile due to a lost facial expression, and the avatar continues to show a facial expression corresponding to the facial expression of the user, as shown in the lower part of FIG.
[0104] Therefore, even if the user is in a lost state as shown in the upper part of Figure 15, the avatar A1 continues to display an expression corresponding to the expression of the user U1 when the information processing device 11 is able to recognize the user's feature points, as shown in the lower part of Figure 15.
[0105] When the user U1 moves quickly, the loss possibility value is considered to be high. When the loss possibility value is high, the movement of the avatar A1 slows down, and it is expected that the user U1 will move more slowly in response to the movement of the avatar A1. Therefore, when the loss possibility value is high, the information processing device 11 can prevent the user U1 from becoming lost by controlling the avatar to move more slowly.
[0106] When the loss possibility value is equal to or greater than the threshold, the information processing device 11 may not only control the speed of the avatar's movement, but may also perform interactions such as causing each sensor to start detecting modal information or notifying the user that the loss possibility value is equal to or greater than the threshold.
[0107] For example, if the user's posture deteriorates over time and it is predicted that the user will become lost if the posture continues to deteriorate, the system may control the speed of the avatar's movement, start measuring modal information using each sensor, and notify the user. The user's posture is detected based on camera images or by an IMU sensor. Misalignment of the HMD relative to the user's face can also be detected as a change in the user's posture.
[0108] Furthermore, when the user grabs a tissue, the information processing device 11 increases the loss possibility value and controls the speed of the avatar's movement so that the avatar does not show a lost expression even if the user sneezes. At the same time as controlling the speed of the avatar's movement, the user's voice may be muted.
[0109] If a predetermined number or more of the facial feature points of the user are lost, not only may the loss probability value be increased, but a notification prompting recalibration may also be presented to the user.
[0110] <Example of User Turning On Automatic Animation Mode> If the user can foresee that the user will become lost, the user can turn on automatic animation mode by operating the information processing device 11. When turning on automatic animation mode, the user can input a predicted time until the lost state is resolved. When the predicted time is input, the information processing device 11 moves the avatar based on the predicted time.
[0111] <Example of User Defining Preset Facial Expressions> FIG. 16 is a diagram showing an example of a screen for defining preset facial expressions.
[0112] On the screen for defining preset facial expressions, for example, images P1 to P3 showing facial expressions previously expressed by an avatar are displayed in a row, as shown in Fig. 16. The information processing device 11 records parameter values for making the avatar express each facial expression. The user can select an image of an expression that the user wants to use as a recovery facial expression from among images P1 to P3. The information processing device 11 sets the facial expression shown in the image selected by the user as a preset facial expression.
[0113] In the example of FIG. 16, the facial expression shown in image P1, as indicated by the gray frame, is selected by the user as a preset facial expression.
[0114] In this way, the user can define a preset facial expression simply by selecting an image of a preferred facial expression from the images displayed on the screen. If the information processing system is used as a system for distributing avatar images, this function may be used to create a preset facial expression for the avatar.
[0115] <When the lost state continues> When the lost state continues for a predetermined period of time or longer, the information processing device 11 suggests to the user that the avatar be changed. For example, if the feature points around the mouth are lost because the user rests their chin on their hand, the information processing device 11 suggests to the user an avatar without a mouth. Also, if the feature points of the eyebrows are lost because the eyebrows are hidden by bangs, the information processing device 11 suggests to the user an avatar without eyebrows or with hidden eyebrows. The user may also be notified that the lost state has occurred due to the user resting their chin on their hand or bangs.
[0116] In this way, when a certain feature point is lost, the information processing device 11 may suggest an avatar that does not express the lost feature point, or automatically change the avatar to an avatar that does not express the lost feature point. The information processing device 11 may also make the avatar wear an accessory such as a mask or glasses so that the lost feature point is not expressed.
[0117] 17 , when the eye feature points of the user U11 are lost because the user U11's eyes are hidden by bangs, the information processing device 11 may estimate the intensity of the eye feature points based on the intensity of the feature points of the user U11's mouth, for example, and may reflect this in the parameter value for the eye feature points of the avatar A11. In this case, the intensity of the feature points of the user U11's mouth may or may not be reflected in the parameter value for the feature points of the mouth of the avatar A11.
[0118] <When feature points differ between user and avatar> There are feature points that the user has but the avatar does not, and conversely, there are feature points that the user does not have but the avatar does. For example, as shown in Fig. 18, the user U12 has a mouth but the avatar A11 does not. Also, the user U12 does not have animal ears but the avatar A11 does.
[0119] In this case, since the user U12 does not have animal ears, the intensity of feature points other than the animal ears, such as the feature point of the mouth of the user U12, is reflected in the parameter value for the animal ear feature point of the avatar A11. Modal information such as gestures and voice volume may also be reflected in the parameter value for the animal ear feature point of the avatar A11.
[0120] Furthermore, since the avatar A11 does not have a mouth, vibrations and sounds corresponding to the intensity of the feature points of the user U12's mouth are output from the speaker 32 of the output device 21.
[0121] <Example of changing the facial expression of the user's avatar based on information from other users> There may be a situation where other users are having a lively conversation, but the user cannot keep up with the conversation. In this situation, the user may not be able to make a happy face or may laugh later than the other users.
[0122] Therefore, when the facial expressions of the other users are happy but the user's facial expression is not happy, the information processing device 11 determines that the user's facial expression has been lost and causes the user's avatar to display a smiling recovery facial expression. Here, the information processing device 11 determines recovery points for each feature point of the user's avatar based on the context of the conversation between the other users, the facial expressions of the other users (avatars), etc.
[0123] The return point for each feature point of the user's avatar may be determined based on the average value of the parameter values for each feature point of the avatars used by multiple other users. Also, an automatic animation mode may be turned on, and the user's avatar may perform an action such as responding to a user's request.
[0124] <Example of correcting the algorithm> When the intensity of a user's feature points is recognized using single-class estimation, the intensity of a certain feature point also affects the intensity of other feature points. Therefore, if a feature point is lost, an incorrect value may be recognized as the intensity of other feature points. Also, if there are feature points that are lost from the beginning, an incorrect value may be recognized as the intensity of other feature points.
[0125] Therefore, for example, the information processing device 11 recognizes the presence or absence of an obstruction based on a camera image, determines whether or not there is a lost feature point based on the presence or absence of an obstruction, or determines whether or not there is a lost feature point based on the movement of the user. If there is a lost feature point, the information processing device 11 determines a return point for each feature point of the avatar while ignoring the intensities of all feature points of the user, and notifies the user that there is a lost feature point.
[0126] <Regarding Calibration> During preliminary calibration, the information processing device 11 sets a higher threshold value used to determine main feature points for feature points that are likely to be lost than for other feature points. For example, if the user has long hair, the information processing device 11 determines that feature points of the eyebrows and eyes are likely to be lost, and makes it less likely that the feature points of the eyebrows and eyes will be set as main feature points.
[0127] Furthermore, during the preliminary calibration, the information processing device 11 recognizes feature points that have already been lost. For example, if the user is wearing a mask, the information processing device 11 recognizes that feature points around the mouth have been lost.
[0128] During the preliminary calibration, the user can input feature points that the user does not want to be expressed by the avatar by operating the information processing device 11. The information processing device 11 considers the feature points input by the user to be lost, and turns on the automatic animation mode for the feature points or uses an avatar in which the feature points are not expressed.
[0129] During the preliminary calibration, other feature points such as the hands and head may be input as feature points that are not desired to be reflected in the avatar.
[0130] Conversely, the user can also input the feature points that he or she wants the avatar to express.
[0131] FIG. 19 is a diagram showing an example of a screen for inputting feature points to be reflected in an avatar.
[0132] 19 displays, for example, an image P21 of the user U21. When the user U21 wants to reflect the feature points of the eyes and nose in the avatar, the user U21 performs an operation such as drawing an ellipse around an area including the feature points of the eyes and nose. The information processing device 11 determines that feature points other than the eyes and nose are lost, and turns on an automatic animation mode for feature points other than the eyes and nose.
[0133] <Response to Lost State Due to Device-Side Factors> When a lost state occurs due to a deviation in the position of the camera 1, the information processing device 11 responds to the lost state by, for example, making the avatar express a recovery expression.
[0134] If the lost state continues for a predetermined period of time or longer and there is another camera that can be used to recognize the feature points of the user, the information processing device 11 may recognize the feature points of the user based on camera images captured by the other camera. If the lost state continues for a predetermined period of time or longer, the information processing device 11 sends a notification to the user urging the user to correct the position of the camera 1. In this case, it is expected that the user who received the notification will approach the camera 1 or shake the camera 1 to correct its position, which will cause further loss of the main feature points, and therefore the information processing device 11 increases the loss probability value while sending the notification to the user.
[0135] Furthermore, if a child, pet, or the like that may cause the position of the camera 1 to shift is approaching the camera 1, the information processing device 11 increases the loss probability value.
[0136] When a lost state occurs due to the HMD being displaced, the information processing device 11 responds to the lost state by, for example, performing an interaction according to the loss possibility value.
[0137] For example, if the feature points of the eyebrows or eyes are lost because the bangs are caught between the HMD and the user's face, the information processing device 11 may prompt the user to put on the HMD again, or estimate the intensity of the feature points of the eyebrows or eyes based on the intensity of feature points other than the eyebrows or eyes. The information processing device 11 may change the appearance of the avatar so that the eyebrows or eyes are not exposed, for example, by having the avatar wear sunglasses.
[0138] FIG. 20 is a diagram showing examples of the facial expressions of the avatar before the user removes the HMD and after the user puts the HMD back on.
[0139] When the user U51 removes and puts on the HMD 51 again, calibration needs to be performed again. Even if calibration is performed, the facial expression of the user U51 may not be correctly reflected on the avatar A1 after the user U51 puts on the HMD 51 again, as shown on the right side of Fig. 20 , even though the facial expression of the user U51 was correctly reflected on the avatar A1 before the user U51 removed the HMD 51, as shown on the left side of Fig. 20 .
[0140] Therefore, the information processing device 11 records the intensities of the feature points of the user U51 when the user U51's facial expression is neutral before removing the HMD 51. When the user U51 puts the HMD 51 back on, the information processing device 11 recognizes the intensities of the feature points of the user U51 again when the user U51's facial expression is neutral, and displays a UI that prompts the user U51 to adjust the wearing position (posture) of the HMD 51 so that the intensity becomes a value close to the intensity before removing the HMD 51.
[0141] FIG. 21 is a diagram showing an example of a UI that prompts the user to adjust the wearing position of the HMD 51.
[0142] 21 , when the HMD 51 is re-mounted, a crosshair L1 indicating the current posture of the HMD 51 and a crosshair L2 indicating the reference posture of the HMD 51 are displayed on the display of the HMD 51. Also, text is displayed on the display of the HMD 51 urging the user to adjust the wearing position of the HMD 51 so that the crosshair L1 overlaps the crosshair L2. The reference posture of the HMD 51 indicated by the crosshair L2 corresponds to the posture of the HMD 51 before the user U51 removes the HMD 51.
[0143] The user U51 may adjust the posture of the HMD 51 to some extent so that the crosshairs L1 and L2 overlap. After the posture of the HMD 51 is adjusted, the information processing device 11 performs software calibration so that the facial expression of the user U51 is correctly reflected on the avatar A1.
[0144] If the user is unable to put the HMD 51 back on in the same way as before, due to injury or the like, the automatic animation mode may be turned on.
[0145] <Response to Lost State Due to Environmental Factors> When a user is lost due to factors in the user's environment, the information processing device 11 responds by, for example, performing an interaction according to the loss possibility value.
[0146] For example, when a person other than the user approaches the camera 1, the information processing device 11 increases the loss probability value, taking into account the possibility that the person other than the user may block the user or the possibility that the facial expression of the person other than the user may be recognized.
[0147] For example, the information processing device 11 recognizes that an event that may result in a lost state, such as an intercom ringtone, a call from another person, or an incoming phone call, has occurred by collecting sounds around the user or by performing app linkage. When an event that may result in a lost state occurs, the information processing device 11 increases the lost possibility value.
[0148] For example, when the user puts on glasses, the information processing device 11 re-determines the main feature points.
[0149] For example, the information processing device 11 detects the brightness of the environment around the user. When the environment around the user becomes dark, the information processing device 11 turns on the lights around the user or increases the brightness of the HMD display to prevent the user from becoming lost.
[0150] <Response to Lost State Due to Software-Side Factors> When an entire frame is lost, such as when a frame is dropped due to an error, the information processing device 11 responds to the lost state by, for example, making the avatar express a recovery facial expression.
[0151] If a frame is dropped due to a network problem, a method for dealing with the lost state is taken depending on whether the avatar image is output from the device on the user's side or from the device on the other user's side.
[0152] For example, when the device on the user's side outputs an image of an avatar, if the network on the user's side is unstable, the output device 21 on the other user's side turns on the automatic animation mode for the user's avatar.
[0153] For example, when a device on the user's side outputs an image of an avatar, if the network on the other user's side is unstable, the output device 21 on the other user's side causes the user's avatar to display, as a return expression, a facial expression or a neutral facial expression that was displayed by the user's avatar when the network was stable. Also, the output device 21 on the other user's side displays an avatar locally held by the other user as the user's avatar, and performs lip-sync to move the avatar in sync with the user's voice.
[0154] For example, when an image of an avatar is output by a device on the side of another user, if the network on the user's side is unstable, the output device 21 on the side of the other user interpolates the parameter values of the feature points of the avatar by predicting the strength of the feature points while frames are being dropped.
[0155] For example, when an image of an avatar is output from a device on the other user's side, if the network on the other user's side is unstable, the output device 21 on the other user's side reduces the number of main feature points and reflects this in the user's avatar. The output device 21 on the other user's side also displays an avatar with fewer feature points as the user's avatar. Furthermore, if the intensity of the main feature points changes suddenly, the output device 21 on the other user's side gradually changes the facial expression of the user's avatar by temporarily displaying a recovery facial expression.
[0156] <Application to Features Other Than Facial Feature Points> The information processing device 11 recognizes not only the intensity of facial feature points but also the intensity of feature points of other body parts, and when a feature point of another body part is lost, determines a return point for that feature point so that the movement of the avatar does not appear unnatural.
[0157] For example, the information processing system can use the movement of muscles and tendons in the arm as a feature point. The information processing system can also use the line of sight as a feature point. For example, if there is a high possibility that the line of sight feature point will be lost, the information processing device 11 increases the loss possibility value.
[0158] 2. Configuration and Operation of Information Processing Apparatus> FIG. 22 is a block diagram showing an example of the functional configuration of the information processing apparatus 11. As shown in FIG.
[0159] As shown in FIG. 22, the information processing device 11 includes a data input unit 101, a recognition unit 102, a loss possibility value calculation unit 103, a personal data acquisition unit 104, a main feature point determination unit 105, and an avatar control unit 106.
[0160] The data input unit 101 acquires information supplied from the camera 1, sensor 2, motion capture system 3, and microphone 4, and supplies the information to the loss possibility value calculation unit 103 and the avatar control unit 106. The data input unit 101 also supplies camera images to the recognition unit 102.
[0161] The recognition unit 102 functions as a feature point acquisition unit that recognizes and acquires the intensities of the feature points of the user himself / herself based on the camera image supplied from the data input unit 101. The recognition unit 102 supplies the recognition results of the intensities of the feature points of the user himself / herself to the main feature point determination unit 105 and the avatar control unit 106.
[0162] The recognition unit 102 determines whether or not a main feature point has been lost based on the recognition result of the strength of the main feature point determined by the main feature point determination unit 105, and supplies a lost flag indicating whether or not the main feature point has been lost to the avatar control unit 106. If it is determined that the main feature point has been lost, the recognition unit 102 sets the lost flag to true.
[0163] The loss possibility value calculation unit 103 calculates a loss possibility value based on the information supplied from the data input unit 101 and supplies the loss possibility value to the avatar control unit 106 .
[0164] The personal data acquisition unit 104 acquires personal data such as a threshold value used to determine the main feature points, the user's evaluation of lost expressions, and preset expressions defined by the user, and supplies this data to the main feature point determination unit 105 and the avatar control unit 106.
[0165] The main feature point determination unit 105 determines a main feature point based on the time series of the recognition results of the feature point intensity by the recognition unit 102. For example, the main feature point determination unit 105 determines a feature point whose intensity over several frames from the past frame to the current frame exceeds a threshold value supplied from the personal data acquisition unit 104 as a main feature point.
[0166] The main feature point determination unit 105 supplies information indicating the main feature points to the recognition unit 102 and the avatar control unit 106 .
[0167] The avatar control unit 106 has an avatar DB (Data Base) in which avatar data is recorded. The avatar control unit 106 controls, for example, the facial expression of the avatar by using, for example, Blend Shape.
[0168] The avatar control unit 106 controls the parameters of the avatar corresponding to the main feature points recognized by the recognition unit 102 so that the parameter values correspond to the intensities of the main feature points, thereby causing the avatar to express a realistic facial expression.
[0169] When the recognition unit 102 determines that a principal feature point has been lost, the avatar control unit 106 determines a return point based on information supplied from the data input unit 101, a parameter value for the principal feature point immediately before the principal feature point was lost, etc. The avatar control unit 106 causes the avatar to express a return facial expression by gradually changing the parameter value for the lost feature point from the parameter value for expressing a lost facial expression to the return point.
[0170] When the recognition unit 102 determines that the lost state of the main feature points has been resolved, the avatar control unit 106 causes the avatar to express a realistic expression by gradually changing the parameter value for the lost feature points from a parameter value for expressing a restored expression (return point) to a parameter value for expressing a realistic expression (a parameter value corresponding to the intensity of the feature points of the user recognized again by the recognition unit 102).
[0171] The avatar control unit 106 controls the on / off of the automatic animation mode. The avatar control unit 106 has an expression DB in which expression data indicating a time series of parameter values of each feature point for automatically changing the avatar's expression is recorded. When the automatic animation mode is on, the avatar control unit 106 acquires the expression data from the expression DB and controls the avatar's parameters corresponding to each feature point so that the parameter values become the parameter values indicated by the expression data, thereby automatically changing the avatar's expression.
[0172] The avatar control unit 106 performs various interactions according to the loss possibility value calculated by the loss possibility value calculation unit 103 .
[0173] The avatar control unit 106 generates (renders) an image including an avatar and supplies it to the output device 21. When the image including an avatar is generated by the output device 21, the avatar control unit 106 controls the avatar by transmitting parameter values for each feature point of the avatar to the output device 21.
[0174] Next, a process in which the information processing device 11 reflects the user's facial expression on an avatar will be described with reference to the flowchart of FIG.
[0175] In step S1, the information processing device 11 performs calibration. Here, for example, the user inputs feature points that the user does not want the avatar to display, or adjusts the position of the camera 1 or the wearing position of the HMD. In addition, the main feature point determination unit 105 determines the main feature points.
[0176] In step S2, the recognition unit 102 recognizes the intensity of the feature points of the user based on the camera image captured by the camera 1.
[0177] In step S3, the main feature point determination unit 105 determines whether or not it is necessary to re-determine the main feature points.
[0178] If it is determined in step S3 that redetermining of the main feature points is not necessary, then in step S4, the recognition unit 102 determines whether the main feature points have been lost.
[0179] If it is determined in step S4 that a main feature point has been lost, the information processing device 11 performs lost-state processing in step S5. The lost-state processing may include turning on automatic animation mode and determining a return point. Details of the lost-state processing will be described later with reference to FIG. 24.
[0180] On the other hand, if it is determined in step S4 that the main feature points are not lost, the recognition unit 102 determines in step S6 whether the lost flag is true.
[0181] If it is determined in step S6 that the lost flag is true, the information processing device 11 performs re-recognition processing in step S7. The re-recognition processing calculates the speed of the avatar's movement when changing the avatar's facial expression from a restored facial expression to a realistic facial expression. Details of the re-recognition processing will be described later with reference to FIG. 25.
[0182] After the processing of step S5 or step S7 is performed, or if it is determined in step S6 that the lost flag is false, in step S8, the avatar control unit 106 updates the parameter values for each feature point of the avatar.
[0183] Here, for example, the avatar control unit 106 updates the parameter value for each feature point of the avatar to a return point, or to a parameter value corresponding to the strength of the feature point recognized by the recognition unit 102. Thereafter, the process returns to step S2, and the subsequent processes are repeated.
[0184] If it is determined in step S3 that redetermining of the principal feature points is necessary, then in step S9 the principal feature point determiner 105 determines whether a sufficient number of frames (number of data) of the time series (log) of feature point intensities have been collected to determine the principal feature points.
[0185] If it is determined in step S9 that the logs of feature point intensities have been collected for a sufficient number of frames to determine the main feature points, the main feature point determiner 105 determines the main feature points in step S10, after which the process proceeds to step S4 and subsequent steps are performed.
[0186] On the other hand, if it is determined in step S9 that the logs of feature point intensities have not been collected for a sufficient number of frames to determine the main feature point, the main feature point determiner 105 continues to collect the logs of the intensities of each feature point in step S11. Then, the process proceeds to step S8, where the parameter values of the main feature points of the avatar are updated to, for example, the parameter values corresponding to the intensities of the main feature points recognized by the recognition unit 102.
[0187] Next, the lost time process performed in step S5 of FIG. 23 will be described with reference to the flowchart of FIG.
[0188] In step S21, the recognition unit 102 determines whether the lost flag is true.
[0189] If it is determined in step S21 that the lost flag is true, the avatar control unit 106 determines in step S22 whether the period during which the main feature points have been lost is longer than a threshold value.
[0190] If it is determined in step S22 that the period during which the main feature points have been lost is longer than the threshold, the avatar control unit 106 turns on the auto-animation mode in step S23. Then, the process returns to step S5 in Fig. 23, where the parameter values for each feature point of the avatar are updated to, for example, the parameter values indicated by the facial expression data recorded in the facial expression DB.
[0191] On the other hand, if it is determined in step S21 that the lost flag is false, or if it is determined in step S22 that the period during which the main feature points have been lost is shorter than the threshold, the process proceeds to step S24. In step S24, the recognition unit 102 determines whether a sufficient number of frames (number of data) of camera images have been input to determine that the main feature points have been lost.
[0192] If it is determined in step S24 that a sufficient number of camera images have been input to determine that the main feature points have been lost, then in step S25, the recognition unit 102 sets the lost flag to true.
[0193] In step S26, the avatar control unit 106 determines whether or not a return point (preset facial expression) has been set in advance, for example, during calibration.
[0194] If it is determined in step S26 that a return point has not been set in advance, the avatar control unit 106 determines a new return point in step S27. Thereafter, the process returns to step S5 in Fig. 23, where, for example, the parameter values for the main feature points of the avatar are updated so that the return point becomes the newly determined return point, and the avatar shows a return facial expression.
[0195] On the other hand, if it is determined in step S26 that a return point has been set in advance, the process returns to step S5 in FIG. 23, where, for example, the parameter values for the main feature points of the avatar are updated so as to reach the previously set return point, and a return facial expression is displayed on the avatar.
[0196] If it is determined in step S24 that the number of camera images input is not sufficient to determine that the main feature points are lost, the process returns to step S5 in FIG. 23, and, for example, the parameter values for the main feature points of the avatar are not updated, and a lost expression is displayed on the avatar.
[0197] Next, the re-recognition process performed in step S7 of FIG. 23 will be described with reference to the flowchart of FIG.
[0198] In step S41, the recognition unit 102 sets the lost flag to false.
[0199] In step S42, the recognition unit 102 determines, based on the camera images, whether there has been a change in the user's appearance before the lost state and after the lost state is resolved. If the user wears a mask, for example, during the period in which the main feature points were lost, it is determined that there has been a change in the user's appearance.
[0200] If it is determined in step S42 that there has been a change in the user's appearance, the avatar control unit 106 changes the appearance of the avatar in step S43. If the user was wearing a mask during the period in which the main feature points were lost, for example, the avatar control unit 106 refers to the avatar DB and displays an avatar without a mouth as the user's avatar or makes the avatar wear a mask.
[0201] On the other hand, if it is determined in step S42 that there has been no change in the user's appearance, the process of step S43 is skipped.
[0202] In step S44, the avatar control unit 106 calculates the speed (animation speed) of the facial movement when changing the avatar's facial expression from the normal facial expression to a realistic facial expression. Then, the process returns to step S7 in Fig. 23, where the parameter values of the main feature points are updated so that the parameter values correspond to the intensities of the feature points recognized by the recognition unit 102, and the realistic facial expression is displayed on the avatar.
[0203] Next, a process in which the information processing device 11 performs an interaction according to a loss possibility value will be described with reference to a flowchart in Fig. 26. The process in Fig. 26 is executed in parallel with the process in Fig. 23, for example.
[0204] In step S61, the loss possibility value calculation unit 103 calculates a loss possibility value.
[0205] In step S62, the avatar control unit 106 determines whether the loss possibility value is equal to or greater than a threshold value.
[0206] If it is determined in step S62 that the loss possibility value is equal to or greater than the threshold value, the information processing device 11 determines in step S63 whether recalibration is necessary.
[0207] If it is determined in step S63 that recalibration is necessary, the information processing apparatus 11 performs recalibration in step S64.
[0208] On the other hand, if it is determined in step S63 that recalibration is not necessary, the avatar control unit 106 determines in step S65 whether or not to change the appearance of the avatar.
[0209] If it is determined in step S65 that the appearance of the avatar should be changed, the avatar control unit 106 changes the appearance of the avatar in step S66.
[0210] On the other hand, if it is determined in step S65 that the appearance of the avatar will not be changed, in step S67, the avatar control unit 106 performs an interaction according to the loss possibility value.
[0211] After the processing of step S64, step S66, or step S67 is performed, the processing returns to step S61, and the subsequent processing is repeated. Also, if it is determined in step S62 that the loss possibility value is less than the threshold value, the processing returns to step S61, and the subsequent processing is repeated.
[0212] As described above, in the information processing device of the present technology, if the recognition (acquisition) of the intensity of the user's feature points is successful, the parameter values for the avatar's feature points are controlled to be values corresponding to the intensity of the user's feature points, and if the recognition of the intensity of the user's feature points is unsuccessful, the parameter values are controlled to be values for the unsuccessful case. This makes it possible for the avatar to display a natural facial expression even in a lost state.
[0213] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware, or into a general-purpose personal computer, etc.
[0214] FIG. 27 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0215] A CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .
[0216] An input / output interface 505 is also connected to the bus 504. An input unit 506 including a keyboard, a mouse, etc., and an output unit 507 including a display, a speaker, etc. are connected to the input / output interface 505. Also connected to the input / output interface 505 are a storage unit 508 including a hard disk, a nonvolatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 that drives removable media 511.
[0217] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program stored in the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.
[0218] The program executed by the CPU 501 is installed in the storage unit 508 by being recorded on, for example, a removable medium 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.
[0219] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0220] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0221] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0222] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0223] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0224] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0225] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0226] <Examples of Combinations of Configurations> The present technology can also have the following configurations.
[0227] (1) An information processing device comprising: an avatar control unit that, when acquiring the intensity of a user's feature point, controls a parameter value for the feature point of an avatar to be a value corresponding to the intensity of the user's feature point, and, when acquiring the intensity of the user's feature point fails, controls the parameter value to be a value used in the event of failure. (2) The information processing device described in (1), in which the avatar control unit determines the value used in the event of failure based on the value of the parameter before failing to acquire the intensity of the user's feature point. (3) The information processing device described in (2), in which the avatar control unit determines the value used in the event of failure based on the intensity of the user's feature point from among a plurality of candidates set at equal intervals within a possible range of the parameter. (4) The information processing device described in (2) or (3), in which, when acquiring the intensity of a certain feature point of the user fails, the avatar control unit determines the value used in the event of failure based on the value of the parameter for another feature point of the avatar. (5) The information processing device according to any one of (2) to (4), further comprising a feature point acquisition unit that acquires intensities of the feature points of the user based on a camera image obtained by photographing the user with a camera. (6) The information processing device according to (5), wherein the avatar control unit determines the failure value based on sensor information about the user detected by a sensor other than the camera. (7) The information processing device according to any one of (2) to (6), wherein the avatar control unit determines the failure value based on context information about communication via the avatar. (8) The information processing device according to (1), wherein the failure value is a fixed value that is set in advance. (9) The information processing device according to any one of (1) to (8), wherein the avatar control unit, when failing to acquire the intensities of the feature points of the user, gradually changes the value of the parameter from the value of the parameter before the failure to acquire the intensities of the feature points of the user to the failure value.(10) The information processing device according to any of (1) to (9), wherein, when acquiring the intensity of the feature point of the user is successful in a state where acquiring the intensity of the feature point of the user has failed, the avatar control unit gradually changes the value of the parameter from a value used in the event of the failure to a value corresponding to the intensity of the feature point of the user. (11) The information processing device according to (10), wherein, when acquiring the intensity of the feature point of the user is successful in a state where acquiring the intensity of the feature point of the user has failed, the avatar control unit gradually changes the value of the parameter after a predetermined delay. (12) The information processing device according to any of (1) to (7) and (9) to (11), further comprising a calculation unit that calculates possibility information indicating a possibility of failing to acquire the intensity of the feature point of the user, and the avatar control unit performs an interaction according to the possibility information. (13) The information processing device according to (12), wherein the avatar control unit performs the interaction by controlling the avatar to move more slowly when a possibility of failing to acquire the intensity of the feature point of the user is equal to or greater than a threshold. (14) The information processing device according to (12) or (13), wherein the avatar control unit performs the interaction by notifying the user that the possibility of failing to acquire the intensity of the feature point of the user is equal to or greater than a threshold. (15) The information processing device according to any of (12) to (14), further comprising a feature point acquisition unit that acquires intensities of the feature points of the user based on a camera image obtained by photographing the user with a camera, wherein the avatar control unit determines the failure value based on sensor information about the user detected by a sensor other than the camera, and performs the interaction by starting detection of the sensor information by the sensor when the possibility of failing to acquire the intensity of the feature point of the user is equal to or greater than a threshold.(16) The information processing device according to any of (1) to (15), further comprising a determination unit that determines a main feature point among the feature points of the user based on intensities of the feature points, wherein the avatar control unit controls values of the parameters for the main feature points of the avatar. (17) The information processing device according to any of (1) to (16), wherein the feature points include AUs or landmarks. (18) The information processing device according to any of (1) to (17), wherein the avatar control unit changes the appearance of the avatar when there is a change in the appearance of the user. (19) An information processing method comprising: when acquisition of intensities of the user's feature points is successful, controlling values of the parameters for the feature points of the avatar to be values corresponding to the intensities of the feature points of the user, and when acquisition of intensities of the user's feature points fails, controlling the values of the parameters to be values for use in the event of failure. (20) A program for causing a computer to execute a process of controlling, when acquisition of the intensity of a user's feature point is successful, a parameter value for the feature point of an avatar to become a value corresponding to the intensity of the user's feature point, and, when acquisition of the intensity of the user's feature point is unsuccessful, controlling the parameter value to become a value for the case of failure.
[0228] REFERENCE SIGNS LIST 1 Camera, 2 Sensor, 3 Motion capture system, 4 Microphone, 11 Information processing device, 21 Output device, 31 Display, 32 Speaker 101 Data input unit, 102 Recognition unit, 103 Loss possibility value calculation unit, 104 Personal data acquisition unit, 105 Main feature point determination unit, 106 Avatar control unit
Claims
1. An information processing device comprising an avatar control unit that, if acquisition of the strength of a user's feature point is successful, controls the value of a parameter for the avatar's feature point to be a value corresponding to the strength of the user's feature point, and, if acquisition of the strength of the user's feature point fails, controls the value of the parameter to be a value used in the event of failure.
2. The information processing device according to claim 1, wherein the avatar control unit determines the value for when the failure occurs based on the value of the parameter before the failure to acquire the intensity of the feature point of the user.
3. The information processing device according to claim 2, wherein the avatar control unit determines the failure value from among a plurality of candidates set at equal intervals within the range of the parameter based on the strength of the feature point of the user.
4. The information processing device according to claim 2, wherein, when the avatar control unit fails to acquire the intensity of one of the feature points of the user, the avatar control unit determines the value for the failure based on the parameter values for other feature points of the avatar.
5. The information processing device according to claim 2, further comprising a feature point acquisition unit that acquires the intensity of the feature points of the user based on a camera image obtained by photographing the user with a camera.
6. The information processing device according to claim 5, wherein the avatar control unit determines the failure value based on sensor information about the user detected by a sensor other than the camera.
7. The information processing device according to claim 2, wherein the avatar control unit determines the failure value based on context information related to communication via the avatar.
8. The information processing device according to claim 1, wherein the value for failure is a preset fixed value.
9. The information processing device according to claim 1, wherein, when the avatar control unit fails to acquire the intensity of the feature point of the user, the avatar control unit gradually changes the value of the parameter from the value of the parameter before the failure to acquire the intensity of the feature point of the user to a value for the failure.
10. The information processing device according to claim 1, wherein, when the avatar control unit has failed to acquire the intensity of the feature point of the user but has succeeded in acquiring the intensity of the feature point of the user, the avatar control unit gradually changes the value of the parameter from the value used in the case of the failure to a value corresponding to the intensity of the feature point of the user.
11. The information processing device according to claim 10, wherein, when the avatar control unit has failed to acquire the intensity of the feature point of the user but has succeeded in acquiring the intensity of the feature point of the user, the avatar control unit delays for a predetermined period and then gradually changes the value of the parameter.
12. The information processing device according to claim 1, further comprising a calculation unit that calculates possibility information indicating a possibility of failing to acquire the intensity of the feature point of the user, wherein the avatar control unit performs an interaction according to the possibility information.
13. The information processing device according to claim 12, wherein the avatar control unit performs the interaction by controlling the avatar to move slowly when the possibility of failing to acquire the intensity of the feature point of the user is equal to or greater than a threshold.
14. The information processing device according to claim 12, wherein, when the possibility of failing to acquire the intensities of the feature points of the user is equal to or greater than a threshold, the avatar control unit performs the interaction by notifying the user that the possibility of failing to acquire the intensities of the feature points of the user is equal to or greater than a threshold.
15. An information processing device as described in claim 12, further comprising a feature point acquisition unit that acquires the intensity of the feature points of the user based on a camera image obtained by photographing the user with a camera, wherein the avatar control unit determines the value for the failure case based on sensor information about the user detected by a sensor other than the camera, and when the possibility of failing to acquire the intensity of the feature points of the user is equal to or greater than a threshold, performs the interaction by starting detection of the sensor information by the sensor.
16. The information processing device according to claim 1, further comprising a determination unit that determines a main feature point among the feature points of the user based on the intensities of the feature points, and the avatar control unit controls the values of the parameters for the main feature point of the avatar.
17. The information processing device according to claim 1, wherein the feature points include AUs or landmarks.
18. The information processing device according to claim 1, wherein the avatar control unit changes the appearance of the avatar when there is a change in the appearance of the user.
19. An information processing method comprising: when acquisition of the intensity of a user's feature point is successful, controlling the value of a parameter for the avatar's feature point to be a value corresponding to the intensity of the user's feature point; and when acquisition of the intensity of the user's feature point fails, controlling the value of the parameter to be a value for when failure occurs.
20. A program for causing a computer to execute a process of controlling, if acquisition of the intensity of a user's feature point is successful, the value of a parameter for the avatar's feature point to become a value corresponding to the intensity of the user's feature point, and, if acquisition of the intensity of the user's feature point fails, controlling the value of the parameter to become a value for when failure occurs.
Citation Information
Patent Citations
Face information transmitting device, face information transmitting method and recording medium with its program recorded
JP2006331065A
Terminal equipment, information communication method, and information communication program
JP2015172883A
Recording and sending emojis
JP2020520030A