Expression prediction using image motion-based metrics

By training the machine learning model, the expression unit value is calculated using the motion metric value of the user's facial part image, the problem of difficult detection of subtle changes in the face in the prior art is solved, and efficient and robust user expression prediction is achieved in AR/VR systems.

CN115515491BActive Publication Date: 2025-08-08MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180029499.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-20
Filing Date
2021-04-19
Publication Date
2025-08-08
Estimated Expiration
2041-04-19

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently predict user expressions based on part of the image of the user's face, especially in AR/VR systems, where subtle changes in the face are difficult to detect and are not robust enough.

Method used

By training a machine learning model, the expression unit value is calculated using the motion metric values of partial images of the user's face (such as eyes), and the model parameters are adjusted by comparing error data to achieve prediction of user expressions.

Benefits of technology

It allows efficient and robust prediction of user expressions based on part of user facial images in AR/VR systems, and is suitable for customized and personalized expression prediction, improving detection accuracy and system adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115515491B_ABST
    Figure CN115515491B_ABST
Patent Text Reader

Abstract

A technique for training a machine learning model to predict user expressions is disclosed. A plurality of images are received, each of the plurality of images containing at least a portion of a user's face. A plurality of values of a motion metric are calculated based on the plurality of images, each of the plurality of values of the motion metric indicating a motion of the user's face. A plurality of values of an expression unit are calculated based on the plurality of values of the motion metric, each of the plurality of values of the expression unit corresponding to a degree to which the user's face is producing the expression unit. A machine learning model is trained using the plurality of images and the plurality of values of the expression unit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63,012,579, filed on April 20, 2020, entitled “EXPRESSION PREDICTION USING IMAGE-BASED MOVEMENT METRIC,” which is incorporated herein by reference in its entirety for all purposes. Background Art

[0003] Modern computing and display technologies have facilitated the development of so-called "virtual reality" or "augmented reality" experience systems, in which digitally reproduced images, or portions thereof, are presented to a user in such a way that they appear or can be perceived as real. Virtual reality or "VR" scenarios generally involve the presentation of digital or virtual image information that is opaque to other actual real-world visual input; augmented reality or "AR" scenarios generally involve the presentation of digital or virtual image information as an enhancement to the visualization of the actual world around the user.

[0004] Despite advances in these display technologies, there remains a need in the art for improved methods, systems, and devices related to augmented reality systems, particularly display systems. Summary of the Invention

[0005] The present disclosure generally relates to techniques for improving the performance and user experience of optical systems. More particularly, embodiments of the present disclosure provide systems and methods for predicting a user's facial expressions based on an image of the user's face. Although the present disclosure is often described with reference to augmented reality (AR) devices, the present disclosure is applicable to a variety of applications.

[0006] A summary of various embodiments of the present invention is provided below as a list of examples. As used below, any reference to a series of examples should be understood as a separate reference to each of these examples (e.g., "Examples 1-4" should be understood as "Examples 1, 2, 3, or 4").

[0007] Example 1 is a method for training a machine learning model to predict user expressions, the method comprising: receiving a plurality of images, each of the plurality of images containing at least a portion of a user's face; calculating a plurality of values of a motion metric based on the plurality of images, each of the plurality of values of the motion metric indicating a movement of the user's face; calculating a plurality of values of an expression unit based on the plurality of values of the motion metric, each of the plurality of values of the expression unit corresponding to the extent to which the user's face is producing the expression unit; and using the plurality of images and the plurality of values of the expression unit, training the machine learning model by: generating training output data by the machine learning model based on the plurality of images; and modifying the machine learning model based on the plurality of values of the expression unit and the training output data.

[0008] Example 2 is a method according to Example 1, wherein the training output data includes multiple output values for each expression unit in a set of expression units, the expression unit being a first expression unit from the set of expression units.

[0009] Example 3 is a method according to Example 2, wherein the set of expression units includes at least one of the following: raising the inner eyebrow, raising the outer eyebrow, lowering the eyebrow, raising the upper eyelid, raising the cheek, tightening the eyelid, wrinkling the nose, closing the eyes, blinking the left eye, or blinking the right eye.

[0010] Example 4 is a method according to Examples 1-3, wherein training the machine learning model using the multiple images and the multiple values of the expression units also includes: performing a comparison of the multiple values of the expression units with the multiple output values of the expression units of the training output data; and generating error data based on the comparison, wherein the machine learning model is modified based on the error data.

[0011] Example 5 is a method according to Examples 1-4, wherein the machine learning model is an artificial neural network with a set of adjustable parameters.

[0012] Example 6 is a method according to Examples 1-5, wherein the motion metric is the number of eye pixels, and wherein calculating the multiple values of the motion metric based on the multiple images includes: segmenting each image of the multiple images so that each image of the multiple images includes eye pixels and non-eye pixels; counting the number of eye pixels in each image of the multiple images; and setting each of the multiple values of the motion metric to be equal to the number of eye pixels in a corresponding image from the multiple images.

[0013] Example 7 is a method according to Examples 1-6, wherein calculating the multiple values of the expression unit based on the multiple values of the motion metric includes: identifying a first extreme value among the multiple values of the motion metric and a first corresponding image for which the first extreme value is identified; setting each of the multiple values of the expression unit associated with the first corresponding image to be equal to one; identifying a second extreme value among the multiple values of the motion metric and a second corresponding image for which the second extreme value is identified; setting each of the multiple values of the expression unit associated with the second corresponding image to be equal to zero; and setting each remaining value of the multiple values by interpolating between zero and one.

[0014] Example 8 is a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: receiving a plurality of images, each of the plurality of images containing at least a portion of a user's face; calculating a plurality of values of a motion metric based on the plurality of images, each of the plurality of values of the motion metric indicating movement of the user's face; calculating a plurality of values of the expression unit based on the plurality of values of the motion metric, each of the plurality of values of the expression unit corresponding to the extent to which the user's face is producing the expression unit; and training a machine learning model using the plurality of images and the plurality of values of the expression unit by: generating training output data by the machine learning model based on the plurality of images; and modifying the machine learning model based on the plurality of values of the expression unit and the training output data.

[0015] Example 9 is a non-transitory computer-readable medium according to example 8, wherein the training output data comprises a plurality of output values for each expression unit in a set of expression units, the expression unit being a first expression unit from the set of expression units.

[0016] Example 10 is a non-transitory computer-readable medium according to Example 9, wherein the set of expression units includes at least one of: raising the inner eyebrow, raising the outer eyebrow, lowering the eyebrow, raising the upper eyelid, raising the cheek, tightening the eyelid, wrinkling the nose, closing the eyes, blinking the left eye, or blinking the right eye.

[0017] Example 11 is a non-transitory computer-readable medium according to Examples 8-10, wherein training the machine learning model using the multiple images and the multiple values of the expression unit also includes: performing a comparison of the multiple values of the expression unit with the multiple output values of the expression unit of the training output data; and generating error data based on the comparison, wherein the machine learning model is modified based on the error data.

[0018] Example 12 is a non-transitory computer-readable medium according to Examples 8-11, wherein the machine learning model is an artificial neural network having a set of adjustable parameters.

[0019] Example 13 is a non-transitory computer-readable medium according to Examples 8-12, wherein the motion metric is a number of eye pixels, and wherein calculating the multiple values of the motion metric based on the multiple images includes: segmenting each image of the multiple images so that each image of the multiple images includes eye pixels and non-eye pixels; counting the number of eye pixels in each image of the multiple images; and setting each of the multiple values of the motion metric to be equal to the number of eye pixels in a corresponding image from the multiple images.

[0020] Example 14 is a non-transitory computer-readable medium according to Examples 8-13, wherein calculating the multiple values of the expression unit based on the multiple values of the motion metric includes: identifying a first extreme value among the multiple values of the motion metric and a first corresponding image for which the first extreme value is identified; setting each of the multiple values of the expression unit associated with the first corresponding image to be equal to one; identifying a second extreme value among the multiple values of the motion metric and a second corresponding image for which the second extreme value is identified; setting each of the multiple values of the expression unit associated with the second corresponding image to be equal to zero; and setting each remaining value of the multiple values by interpolating between zero and one.

[0021] Example 15 is a system comprising: one or more processors; and a non-transitory computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving a plurality of images, each of the plurality of images containing at least a portion of a user's face; calculating a plurality of values of a motion metric based on the plurality of images, each of the plurality of values of the motion metric indicating movement of the user's face; calculating a plurality of values of an expression unit based on the plurality of values of the motion metric, each of the plurality of values of the expression unit corresponding to the extent to which the user's face is producing the expression unit; and training a machine learning model using the plurality of images and the plurality of values of the expression units by: generating training output data by the machine learning model based on the plurality of images; and modifying the machine learning model based on the plurality of values of the expression unit and the training output data.

[0022] Example 16 is the system of example 15, wherein the training output data comprises a plurality of output values for each expression unit in a set of expression units, the expression unit being a first expression unit from the set of expression units.

[0023] Example 17 is a system according to Example 16, wherein the set of expression units includes at least one of the following: raising the inner eyebrow, raising the outer eyebrow, lowering the eyebrow, raising the upper eyelid, raising the cheek, tightening the eyelid, wrinkling the nose, closing the eyes, blinking the left eye, or blinking the right eye.

[0024] Example 18 is a system according to Examples 15-17, wherein training the machine learning model using the multiple images and the multiple values of the expression units also includes: performing a comparison of the multiple values of the expression units with the multiple output values of the expression units of the training output data; and generating error data based on the comparison, wherein the machine learning model is modified based on the error data.

[0025] Example 19 is a system according to Examples 15-18, wherein the motion metric is a number of eye pixels, and wherein calculating the multiple values of the motion metric based on the multiple images includes: segmenting each image of the multiple images so that each image of the multiple images includes eye pixels and non-eye pixels; counting the number of eye pixels in each image of the multiple images; and setting each of the multiple values of the motion metric to be equal to the number of eye pixels in a corresponding image from the multiple images.

[0026] Example 20 is a system according to Examples 15-19, wherein calculating the multiple values of the expression unit based on the multiple values of the motion metric includes: identifying a first extreme value among the multiple values of the motion metric and a first corresponding image for which the first extreme value is identified; setting each of the multiple values of the expression unit associated with the first corresponding image to be equal to one; identifying a second extreme value among the multiple values of the motion metric and a second corresponding image for which the second extreme value is identified; setting each of the multiple values of the expression unit associated with the second corresponding image to be equal to zero; and setting each remaining value of the multiple values by interpolating between zero and one.

[0027] Compared to conventional techniques, many benefits are achieved through the presently disclosed approaches. For example, the embodiments described herein allow prediction of a user's expression using only a portion of the user's face, which has useful applications in head-mounted systems such as AR systems. The embodiments described herein further allow training of machine learning models to predict user expressions that can be customized to be user-specific or used by any user. For example, a machine learning model can be trained for all users first, and then the end user can further calibrate and fine-tune the training upon receiving the device, before each use of the device, and / or periodically based on the user's needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated into and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the detailed description, serve to explain the principles of the present disclosure. No attempt is made to show in greater detail the structural details of the disclosure, except as necessary for a basic understanding of the present disclosure and various ways in which it may be implemented.

[0029] Figure 1 Shown are example instructions and corresponding motion metrics that may be detected for training a machine learning model to predict user expressions.

[0030] Figure 2A and Figure 2B An example calculation of expression unit values based on motion metric values is shown.

[0031] Figure 3A and Figure 3B An example calculation of expression unit values based on motion metric values is shown.

[0032] Figure 4A and Figure 4B An example calculation of expression unit values based on motion metric values is shown.

[0033] Figure 5A An example system is shown in which a machine learning model operates in training mode.

[0034] Figure 5B An example system is shown in which a machine learning model operates in runtime mode.

[0035] Figure 6 An example implementation is shown in which the number of eye pixels in an image is used as a motion metric.

[0036] Figure 7A Show Figure 6 Example motion metric values for an example implementation of .

[0037] Figure 7B Show Figure 7A Example expression unit values for the motion metrics shown in .

[0038] Figure 8 A method for training a machine learning model to predict user expressions is shown.

[0039] Figure 9 A schematic diagram illustrating an example wearable system is shown.

[0040] Figure 10 A simplified computer system is shown. DETAILED DESCRIPTION

[0041] Predicting user expressions is useful in a variety of applications. For example, the ability to detect user expressions (and corresponding user emotional states) can allow a computing system to communicate with a user based on perceived needs, thereby enabling the computing system to provide relevant information to the user. In augmented reality (AR) or virtual reality (VR) environments, detecting user expressions can facilitate the animation of avatars and other digital characters. For example, the expressions produced by a user's avatar in the digital world may immediately react to the user's expressions in the real world.

[0042] While many previous studies have predicted a user's expression based on images of the user's entire face, prediction based on images of only a portion of the user's face, such as the user's eyes, is much more complex. For example, some facial expressions may only cause subtle changes to the eyes, while changes to other parts of the user's face, such as the user's mouth, may be more obvious. These subtle changes can be difficult to detect and difficult to associate with specific user expressions. Given that in many applications, especially AR / VR applications that use eye-tracking cameras, the camera's field of view is limited, there is a great need for robust methods to predict user expressions based on images of only a portion of the user's face.

[0043] The embodiments described herein provide systems and methods for training a machine learning model to predict user expressions. Specifically, when an input image of a user's face (e.g., the user's eyes) is provided, the machine learning model can be trained to generate a set of fractional values representing different facial movements. The different facial movements can be referred to as expression units, and the value generated by the machine learning model for each expression unit can be referred to as an expression unit value. In some cases, each expression unit value can range between zero and one, where zero corresponds to the least degree to which the user generates the expression unit and one corresponds to the most degree to which the user generates the expression unit.

[0044] In some embodiments, different expression units may be Facial Action Coding System (FACS) action units, which are a widely used classification method for facial movements. Each FACS action unit in the FACS action unit corresponds to a different contraction or relaxation of one or more muscles in the user's face. The combination of action units may help the user express a specific emotion. For example, when the user produces a cheek raise (action unit 6) and a lip corner pull (action unit 12), the user may show a "happy" emotion. As another example, when the user produces an inner brow raiser (action unit 1), a brow lowerer (action unit 4), and a lip corner press (action unit 15), the user may show a "sad" emotion.

[0045] To train a machine learning model using a series of images, a set of expression unit values generated for each image is compared to ground truth data, which may include a different set of expression unit values (or a single expression unit value) calculated based on a motion metric for the series of images. To distinguish between the two sets of expression unit values, the values generated by the machine learning model may be referred to as output values. For each image, error data may be generated by comparing the output value with the expression unit value calculated using the motion metric value for the image. The error data is then used to modify the machine learning model, for example by adjusting weights associated with the machine learning model, so as to generate more accurate output values during subsequent inference.

[0046] Figure 1 Example instructions 101 and corresponding motion metric values 102 for training a machine learning model to predict user expressions according to some embodiments of the present invention are shown. Instructions 101 may be provided to a user to instruct the user to produce one or more expression units. While the user is producing the expression units, a camera captures an image of the user's face (or a portion thereof, such as the user's eyes). The captured image is analyzed to extract motion metric values 102 for specific motion metrics associated with the user's face.

[0047] In some examples, the user wears an AR / VR headset. The headset may include a camera whose field of view includes at least a portion of the user's face, such as one or both of the user's eyes. Such a camera may be referred to as an eye-tracking camera, which is often used in AR / VR headsets. In some examples, the camera may capture an image of the user's entire face while providing instructions to the user, and the image may be cropped to reduce the image to a desired area (such as the user's eyes). Alternatively, the camera may capture an image of the desired area directly by focusing or zooming in on the user's eyes. Accordingly, embodiments of the present invention may include scenarios in which the user is or is not wearing a head-mounted device.

[0048] Although Instruction 101 Figure 1 Instructions 101 are shown as written instructions, but instructions 101 may include auditory instructions played through a speaker in the AR / VR headset or through a remote speaker, and visual instructions displayed at the AR / VR headset or on a remote display device, among other possibilities. For example, during a calibration step of the AR / VR headset, the headset may generate virtual content showing written instructions or examples of a virtual character demonstrating different expression units. The user may see these visual instructions and subsequently produce the indicated expression units.

[0049] In the example shown, the user is first provided with an instruction to "repeat expression unit 1." In this example, "expression unit 1" may correspond to raising an inner eyebrow. Then, while images of the user's face are captured, the user repeatedly produces an inner eyebrow raise. These images are analyzed to detect motion metrics 102 indicating the movement of the user's face when producing an inner eyebrow raise. The motion metrics 102 may be analyzed to identify maximum and minimum values (and their respective corresponding timestamps T max and T min ).

[0050] Timestamp T max and T min It can be used to identify images of interest and generate true values for training machine learning models. For example, the motion metric value 102 is at a relative maximum (at timestamp T max ) may be when the user fully produces the inner eyebrow lift, the motion metric value 102 is at a relative minimum value (at time stamp T min ) can be when the user produces an inner brow raise by a minimal amount, and images in between can be when the user produces a partial inner brow raise. Thus, different expression unit values can be calculated based on the motion metric value 102. For example, an expression unit value of one can be calculated for a relative maximum motion metric value, an expression unit value of zero can be calculated for a relative minimum motion metric value, and expression unit values between zero and one can be interpolated (e.g., linearly) between the maximum and minimum motion metric values.

[0051] Continuing with the example shown, the user is then provided with an instruction to "repeat expression unit 2." In this example, "expression unit 2" may correspond to lowering the eyebrows. An image of the user's face is then captured, and the user repeatedly lowers the eyebrows multiple times. The image is analyzed to detect motion metric values 102, from which the maximum and minimum values and corresponding timestamps T are identified. max and T min Compared with the inner eyebrow lift, the motion metric value 102 is at a relative minimum (at time stamp T min) may be when the user fully lowers his eyebrows, and the motion metric value 102 is at a relative maximum (at time stamp T max ) may be when the user lowers their eyebrows with the smallest amount. Thus, the expression unit value may be calculated as one for the relative minimum motion metric value, and zero for the relative maximum motion metric value.

[0052] Next, the user is provided with an instruction to "repeat expression unit 3". In this example, "expression unit 3" may correspond to tightening the eyelids. Then, while an image of the user's face is captured, the user repeatedly tightens the eyelids multiple times, and the image is analyzed to detect motion metric values 102, from which the maximum and minimum values and corresponding timestamps T are identified. max and T min Finally, the user is provided with an instruction to "repeat expression unit 4". In this example, "expression unit 4" may correspond to raising the upper eyelid. Then, while an image of the user's face is captured, the user repeatedly raises the upper eyelid multiple times, and the image is analyzed to detect motion metric values 102, from which the maximum and minimum values and corresponding timestamps T are identified. max and T min .

[0053] Figure 2A and Figure 2B 2 shows an example calculation of an expression unit value 204 based on a motion metric value 202 according to some embodiments of the present invention. Figure 2A As shown in FIG, firstly, the motion metric value 202 and its corresponding time stamp (T1, T4, T7, T 10 、T 13 、T 16 and T 19 ) relative maximum 208 and relative minimum 210. To avoid over-identification of relative extremes, a constraint can be imposed that sequential extremes have at least a certain spacing (e.g., an amount of time or a number of frames). Next, an upper threshold 212 can be set at a predetermined distance below each relative maximum 208, and a lower threshold 214 can be set at a predetermined distance above each relative minimum 210. The motion metric value 202 crosses the upper threshold 212 (T3, T5, T9, T 11 、T 15 and T 17 ) and lower threshold 214 (T2, T6, T8, T 12 、T 14 and T 18 ) timestamp can be identified.

[0054] like Figure 2BAs shown in , the expression unit value 204 can then be calculated by setting the value equal to one at the time stamp of identifying the relative maximum value 208 and / or the time stamp of the motion metric value 202 crossing the upper threshold value 212. The expression unit value 204 can be set equal to zero at the time stamp of identifying the relative minimum value 210 and / or the time stamp of the motion metric value 202 crossing the lower threshold value 214. The remaining values of the expression unit value 204 can be linearly interpolated. For example, the expression unit value 204 between T2 and T3 can be linearly interpolated between zero and one, the expression unit value 204 between T5 and T6 can be linearly interpolated between one and zero, and so on.

[0055] In some embodiments, interpolation schemes other than linear interpolation may be used. For example, a nonlinear interpolation scheme may be used, where the expression unit value is calculated based on the nearest motion metric value as shown below. i ) and E(T i ) are T i The metric motion value and expression unit value of time, then the expression unit value between T2 and T3 can be interpolated between zero and one, as defined by the following equation:

[0056]

[0057] Likewise, the expression unit values between T5 and T6 can be interpolated between zero and one, as defined by the following equation:

[0058]

[0059] Figure 3A and Figure 3B 3 shows an example calculation of an expression unit value 304 based on a motion metric value 302 according to some embodiments of the present invention. Figure 2A and Figure 2B compared to, Figure 3A and Figure 3B An expression unit in is one, where a minimum motion metric value occurs when the user fully produces the expression unit, and a maximum motion metric value occurs when the user produces the expression unit with a minimum amount.

[0060] like Figure 3A As shown in FIG, first, the motion metric value 302 and its corresponding time stamp (T1, T4, T7, T 10 、T 13 、T 16 and T 19 ) identifies the relative maximum 308 and the relative minimum 310. Figure 2ASimilarly to that described in , an upper threshold 312 may be set at a predetermined distance below each relative maximum value 308, and a lower threshold 314 may be set at a predetermined distance above each relative minimum value 310. 12 、T 14 and T 18 ) and lower threshold 314 (T3, T5, T9, T 11 、T 15 and T 17 ) timestamp can be identified.

[0061] like Figure 3B , the expression unit value 304 may then be calculated by setting a value equal to zero at the timestamp at which the relative maximum value 308 is identified and / or the timestamp at which the motion metric value 302 crosses the upper threshold 312. The expression unit value 304 may be set equal to one at the timestamp at which the relative minimum value 310 is identified and / or the timestamp at which the motion metric value 302 crosses the lower threshold 314. The remaining values of the expression unit value 304 may be linearly interpolated.

[0062] Figure 4A and Figure 4B An example calculation of an expression unit value 404 based on a motion metric value 402 is shown according to some embodiments of the present invention. Figure 4A The approach adopted in this paper is a simplified approach in which no Figure 2A and Figure 3A The threshold value described in . Similar to Figure 2A and Figure 2B , Figure 4A and Figure 4B The expression unit in is one, where the maximum motion metric value occurs when the user fully produces the expression unit, and the minimum motion metric value occurs when the user produces the expression unit with the minimum amount. Figure 4A As shown in FIG, for the motion metric value 402 and its corresponding time stamp (T1, T4, T7, T 10 、T 13 、T 16 and T 19 ) identifies a relative maximum 408 and a relative minimum 410.

[0063] like Figure 4B As shown in FIG, the expression unit value 404 can be calculated by setting the value equal to one at the time stamp of identifying the relative maximum value 408 and setting the value equal to zero at the time stamp of identifying the relative minimum value 410. The remaining value of the expression unit value 404 is calculated by setting the value equal to one at T1 and T4, T7 and T8. 10 and T 13 and T 16 Between zero and one, and between T4 and T7, T10 and T 13 and T 16 and T 19 The values are calculated by linear or nonlinear interpolation between one and zero.

[0064] Figure 5A An example system 500A is shown in which a machine learning model 550 operates in a training mode according to some embodiments of the present invention. System 500A includes an image capture device 505 configured to capture an image 506 of a user's face. Image 506 is received and processed by image processors 508A and 508B. Image processor 508A calculates a value 502 of a motion metric 510. Motion metric 510 may be constant during the training process or may vary for different expression units. Value 502 of motion metric 510 is sent from image processor 508A to image processor 508B, which calculates value 504 of expression unit 514 based on value 502 of motion metric 510.

[0065] The image 506 and the value 504 of the expression unit 514 can form training input data 518. During the training process, each image 506 can be sequentially input to the machine learning model 550 along with the expression unit value corresponding to the image from the value 504. After receiving the image, the machine learning model 550 can generate an output value 522 for each expression unit 520 in a set of N expression units 520. The output value of the expression unit that is the same as the expression unit 514 is compared with the corresponding value from the value 504 to generate error data 524. The weights associated with the machine learning model 550 are then modified (e.g., adjusted) based on the error data 524.

[0066] As an example, during a first training iteration, a first image from images 506 may be provided to a machine learning model 550, which may generate N output values 522 (one for each of the N expression units 520). In some embodiments, each of the N output values may be a fractional value between zero and one. The output value 522 of an expression unit 520 that is identical to the expression unit 514 is compared with a first value from the values 504 (representing true data) to generate error data 524, which corresponds to the first image. In some embodiments, the output values 522 of the remaining expression units 520 are also used to generate the error data 524, thereby allowing the machine learning model 550 to learn that these output values 522 should be zero. The weights associated with the machine learning model 550 are then modified based on the error data 524.

[0067] Continuing with this example, during a second training iteration following the first training iteration, a second image from image 506 can be provided to a machine learning model 550, which can generate N output values 522 (one for each of the N expression units 520). The output value 522 of the expression unit 520 that is the same as the expression unit 514 is compared with the second value from value 504 to generate error data 524 corresponding to the second image (optionally, the output values 522 of the remaining expression units 520 are also used to generate the error data 524). The weights associated with the machine learning model 550 are then modified based on the error data 524.

[0068] This process continues until all images 506 have been used in the training process. During the training process, expression units 514 may be changed as needed, resulting in different output values 522 being selected and used in the generation of error data 524. Thus, machine learning model 550 may "learn" to predict how well a user will produce each of N expression units 520 based on a single image.

[0069] Figure 5B An example system 500B is shown in which a machine learning model 550 operates in a run mode according to some embodiments of the present invention. During run time, an image capture device 505 captures an image 506 and provides the image 506 to the machine learning model 550, which generates an output value 522 for each of the expression units 520, resulting in N output values 522. Figure 5B 506, in some embodiments, multiple input images may be provided to improve the accuracy of the machine learning model 550. For example, one or more previous or subsequent images of image 506 may be provided to the machine learning model 550 along with image 506 when generating a single set of N values 522. In such embodiments, the training process may similarly utilize multiple input images during each training iteration.

[0070] Figure 6An example implementation is shown in which the number of eye pixels in an image is used as a motion metric, according to some embodiments of the present invention. In the illustrated example, a left image 602A and a right image 602B of a user's eyes are captured using an image capture device. Each of images 602 is segmented into eye pixels 606 and non-eye pixels 608 (or alternatively, non-skin pixels and skin pixels, respectively), as shown in eye segmentation 604. Eye pixels 606 can be further segmented into different regions of the eye, including the sclera, iris, and pupil. In some embodiments, an additional machine learning model can be used to generate eye segmentation 604. Such a machine learning model can be trained using labeled images prepared by a user, in which the user manually identifies eye pixels 606 and non-eye pixels 608, as well as different regions of the eye.

[0071] Figure 7A FIG. 1 shows a method according to some embodiments of the present invention in which the number of eye pixels in an image is used as a motion metric. Figure 6 Example motion metric values for an example implementation of . In the example shown, data for the left eye and the right eye are superimposed. The curve shows the number of eye pixels (or non-skin pixels) over time. In some embodiments, motion metric values corresponding to "strong expressions" (e.g., the user is generating expression units to the greatest extent) can be automatically or manually identified. Automatic identification can be achieved by identifying reference Figures 1 to 4B The motion metric values corresponding to "neutral expressions" (e.g., minimal user-generated expression units) can be automatically or manually identified.

[0072] Figure 7B Show Figure 7A Example expression unit values for the motion metric values shown in . The expression unit values are calculated by setting a value equal to one for frames (images) for which a strong expression is identified (and optionally frames for which the motion metric value is within a threshold distance), and setting a value equal to zero for frames for which a neutral expression is identified (and optionally frames for which the motion metric value is within a threshold distance). As shown with reference to FIG. Figure 4B As described above, the remaining expression unit values are linearly or nonlinearly interpolated between zero and one.

[0073] Figure 8 A method 800 is shown for training a machine learning model (e.g., machine learning model 550) to predict user expressions according to some embodiments of the present invention. During execution of method 800, one or more steps of method 800 may be omitted, and the steps of method 800 do not necessarily need to be performed in the order shown. One or more steps of method 800 may be performed or facilitated by one or more processors.

[0074] At step 802, a plurality of images (e.g., images 506, 602) are received. The plurality of images may be received from an image capture device (e.g., image capture device 505), which may capture the plurality of images and send the plurality of images to a processing module. One or more of the plurality of images may be grayscale images, multi-channel images (e.g., RGB images), and other possibilities. In some embodiments, the image capture device may be an eye-tracking camera mounted to a wearable device. Each of the plurality of images may include at least a portion of a user's face. For example, each of the plurality of images may include an eye of the user.

[0075] At step 804, a plurality of values (e.g., values 102, 202, 302, 402, 502) of a motion metric (e.g., motion metric 510) are calculated based on the plurality of images. The motion metric can be some metric that indicates motion of the user's face (or that, through analysis thereof, can indicate motion of the user's face). For example, the motion metric can be the number of eye pixels in the image, the number of non-eye pixels in the image, the distance between the top and bottom of the eye, the distance between the left and right sides of the eye, the position of a point along the eye within the image, the gradient of the image, and other possibilities.

[0076] For embodiments in which the motion metric is a number of eye pixels, calculating the plurality of values of the motion metric may include segmenting each of the plurality of images such that each of the plurality of images includes eye pixels (e.g., eye pixels 606) and non-eye pixels (e.g., non-eye pixels 608), counting the number of eye pixels in each of the plurality of images, and setting each of the plurality of values of the motion metric equal to the number of eye pixels in a corresponding image from the plurality of images. Segmenting the image from the plurality of images may result in an eye segmentation (e.g., eye segmentation 604).

[0077] At step 806, multiple values (e.g., values 204, 304, 404, 504) of an expression unit (e.g., expression unit 514) are calculated based on the multiple values of the motion metric. Each of the multiple values of the expression unit may correspond to a degree to which the user (e.g., the user's face) produced the expression unit. In some embodiments, a larger value may correspond to a greater degree to which the user produced the expression unit.

[0078] In some embodiments, calculating the multiple values of the expression unit may include identifying extreme values (maximum and / or minimum values) among the multiple values of the motion metric. In one example, a first extreme value (e.g., a maximum value) among the multiple values of the motion metric is identified and a first corresponding image for the first extreme value is identified. Each of the multiple values of the expression unit associated with the first corresponding image can be set equal to one. In addition, a second extreme value (e.g., a minimum value) among the multiple values of the motion metric is identified and a second corresponding image for the second extreme value is identified. Each of the multiple values of the expression unit associated with the second corresponding image can be set equal to zero. In addition, each remaining value of the multiple values can be set to a value between zero and one by interpolation.

[0079] At step 808, a machine learning model is trained using the plurality of images and the plurality of values of the expression unit. In some embodiments, step 808 includes one or both of steps 810 and 812.

[0080] At step 810, training output data (e.g., training output data 526) is generated based on the plurality of images. The training output data may include a plurality of output values (e.g., output value 522) for each expression unit in a set of expression units (e.g., expression unit 520). The expression unit may be one of a set of expression units. The set of expression units may include one or more of: inner eyebrow raise, outer eyebrow raise, lower eyebrow, upper eyelid raise, cheek raise, eyelid tuck, nose wrinkle, eye closed, left eye blink, or right eye blink. The set of expression units may be FACS action units, such that the expression unit may be one of the FACS action units.

[0081] In some embodiments, training a machine learning model using multiple images and multiple values of an expression unit further includes performing a comparison of the multiple values of the expression unit with the multiple output values of the expression unit. In some embodiments, error data (e.g., error data 524) can be generated based on the comparison. For example, the error data can be generated by subtracting the multiple values of the expression unit from the multiple output values of the expression unit (or vice versa). The error data can be set to be equal to the magnitude of the difference, the sum of the magnitudes of the differences, the sum of the squares of the differences, and other possibilities. In general, the error data can indicate the difference between the multiple values of the expression unit and the multiple output values of the expression unit.

[0082] At step 812, the machine learning model is modified based on the plurality of values of the expression units and the training output data. Modifying the machine learning model may include adjusting one or more parameters associated with the machine learning model (e.g., weights and / or biases). For example, the machine learning model may be an artificial neural network having a plurality of adjustable parameters for calculating a set of output values for a set of expression units based on an input image.

[0083] In some embodiments, the machine learning model can be modified based on the error data. In some embodiments, the degree of adjustment of parameters associated with the machine learning model can be related to (e.g., proportional to) the magnitude of the error data, such that a larger difference between the multiple values of the expression unit and the multiple output values of the expression unit will result in a larger modification to the machine learning model. In some embodiments, the machine learning model can be modified for each of a plurality of training iterations. For example, each training iteration can include training the machine learning model using a single input image from a plurality of images and a corresponding value of the expression unit from a plurality of values of the expression unit.

[0084] Figure 9 A schematic diagram of an example wearable system 900 that can be used in one or more of the above embodiments according to some embodiments of the present invention is shown. Wearable system 900 may include a wearable device 901 and at least one remote device 903 that is remote from wearable device 901 (e.g., separate hardware but communicatively coupled). While wearable device 901 is worn by a user (typically as a head-mounted device), remote device 903 can be handheld by the user (e.g., as a handheld controller) or mounted in various configurations, such as fixedly attached to a frame, fixedly attached to a helmet or hat worn by the user, embedded in headphones, or otherwise removably attached to the user (e.g., in a backpack configuration, a belt-coupled configuration, etc.).

[0085] Wearable device 901 may include a left eyepiece 902A and a left lens assembly 905A arranged in a side-by-side configuration and forming a left optical stack. Left lens assembly 905A may include an accommodation lens located on the user side of the left optical stack and a compensation lens located on the world side of the left optical stack. Similarly, wearable device 901 may include a right eyepiece 902B and a right lens assembly 905B arranged in a side-by-side configuration and forming a right optical stack. Right lens assembly 905B may include an accommodation lens located on the user side of the right optical stack and a compensation lens located on the world side of the right optical stack.

[0086] In some embodiments, the wearable device 901 includes one or more sensors, including but not limited to: a left front-facing world camera 906A directly attached to or near the left eyepiece 902A, a right front-facing world camera 906B directly attached to or near the right eyepiece 902B, a left-facing world camera 906C directly attached to or near the left eyepiece 902A, a right-facing world camera 906D directly attached to or near the right eyepiece 902B, a left-eye tracking camera 926A pointed toward the left eye, a right-eye tracking camera 926B pointed toward the right eye, and a depth sensor 928 attached between the eyepieces 902. The wearable device 901 may include one or more image projection devices, such as a left projector 914A optically connected to the left eyepiece 902A and a right projector 914B optically connected to the right eyepiece 902B.

[0087] The wearable system 900 may include a processing module 950 for collecting, processing, and / or controlling data within the system. Components of the processing module 950 may be distributed between the wearable device 901 and the remote device 903. For example, the processing module 950 may include a local processing module 952 on the wearable portion of the wearable system 900 and a remote processing module 956 that is physically separate from the local processing module 952 and communicatively connected to the local processing module 952. Each of the local processing module 952 and the remote processing module 956 may include one or more processing units (e.g., a central processing unit (CPU), a graphics processing unit (GPU), etc.) and one or more storage devices, such as non-volatile memory (e.g., flash memory).

[0088] The processing module 950 can collect data captured by various sensors of the wearable system 900, such as the camera 906, the eye-tracking camera 926, the depth sensor 928, the remote sensor 930, the ambient light sensor, the microphone, the inertial measurement unit (IMU), the accelerometer, the compass, the global navigation satellite system (GNSS) unit, the radio device, and / or the gyroscope. For example, the processing module 950 can receive images 920 from the camera 906. Specifically, the processing module 950 can receive a left front image 920A from the left front-facing world camera 906A, a right front image 920B from the right front-facing world camera 906B, a left side image 920C from the left-facing world camera 906C, and a right side image 920D from the right-facing world camera 906D. In some embodiments, the images 920 can include a single image, a pair of images, a video containing an image stream, a video containing a pair of image streams, and the like. Images 920 may be generated and sent to processing module 950 periodically when wearable system 900 is powered on, or may be generated in response to instructions sent by processing module 950 to one or more cameras.

[0089] Cameras 906 can be configured in different positions and orientations along the outer surface of wearable device 901 to capture images around the user. In some cases, cameras 906A and 906B can be positioned to capture images that substantially overlap the FOVs of the user's left and right eyes, respectively. Thus, camera 906 can be placed close to the user's eyes, but not so close as to obscure the user's FOV. Alternatively or in addition, cameras 906A and 906B can be positioned to align with the incoupling locations of virtual image light 922A and 922B, respectively. Cameras 906C and 906D can be positioned to capture images to one side of the user, for example, within or outside the user's peripheral field of view. Images 920C and 920D captured using cameras 906C and 906D do not necessarily overlap with images 920A and 920B captured using cameras 906A and 906B.

[0090] In some embodiments, the processing module 950 may receive ambient light information from an ambient light sensor. The ambient light information may indicate a brightness value or a range of spatially resolved brightness values. The depth sensor 928 may capture a depth image 932 in the forward-facing direction of the wearable device 901. Each value of the depth image 932 may correspond to the distance between the depth sensor 928 and the nearest detected object in a particular direction. As another example, the processing module 950 may receive eye tracking data 934 from the eye tracking camera 926, which may include images of the left eye and the right eye. As another example, the processing module 950 may receive projected image brightness values from one or both of the projectors 914. The remote sensor 930 located within the remote device 903 may include any of the above-mentioned sensors with similar functionality.

[0091] Virtual content is delivered to the user of the wearable system 900 using the projector 914 and the eyepiece 902, as well as other components in the optical stack. For example, the eyepieces 902A, 902B may include transparent or translucent waveguides configured to guide and decouple light generated by the projectors 914A, 914B, respectively. Specifically, the processing module 950 may cause the left projector 914A to output left virtual image light 922A onto the left eyepiece 902A, and may cause the right projector 914B to output right virtual image light 922B onto the right eyepiece 902B. In some embodiments, the projector 914 may include a microelectromechanical system (MEMS) spatial light modulator (SLM) scanning device. In some embodiments, each of the eyepieces 902A, 902B may include multiple waveguides corresponding to different colors. In some embodiments, the lens assemblies 905A, 905B may be coupled to and / or integrated with the eyepieces 902A, 902B. For example, lens assemblies 905A, 905B may be incorporated into a multi-layer eyepiece and may form one or more layers that make up one of the eyepieces 902A, 902B.

[0092] Figure 10 1 shows a simplified computer system 1000 according to the described embodiment. Figure 10 The computer system 1000 shown in FIG. 1 may be incorporated into the devices described herein. Figure 10 A schematic diagram of one embodiment of a computer system 1000 is provided, which can perform some or all of the steps of the methods provided by various embodiments. It should be noted that Figure 10 It is meant only to provide a generalized description of the various components, any or all of which may be used as appropriate. Figure 10 It broadly illustrates how individual system elements may be implemented in a relatively discrete or relatively more integrated manner.

[0093] The illustrated computer system 1000 includes hardware elements that may be electrically coupled via a bus 1005, or may communicate in other ways as appropriate. The hardware elements may include one or more processors 1010, including but not limited to one or more general-purpose processors and / or one or more specialized processors, such as digital signal processing chips, graphics acceleration processors, etc.; one or more input devices 1015, which may include but are not limited to a mouse, keyboard, camera, etc.; and one or more output devices 1020, which may include but are not limited to a display device, printer, etc.

[0094] The computer system 1000 may further include and / or communicate with one or more non-transitory storage devices 1025, which may include, but are not limited to, local and / or network accessible storage, and / or may include, but are not limited to, disk drives, drive arrays, optical storage devices, solid-state storage devices, such as random access memory ("RAM") and / or read-only memory ("ROM"), which may be programmable, flash-updatable, etc. Such storage devices may be configured to implement any suitable data storage, including, but not limited to, various file systems, database structures, etc.

[0095] The computer system 1000 may also include a communication subsystem 1019, which may include but is not limited to a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a Bluetooth TMDevices, 802.11 devices, WiFi devices, WiMax devices, chipsets for cellular communication facilities, etc. The communication subsystem 1019 may include one or more input and / or output communication interfaces to allow data to be exchanged with a network as described below (for example), other computer systems, televisions, and / or any other devices described herein. Depending on the desired functionality and / or other implementation considerations, a portable electronic device or similar device may communicate images and / or other information via the communication subsystem 1019. In other embodiments, a portable electronic device (e.g., a first electronic device) may be incorporated into the computer system 1000, for example, as an electronic device for the input device 1015. In some embodiments, the computer system 1000 will further include a working memory 1035, which may include a RAM or ROM device, as described above.

[0096] The computer system 1000 may also include software elements, shown as currently located within working memory 1035, including an operating system 1040, device drivers, executable libraries, and / or other code, such as one or more application programs 1045, which may include computer programs provided by various embodiments and / or may be designed to implement methods and / or configure systems provided by other embodiments, as described herein. By way of example only, one or more processes regarding the methods discussed above may be implemented as code and / or instructions executable by a computer and / or a processor within a computer; in one aspect, such code and / or instructions may then be used to configure and / or adjust a general-purpose computer or other device to perform one or more operations according to the described methods.

[0097] A set of these instructions and / or codes may be stored on a non-transitory computer-readable storage medium, such as the storage device 1025 described above. In some cases, the storage medium may be incorporated into a computer system, such as the computer system 1000. In other embodiments, the storage medium may be separate from the computer system, such as a removable medium, such as an optical disc, and / or provided in an installation package so that the storage medium can be used to program, configure, and / or adjust a general-purpose computer using the instructions / codes stored thereon. These instructions may be in the form of executable code that is executed by the computer system 1000, and / or may be in the form of source and / or installable code that is compiled and / or installed on the computer system 1000 using, for example, various general-purpose compilers, installers, compression / decompression utilities, and the like and then in the form of executable code.

[0098] It will be apparent to those skilled in the art that substantial variations may be made depending on specific requirements. For example, customized hardware may be used, and / or specific elements may be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connections to other computing devices (such as network input / output devices) may also be employed.

[0099] As described above, in one aspect, some embodiments may employ a computer system (such as computer system 1000) to perform methods according to various embodiments of the technology. According to one set of embodiments, some or all of the processes of such methods are performed by computer system 1000 in response to processor 1010 executing one or more sequences of one or more instructions, which may be incorporated into operating system 1040 and / or other code contained in working memory 1035, such as application program 1045. Such instructions may be read into working memory 1035 from another computer-readable medium (such as one or more storage devices 1025). By way of example only, execution of the sequences of instructions contained in working memory 1035 may cause processor 1010 to perform one or more processes of the methods described herein. Additionally or alternatively, portions of the methods described herein may be performed by dedicated hardware.

[0100] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any medium that participates in providing data to cause a machine to operate in a specific manner. In embodiments implemented using computer system 1000, various computer-readable media may be involved in providing instructions / code to processor 1010 for execution and / or may be used to store and / or carry such instructions / code. In many implementations, computer-readable media are physical and / or tangible storage media. Such media can take the form of non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as storage device 1025. Volatile media include, but are not limited to, dynamic memory, such as working memory 1035.

[0101] Common forms of physical and / or tangible computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape or any other magnetic medium, CD-ROMs, any other optical medium, punched cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.

[0102] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processor 1010 for execution. By way of example only, the instructions may initially be carried on a magnetic disk and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions as signals over a transmission medium for computer system 1000 to receive and / or execute.

[0103] The communication subsystem 1019 and / or its components will typically receive the signal, and the bus 1005 may then carry the signal and / or the data, instructions, etc. carried by the signal to the working memory 1035, from which the processor 1010 retrieves and executes the instructions. The instructions received by the working memory 1035 may optionally be stored on the non-transitory storage device 1025 before or after execution by the processor 1010.

[0104] The methods, systems, and devices discussed above are examples. Various configurations may omit, substitute, or add various processes or components as appropriate. For example, in alternative configurations, the methods may be performed in an order different from that described, and / or various stages may be added, deleted, and / or combined. Furthermore, features described with respect to certain configurations may be combined in various other configurations. Different aspects and elements of the configurations may be combined in similar manners. Furthermore, technology is evolving, and therefore, many of the elements are examples and do not limit the scope of the disclosure or claims.

[0105] Specific details are given in the description to provide a comprehensive understanding of the example configurations, including implementations. However, the configurations can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been presented without unnecessary detail to avoid obscuring the configurations. This description provides only example configurations and does not limit the scope, applicability, or configurations of the claims. On the contrary, the description of the above configurations will provide those skilled in the art with an enabling description for implementing the described technology. Various changes may be made to the functions and arrangements of the elements without departing from the spirit or scope of the disclosure.

[0106] In addition, a configuration can be described as a process, which is described as a schematic flowchart or block diagram. Although each configuration can describe the operation as a sequential process, many operations can be performed in parallel or concurrently. In addition, the order of the operations can be rearranged. The process may have additional steps not included in the figure. In addition, examples of methods can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments that perform the necessary tasks can be stored in a non-transitory computer-readable medium (such as a storage medium). The processor can perform the described tasks.

[0107] Having described several example configurations, various modifications, alternative structures, and equivalents may be used without departing from the spirit of the present disclosure. For example, the above elements may be components of a larger system, in which other rules may take precedence over or otherwise modify the application of the technology. In addition, some steps may be taken before, during, or after considering the above elements. Therefore, the above description does not restrict the scope of the claims.

[0108] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "a user" includes a plurality of such users and reference to "a processor" includes reference to one or more processors and equivalents thereof known to those skilled in the art, and so forth.

[0109] Furthermore, the words “comprises,” “comprising,” “containing,” “included,” “included,” and “including,” when used in this specification and the following claims, are intended to specify the presence of stated features, integers, components or steps, but they do not preclude the presence or addition of one or more other features, integers, components, steps, acts or groups.

[0110] It should also be understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or alterations based on these examples and embodiments will be suggested to those skilled in the art and will be included within the spirit and purview of this application and the scope of the appended claims.

Claims

1. A method for training a machine learning model to predict user expressions, the method comprising: receiving a plurality of images, each image of the plurality of images containing at least a portion of a user's face; calculating, based on the plurality of images, a plurality of values of a motion metric, each value of the plurality of values of the motion metric indicating motion of the user's face; calculating a plurality of values of an expression unit based on the plurality of values of the motion metric, each value of the plurality of values of the expression unit corresponding to a degree to which the user's face is producing the expression unit; as well as Using the plurality of images and the plurality of values of the expression units, the machine learning model is trained by: generating, by the machine learning model, training output data based on the plurality of images; and The machine learning model is modified based on the multiple values of the expression unit and the training output data.

2. The method according to claim 1, wherein The training output data includes a plurality of output values for each expression unit in a set of expression units, the expression unit being a first expression unit from the set of expression units.

3. The method according to claim 2, wherein: The set of expression units includes at least one of the following: raising inner eyebrows, raising outer eyebrows, lowering eyebrows, raising upper eyelids, raising cheeks, tightening eyelids, wrinkling nose, closing eyes, blinking left eyes, or blinking right eyes.

4. The method according to claim 1, wherein Training the machine learning model using the plurality of images and the plurality of values of the expression units further comprises: performing a comparison of the plurality of values of the expression unit with a plurality of output values of the expression unit of the training output data; and Error data is generated based on the comparison, wherein the machine learning model is modified based on the error data.

5. The method according to claim 1, wherein The machine learning model is an artificial neural network with a set of adjustable parameters.

6. The method according to claim 1, wherein The motion metric is a number of eye pixels, and wherein calculating the plurality of values of the motion metric based on the plurality of images comprises: segmenting each of the plurality of images so that each of the plurality of images includes eye pixels and non-eye pixels; counting the number of eye pixels in each of the plurality of images; and Each of the plurality of values of the motion metric is set equal to the number of eye pixels in a corresponding image from the plurality of images.

7. The method according to claim 1, wherein Calculating the plurality of values of the expression unit based on the plurality of values of the motion metric includes: identifying a first extreme value among the plurality of values of the motion metric and a first corresponding image for which the first extreme value is identified; setting each of the plurality of values of the expression element associated with the first corresponding image equal to one; identifying a second extreme value among the plurality of values of the motion metric and a second corresponding image for which the second extreme value is identified; setting each of the plurality of values of the expression element associated with the second corresponding image equal to zero; and Each remaining value of the plurality of values is set by interpolating between zero and one.

8. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: receiving a plurality of images, each image of the plurality of images containing at least a portion of a user's face; calculating, based on the plurality of images, a plurality of values of a motion metric, each value of the plurality of values of the motion metric indicating motion of the user's face; calculating a plurality of values of an expression unit based on the plurality of values of the motion metric, each value of the plurality of values of the expression unit corresponding to a degree to which the user's face is producing the expression unit; as well as Using the plurality of images and the plurality of values of the expression units, a machine learning model is trained by: generating, by the machine learning model, training output data based on the plurality of images; and The machine learning model is modified based on the multiple values of the expression unit and the training output data.

9. The non-transitory computer readable medium of claim 8, wherein: The training output data includes a plurality of output values for each expression unit in a set of expression units, the expression unit being a first expression unit from the set of expression units.

10. The non-transitory computer readable medium of claim 9, wherein: The set of expression units includes at least one of the following: raising inner eyebrows, raising outer eyebrows, lowering eyebrows, raising upper eyelids, raising cheeks, tightening eyelids, wrinkling nose, closing eyes, blinking left eyes, or blinking right eyes.

11. The non-transitory computer readable medium of claim 8, wherein: Training the machine learning model using the plurality of images and the plurality of values of the expression units further comprises: performing a comparison of the plurality of values of the expression unit with a plurality of output values of the expression unit of the training output data; and Error data is generated based on the comparison, wherein the machine learning model is modified based on the error data.

12. The non-transitory computer readable medium of claim 8, wherein: The machine learning model is an artificial neural network with a set of adjustable parameters.

13. The non-transitory computer readable medium of claim 8, wherein: The motion metric is a number of eye pixels, and wherein calculating the plurality of values of the motion metric based on the plurality of images comprises: segmenting each of the plurality of images so that each of the plurality of images includes eye pixels and non-eye pixels; counting the number of eye pixels in each of the plurality of images; and Each of the plurality of values of the motion metric is set equal to the number of eye pixels in a corresponding image from the plurality of images.

14. The non-transitory computer readable medium of claim 8, wherein: Calculating the plurality of values of the expression unit based on the plurality of values of the motion metric includes: identifying a first extreme value among the plurality of values of the motion metric and a first corresponding image for which the first extreme value is identified; setting each of the plurality of values of the expression element associated with the first corresponding image equal to one; identifying a second extreme value among the plurality of values of the motion metric and a second corresponding image for which the second extreme value is identified; setting each of the plurality of values of the expression element associated with the second corresponding image equal to zero; and Each remaining value of the plurality of values is set by interpolating between zero and one.

15. A system for training a machine learning model to predict user expressions, the system comprising: one or more processors; as well as A non-transitory computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving a plurality of images, each image of the plurality of images containing at least a portion of a user's face; calculating, based on the plurality of images, a plurality of values of a motion metric, each value of the plurality of values of the motion metric indicating motion of the user's face; calculating a plurality of values of an expression unit based on the plurality of values of the motion metric, each value of the plurality of values of the expression unit corresponding to a degree to which the user's face is producing the expression unit; as well as Using the plurality of images and the plurality of values of the expression units, the machine learning model is trained by: generating, by the machine learning model, training output data based on the plurality of images; and The machine learning model is modified based on the multiple values of the expression unit and the training output data.

16. The system according to claim 15, wherein: The training output data includes a plurality of output values for each expression unit in a set of expression units, the expression unit being a first expression unit from the set of expression units.

17. The system according to claim 16, wherein: The set of expression units includes at least one of the following: raising inner eyebrows, raising outer eyebrows, lowering eyebrows, raising upper eyelids, raising cheeks, tightening eyelids, wrinkling nose, closing eyes, blinking left eyes, or blinking right eyes.

18. The system according to claim 15, wherein: Training the machine learning model using the plurality of images and the plurality of values of the expression units further comprises: performing a comparison of the plurality of values of the expression unit with a plurality of output values of the expression unit of the training output data; and Error data is generated based on the comparison, wherein the machine learning model is modified based on the error data.

19. The system of claim 15, wherein: The motion metric is a number of eye pixels, and wherein calculating the plurality of values of the motion metric based on the plurality of images comprises: segmenting each of the plurality of images so that each of the plurality of images includes eye pixels and non-eye pixels; counting the number of eye pixels in each of the plurality of images; and Each of the plurality of values of the motion metric is set equal to the number of eye pixels in a corresponding image from the plurality of images.

20. The system of claim 15, wherein: Calculating the plurality of values of the expression unit based on the plurality of values of the motion metric includes: identifying a first extreme value among the plurality of values of the motion metric and a first corresponding image for which the first extreme value is identified; setting each of the plurality of values of the expression element associated with the first corresponding image equal to one; identifying a second extreme value among the plurality of values of the motion metric and a second corresponding image for which the second extreme value is identified; setting each of the plurality of values of the expression element associated with the second corresponding image equal to zero; and Each remaining value of the plurality of values is set by interpolating between zero and one.

Citation Information

Patent Citations

  • Tracking a range of body movement based on 3D captured image streams of a user

    CN101238981A

  • Method, device, terminal and readable storage medium for predicting stock price based on microexpression

    CN109241873A