Information processing method and device, head-mounted equipment and medium

By using eyebrow and eye tracking cameras in AR headsets, combined with expression detection networks, extracting facial feature information and generating anthropomorphic images, the problem of device movement and lighting changes affecting facial expression capture is solved, and the user experience and display effect is improved.

CN120183008APending Publication Date: 2025-06-20BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311760722.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

During use, due to movement and lighting changes, it is difficult to fix the relative position of the user's eyes, affecting the capture and display of facial expressions.

Method used

Eyebrow tracking camera and eye tracking camera are used to obtain facial images, combine expression detection networks to extract eyebrow and eye feature information, and generate anthropomorphic images to simulate user expressions.

Benefits of technology

It improves the robustness of user eye tracking, and the generated anthropomorphic images truly simulate user expressions, enhances the fun and vividness of the display, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183008A_ABST
    Figure CN120183008A_ABST
Patent Text Reader

Abstract

The invention relates to an information processing method and device, head-mounted equipment and a medium, the information processing method is applied to the head-mounted equipment, the method comprises the steps that at least one frame of face image is acquired, the face image comprises an eyebrow image and an eye image, the eyebrow image is acquired by an eyebrow tracking camera, and the eye image is acquired by an eye tracking camera; determining expression information of the user based on the facial image and an expression detection network; and obtaining and displaying an anthropomorphic image, wherein the expression of the anthropomorphic image is related to the expression information. The head-mounted device is provided with the eyebrow tracking camera and the eye tracking camera, it is guaranteed that even if the position and distance between the head-mounted device and the face of the user are changed, the complete face image of the user can be obtained, the eye tracking robustness of the user is effectively improved, and the user experience is improved. The anthropomorphic image generated based on the expression information realizes the real simulation of the expression of the user, enhances the interestingness and vividness of display, and improves the use experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of electronic devices, and particularly to an information processing method, apparatus, head-mounted device, and medium. Background Art

[0002] With the rapid development of technology, head-mounted devices are widely used in people's daily lives. Head-mounted devices use a camera array or a depth camera to capture subtle movements and morphological changes of the face, and convert them into digital data, extracting key facial features from the collected data for subsequent analysis and constructing facial expressions. AR (Augmented Reality) head-mounted devices are lightweight and foldable, similar to traditional glasses, with multi-point force support. During use, the AR head-mounted device will move, and its relative position to the user's eyes cannot be fixed. Moreover, the AR head-mounted device has light transmissibility, and both light changes and changes in the wearing position will affect the capture of facial expressions by the AR head-mounted device. Summary of the Invention

[0003] To overcome the problems in the related art, the present disclosure provides an information processing method, apparatus, head-mounted device, and medium.

[0004] According to a first aspect of an embodiment of the present disclosure, an information processing method is provided, which is applied to a head-mounted device. The head-mounted device includes an eyebrow tracking camera and an eye tracking camera. The eyebrow tracking camera is used to collect eyebrow images, and the eye tracking camera is used to collect eye images. The information processing method includes:

[0005] Obtain at least one frame of facial image, where the facial image includes the eyebrow image and the eye image;

[0006] Based on the facial image and an expression detection network, determine the expression information of the user;

[0007] Obtain and display an anthropomorphic image, where the expression of the anthropomorphic image is related to the expression information.

[0008] In some exemplary embodiments of the present disclosure, the method further includes:

[0009] Based on the eyebrow image and a pre-stored eyebrow feature extraction network, obtain first eyebrow feature information;

[0010] Based on the eye image and a pre-stored eye feature extraction network, obtain first eye feature information;

[0011] The determining the expression information of the user based on the facial image and the expression detection network includes:

[0012] Determine the expression information of the user based on the expression detection network, the first eyebrow feature information, and the first eye feature information.

[0013] In some exemplary embodiments of the present disclosure, the determining the expression information of the user based on the expression detection network, the first eyebrow feature information, and the first eye feature information includes:

[0014] Input the first eyebrow feature information and the first eye feature information into the expression detection network, and the expression detection network outputs a first weight value;

[0015] Determine the expression information based on the first weight value.

[0016] In some exemplary embodiments of the present disclosure, the information processing method further includes:

[0017] Optimize the first eyebrow feature information and the first eye feature information to obtain second eyebrow feature information and second eye feature information;

[0018] Determine the expression information of the user based on the expression detection network, the second eyebrow feature information, and the second eye feature information.

[0019] In some exemplary embodiments of the present disclosure, the optimizing the first eyebrow feature information and the first eye feature information to obtain second eyebrow feature information and second eye feature information includes:

[0020] Based on a moving window, calculate a first moving average of the first eyebrow feature information of the facial image and a second moving average of the first eye feature information. The moving window slides in the first eyebrow feature information and the first eye feature information respectively, and the number of frames of the facial image is greater than or equal to the length of the moving window;

[0021] Obtain the second eyebrow feature information and the second eye feature information based on the first moving average and the second moving average.

[0022] In some exemplary embodiments of the present disclosure, the determining the expression information of the user based on the expression detection network, the second eyebrow feature information, and the second eye feature information includes:

[0023] Input the second eyebrow feature information and the second eye feature information into the expression detection network, and the expression detection network outputs a second weight value;

[0024] Determine the expression information based on the second weight value.

[0025] In some exemplary embodiments of the present disclosure, the facial image includes an initial facial image and a current facial image, and the information processing method further includes:

[0026] Using a first weight value of the initial facial image to correct a first weight value or a second weight value of the current facial image, and obtaining a corrected third weight value of the current facial image;

[0027] Based on the third weight value, determining the expression information.

[0028] In some exemplary embodiments of the present disclosure, the first eyebrow feature information and the second eyebrow feature information include coordinate values of a plurality of eyebrow feature points, and the first eye feature information and the second eye feature information include coordinate values of a plurality of eye feature points.

[0029] In some exemplary embodiments of the present disclosure, the anthropomorphic image is generated by an output end based on the expression information.

[0030] According to a second aspect of the embodiments of the present disclosure, there is provided an information processing device applied to a head-mounted device. The head-mounted device includes an eyebrow tracking camera and an eye tracking camera. The eyebrow tracking camera is configured to collect eyebrow images, and the eye tracking camera is configured to collect eye images. The information processing device includes:

[0031] An acquisition module, configured to acquire at least one frame of facial image, where the facial image includes the eyebrow image and the eye image;

[0032] A processing module, configured to determine user's expression information based on the facial image and an expression detection network;

[0033] A display module, configured to acquire and display an anthropomorphic image, where an expression of the anthropomorphic image is related to the expression information.

[0034] According to a third aspect of the embodiments of the present disclosure, there is provided a head-mounted device, including:

[0035] A fill light; an eyebrow tracking camera; an eye tracking camera;

[0036] A processor;

[0037] A memory for storing executable instructions executable by the processor;

[0038] Wherein, the processor is configured to execute the executable instructions in the memory to implement the information processing method provided in the first aspect of the present disclosure.

[0039] In some exemplary embodiments of the present disclosure, the head-mounted device has a frame. The eye tracking camera includes an upper eye tracking camera and a lower eye tracking camera. The eyebrow tracking camera is disposed on the upper edge of the frame, the upper eye tracking camera is disposed on the side edge of the frame, and the lower eye tracking camera is disposed on the lower edge of the frame.

[0040] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium having executable instructions stored thereon. When the executable instructions are executed by a processor, the information processing method provided in the first aspect of the present disclosure is implemented.

[0041] Adopting the above method of the present disclosure has the following beneficial effects: The head-mounted device is provided with an eyebrow tracking camera and an eye tracking camera, ensuring that even if the position and distance of the head-mounted device change, a complete facial image of the user can be obtained, effectively improving the robustness of eye tracking of the user. The anthropomorphic image generated based on the expression information realizes a real simulation of the user's expression, enhances the interest and vividness of the display, and improves the user experience.

[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0044] Figure 1 is a flowchart of an information processing method shown according to an exemplary embodiment.

[0045] Figure 2 is a flowchart of an information processing method shown according to an exemplary embodiment.

[0046] Figure 3 is a flowchart of an information processing method shown according to an exemplary embodiment.

[0047] Figure 4 is a schematic diagram of a head-mounted device shown according to an exemplary embodiment.

[0048] Figure 5a is a schematic diagram of eyebrow feature points shown according to an exemplary embodiment.

[0049] Figure 5b is a schematic diagram of eye feature points shown according to an exemplary embodiment.

[0050] Figure 6 is a block diagram of an information processing device shown according to an exemplary embodiment.

[0051] Figure 7 It is a block diagram of a head-mounted device shown according to an exemplary embodiment. Detailed implementation

[0052] Here, the exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0053] AR (Augmented Reality) and VR (Virtual Reality) are two different virtual reality technologies. AR is a technology that superimposes virtual objects on the real world, while VR is a technology that provides users with an immersive experience by simulating a virtual environment. Both VR head-mounted devices and AR head-mounted devices can use cameras to capture subtle movements and morphological changes of the face, convert them into digital data, extract key facial features from the collected data, and reproduce facial expression movements.

[0054] Since VR head-mounted devices are relatively large in size and need to be fixed using straps, users adjust the tightness of the straps according to their head circumference to make the VR head-mounted device fit closely to the face. During use, the relative position between the VR head-mounted device and the user's eyes is fixed, which will not affect the capture of facial expressions by the VR head-mounted device. However, because the VR head-mounted device fits closely to the face, the VR head-mounted device cannot use the camera to collect the eyebrow information of the user's face. On the other hand, AR head-mounted devices are similar to traditional glasses and do not have straps for fixation. During use, the AR head-mounted device may move, and the relative position with the user's eyes cannot be fixed. In addition, there are differences in the height of each user's eyes. Some users have high eye positions, while some users have low eye positions, and the specific wearing positions are also different. Moreover, AR head-mounted devices have light transmittance, and both light changes and changes in the wearing position will affect the capture of facial expressions by the AR head-mounted device.

[0055] To solve the above problems, the present disclosure provides an information processing method applied to a head-mounted device. The method uses the eyebrow tracking camera and eye tracking camera of the head-mounted device to obtain at least one frame of facial image. The facial image includes an eyebrow image and an eye image. Therefore, the expression detection network can determine the user's expression information based on the eyebrow feature information and eye feature information in the facial image, obtain and display an anthropomorphic image related to the expression information. The anthropomorphic image has the user's personal characteristics, can truly simulate the user's expression, enrich the display content of the head-mounted device, enhance the interest and vividness of the display, and improve the user's experience.

[0056] Exemplary embodiments of the present disclosure provide an information processing method, which is applied to a head-mounted device, specifically an AR head-mounted device. As Figure 4 shown, the head-mounted device is similar to traditional glasses. A brow tracking camera 43 and an eye tracking camera 44 are provided on the head-mounted device. The brow tracking camera 43 is disposed on the upper edge of the frame 41, and the eye tracking camera 44 is disposed on the lower edge and side edges of the frame 41. The number and lens types of the brow tracking camera and the eye tracking camera on the head-mounted device are not limited. The number of cameras can be adjusted according to the lens type of the cameras, as long as it is ensured that the brow tracking camera and the eye tracking camera can collect complete facial images. In addition, the head-mounted device further includes a fill light 42, and the fill light 42 can automatically adjust the brightness according to the ambient light brightness, so that the brow tracking camera and the eye tracking camera can obtain clear brow images and eye images.

[0057] As Figure 1 shown, the information processing method shown in the present disclosure includes:

[0058] S101. Obtain at least one frame of facial image, where the facial image includes a brow image and an eye image;

[0059] S102. Determine the expression information of the user based on the facial image and the expression detection network;

[0060] S103. Obtain and display an anthropomorphic image, where the expression of the anthropomorphic image is related to the expression information.

[0061] In step S101, the head-mounted device is provided with a brow tracking camera and an eye tracking camera. Among them, the brow tracking camera is used to collect the user's brow image, and the eye tracking camera is used to collect the user's eye image. A frame of facial image is composed of a frame of brow image and a frame of eye image. After the user starts using the head-mounted device, the brow tracking camera and the eye tracking camera are turned on synchronously to collect the user's facial image in real time. The facial image can reflect the user's brow state and eye state, such as the user's brow raising, frowning, blinking, pupil movement, etc. Since the head-mounted device is not fixed by a strap, the position and distance between the head-mounted device and the user's eyes may change during use. To ensure that the cameras on the head-mounted device can collect the user's complete brow image and eye image, the brow tracking camera and the eye tracking camera can use wide-angle lenses to increase the viewing angle range of the cameras, so that the cameras can collect images with a larger range. In addition, the number of the brow tracking camera and the eye tracking camera can also be increased. For example, 2 brow tracking cameras and 3 eye tracking cameras are respectively disposed on the left frame and the right frame of the head-mounted device.

[0062] The head-mounted device acquires at least one frame of facial image, which at least includes the initial facial image of the user when using the head-mounted device. When the user just starts using the head-mounted device, the head-mounted device can remind the user to keep a natural state, look straight ahead, and relax the eyebrows and eyes in the form of sound, vibration, etc. The initial facial image of the user when using the head-mounted device is collected by the eyebrow tracking camera and the eye tracking camera to record the relative position between the head-mounted device and the user's face, which is used for comparison with the facial images obtained during subsequent use. In this way, by comparing the current facial image with the initial facial image, it is possible to determine whether there are changes in the position and distance between the head-mounted device and the user's eyes, and it is also possible to determine the user's expression changes.

[0063] In step S102, a frame of facial image contains the user's eyebrow image and eye image. To facilitate the processing of the facial image by the expression detection network, the feature information that can reflect the user's expression in the facial image can be extracted, and only the feature information is input into the expression detection network to reduce the processing time of the expression detection network. The expression detection network can be pre-trained by a small neural network MLPMixer and stored in the head-mounted device. The expression detection network can detect a variety of eyebrow expressions and eye expressions. For example, eyebrow expressions include frowning, left eyebrow moving left-upward, right eyebrow moving right-upward, etc., and eye expressions include left eye blinking, right eye blinking, left eye opening wide, right eye opening wide, etc. Therefore, the expression detection network can output the weight values of the eyebrow expression and the eye expression that make up the user's facial expression based on the feature information of the eyebrows and eyes in the facial image. The expression information is the eyebrow expression that makes up the facial expression and the corresponding weight value, as well as the eye expression and the corresponding weight value. For example, the initial facial image is obtained based on the user being in a relaxed state. The feature information of the eyebrows and eyes in the initial facial image is input into the expression detection network. The expression detection network can detect the eyebrow expressions of the left eyebrow moving left-upward and the right eyebrow moving right-upward, and the eye expressions of the left eye opening wide and the right eye opening wide. Therefore, the expression detection network outputs four weight values based on the feature information, which are 0.5, 0.5, 0.7, and 0.7 respectively. That is, the expression information is that the weight value of the left eyebrow moving left-upward is 0.5, the weight value of the right eyebrow moving right-upward is 0.5, the weight value of the left eye opening wide is 0.7, and the weight value of the right eye opening wide is 0.7. The expression information can reflect the user's facial expression state.

[0064] In step S103, the head-mounted device can send the obtained expression information to game engines at the output end, such as Unity, Ureal, etc. The output end can use Blendshapes technology to generate anthropomorphic images based on the expression information. BlendShapes refers to a deformation method based on linear weighting, which can be used to change the specific shape of a human body model and is usually used in applications such as virtual character animation production, game design, and virtual fitting. The principle of BlendShapes is to add a series of shape deformations (Shape Keys) on the basis of a deformed mesh to simulate facial expressions. Each shape deformation can be regarded as an independent deformation controller, and these controllers can be combined through weight values to create various facial expressions. That is, the output end can use Blendshapes technology to fuse each expression according to the weight values in the expression information to obtain an anthropomorphic image with an expression consistent with the user's current expression. For example, if the expression information reflects that the user is currently frowning and the eyes are slightly open, the output end can generate an anthropomorphic image with a frowning expression and slightly open eyes based on the obtained expression information.

[0065] The head-mounted device can obtain the anthropomorphic image generated by the output end through an application program and display it in the head-mounted device. In this way, the user can view an anthropomorphic image consistent with their current facial expression through the head-mounted device. Among them, the prototype of the anthropomorphic image can be a model consistent with the user's image, or a model inconsistent with the user's image. For example, the anthropomorphic image is the image of a cartoon cat or dog. The anthropomorphic image can include only the face, or it can include the upper body or the whole body, as long as the anthropomorphic image can simulate the user's facial expression. In addition, there are no restrictions on the scene and position where the head-mounted device displays the anthropomorphic image. For example, the anthropomorphic image can be displayed in the blank area of the menu bar, and the user can also change the display position of the anthropomorphic image according to their personal visual habits. For example, if the default display position of the anthropomorphic image is the lower right corner of the display interface, the user can move the anthropomorphic image to the upper right corner of the display interface.

[0066] In the present disclosure, the head-mounted device is provided with an eyebrow tracking camera and an eye tracking camera, which can ensure that even if the position and distance of the head-mounted device change, a complete facial image of the user can be obtained, effectively improving the robustness of eye tracking of the user. The anthropomorphic image generated based on the expression information can truly simulate the user's expression, enhancing the fun and vividness of the display and improving the user experience.

[0067] According to an exemplary embodiment, as Figure 2 shown, the information processing method in this embodiment includes:

[0068] S201. Obtain at least one frame of facial image, where the facial image includes an eyebrow image and an eye image;

[0069] S202. Obtain the first eyebrow feature information based on the eyebrow image and the pre-stored eyebrow feature extraction network;

[0070] S203. Obtain the first eye feature information based on the eye image and the pre-stored eye feature extraction network;

[0071] S204. Determine whether the number of frames of the facial image is greater than or equal to the length of the moving window. If the number of frames of the facial image is less than the length of the moving window, execute step S205; if the number of frames of the facial image is greater than or equal to the length of the moving window, execute step S207;

[0072] S205. Input the first eyebrow feature information and the first eye feature information into the expression detection network, and the expression detection network outputs the first weight value;

[0073] S206. Determine the expression information based on the first weight value;

[0074] S207. Optimize the first eyebrow feature information and the first eye feature information to obtain the second eyebrow feature information and the second eye feature information;

[0075] S208. Input the second eyebrow feature information and the second eye feature information into the expression detection network, and the expression detection network outputs the second weight value;

[0076] S209. Determine the expression information based on the second weight value;

[0077] S210. Obtain and display the anthropomorphic image, and the expression of the anthropomorphic image is related to the expression information.

[0078] Among them, steps S201 and S210 are the same as steps S101 and S103 in the implementation manner in the above embodiment, and will not be elaborated here.

[0079] In step S202, the eyebrow feature extraction network is a lightweight neural network pre-stored in the head-mounted device after training. The neural network can be EfficientNetv2, MobileNetv2, etc. During the training process of the eyebrow feature extraction network, the set of eyebrow sample images marked with eyebrow feature points is input into the neural network for training. The set of eyebrow sample images used to train the eyebrow feature extraction network needs to be pre-processed, such as cropping, rotating, and enhancing, to ensure that each eyebrow sample image has the same size and can clearly present the complete eyebrow shape. In addition, to improve the robustness of the eyebrow feature extraction network, some of the eyebrow sample images in the set of eyebrow sample images can be partially occluded, so that the trained eyebrow feature extraction network can output complete first eyebrow feature information even based on incomplete eyebrow images. The first eyebrow feature information is the coordinate values of each eyebrow feature point. Through the coordinate relationship between each eyebrow feature point, the specific shape of the eyebrow can be determined. Among them, the coordinate values can be presented in u-v coordinates or x-y coordinates.

[0080] For example, the set of eyebrow sample images consists of 100,000 224*224*3 RGB images with a size of 640x480. Each eyebrow sample image is marked with 20 eyebrow feature points, as Figure 5a shown. The 20 eyebrow feature points 4 are evenly distributed around the left and right eyebrows 3. Taking the lower left corner of the eyebrow sample image as (0,0) and the upper right corner of the eyebrow sample image as (1,1), the u-v coordinates of each eyebrow feature point are recorded in this way. The set of eyebrow sample images is input into EfficientNetv2, and the Adam optimizer is used to optimize the neural network during the training process. The loss function during the training process is set as the sum of squared errors, the initial learning rate is 0.001, and the convergence round is 105, so that EfficientNetv2 has a better training effect.

[0081] In step 203, the eye feature extraction network is a lightweight neural network pre-stored in the head-mounted device after training. The training process of the eye feature extraction network is similar to that of the eyebrow feature extraction network. The set of eye sample images marked with eye feature points is input into the neural network for training. The set of eye sample images also needs to be pre-processed so that the set of eye sample images has the same size and can clearly present the complete eye shape. Since the eye includes the eye socket, eyelid, and pupil, to improve the training effect of the eye feature extraction network, the marking of eye feature points on the eye sample images is increased. As Figure 5b shown, there are a total of 34 eye feature points 5, which are evenly distributed on the left and right eyes. Among them, there are 7 eye feature points 5 at one side of the pupil 2, and 10 eye feature points 5 at one side of the eye socket 1.

[0082] During the training process of the eye feature extraction network, the settings of various parameters can be the same as those of the eyebrow feature extraction network, so that the trained eye feature extraction network can output first eye feature information reflecting the user's eye state based on the eye image. The first eye feature information includes first pupil feature information and first eyelid feature information. The coordinate values of the eye feature points at the pupil can reflect the movement of the pupil, and the coordinate values of the eye feature points at the eye socket can reflect the opening and closing of the eyes. For the convenience of subsequent data processing, the form of the coordinate values of the first eye feature information is the same as that of the first eyebrow feature information, that is, presented in u-v coordinates or x-y coordinates.

[0083] As Figure 4 shown, the eye tracking camera 44 includes a middle eye tracking camera 441 and a lower eye tracking camera 442. The middle eye tracking camera 441 is disposed on the side edge of the frame 41, and the lower eye tracking camera 442 is disposed on the lower edge of the frame 41. Based on the positions of the middle eye tracking camera and the lower eye tracking camera on the head-mounted device, it can be known that the middle eye tracking camera can better sense the up and down movement of the pupil, and the lower eye tracking camera can better sense the left and right movement of the pupil. Therefore, the horizontal coordinate value (i.e., the u value or the x value) of the eye feature points at the pupil collected by the lower eye tracking camera in the first eye feature information and the vertical coordinate value (i.e., the v value or the y value) of the eye feature points at the pupil collected by the middle eye tracking camera can be taken and combined as the coordinate value of the eye feature points at the pupil to improve the accuracy of the first pupil feature information.

[0084] In step S204, during the process of the user using the head-mounted device, the user's facial state is in a changing state. For example, when playing a game with the head-mounted device, when encountering a scary scene, the user will close his eyes tightly and frown, and when encountering a surprising scene, the user will open his eyes wide and raise his eyebrows. Since the eyebrow tracking camera and the eye tracking camera are always in the on state in real time, the head-mounted device can obtain multiple frames of facial images. Each frame of image is equivalent to being extracted from a video stream. The facial images obtained near the frame value may jitter and there is Gaussian noise. In order to improve the accuracy of the output of the subsequent expression detection network, it is necessary to perform noise optimization on the first eyebrow feature information and the first eye feature information to filter out high-frequency disturbances and achieve a smoothing effect. The present disclosure adopts a moving window to process data, that is, sliding a moving window on the first eyebrow feature information and the first eye feature information of multiple frames of facial images, and performing a series of calculations or operations on the coordinate values of the moving window, such as calculating the average value, maximum value, minimum value, etc. of the first eyebrow feature information and the first eye feature information within the moving window.

[0085] If it is determined that the number of frames of the facial image is less than the length of the moving window, that is, the current head-mounted device has not yet obtained a large amount of first eyebrow feature information and first eye feature information, and it is not sufficient to use the moving window for data processing. Therefore, step S205 is executed, and the first eyebrow feature information and the first eye feature information are not processed and are directly input into the expression detection network. If it is determined that the number of frames of the facial image is greater than or equal to the length of the moving window, then step S207 can be executed to optimize the first eyebrow feature information and the first eye feature information, and then the optimized second eyebrow feature information and second eye feature information are used and input into the expression detection network. Among them, the length of the moving window is determined based on the scale of the data to be processed. If the length of the moving window is too large, a lag effect will occur; if the length of the moving window is too small, the noise optimization effect cannot be achieved. For example, the first eyebrow feature information of each frame of the facial image includes the coordinate values of 20 eyebrow feature points, and the first eye feature information includes the coordinate values of 34 eye feature points. The length of the moving window can be set to 8.

[0086] In step S205, the first eyebrow feature information and the first eye feature information are input into the expression detection network, and the expression detection network pre-stored in the head-mounted device outputs a first weight value based on the coordinate values of multiple eyebrow feature points and eye feature points in the first eyebrow feature information and the first eye feature information. Since the expression detection network can be used to detect multiple eye expressions and multiple eyebrow expressions, the output first weight value includes multiple expression weight values, and the range of the weight value of each expression is 0-1.

[0087] In one example, the expression detection network can detect 5 eyebrow expressions and 18 eye expressions. The 5 eyebrow expressions include: BrowDownLeft, BrowDownRight, BrowInnerUp, BrowOuterUpLeft, and BrowOuterUpRight. The 18 eye expressions include: EyeBlinkLeft, EyeLookDownLeft, EyeLookInLeft, EyeLookOutLeft, EyeLookUpLeft, EyeSquintLeft, EyeWideLeft, EyeBlinkRight, EyeLookDownRight, EyeLookInRight, EyeLookOutRight, EyeLookUpRight, EyeSquintRight, and EyeWideRight. The first eyebrow feature information and the first eye feature information are input into the expression detection network. The expression detection network determines the weight values of each eyebrow expression and each eye expression based on the coordinate values of each feature point. That is, the first weight values output by the expression detection network include 23 expression weight values, including 5 eyebrow expression weight values and 18 eye expression weight values.

[0088] In step S206, each eyebrow expression and eye expression and their corresponding expression weight values constitute expression information. For example, if the first weight values of the expression detection network include 23 expression weight values, then the expression information is composed of 23 expressions and 23 expression weight values. The head-mounted device sends the expression information to the game engine at the output end. The output end uses the Blendshapes technology to fuse each expression according to the expression weight values in the expression information to generate an anthropomorphic image.

[0089] In some embodiments, step S207 optimizes the first eyebrow feature information and the first eye feature information to obtain the second eyebrow feature information and the second eye feature information, including:

[0090] Based on a moving window, calculate the first moving average of the first eyebrow feature information of the facial image and the second moving average of the first eye feature information.

[0091] Based on the first moving average and the second moving average, obtain the second eyebrow feature information and the second eye feature information.

[0092] The moving window slides in the first eyebrow feature information and the first eye feature information respectively, and calculates the first moving average value of the first eyebrow feature information and the second moving average value of the first eye feature information within the moving window. For example, if the length of the moving window is 8 and a total of 20 facial images are obtained, and each facial image includes the coordinate values of 20 eyebrow feature points and 34 eye feature points, then each feature point has 20 coordinate values. The moving window moves among the 20 coordinate values corresponding to each feature point, and calculates the moving average value of 8 coordinate values of each feature point in the current facial image and the 7 facial images obtained before the current moment, that is, the moving window processes the 20 coordinate values corresponding to 54 feature points in groups, and only processes 8 coordinate values corresponding to one feature point each time.

[0093] In one example, Herd moving average is used for noise optimization, that is, Herd moving average is used to calculate the first moving average value and the second moving average value within the moving window. The specific processing process can adopt the following formula:

[0094]

[0095] Among them, WMA t represents the moving average value, T represents the length of the moving window, and y t represents the first eyebrow feature information or the first eye feature information of the currently obtained facial image, and y t-T+i represents the first eyebrow feature information or the first eye feature information of the first T - i frames.

[0096] The following formula is used to calculate the first moving average value and the second moving average value, and further obtain the second eyebrow feature information and the second eye feature information.

[0097]

[0098] After noise optimization of the first eyebrow feature information and the first eye feature information by the above method, high - frequency perturbations can be effectively filtered, avoiding the influence of facial image jitter on the confirmation of expression information, and the second eyebrow feature information and the second eye feature information achieve a good smoothing effect. For example, if there is jitter in the currently obtained facial image, and the coordinate value of a certain eye feature point is (0.8, 0.5), after noise optimization, the coordinate value of this eye feature point is (0.5, 0.5), then the coordinate value of this feature point in the second eye feature information is (0.5, 0.5).

[0099] Step S208 is similar to step S205 and will not be elaborated here. Since the second eyebrow feature information and the second eye feature information are obtained through noise optimization processing, the second weight value obtained by the expression detection network based on the second eyebrow feature information and the second eye feature information is more accurate.

[0100] Step S209 is similar to step S206 and will not be elaborated here. It should be noted that the second weight value is the same as the first weight value, both including the weight values of each expression. The first weight value is output by the expression detection network based on the first eyebrow feature information and the first eye feature information without noise optimization, and the second weight value is output by the expression detection network based on the second eyebrow feature information and the second eye feature information after noise optimization.

[0101] In the present disclosure, the head-mounted device performs noise optimization on the coordinate values of the eyebrow feature points and eye feature points of the facial image, which can effectively filter out high-frequency perturbations, achieve a smoothing effect, avoid affecting the determination of expression information due to the jitter of the facial image, improve the accuracy of expression information, enable the head-mounted device to display an anthropomorphic image consistent with the user's true expression, and increase the playability during the user's use of the head-mounted device.

[0102] According to an exemplary embodiment, as Figure 3 shown, the information processing method in this embodiment includes:

[0103] S301. Obtain at least one frame of facial image, where the facial image includes an eyebrow image and an eye image;

[0104] S302. Based on the eyebrow image and a pre-stored eyebrow feature extraction network, obtain first eyebrow feature information;

[0105] S303. Based on the eye image and a pre-stored eye feature extraction network, obtain first eye feature information;

[0106] S304. Determine whether the number of frames of the facial image is greater than or equal to the length of the moving window. If the number of frames of the facial image is less than the length of the moving window, execute step S305; if the number of frames of the facial image is greater than or equal to the length of the moving window, execute step S306;

[0107] S305. Input the first eyebrow feature information and the first eye feature information into the expression detection network, and the expression detection network outputs the first weight value;

[0108] S306. Optimize the first eyebrow feature information and the first eye feature information to obtain second eyebrow feature information and second eye feature information;

[0109] S307. Input the second eyebrow feature information and the second eye feature information into the expression detection network, and the expression detection network outputs the second weight value;

[0110] S308. Use the first weight value of the initial facial image to correct the first weight value or the second weight value of the current facial image to obtain the corrected third weight value of the current facial image;

[0111] S309. Determine the expression information based on the third weight value;

[0112] S310. Obtain and display an anthropomorphic image, where the expression of the anthropomorphic image is related to the expression information.

[0113] Among them, steps S301 - S307 and S310 are the same as the implementation manners in the above embodiments and will not be elaborated here.

[0114] In step S308, the facial image includes an initial facial image and a current facial image. The initial facial image is obtained by the eyebrow tracking camera and the eye tracking camera when the user is in a relaxed state at the beginning of using the head-mounted device, and can reflect the user's regular usage state. Using the expression information determined by the initial facial image as a reference, the change in the user's current facial expression can be judged. Since the first weight value or the second weight value of the current facial image is directly output by the expression detection network, there may be a problem that it does not fit the actual situation of the user well. Taking the eye expression of "left eye squinting" as an example, since the eyes of each user are of different sizes, the degree of eye opening is different. For a user with larger eyes, the coordinate values of the eye feature points when squinting may be the same as the coordinate values of the eye feature points when a user with smaller eyes opens their eyes. Therefore, directly judging the eye expression of a user with smaller eyes based on the expression weight value output by the expression detection network will result in an error.

[0115] To ensure the accuracy of the expression information and make it fit the actual situation of the user, the first weight value of the initial facial image can be used to correct the first weight value or the second weight value of the current facial image. It should be noted that since the initial facial image is the first frame of the facial image obtained by the head-mounted device and the number of frames is less than the length of the moving window, the first eyebrow feature information and the first eye feature information of the initial facial image are not optimized for noise, and the first eyebrow feature information and the first eye feature information of the initial facial image are directly input into the expression detection network, and the expression detection network outputs the first weight value of the initial facial image. The third weight value is obtained using the following formula:

[0116]

[0117] For the i-th expression, BS′ i represents the corrected weight value of the facial image corresponding to the i-th expression, that is, the third weight value, and BS i represents the first weight value of the facial image corresponding to the i-th expression or the optimized second weight value, represents the first weight value of the initial facial image.

[0118] For example, for an expression of closed eyes, the expression weight value output by the expression detection network is 0.7. If the expression weight value of the closed eyes is greater than 0.7, the output end will output an anthropomorphic image of closed eyes. For a user with relatively small eyes, among the first weight values of the initial facial image obtained in a relaxed state, the expression weight value representing "left eye closed" is 0.9, that is, the expression weight value when the user's eyes are fully closed is 0.9. The expression weight value representing "left eye closed" in the first weight value of the current facial image is 0.85, which means that the user's left eye is considered to be fully closed, but in fact, the user's eyes are not fully closed, and may only be half-closed. Therefore, the expression weight value in the first weight value of the initial facial image is used for correction, that is, the expression weight value in the third weight value is abs(0.85 - 0.9) / (1 - 0.9) = 0.5. The corrected expression weight value is closer to the actual state of the user.

[0119] Step S309 is similar to step S206 in the above embodiment, and will not be described in detail here.

[0120] In the present disclosure, by using the first weight value of the initial facial image to correct the first weight value or the second weight value of the current facial image, the finally determined expression information can conform to the real situation of the user, ensuring that the obtained and displayed anthropomorphic image is consistent with the user's current expression, and improving the user experience.

[0121] An exemplary embodiment of the present disclosure provides an information processing device, which is applied to a head-mounted device. The head-mounted device includes an eyebrow tracking camera and an eye tracking camera. The eyebrow tracking camera is used to collect eyebrow images, and the eye tracking camera is used to collect eye images. As Figure 6 shown, a block diagram of an information processing device shown in the present disclosure.

[0122] The block diagram includes: an acquisition module 61, a processing module 62, and a display module 63. The acquisition module 61 is used to acquire at least one frame of facial image, and the facial image includes an eyebrow image and an eye image; the processing module 62 is used to determine the expression information of the user based on the facial image and the expression detection network; the display module 63 is used to acquire and display an anthropomorphic image, and the expression of the anthropomorphic image is related to the expression information.

[0123] In an exemplary embodiment of the present disclosure, based on the eyebrow image and a pre-stored eyebrow feature extraction network, first eyebrow feature information is obtained; based on the eye image and a pre-stored eye feature extraction network, first eye feature information is obtained; the processing module 62 is further used to: determine the expression information of the user based on the expression detection network, the first eyebrow feature information, and the first eye feature information.

[0124] In an exemplary embodiment of the present disclosure, the processing module 62 is further configured to: input the first eyebrow feature information and the first eye feature information into an expression detection network, and the expression detection network outputs a first weight value; determine the expression information based on the first weight value.

[0125] In an exemplary embodiment of the present disclosure, the processing module 62 is further configured to: optimize the first eyebrow feature information and the first eye feature information to obtain second eyebrow feature information and second eye feature information; determine the expression information of the user based on the expression detection network, the second eyebrow feature information, and the second eye feature information.

[0126] In an exemplary embodiment of the present disclosure, the processing module 62 is further configured to: calculate a first moving average value of the first eyebrow feature information and a second moving average value of the first eye feature information of the facial image based on a moving window, the moving window slides in the first eyebrow feature information and the first eye feature information respectively, and the number of frames of the facial image is greater than or equal to the length of the moving window; obtain the second eyebrow feature information and the second eye feature information based on the first moving average value and the second moving average value.

[0127] In an exemplary embodiment of the present disclosure, the processing module 62 is further configured to: input the second eyebrow feature information and the second eye feature information into an expression detection network, and the expression detection network outputs a second weight value; determine the expression information based on the second weight value.

[0128] In an exemplary embodiment of the present disclosure, the facial image includes an initial facial image and a current facial image, and the processing module 62 is further configured to: correct the first weight value or the second weight value of the current facial image by using the first weight value of the initial facial image to obtain a third weight value after correction of the current facial image; determine the expression information based on the third weight value.

[0129] In an exemplary embodiment of the present disclosure, the first eyebrow feature information and the second eyebrow feature information include coordinate values of multiple eyebrow feature points, and the first eye feature information and the second eye feature information include coordinate values of multiple eye feature points.

[0130] In an exemplary embodiment of the present disclosure, the anthropomorphic image is generated by the output end based on the expression information.

[0131] Regarding the information processing device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0132] Figure 7It is a block diagram of a head-mounted device 700 shown according to an exemplary embodiment. For example, the head-mounted device 700 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0133] Referring to Figure 7 , the head-mounted device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 716, and a communication component 716.

[0134] The processing component 702 generally controls the overall operation of the head-mounted device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.

[0135] The memory 704 is configured to store various types of data to support the operation of the head-mounted device 700. Examples of these data include instructions for any application or method operating on the head-mounted device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0136] The power supply component 706 provides power to various components of the head-mounted device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the head-mounted device 700.

[0137] The multimedia component 708 includes a screen that provides an output interface between the head-mounted device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the head-mounted device 700 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0138] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the head-mounted device 700 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0139] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0140] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the head-mounted device 700. For example, the sensor component 714 can detect the on / off state of the head-mounted device 700, the relative positioning of components, such as the display and keypad of the head-mounted device 700. The sensor component 714 can also detect a change in the position of the head-mounted device 700 or a component of the head-mounted device 700, the presence or absence of user contact with the head-mounted device 700, the orientation or acceleration / deceleration of the head-mounted device 700, and the temperature change of the head-mounted device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0141] The communication component 716 is configured to facilitate communication between the head-mounted device 700 and other devices in a wired or wireless manner. The head-mounted device 700 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0142] In an exemplary embodiment, the head-mounted device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0143] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the above instructions can be executed by the processor 720 of the head-mounted device 700 to complete the above information processing method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0144] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the head-mounted device, enables the processing device of the head-mounted device to execute the information processing method provided by the exemplary embodiments of the present disclosure.

[0145] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0146] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. An information processing method, characterized in that, Applied to a head-mounted device, the head-mounted device includes an eyebrow tracking camera and an eye tracking camera. The eyebrow tracking camera is used to collect eyebrow images, and the eye tracking camera is used to collect eye images. The information processing method includes: Obtain at least one frame of facial image, where the facial image includes the eyebrow image and the eye image; Based on the facial image and an expression detection network, determine the user's expression information; Obtain and display an anthropomorphic image, where the expression of the anthropomorphic image is related to the expression information.

2. The information processing method according to claim 1, characterized in that, The method further includes: Based on the eyebrow image and a pre-stored eyebrow feature extraction network, obtain first eyebrow feature information; Based on the eye image and a pre-stored eye feature extraction network, obtain first eye feature information; The step of determining the user's expression information based on the facial image and the expression detection network includes: Based on the expression detection network, the first eyebrow feature information, and the first eye feature information, determine the user's expression information.

3. The information processing method according to claim 2, characterized in that, The step of determining the user's expression information based on the expression detection network, the first eyebrow feature information, and the first eye feature information includes: Input the first eyebrow feature information and the first eye feature information into the expression detection network, and the expression detection network outputs a first weight value; Based on the first weight value, determine the expression information.

4. The information processing method according to claim 2, characterized in that, The information processing method further includes: Optimize the first eyebrow feature information and the first eye feature information to obtain second eyebrow feature information and second eye feature information; Based on the expression detection network, the second eyebrow feature information, and the second eye feature information, determine the user's expression information.

5. The information processing method according to claim 4, characterized in that, The step of optimizing the first eyebrow feature information and the first eye feature information to obtain second eyebrow feature information and second eye feature information includes: Based on a moving window, calculate a first moving average of the first eyebrow feature information of the facial image and a second moving average of the first eye feature information. The moving window slides in the first eyebrow feature information and the first eye feature information respectively, and the number of frames of the facial image is greater than or equal to the length of the moving window; Based on the first moving average and the second moving average, obtain the second eyebrow feature information and the second eye feature information.

6. The information processing method according to claim 5, characterized in that, The step of determining the user's expression information based on the expression detection network, the second eyebrow feature information, and the second eye feature information includes: Input the second eyebrow feature information and the second eye feature information into the expression detection network, and the expression detection network outputs a second weight value; Based on the second weight value, determine the expression information.

7. The information processing method according to claim 6, characterized in that, The facial image includes an initial facial image and a current facial image. The information processing method further includes: Use the first weight value of the initial facial image to correct the first weight value or the second weight value of the current facial image to obtain a corrected third weight value of the current facial image; Based on the third weight value, determine the expression information.

8. The information processing method according to any one of claims 2-7, characterized in that, The first eyebrow feature information and the second eyebrow feature information include coordinate values of a plurality of eyebrow feature points, and the first eye feature information and the second eye feature information include coordinate values of a plurality of eye feature points.

9. The information processing method according to claim 1, characterized in that, The anthropomorphic image is generated by an output end based on the expression information.

10. An information processing device, characterized in that, Applied to a head-mounted device, the head-mounted device includes an eyebrow tracking camera and an eye tracking camera. The eyebrow tracking camera is used to collect eyebrow images, and the eye tracking camera is used to collect eye images. The information processing device includes: An acquisition module, configured to acquire at least one frame of facial image, where the facial image includes the eyebrow image and the eye image; A processing module, configured to determine the expression information of a user based on the facial image and an expression detection network; A display module, configured to acquire and display an anthropomorphic image, where the expression of the anthropomorphic image is related to the expression information.

11. A head-mounted device, characterized in that, The head-mounted device includes: A fill light; an eyebrow tracking camera; an eye tracking camera; A processor; A memory for storing executable instructions executable by the processor; Wherein, the processor is configured to execute the executable instructions in the memory to implement the information processing method according to any one of claims 1 to 9.

12. The head-mounted device according to claim 11, characterized in that, The head-mounted device has a frame. The eye tracking camera includes a middle eye tracking camera and a lower eye tracking camera. The eyebrow tracking camera is arranged on the upper edge of the frame, the middle eye tracking camera is arranged on the side edge of the frame, and the lower eye tracking camera is arranged on the lower edge of the frame.

13. A non-transitory computer-readable storage medium, on which executable instructions are stored, characterized in that, When the executable instructions are executed by the processor, the information processing method according to any one of claims 1 to 9 is implemented.