Facial data entry method, device, equipment, medium and program product

By using images with color and depth information combined with a pre-trained model to generate head pose and facial information, the problem of difficulty in obtaining three-dimensional geometric features and light sensitivity in robot facial recognition technology is solved, and efficient and accurate input and recognition of facial data under multiple poses is achieved.

CN120976995APending Publication Date: 2025-11-18JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511476694.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Robotic facial recognition technology relies on two-dimensional visible light RGB images, making it difficult to obtain three-dimensional geometric features of the face. It is also sensitive to lighting conditions, resulting in low recognition accuracy, especially when the face is not directly facing the image.

Method used

Facial data is recorded using images that include color and depth information. A pre-trained facial detection and feature extraction model is used to generate head pose and facial information. The facial data is stored under pose conditions that meet the robot's perspective settings, thus realizing facial information recording under multiple poses.

Benefits of technology

This improves the accuracy and efficiency of robot facial recognition, enabling accurate recording of facial data under different head postures, reducing the impact of lighting conditions on recognition, and ensuring the efficiency and accuracy of subsequent facial recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976995A_ABST
    Figure CN120976995A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a face data entry method and device, equipment, a medium and a program product. A specific embodiment of the method comprises the following steps: in response to currently executed face data entry for a face registration object, obtaining an image of currently shooting the face registration object; determining facial feature information corresponding to the image according to the color information and the depth information; according to the facial feature information, generating a head posture and facial information corresponding to the image; and in response to determining that the head posture meets a posture condition set based on the visual angle of the robot, storing face data corresponding to the face registration object to a face database corresponding to the robot according to the face information so as to realize face information input under the head posture. The implementation mode is related to an intelligent robot, the face data under different head postures can be accurately and efficiently input into the robot, and efficient interaction between a subsequent object and the robot is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and specifically to facial data entry methods, apparatus, devices, media, and program products. Background Technology

[0002] Currently, with the advent of the intelligent era, robots are gradually becoming an important part of daily life. How robots can accurately identify objects is crucial. The common method for inputting facial data into robots is as follows: First, an RGB image of the object's face is captured. Then, the RGB image is entered into the robot's corresponding facial database.

[0003] However, the inventors discovered that the following technical problems often arise when using the above method: Facial recognition technology in robots typically relies on two-dimensional visible light RGB images. However, this method suffers from low dimensionality, making it difficult to effectively capture the three-dimensional geometric features of the face. Furthermore, it is extremely sensitive to lighting conditions; strong light, weak light, backlighting, and uneven lighting can all significantly reduce recognition accuracy. In addition, RGB images of the face are often head-on, which can lead to ineffective facial recognition by robots, resulting in inefficient interactions.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide methods, apparatus, devices, media, and program products for facial data entry to address the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a facial data entry method, comprising: in response to currently performing facial data entry for a facial registration object, acquiring an image currently captured of the facial registration object, wherein the image is an image including color information and depth information; determining facial feature information corresponding to the image based on the color information and depth information; generating a head pose and facial information corresponding to the image based on the facial feature information; and in response to determining that the head pose satisfies pose conditions set based on a robot's perspective, storing the facial data corresponding to the facial registration object into a facial database corresponding to the robot based on the facial information, thereby realizing facial information entry under the head pose.

[0008] Optionally, the robot performs real-time facial tracking during the multi-pose facial data entry process; and the method further includes: in response to a facial tracking interruption during the facial data entry process and the re-detection of the registered facial object within a preset interruption time, generating re-detected facial feature information of the object; comparing the object's facial feature information with the facial data corresponding to the registered facial object stored in the facial database to obtain comparison information; and in response to determining that the comparison information is correct, continuing to perform multi-pose facial data entry for the registered facial object.

[0009] Optionally, the method further includes: in response to performing facial data addition to the target facial database, during the interaction with the facial registration object, obtaining an interactive image set corresponding to the facial registration object, wherein the target facial database is a facial database in which the facial registration object has completed multi-pose facial data recording; selecting facial data enhancement images from the interactive image set based on at least one of facial distance, shooting illumination, head posture, and shooting time corresponding to the interactive images, to obtain a facial data enhancement image set; and adding the facial data enhancement image set to the facial data corresponding to the facial registration object in the target facial database.

[0010] Optionally, determining the facial feature information corresponding to the image based on the color information and depth information includes: inputting the color information into a pre-trained facial detection model to obtain facial detection information; filtering out the depth information corresponding to the facial region from the depth information based on the facial detection information to obtain facial depth information; combining the facial detection information and the facial depth information to obtain facial combination information; and inputting the facial combination information into a pre-trained facial feature extraction model to obtain facial feature information.

[0011] Optionally, generating the head pose and facial information corresponding to the image based on the facial feature information includes: inputting the facial feature information into a pre-trained head pose generation model to obtain the head pose; and training the head pose generation model through the following steps: acquiring a first training image dataset and the second training image dataset, wherein the first training image dataset is an image including color information, and the second training image dataset is an image including color information and depth information; performing data augmentation on the first training image dataset to generate training image data under various target poses, obtaining an augmented image dataset; training the initial head pose generation model based on the augmented image dataset to obtain a trained head pose generation model; and retraining the trained head pose generation model based on the second training image dataset to obtain a head pose generation model.

[0012] Optionally, the method further includes: in response to detecting a first object during the interaction, acquiring a first facial image corresponding to the first object; in response to the head pose corresponding to the first facial image conforming to the pose condition, determining whether there exists a facial input object in the target facial database whose feature similarity with the first facial image meets the target similarity condition; in response to determining that there is a target facial input object, determining that the first object is the information of the target facial input object; and in the interaction, performing object tracking for the target facial input object.

[0013] Optionally, the method further includes: in response to the object tracking for the target facial object being disconnected and a second object being detected again, acquiring a second facial image corresponding to the second object; re-determining whether there is a facial object in the target facial database whose feature similarity with the second facial image meets the target similarity condition; and in response to determining that there is a target facial object, continuing to perform object tracking for the target facial object.

[0014] Optionally, the robot moves around the face registration object to record facial data in real time in multiple poses.

[0015] Secondly, some embodiments of this disclosure provide a facial data entry device, including: an acquisition unit configured to acquire an image of the facial registration object currently being captured in response to performing facial data entry for a facial registration object, wherein the image is an image including color information and depth information; a determination unit configured to determine facial feature information corresponding to the image based on the color information and depth information; a generation unit configured to generate a head pose and facial information corresponding to the image based on the facial feature information; and a storage unit configured to store the facial data corresponding to the facial registration object in a facial database corresponding to the robot based on the facial information in response to determining that the head pose satisfies the pose conditions set based on the robot's perspective, thereby realizing facial information entry under the head pose.

[0016] Optionally, the robot performs real-time facial tracking during the multi-pose facial data entry process; and the device further includes: in response to a facial tracking disconnection during the facial data entry process and the re-detection of the registered facial object within a preset disconnection time, generating re-detected facial feature information of the object; comparing the object's facial feature information with the facial data corresponding to the registered facial object stored in the facial database to obtain comparison information; and in response to determining that the comparison information is correct, continuing to perform multi-pose facial data entry for the registered facial object.

[0017] Optionally, the apparatus further includes: in response to performing facial data addition to a target facial database, during interaction with the facial registration object, acquiring an interactive image set corresponding to the facial registration object, wherein the target facial database is a facial database in which the facial registration object has completed multi-pose facial data recording; selecting facial data enhancement images from the interactive image set based on at least one of facial distance, shooting illumination, head posture, and shooting time corresponding to the interactive images, to obtain a facial data enhancement image set; and adding the facial data enhancement image set to the facial data corresponding to the facial registration object in the target facial database.

[0018] Optionally, the determining unit is configured to: input the aforementioned color information into a pre-trained face detection model to obtain face detection information; based on the aforementioned face detection information, filter out the depth information corresponding to the face region from the aforementioned depth information to obtain face depth information; combine the aforementioned face detection information and the aforementioned face depth information to obtain face combination information; and input the aforementioned face combination information into a pre-trained face feature extraction model to obtain face feature information.

[0019] Optionally, the generation unit is configured to: input the aforementioned facial feature information into a pre-trained head pose generation model to obtain a head pose; and train the aforementioned head pose generation model through the following steps: acquiring a first training image dataset and the aforementioned second training image dataset, wherein the first training image dataset is an image including color information, and the second training image dataset is an image including color information and depth information; performing data augmentation on the aforementioned first training image dataset to generate training image data under various target poses, thereby obtaining an augmented image dataset; training the initial head pose generation model based on the aforementioned augmented image dataset to obtain a trained head pose generation model; and retraining the trained head pose generation model based on the aforementioned second training image dataset to obtain a head pose generation model.

[0020] Optionally, the apparatus further includes: in response to detecting a first object during the interaction, acquiring a first facial image corresponding to the first object; in response to the head posture corresponding to the first facial image conforming to the posture condition, determining whether there exists a facial input object in the target facial database whose feature similarity with the first facial image meets the target similarity condition; in response to determining that there is a target facial input object, determining that the first object is the information of the target facial input object; and in the interaction, performing object tracking for the target facial input object.

[0021] Optionally, the apparatus further includes: in response to the disconnection of object tracking for the target facial object and the detection of a second object again, acquiring a second facial image corresponding to the second object; re-determining whether there is a facial object in the target facial database whose feature similarity with the second facial image meets the target similarity condition; and in response to determining that there is a target facial object, continuing to perform object tracking for the target facial object.

[0022] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0023] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0024] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0025] The above-described embodiments of this disclosure have the following beneficial effects: Through the facial data entry methods of some embodiments of this disclosure, facial data under different head postures can be accurately and efficiently entered into the robot, facilitating efficient interaction between the object and the robot. Specifically, the reason for the insufficient accuracy and efficiency of interaction with the robot is that facial recognition technology in robots typically relies on two-dimensional visible light RGB images. However, this method suffers from low dimensionality, making it difficult to effectively acquire the three-dimensional geometric features of the face. Furthermore, it is extremely sensitive to lighting conditions; strong light, weak light, backlight, and uneven lighting can all significantly reduce the accuracy of recognition. In addition, RGB images of the face are often head-on facial images, which often lead to the robot's inability to effectively recognize faces, resulting in inefficient interaction. Therefore, the facial data entry method of some embodiments of this disclosure first, in response to the current execution of facial data entry for a registered facial object, acquires an image of the currently captured face. This image includes color information and depth information. Here, by capturing images that include depth information, three-dimensional facial data corresponding to the registered facial object can be effectively recorded during subsequent facial data entry, avoiding the problem of low dimensionality and effectively acquiring the three-dimensional geometric features of the face. Furthermore, by setting facial data that includes depth information, the problem of low recognition accuracy under sensitive lighting conditions can be effectively avoided. Then, based on the aforementioned color and depth information, the facial feature information corresponding to the image can be accurately determined to obtain the semantic content of the facial features corresponding to the registered facial object. Next, based on the aforementioned facial feature information, the head pose and facial information corresponding to the image are generated. Here, by generating the head pose, facial images that meet the requirements for accurate facial recognition by the robot are selected, enabling the robot to accurately and efficiently extract more effective facial feature information during subsequent facial recognition, making the subsequent recognition process more accurate and efficient. In addition, by generating facial information, the robot can effectively understand the facial identity information corresponding to the registered facial object, thereby achieving accurate generation of subsequent facial data. Finally, in response to the determination that the head posture satisfies the posture conditions set based on the robot's perspective, the facial data corresponding to the registered facial object is stored in the robot's facial database according to the facial information, thereby realizing the facial information input under the head posture. Here, by inputting facial information under the head posture, the robot can subsequently achieve accurate recognition of the registered facial object under that head posture. Furthermore, by inputting facial data under multiple head postures through the robot, the robot can achieve accurate recognition under multiple head postures during the subsequent recognition process, greatly improving recognition efficiency and accuracy.In summary, by including depth information in the corresponding facial images and recording facial information under various head poses, the accuracy and efficiency of subsequent facial recognition by the robot can be improved. Attached Figure Description

[0026] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0027] Figure 1 This is a schematic diagram illustrating an application scenario of a facial data entry method according to some embodiments of the present disclosure; Figure 2 This is a flowchart of some embodiments of the facial data entry method according to the present disclosure; Figure 3 This is a schematic diagram of the face data entry process according to some embodiments of the face data entry method of this disclosure; Figure 4 This describes the face recognition process according to some embodiments of the face data entry method disclosed herein; Figure 5 This is a flowchart of some other embodiments of the facial data entry method according to the present disclosure; Figure 6 These are schematic diagrams illustrating the structure of some embodiments of the facial data entry device according to the present disclosure; Figure 7 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0029] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] Before performing any of the operations involving the collection, storage, or use of user personal information (such as facial data) disclosed in this disclosure, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, informing personal information subjects, and obtaining prior authorization and consent from personal information subjects.

[0034] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0035] Figure 1 This is a schematic diagram illustrating an application scenario of a facial data entry method according to some embodiments of the present disclosure.

[0036] exist Figure 1 In this application scenario, firstly, in response to the current facial data entry for the registered facial object, robot 101 can acquire an image 102 currently captured on the registered facial object. Image 102 includes color information 1031 and depth information 1032. Then, robot 101 can determine the facial feature information 104 corresponding to image 102 based on the color information 1031 and depth information 1032. Next, robot 101 can generate a head pose 105 and facial information 106 corresponding to image 102 based on the facial feature information 104. In this application scenario, head pose 105 can be "pitch angle: angle A, yaw angle: angle B, roll angle: angle C". Facial information 106 can be an object identity description. Finally, in response to determining that the head posture 105 satisfies the posture conditions set based on the robot's perspective, the robot 101 can store the facial data 107 corresponding to the facial registration object into the facial database 108 corresponding to the robot 101 based on the facial information 106, so as to realize the facial information input under the head posture 105.

[0037] Continue to refer to Figure 2The diagram illustrates a flow 200 of some embodiments of a facial data entry method according to the present disclosure. This facial data entry method, applied to a robot, includes the following steps: Step 201: In response to the current execution of facial data entry for the facial registration object, obtain the image currently being captured on the aforementioned facial registration object.

[0038] In some embodiments, in response to the current execution of facial data entry for a facial registration object, the executing entity of the aforementioned facial data entry method (e.g., Figure 1 The robot 101 shown can acquire images of the currently registered facial objects via wired or wireless connections. The registered facial object can be an object whose facial data is entered based on facial features. The robot can be a robot that supports facial recognition and interaction functions. In practice, the robot can be an intelligent robot that supports facial recognition and various interactive functions. For example, the robot can be an intelligent home robot that supports facial recognition. In a home setting, the robot can be a home robot used for health monitoring. The registered facial object can be a person entering facial data based on facial features, or an animal entering facial data based on pet facial features. For example, in a factory setting, the robot can be a monitoring robot. The registered facial object can be an employee entering facial data based on facial features. Facial data entry can involve inputting facial data into the robot so that the robot can clearly understand the identity information of the entered object. Facial data can be data related to facial features. In practice, facial data can include, but is not limited to, at least one of the following: the object's identity information, face size, and facial pose. The image can be an image captured at the current time. The image includes color information and depth information. Color information can be a channel matrix corresponding to each color channel. In practice, each color channel can include: red channel, green channel, and blue channel. Depth information can be a channel matrix under the depth channel. Depth information can represent the distance between the object displayed in the image and the camera device.

[0039] As an example, the robot can use an RGBD camera to capture images of registered facial objects.

[0040] Step 202: Based on the color information and depth information mentioned above, determine the facial feature information corresponding to the above image.

[0041] In some embodiments, the executing entity can determine the facial feature information corresponding to the image based on the color information and depth information. The facial feature information can characterize the semantic content of the facial content displayed in the image. In practice, for images including faces, the corresponding facial feature information can be the semantic content of the face in the image. In practice, facial feature information can be in vector form.

[0042] As an example, firstly, the aforementioned execution entity can concatenate the color matrix corresponding to the color information and the depth matrix corresponding to the depth information to obtain a concatenated matrix. Then, the concatenated matrix is ​​input into a pre-trained facial feature extraction model to obtain facial feature information. The facial feature extraction model can be a deep learning model for extracting facial feature information. For example, the facial feature extraction model could be a lightweight convolutional layer-based extraction model.

[0043] In some optional implementations of certain embodiments, the execution entity can determine the facial feature information corresponding to the image based on the color information and depth information, including the following steps: The first step involves inputting the aforementioned color information into a pre-trained face detection model to obtain face detection information. This face detection model can be a neural network model for face detection. The face detection information can be detection results related to corresponding facial features. This information can include: facial identity information, facial landmark information, and facial bounding boxes. Facial identity information can be the identity of the object corresponding to the face. Facial landmark information can include 12 facial landmarks. For example, for an image containing a face, the corresponding facial landmark information could be 12 facial landmarks. In practice, the face detection model can be a RetinaFace model or an anchor-free CenterFace model.

[0044] The second step involves filtering the depth information corresponding to the facial regions from the aforementioned facial detection information to obtain the facial depth information. This facial depth information can be the depth of the facial regions displayed in the image. The facial depth information can be in matrix form.

[0045] As an example, firstly, the aforementioned execution entity can generate a facial contour based on the various facial key points included in the facial detection information. Then, it filters the depth information corresponding to the facial contour from the aforementioned depth information and uses it as the facial depth information.

[0046] The third step involves combining the aforementioned facial detection information and facial depth information to obtain combined facial information. This combined facial information can be a combination of color and depth information related to the facial region displayed in the image. In practice, the combined facial information can be in matrix form.

[0047] As an example, firstly, the aforementioned execution entity can determine the facial contour based on the facial detection information. Then, it filters the color information corresponding to the facial contour from the color information to obtain the facial color information. This facial color information can be in matrix form. Next, the aforementioned facial detection information and the aforementioned facial depth information are matrix-concatenated to obtain the combined facial information.

[0048] The fourth step involves inputting the aforementioned facial combination information into a pre-trained facial feature extraction model to obtain facial feature information. This facial feature extraction model can be a deep learning model for extracting facial feature information. Specifically, it can be a lightweight deep learning model for extracting semantic content from facial features. For example, the facial feature extraction model could be a MobileNet model.

[0049] Optionally, after obtaining facial depth information by filtering out the depth information corresponding to the facial region from the depth information based on the aforementioned facial detection information, the method further includes: The first step is to normalize each element in the matrix corresponding to the facial depth information to obtain normalized depth information.

[0050] The second step is to determine the normalized depth information as facial depth information.

[0051] Alternatively, the facial feature extraction model can be trained through the following steps: The first step is to obtain the first training dataset and the second training dataset. The first training dataset consists of images in RGB format. The second training dataset consists of images in RGBD format.

[0052] The second step involves training the initial facial feature extraction model using the first training dataset mentioned above, resulting in a trained facial feature extraction model. The initial facial feature extraction model may be a model that has not yet completed its training.

[0053] As an example, the aforementioned execution entity can train the initial facial feature extraction model using backpropagation based on the first training dataset to obtain the trained facial feature extraction model.

[0054] The third step is to retrain the trained facial feature extraction model based on the second training dataset to obtain the facial feature extraction model.

[0055] Here, a facial detection model can accurately determine relevant facial information, such as facial landmarks. These landmarks allow for further refinement of color and depth information in facial regions of the image. Based on this, a facial feature extraction model can accurately extract facial features.

[0056] Step 203: Based on the above facial feature information, generate the head pose and facial information corresponding to the above image.

[0057] In some embodiments, the executing entity can generate head pose and facial information corresponding to the image based on the facial feature information. The head pose can be the pose of the head displayed in the image. In practice, head pose can be represented by pitch, yaw, and roll angles. The facial information can be the identity information of the facial content displayed in the image. In practice, facial information can be the object identity information corresponding to a registered facial object. For example, facial information can be an object identifier.

[0058] As an example, firstly, the aforementioned execution entity can input facial feature information into a facial keypoint detection model to obtain facial keypoint information. Then, based on the facial keypoint information, a head pose is generated. Next, the facial feature information is input into an identity recognition model to obtain facial information. The facial keypoint detection model can be a neural network model for detecting facial keypoints. The identity recognition model can be a model for recognizing the identity information of the object displayed in the image. For example, the facial keypoint detection model can be a keypoint detection model based on an attention mechanism. The identity recognition model can be a network layer based on multiple concatenated convolutional layers.

[0059] In some optional implementations of certain embodiments, generating the head pose and facial information corresponding to the image based on the facial feature information includes: The aforementioned execution entity can input the facial feature information into a pre-trained head pose generation model to obtain the head pose. The head pose generation model can be a deep learning model that generates the head pose displayed in the image. That is, the head pose generation model can be a headpose model. For example, the head pose generation model can be a WHENet model.

[0060] Optionally, the above head pose generation model is trained through the following steps: The first step involves obtaining the first training image dataset and the second training image dataset mentioned above. The first training image dataset consists of images including color information. The second training image dataset consists of images including both color and depth information. Both the first and second training image datasets are used subsequently for training the head pose generation model. The first training image dataset is in RGB format. The second training image dataset is in RGBD format.

[0061] The second step involves data augmentation of the first training image dataset to generate enhanced image datasets for each target pose. The target pose can be the pose set for the robot to accurately recognize objects in that pose. By augmenting the target pose, the learning strength of the subsequent initial head pose generation model in that pose can be increased, improving the accuracy of object recognition in that pose. In practice, the target pose can be the pose corresponding to a side profile of a face.

[0062] As an example, the aforementioned execution entity can, in order to adapt to multi-pose and large-pose facial information, simulate the face in the lateral position (affine) through data augmentation, increase the side face weight during training, and thus enhance the training image data to obtain enhanced image data.

[0063] The third step involves training the initial head pose generation model using the aforementioned enhanced image dataset to obtain the trained head pose generation model. The initial head pose generation model may be a head pose generation model that has not yet completed training.

[0064] The fourth step is to retrain the trained head pose generation model based on the second training image dataset to obtain the head pose generation model.

[0065] As an example, the aforementioned execution entity can retrain the trained head pose generation model based on the second training image dataset using the Fewshot Training method to obtain the head pose generation model.

[0066] Step 204: In response to determining that the head posture satisfies the posture conditions set based on the robot's perspective, the facial data corresponding to the registered facial object is stored in the facial database corresponding to the robot based on the facial information, so as to realize the facial information input under the head posture.

[0067] In some embodiments, in response to determining that the head posture meets the posture conditions set based on the robot's perspective, the execution entity can store the facial data corresponding to the registered facial object into the robot's facial database based on the facial information, thereby realizing the input of facial information under the head posture. The robot's perspective can be the range of view of the object being photographed by the robot. In specific scenarios, for cases where the robot is shorter than the object, the corresponding robot perspective can be an upward angle range. In practice, the robot perspective can be preset based on the difference between the robot's height and the object's height. In practice, by setting the robot perspective, accurate object recognition under the perspective can be achieved in subsequent object recognition processes. The posture conditions can be the required posture range for the object being photographed. In practice, the posture range can be flexibly set based on the robot's perspective. For example, for a robot perspective of an upward angle range, the corresponding posture range can be the posture range that is exactly within the upward angle range when photographed from the robot's perspective. That is, the posture range can be set based on the robot's perspective range. This ensures that the robot can clearly obtain the facial content of the object. The facial database can be a database storing facial data corresponding to each object. Facial data can include: object identity information, image, and head posture corresponding to the registered facial object.

[0068] Optionally, based on the aforementioned facial information, the facial data corresponding to the aforementioned registered facial object is stored in the facial database corresponding to the aforementioned robot, so as to realize the facial information input under the aforementioned head posture, including: The first step is to determine the head pose corresponding to the above facial data, which will be used as the target head pose.

[0069] The second step is to store the aforementioned facial data into the facial data group corresponding to the aforementioned facial database. The head pose range corresponding to the facial data group includes the target head pose.

[0070] Here, facial data is grouped and stored according to the head posture range corresponding to the head posture, so as to facilitate the subsequent management of facial data and the filling of facial data.

[0071] As an example, firstly, the aforementioned executing entity can package the object's identity information, head pose, and image to obtain packaged data. Then, the packaged data is added to the data with the same object identity information in the corresponding facial database to obtain the latest facial data corresponding to the registered facial object.

[0072] like Figure 3 The diagram shown illustrates the process of facial data entry.

[0073] like Figure 3As shown, firstly, an image of the currently registered face is captured from an image data source. This image includes RGB channels and a Depth channel (i.e., the channel corresponding to depth information). Then, face detection and keypoint detection are performed based on the channel information corresponding to the RGB channels in the image. Next, based on the face detection and keypoint detection results, face depth information related to the face is filtered from the channel information corresponding to the Depth channel. Then, the face depth information and the face RGB information corresponding to the face detection results are stitched together to obtain face stitching information. Next, based on the above face stitching information, a head pose (pitch, raw, roll) is generated. Then, face feature information is extracted from the face stitching information. Finally, in response to determining that the above head pose satisfies the pose conditions set based on the robot's perspective, the face feature information corresponding to the registered face is stored in the robot's corresponding face database, thereby realizing the face information input under the above head pose. Here, the pose conditions can be "pitch>20||abs(raw)>90". The attitude ranges include: the attitude range corresponding to BigLeft raw < -54 (i.e., the maximum left offset of the yaw angle is less than -54), the attitude range corresponding to Left raw >= -54 && raw < -18 (i.e., the maximum left offset of the yaw angle is greater than or equal to -54 and the left yaw angle is less than -18), the attitude range corresponding to Front -18 <= raw <= 18 (i.e., the yaw angle is greater than -18 forward and less than or equal to 18 forward), the attitude range corresponding to Right raw -18 <= 54 && raw > 18 (i.e., the yaw angle is greater than the right and less than or equal to -18 and the right yaw angle is greater than 18), and the attitude range corresponding to BigRight raw > 54 (i.e., the maximum right yaw angle is greater than 54).

[0074] In some optional implementations of certain embodiments, the robot performs real-time facial tracking during the multi-pose facial data acquisition process. Specifically, during facial registration, the robot tracks the face of the registered subject in real time to ensure that no dummies are encountered during the entire multi-pose acquisition process. The multi-pose facial data acquisition process can involve acquiring facial data under different head postures. Through multi-pose facial data acquisition, the robot can achieve faster and more accurate identity recognition when subsequently performing object identification.

[0075] As an example, in the process of capturing facial data in multiple poses, object tracking algorithms can be used to achieve facial tracking. In practice, object tracking algorithms can be Mean Shift algorithms or Siamese networks.

[0076] Optionally, after step 204, the steps further include: The first step involves generating re-detected facial feature information in response to a facial tracking interruption during facial data entry, provided that the facial registration object is detected again within a preset interruption duration. Facial tracking interruption can occur when the corresponding face is not tracked. The preset interruption duration is the maximum allowed duration of tracking interruption during facial tracking. If the facial tracking interruption duration exceeds the preset duration, multi-pose facial recording for the facial registration object is interrupted. If the facial tracking interruption duration does not exceed the preset duration, multi-pose facial recording for the facial registration object is paused to observe whether the facial registration object is re-detected within the preset interruption duration. The re-detected facial feature information can be the semantic content of the facial features in the newly acquired image of the facial registration object. This facial feature information can be in vector form.

[0077] As an example, firstly, in response to the re-detection of the aforementioned registered facial object within a preset disconnection period, an image corresponding to the re-detected registered facial object is acquired. Then, facial feature information of the object corresponding to the captured image is generated. Here, the method for generating the object's facial feature information is the same as the method for generating facial feature information.

[0078] The second step involves comparing the facial feature information of the aforementioned object with the facial data corresponding to the aforementioned registered facial object stored in the aforementioned facial database to obtain comparison information. The comparison information can be one of the following: indicating no error in the comparison, or indicating an error in the comparison.

[0079] As an example, firstly, the aforementioned execution entity can determine the vector similarity between the facial feature information of the object and the facial feature information in the corresponding facial data. Then, in response to determining that the vector similarity is greater than the target similarity, it generates alignment information indicating that the comparison is correct. In response to determining that the vector similarity is not greater than the target similarity, it generates alignment information indicating that the comparison is incorrect.

[0080] The third step is to continue the multi-pose facial data entry for the above-mentioned facial registration object after confirming that the above-mentioned comparison information is correct.

[0081] As an example, following the next input pose, continue to perform multi-pose facial data input for the aforementioned facial registration object.

[0082] Optionally, in response to determining that the above comparison information representation is incorrect, the multi-pose facial data entry for the above-mentioned facial registration object is stopped.

[0083] Here, by performing facial tracking during the multi-pose facial data entry process, the authenticity of the subject being entered can be ensured, and the situation of entering fake facial information can be avoided.

[0084] In some optional implementations of certain embodiments, after step 204, the steps further include: The first step is to acquire a first facial image corresponding to the first object upon detection during the interaction. The interaction process can be between a robot and an object. For example, if the object is a person, the interaction process can be between a robot and a person. The interaction process can be a voice interaction. If the object is a pet, the interaction process can be between a robot and a pet. The interaction process can be the robot feeding the pet. The first facial image can be an image captured by the robot of the first object during the interaction. The first object can be the object to be identified for the first time during the interaction.

[0085] The second step involves determining whether a facial object exists in the target facial database that satisfies the target similarity condition in terms of feature similarity to the first facial image, given that the head pose corresponding to the first facial image meets the aforementioned pose condition. The head pose corresponding to the first facial image can be determined using a head pose generation model. Feature similarity can be the vector similarity between the facial feature information corresponding to the first facial image and the various facial feature information stored in the target facial database. In practice, vector similarity can be cosine similarity. The target similarity condition can be the facial object with the highest corresponding feature similarity. The target facial database can be a database storing complete facial data for each object in various poses.

[0086] Thirdly, in response to determining the existence of a target facial recognition object, the information that the first object is the target facial recognition object is confirmed, and during the interaction process, object tracking is performed for the target facial recognition object. The target facial recognition object can be a facial recognition object that meets the target similarity condition.

[0087] In some optional implementations of certain embodiments, the steps further include: In the first step, in response to the disconnection of object tracking for the target facial input object and the subsequent detection of a second object, the executing entity can acquire a second facial image corresponding to the second object. The second object can be the object detected for the first time after the first object's tracking was disconnected. The second facial image can be a facial image captured by the robot for the second object.

[0088] The second step is for the executing entity to re-determine whether there exists a facial input object in the target facial database whose feature similarity with the second facial image meets the target similarity condition.

[0089] Third, in response to the determination that the aforementioned target facial recognition object exists, the aforementioned execution entity can continue to perform object tracking for the aforementioned target facial recognition object.

[0090] like Figure 4 The image shows the facial recognition process.

[0091] like Figure 4 As shown, when the upper-layer application wakes up to perform face recognition, the target face image is acquired. The target face image can be the image to be recognized. The target face image includes RGB channels and Depth channels. Then, face stitching information is generated through face detection and key point detection (see details in [link to documentation]). Figure 3 Then, facial identity information is retrieved from the face database (i.e., the face identification database) where the similarity between the captured facial information and the target face is higher than the target similarity. Next, after confirming the facial identity information, a face tracking algorithm is used to track the face in real time to achieve synchronous face tracking. During tracking, if tracking is lost or congested, the similarity calculation needs to be re-executed to re-verify the facial identity information. Thus, face recognition and facial identity tracking are achieved. In addition, during interaction with the corresponding face, the system supports filtering facial data and storing the filtered facial data as enhanced facial data in the face database to update the face database.

[0092] In some alternative implementations of certain embodiments, the robot moves around the face registration object to record facial data in real time in multiple poses.

[0093] The above-described embodiments of this disclosure have the following beneficial effects: Through the facial data entry methods of some embodiments of this disclosure, facial data under different head postures can be accurately and efficiently entered into the robot, facilitating efficient interaction between the object and the robot. Specifically, the reason for the insufficient accuracy and efficiency of interaction with the robot is that facial recognition technology in robots typically relies on two-dimensional visible light RGB images. However, this method suffers from low dimensionality, making it difficult to effectively acquire the three-dimensional geometric features of the face. Furthermore, it is extremely sensitive to lighting conditions; strong light, weak light, backlight, and uneven lighting can all significantly reduce the accuracy of recognition. In addition, RGB images of the face are often head-on facial images, which often lead to the robot's inability to effectively recognize faces, resulting in inefficient interaction. Therefore, the facial data entry method of some embodiments of this disclosure first, in response to the current execution of facial data entry for a registered facial object, acquires an image of the currently captured face. This image includes color information and depth information. Here, by capturing images that include depth information, three-dimensional facial data corresponding to the registered facial object can be effectively recorded during subsequent facial data entry, avoiding the problem of low dimensionality and effectively acquiring the three-dimensional geometric features of the face. Furthermore, by setting facial data that includes depth information, the problem of low recognition accuracy under sensitive lighting conditions can be effectively avoided. Then, based on the aforementioned color and depth information, the facial feature information corresponding to the image can be accurately determined to obtain the semantic content of the facial features corresponding to the registered facial object. Next, based on the aforementioned facial feature information, the head pose and facial information corresponding to the image are generated. Here, by generating the head pose, facial images that meet the requirements for accurate facial recognition by the robot are selected, enabling the robot to accurately and efficiently extract more effective facial feature information during subsequent facial recognition, making the subsequent recognition process more accurate and efficient. In addition, by generating facial information, the robot can effectively understand the facial identity information corresponding to the registered facial object, thereby achieving accurate generation of subsequent facial data. Finally, in response to the determination that the head posture satisfies the posture conditions set based on the robot's perspective, the facial data corresponding to the registered facial object is stored in the robot's facial database according to the facial information, thereby realizing the facial information input under the head posture. Here, by inputting facial information under the head posture, the robot can subsequently achieve accurate recognition of the registered facial object under that head posture. Furthermore, by inputting facial data under multiple head postures through the robot, the robot can achieve accurate recognition under multiple head postures during the subsequent recognition process, greatly improving recognition efficiency and accuracy.In summary, by including depth information in the corresponding facial images and recording facial information under various head poses, the accuracy and efficiency of subsequent facial recognition by the robot can be improved.

[0094] Further reference Figure 5 The diagram illustrates flow 500 of some other embodiments of the facial data entry method according to the present disclosure. This facial data entry method, applied to a robot, includes the following steps: Step 501: In response to the current execution of facial data entry for the facial registration object, obtain the image currently being captured on the aforementioned facial registration object.

[0095] Step 502: Determine the facial feature information corresponding to the above image based on the color information and depth information.

[0096] Step 503: Based on the above facial feature information, generate the head pose and facial information corresponding to the above image.

[0097] Step 504: In response to determining that the head posture satisfies the posture conditions set based on the robot's perspective, the facial data corresponding to the registered facial object is stored in the facial database corresponding to the robot based on the facial information, so as to realize the facial information input under the head posture.

[0098] In some embodiments, the specific implementation of steps 501-504 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 201-204 in the corresponding embodiments will not be repeated here.

[0099] Step 505: In response to performing facial data addition to the target facial database, during the interaction with the aforementioned facial registration object, obtain the set of interactive captured images corresponding to the aforementioned facial registration object.

[0100] In some embodiments, in response to performing facial data addition to a target facial database, the executing entity (e.g., Figure 1 The robot 101 shown can acquire a set of interactive images corresponding to the aforementioned registered facial object during interaction. The target facial database is a database of facial data for which the registered facial object has completed multi-pose facial data entry. Adding facial data can be an operation of adding a new image corresponding to the entered facial object to the target facial database. Interactive images can be images captured by the robot on the registered facial object during the interaction. For example, if the interaction process involves the robot playing a game with the registered facial object, the set of interactive images can be acquired during the game.

[0101] Step 506: Based on at least one of the following factors—facial distance, shooting lighting, head posture, and shooting time—corresponding to the interactively captured image, select facial data-enhanced images from the aforementioned set of interactively captured images to obtain a set of facial data-enhanced images.

[0102] In some embodiments, the aforementioned execution entity may select facial data-enhanced images from the aforementioned interactive image set based on at least one of the facial distance, shooting illumination, head posture, and shooting time corresponding to the interactively captured image, thereby obtaining a facial data-enhanced image set.

[0103] As an example, the aforementioned execution entity can filter facial data-enhanced images from the aforementioned interactive image set based on at least one of the following filtering conditions: facial distance, shooting lighting, head pose, and shooting time. For instance, the execution entity can filter interactive images from the interactive image set that correspond to facial distances within the target distance range, shooting lighting within the target lighting range, head poses different from the stored head poses, and shooting times within the past week, thus obtaining a facial data-enhanced image set.

[0104] As another example, firstly, based on the lighting conditions corresponding to the interactive captured images, and according to the lighting histogram corresponding to the interactive captured images, images with brightness information higher than the target brightness information and darkness information lower than the target darkness information are removed from the set of interactive captured images, resulting in a set of images after removal. Then, a pre-set set of head pose groups is determined. Next, for each head pose group, the first step is to filter out a first subset of captured images corresponding to the head pose group from the set of images after removal. The second step is to determine a second subset of captured images corresponding to the head pose group in the target face database. The third step is to determine the number of identical captured images between the first and second subsets. The fourth step, in response to determining that the number of images is less than a target value, is to determine the first subset of captured images as a subset of facial data augmentation images. Finally, each of the resulting subsets of facial data augmentation images is determined as a set of facial data augmentation images.

[0105] Step 507: Add the above-mentioned enhanced facial data image set to the facial data corresponding to the above-mentioned registered facial object in the above-mentioned target facial database.

[0106] In some embodiments, the execution entity may add the facial data enhancement image set to the facial data corresponding to the facial registration object in the target facial database.

[0107] from Figure 5 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 5In some corresponding embodiments, the facial data entry method process 500 enriches the facial data of the objects in the database by adding images of the interactive activities corresponding to the facial objects to the target facial database, thereby greatly improving the recognition efficiency of the robot in the subsequent object recognition process.

[0108] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a facial data entry device, which are similar to... Figure 2 Corresponding to the method embodiments shown, this facial data entry device can be specifically applied to various electronic devices.

[0109] like Figure 6 As shown, a facial data entry device 600 includes: an acquisition unit 601, a determination unit 602, a generation unit 603, and a storage unit 604. The acquisition unit 601 is configured to acquire an image of the facial registration object currently being captured, in response to performing facial data entry for a registered facial object. The image includes color information and depth information. The determination unit 602 is configured to determine facial feature information corresponding to the image based on the color and depth information. The generation unit 603 is configured to generate a head pose and facial information corresponding to the image based on the facial feature information. The storage unit 604 is configured to, in response to determining that the head pose satisfies pose conditions set based on a robot's perspective, store the facial data corresponding to the registered facial object in the robot's facial database based on the facial information, thereby realizing facial information entry under the aforementioned head pose.

[0110] In some optional implementations of some embodiments, the robot performs real-time facial tracking during the multi-pose facial data entry process; and the facial data entry device 600 may further include: in response to a facial tracking disconnection during the facial data entry process and the re-detection of the facial registration object within a preset disconnection time, generating re-detected object facial feature information; comparing the object facial feature information with the facial data corresponding to the facial registration object stored in the facial database to obtain comparison information; and in response to determining that the comparison information is correct, continuing to perform multi-pose facial data entry for the facial registration object.

[0111] In some optional implementations of certain embodiments, the facial data entry device 600 may further include: in response to performing facial data addition to a target facial database, during interaction with the facial registration object, acquiring an interactive image set corresponding to the facial registration object, wherein the target facial database is a facial database in which the facial registration object has completed multi-pose facial data entry; filtering facial data enhancement images from the interactive image set based on at least one of facial distance, shooting illumination, head posture, and shooting time corresponding to the interactive image, to obtain a facial data enhancement image set; and adding the facial data enhancement image set to the facial data corresponding to the facial registration object in the target facial database.

[0112] In some optional implementations of some embodiments, the determining unit 602 may be further configured to: input the color information into a pre-trained face detection model to obtain face detection information; based on the face detection information, filter out the depth information corresponding to the face region from the depth information to obtain face depth information; combine the face detection information and the face depth information to obtain face combination information; and input the face combination information into a pre-trained face feature extraction model to obtain face feature information.

[0113] In some optional implementations of certain embodiments, the generation unit 603 may be further configured to: input the aforementioned facial feature information into a pre-trained head pose generation model to obtain a head pose; and train the aforementioned head pose generation model through the following steps: acquiring a first training image dataset and the aforementioned second training image dataset, wherein the first training image dataset is an image including color information, and the second training image dataset is an image including color information and depth information; performing data augmentation on the aforementioned first training image dataset to generate training image data under various target poses, thereby obtaining an augmented image dataset; training the initial head pose generation model based on the aforementioned augmented image dataset to obtain a trained head pose generation model; and retraining the trained head pose generation model based on the aforementioned second training image dataset to obtain a head pose generation model.

[0114] In some optional implementations of some embodiments, the determining unit 602 may be further configured to: in response to detecting a first object during the interaction, acquire a first facial image corresponding to the first object; in response to the head pose corresponding to the first facial image conforming to the pose condition, determine whether there exists a facial input object in the target facial database whose feature similarity with the first facial image meets the target similarity condition; in response to determining that there is a target facial input object, determine that the first object is the information of the target facial input object, and in the interaction, perform object tracking for the target facial input object.

[0115] In some optional implementations of some embodiments, the determining unit 602 may be further configured to: in response to the object tracking for the target facial recognition object being disconnected and the second object being detected again, acquire the second facial image corresponding to the second object; re-determine whether there is a facial recognition object in the target facial database whose feature similarity with the second facial image meets the target similarity condition; and in response to determining that there is a target facial recognition object, continue to perform object tracking for the target facial recognition object.

[0116] It is understandable that the units recorded in the facial data entry device 600 are related to the reference. Figure 2 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the facial data entry device 600 and the units contained therein, and will not be repeated here.

[0117] The following is for reference. Figure 7 It shows a schematic diagram of the structure of an electronic device 700 suitable for implementing some embodiments of the present disclosure. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0118] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory 702 or a program loaded from a storage device 708 into a random access memory 703. The random access memory 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, the read-only memory 702, and the random access memory 703 are interconnected via a bus 704. An input / output interface 705 is also connected to the bus 704.

[0119] Typically, the following devices can be connected to the input / output interface 705: input devices 706 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 708 including, for example, magnetic tape, hard disk, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 7 Each box shown can represent a device or multiple devices as needed.

[0120] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 709, or installed from a storage device 708, or installed from a read-only memory 702. When the computer program is executed by the processing device 701, it performs the functions defined in the methods of some embodiments of this disclosure.

[0121] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to currently performing facial data recording for a facial registration object, acquire an image currently captured of the facial registration object, wherein the image includes color information and depth information; determine facial feature information corresponding to the image based on the color information and depth information; generate a head pose and facial information corresponding to the image based on the facial feature information; and, in response to determining that the head pose satisfies pose conditions set based on the robot's perspective, store the facial data corresponding to the facial registration object in the facial database corresponding to the robot based on the facial information, thereby realizing facial information recording under the aforementioned head pose.

[0124] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0126] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a determination unit, a generation unit, and a storage unit. The names of these units do not necessarily limit the unit itself; for example, the generation unit may also be described as "a unit that generates the head pose and facial information corresponding to the above-mentioned image based on the above-mentioned facial feature information."

[0127] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0128] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the facial data entry methods described above.

[0129] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A facial data entry method, applied to a robot, comprising: In response to the current execution of facial data entry for a facial registration object, an image currently captured on the facial registration object is obtained, wherein the image is an image including color information and depth information; Based on the color and depth information, the facial feature information corresponding to the image is determined; Based on the facial feature information, generate the head pose and facial information corresponding to the image; In response to determining that the head posture meets the posture conditions set based on the robot's perspective, the facial data corresponding to the facial registration object is stored in the facial database corresponding to the robot according to the facial information, so as to realize the facial information input under the head posture.

2. The method according to claim 1, wherein, The robot performs real-time facial tracking during the multi-pose facial data acquisition process; and The method further includes: In response to a facial tracking interruption during facial data entry and the re-detection of the registered facial object within a preset interruption time, facial feature information of the re-detected object is generated. The facial feature information of the object is compared with the facial data corresponding to the registered facial object stored in the facial database to obtain comparison information; In response to the determination that the comparison information is correct, the multi-pose facial data entry for the facial registration object continues.

3. The method according to claim 1, wherein, The method further includes: In response to adding facial data to the target facial database, during the interaction with the facial registration object, an interactive image set corresponding to the facial registration object is obtained, wherein the target facial database is a facial database in which the facial registration object has completed the recording of facial data in multiple poses; Based on at least one of the following factors—facial distance, shooting lighting, head posture, and shooting time—corresponding to the interactively captured images, facial data-enhanced images are selected from the set of interactively captured images to obtain a set of facial data-enhanced images. The facial data augmentation image set is added to the facial data corresponding to the facial registration object in the target facial database.

4. The method according to claim 1, wherein, The step of determining the facial feature information corresponding to the image based on the color information and depth information includes: The color information is input into a pre-trained face detection model to obtain face detection information; Based on the facial detection information, the depth information corresponding to the facial region is filtered out from the depth information to obtain the facial depth information; The facial detection information and the facial depth information are combined to obtain combined facial information; The facial combination information is input into a pre-trained facial feature extraction model to obtain facial feature information.

5. The method according to claim 1, wherein, The step of generating the head pose and facial information corresponding to the image based on the facial feature information includes: The facial feature information is input into a pre-trained head pose generation model to obtain the head pose; and The head pose generation model is trained through the following steps: Obtain a first training image dataset and a second training image dataset, wherein the first training image dataset consists of images including color information, and the second training image dataset consists of images including color information and depth information; Data augmentation is performed on the first training image dataset to generate training image data under each target pose, resulting in an augmented image dataset. Based on the enhanced image dataset, the initial head pose generation model is trained to obtain the trained head pose generation model. Based on the second training image dataset, the trained head pose generation model is retrained to obtain a head pose generation model.

6. The method according to claim 1, wherein, The method further includes: In response to detecting a first object during the interaction, a first facial image corresponding to the first object is acquired; In response to the head pose corresponding to the first facial image conforming to the pose condition, it is determined whether there is a facial input object in the target facial database whose feature similarity with the first facial image meets the target similarity condition; In response to determining the existence of a target facial recognition object, determining the information that the first object is the target facial recognition object, and performing object tracking for the target facial recognition object during the interaction.

7. The method according to claim 6, wherein, The method further includes: In response to the disconnection of object tracking for the target facial input object and the detection of a second object again, a second facial image corresponding to the second object is acquired; Re-determine whether there exists a facial input object in the target facial database whose feature similarity with the second facial image meets the target similarity condition; In response to the determination that the target facial recognition object exists, object tracking for the target facial recognition object continues.

8. The method according to claim 1, wherein, The robot moves around the face registration object to record facial data in real time in multiple poses.

9. A facial data input device, applied to a robot, comprising: The acquisition unit is configured to acquire an image of the currently captured face registration object in response to the current execution of facial data recording for the face registration object, wherein the image is an image including color information and depth information; The determining unit is configured to determine facial feature information corresponding to the image based on the color information and depth information; The generation unit is configured to generate head pose and facial information corresponding to the image based on the facial feature information; The storage unit is configured to, in response to determining that the head posture satisfies the posture conditions set based on the robot's perspective, store the facial data corresponding to the facial registration object into the facial database corresponding to the robot, so as to realize the facial information input under the head posture.

10. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.

11. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.

12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Face recognition method and device, robot and storage medium

    CN112488053A

  • Sight tracking data acquisition method and device, computer storage medium and terminal

    CN117666774A

  • Face identification system

    US20040197014A1

  • Methods, systems, and media for swapping faces in images

    US20110123118A1