Nonlinear Mapping-Based Facial Robot Expression Cooperative Control Method and System

Through the method of collaborative control of nonlinear mapping and multi-server, the facial key points of the real face image are obtained, the pixel deviation is calculated and the server angle deviation is generated, which solves the problem of insufficient realistic facial expressions of humanoid robots, and achieves high accuracy and authentic facial expression simulation.

CN120126201BActive Publication Date: 2025-07-25HUAZHONG UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510599607.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-25
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the prior art, the facial expression simulation of humanoid robots has insufficient realism, the expression changes are stiff, and the errors are high when driving different faces, making it difficult to achieve high accuracy anthropomorphic expression.

Method used

Using a method based on nonlinear mapping and multi-server collaborative control, by obtaining facial key points of real face images, calculating pixel deviations and performing nonlinear mappings, generating server angle deviations, and configuring multiple servers to drive key parts of the face to perform expression anthropomorphic expressions.

Benefits of technology

It improves the accuracy and generalization of facial expression simulation, adapts to different faces, reduces simulation deviations, and achieves a more realistic facial expression expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126201B_ABST
    Figure CN120126201B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field related to robot facial expressions, and specifically relates to a method and system for anthropomorphic expression of robot facial expressions based on non-linear mapping and multi-server collaborative control, including: obtaining a plurality of facial key points of a real human face image corresponding to each preset key part of the facial robot, calculating the pixel deviation between the plurality of facial key points and the reference key points preset for quantifying the expression amplitude of the key part, subtracting the preset original state value from the pixel deviation to obtain a pixel deviation representing the relative preset original state of the key part; mapping the pixel deviation to an angular deviation representing the relative original position of each servo for driving the key part through non-linear mapping, and realizing anthropomorphic expression of facial expressions based on the angular deviations; the non-linear mapping method for each part depends on the positions of the servos driving the part; when mapping, the pixel deviation is first normalized based on the size of the human face. The present invention can improve the expression accuracy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to robot facial expressions, and more specifically, relates to a method for anthropomorphic expression of robot facial expressions based on non - linear mapping and multi - server collaborative control. Background Art

[0002] In humanoid embodied robots, emotional expression is the core. As the carrier of emotional output, the face, a facial expression robot can simulate emotions such as smiling and frowning, improving the naturalness of human - robot interaction.

[0003] One of the reasons for designing humanoid robots to be human - shaped is to make them more easily received and understood. The current design of humanoid robots mainly focuses on the hands and legs, while the face, as a key part of emotional expression, has relatively less research. Currently, one of the most important ways to achieve facial expression is through an electronic screen display. The electronic screen can directly generate various expressions and can express various emotions for the robot. However, the electronic screen does not have the three - dimensional sense of the human face, so it will create a sense of distance.

[0004] Therefore, designing a realistic three - dimensional facial expression robot can effectively solve the problem of insufficient robot realism. By using a mechanical structure design to simulate the movement of facial muscles and inputting control signals into the control units of facial muscles, the movement of facial muscles can be achieved, thereby realizing the generation of facial robot expressions. If a fixed linear signal is output to the facial robot, although the target expression can be achieved, the expression change of the facial robot will be very rigid, mainly reflected in the fixed speed of the facial movement unit (AU) controlled by the control unit.

[0005] In an actual scenario, for a facial robot driven by imitating a real human face, the error determination of its joint points is very important. Although the given value can be used to achieve the follow - up movement of the facial robot's AU, the tiny jitter error of the human face detecting the facial AU will also affect the movement of the AU. And robots driven by different human faces cannot be achieved through a single linear mapping. The size of the face and the position between different facial AUs will introduce different errors, resulting in a relatively high error rate in imitation driving. Summary of the Invention

[0006] In view of the above - mentioned defects or improvement requirements of the prior art, the present invention provides a method and system for anthropomorphic expression of robot facial expressions based on non - linear mapping and multi - server collaborative control, aiming to improve the accuracy of imitation driving.

[0007] To achieve the above object, according to one aspect of the present invention, a method for anthropomorphic expression of robot facial expressions based on non - linear mapping and multi - server collaborative control is provided, including:

[0008] Obtain multiple facial key points corresponding to each preset key part of the facial robot, calculate the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part, subtract the preset original state value of the key part from the pixel deviation to obtain the pixel deviation representing the relative preset original state of the key part; through non-linear mapping, map the pixel deviation relative to its preset original state to the angular deviation representing the relative original position of each servo for driving the key part, and generate driving data by synthesizing the angular deviations corresponding to each servo, so as to realize anthropomorphic expression of the robot's facial expression through multiple servos; wherein, the non-linear mapping method corresponding to each key part depends on the setting position of each servo for driving the key part relative to the key part; and when performing non-linear mapping, first normalize the pixel deviation relative to the preset original state based on the face size.

[0009] Further, the preset key parts of the facial robot include the left pupil and the right pupil;

[0010] By taking the average of the two corner positions of each eye as the orbital center position of the eye, and taking the orbital center positions of the two eyes as reference points, calculate the deviation of the pupil of each eye relative to its own orbital center position, and the sum of the two deviations is used as the pixel deviation shared by the two pupils representing the relative original state of each pupil; there are two servos for driving each pupil, which are arranged side by side at the rear position corresponding to the pupil inside the facial robot, then the non-linear mapping method corresponding to each pupil shared by the two pupils is:

[0011]

[0012] In the formula, is the angular deviation corresponding to the pixel deviation of each pupil in the horizontal direction; is the angular deviation corresponding to the pixel deviation of each pupil in the vertical direction; is the horizontal pixel value of the left eye pupil determined by calculation in the real human face image; is the horizontal pixel value of the right eye pupil determined by calculation in the real human face image; 、 、 、 are the horizontal pixel coordinates of the dlib detection key points 40, 37, 46, and 43 in the real human face image respectively; represents the width of the human face in the real human face image; 、 、 、 They are the vertical pixel coordinates of dlib-detected key points 40, 37, 46, and 43 in a real face image, respectively; represents the length of the face in a real face image; is the vertical pixel value of the left eye pupil determined by calculation in a real face image, is the vertical pixel value of the right eye pupil determined by calculation in a real face image;

[0013] The preset key parts of the facial robot include the upper eyelids and the lower eyelids;

[0014] The sum of the pixel differences between the two upper eyelids and the two lower eyelid key points in the vertical direction is used as the pixel deviation representing the relative state of the eyelids from their original state;

[0015] One servo for driving the upper eyelids and one for driving the lower eyelids are respectively arranged at the rear positions corresponding to the two eyes inside the facial robot. Then, the non-linear mapping methods corresponding to each eyelid are as follows:

[0016]

[0017] Among them, is the angular deviation corresponding to the pixel deviation of the eyelid in the vertical direction, , , , They are the vertical pixel coordinates of dlib-detected key points 39, 44, 41, and 48 in a real face image, respectively; is the length of the face in a real face image.

[0018] Furthermore, the preset key part of the facial robot includes the upper lip;

[0019] Taking the nose tip key point as a reference point, calculate the pixel deviation between the nose tip key point and the upper lip key point in the vertical direction, and subtract the original state value of the upper lip as the pixel deviation representing the relative state of the upper lip from its original state; One servo for driving the upper lip is arranged at the upper rear position corresponding to the upper lip inside the facial robot. Then, the non-linear mapping method corresponding to the upper lip is as follows:

[0020]

[0021] In the formula, is the angular deviation corresponding to the pixel deviation of the upper lip in the vertical direction; , , , They are the vertical pixel coordinates of dlib-detected key points 33, 35, 62, and 64 in a real face image, respectively. is the length of the face in a real face image. is a preset original state value of the upper lip.

[0022] The preset key parts of the facial robot further include the lower lip.

[0023] Taking the chin key point as a reference point, calculate the pixel deviation between the chin key point and the lower lip key point in the vertical direction, and subtract the original state value of the lower lip from it as the pixel deviation representing the relative state of the lower lip from its original state. There is one servo for driving the lower lip, which is set at the lower rear position corresponding to the lower lip inside the facial robot. Then the non-linear mapping method corresponding to the lower lip is:

[0024]

[0025] In the formula, is the angular deviation corresponding to the pixel deviation of the lower lip in the vertical direction. and and and and and They are the vertical pixel coordinates of dlib-detected key points 8, 9, 10, 59, 58, and 57 in a real face image, respectively. is the preset original state value of the lower lip. is the length of the face in a real face image.

[0026] Furthermore, the preset key parts of the facial robot include the left corner of the mouth and the right corner of the mouth.

[0027] Calculate the pixel difference between the left corner of the mouth key point and the right corner of the mouth key point as the horizontal pixel deviation representing the relative state of the corners of the mouth from their original state. Taking the two inner corner key points as reference points, calculate the pixel difference between the left corner of the mouth key point and the inner corner key point above it, and the pixel difference between the right corner of the mouth key point and the inner corner key point above it. The sum of the two pixel differences is used as the pixel deviation representing the relative state of the corners of the mouth from their original state. There are two servos for driving the left corner of the mouth and the right corner of the mouth respectively. The two servos corresponding to each corner of the mouth are set at the rear position corresponding to this corner of the mouth inside the facial robot. Then the non-linear mapping method shared by the left corner of the mouth and the right corner of the mouth corresponding to each corner of the mouth is:

[0028]

[0029] In the formula, is the angular deviation corresponding to the pixel deviation of the corners of the mouth in the horizontal direction. is the angular deviation corresponding to the pixel deviation of the corners of the mouth in the vertical direction; , , , are respectively the horizontal pixel coordinates of dlib detected key points 55, 65, 49 and 61 in the real face image; , , , are respectively the vertical pixel coordinates of dlib detected key points 40, 43, 49 and 55 in the real face image; is the width of the face in the real face image.

[0030] Furthermore, the preset key parts of the facial robot include the mouth;

[0031] Calculate the pixel difference between the key point at the lower edge of the nose and the key point of the chin in the vertical direction as the pixel deviation representing the relative state of the mouth to its original state; There is one servo for driving the opening and closing of the mouth, which is set at the rear position corresponding to the mouth inside the facial robot, and the non-linear mapping method corresponding to the mouth is:

[0032]

[0033] In the formula, is the angular deviation corresponding to the pixel deviation of the mouth in the horizontal direction; , , , , , are respectively the vertical pixel coordinates of dlib detected key points 33, 34, 35, 8, 9 and 10 in the real face image; is the width of the face in the real face image.

[0034] Furthermore, the preset key parts of the facial robot include the center of the eyebrows and the eyebrows;

[0035] By taking the key point of the eye corner as the reference key point, calculate the pixel deviation between the key point of the center of the eyebrows and the key point of the eye corner below it in the vertical direction, subtract the preset original state value of the eyebrows as the pixel deviation representing the relative state of the center of the eyebrows to its original state, and calculate the pixel deviation between the key point of the end of the eyebrow and the key point of the eye corner below it as the pixel deviation representing the relative state of the end of the eyebrow to its original state; There is one servo for driving the center of the eyebrows and the end of the eyebrows respectively, which are set at the rear positions corresponding to the two eyebrows inside the facial robot, and the non-linear mapping methods corresponding to the center of the eyebrows and the end of the eyebrows are:

[0036]

[0037] In the formula, represents the angular deviation corresponding to the pixel deviation between the center of the eyebrows and the eye corners in the vertical direction; represents the angular deviation corresponding to the pixel deviation between the end of the eyebrows and the eye corners in the vertical direction; represents the preset original state value of the eyebrows; , , , , , , and are respectively the vertical pixel coordinates of the dlib detected key points 22, 23, 40, 43, 18, 27, 37 and 46 in the real face image; represents the length of the face in the real face image.

[0038] Furthermore, the preset key part of the facial robot includes the nose;

[0039] Taking the two inner eye corner key points as reference points, calculate the pixel deviation between the two inner eye corner key points and the nose tip key point in the vertical direction, and subtract the preset original state value of the nose as the pixel deviation representing the nose relative to its original state; there is one servo for driving the nose, which is set at the rear position corresponding to the nose inside the facial robot, then the non - linear mapping method corresponding to the nose is:

[0040]

[0041] In the formula, is the angular deviation corresponding to the pixel deviation of the nose in the vertical direction, is the preset original state value of the nose; , , are the vertical pixel coordinates of the dlib detected key points 40, 43 and 31 in the real face image; is the length of the face in the real face image.

[0042] According to another aspect of the present invention, there is provided a robot facial expression anthropomorphic expression system based on non - linear mapping and multi - servo collaborative control, which is characterized in that it is used to execute a robot facial expression anthropomorphic expression method based on non - linear mapping and multi - servo collaborative control as described above, including a PC side and a facial robot; wherein, the PC side is used to obtain and generate driving data based on the real face image, and the facial robot is used to realize the anthropomorphic expression of the robot facial expression based on the driving data through multiple servos.

[0043] Furthermore, the facial robot includes a nose muscle control structure;

[0044] Among them, a screw hole is provided at the bottom end of the nose muscle control structure for fixing the entire nose muscle control structure; a forward protruding link is provided at the front end of the nose muscle control structure for connecting the nose muscle through the screw hole; the servo for controlling the nose is horizontally placed directly behind the forward protruding link, and the part of the nose muscle control structure on one side of the robot is used to connect the servo rotating link on one side of the servo; the part of the nose muscle control structure on the other side of the robot is movably connected to the other side of the servo through a circular hole structure for strengthening the rotating shaft of the servo rotating link; the servo rotation plane is a vertical plane in space.

[0045] Furthermore, the facial robot includes an eyebrow control structure; the eyebrow control structure is composed of a servo bearing structure, a center of the eyebrows movement link structure, and an end of the eyebrows movement link structure.

[0046] Among them, the servo bearing structure is used to fix two servos that are horizontally placed behind the eyebrows and used to control the movement of the eyebrows; one side of the center of the eyebrows movement link structure is connected to the servo rotating link on one side of one of the servos, and the other side of the center of the eyebrows movement link structure is movably connected to the other side of the other servo through a circular hole structure for strengthening the rotating shaft of the servo rotating link; one side of the end of the eyebrows movement link structure is connected to the servo rotating link on one side of the other servo, and the other side of the end of the eyebrows movement link structure is movably connected to the other side of one of the above servos through a circular hole structure for strengthening the rotating shaft of the servo rotating link; the rotation plane of each servo is a vertical plane in space.

[0047] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the technical solutions provided by the present invention mainly have the following beneficial effects:

[0048] 1. The present invention proposes a method for anthropomorphic expression of robot facial expressions based on nonlinear mapping and multi-server collaborative control, wherein multiple facial key points of a real face image corresponding to each preset key part of the facial robot are selected, and the pixel deviation between the multiple facial key points and the preset reference key points used to quantify the expression amplitude of the key part is calculated as the amplitude of the expression of the key part, wherein the reference key points are selected and normalized based on the face size during nonlinear mapping, thereby reducing the simulation deviation caused by the difference between different faces, adapting to different faces, and ensuring that the accuracy of simulation driving of different faces is high and the generalization is high. In addition, nonlinear mapping is used to convert the pixel information of multiple key points into the angle deviation of the server relative to its original position, and different key parts are configured with their own servers. One key part is driven by one or more servers, which can avoid the error introduced by the size of the face and the relative position between different facial AUs. Therefore, in general, the method of this embodiment can improve the accuracy of imitation driving.

[0049] 2. The present invention further preferably designs the number, position, reference key points and nonlinear mapping formula of the servos corresponding to different key parts, and more specifically realizes the anthropomorphic expression of facial expressions applicable to any human face.

[0050] 3. The present invention innovatively designs a facial robot structure, and sets up the control structure of the key parts of the facial robot by simulating muscle pulling, including the nose muscle control structure and the eyebrow control structure, and combines it with the eyebrow control structure of the driving data obtained by the PC based on the mapping function, so that fewer servers can be used to achieve more realistic facial expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 The present invention provides a flowchart of a method for anthropomorphic expression of robot facial expressions based on nonlinear mapping and multi-server collaborative control.

[0052] Figure 2 It is a comparison diagram of dlib detection key points and facial robot key parts provided by an embodiment of the present invention.

[0053] Figure 3 It is a flowchart of the facial robot software implementation provided by an embodiment of the present invention.

[0054] Figure 4 1 is a structural diagram of a facial shell provided by an embodiment of the present invention.

[0055] Figure 5 It is a structural diagram of a cheek and nose shell provided by an embodiment of the present invention.

[0056] Figure 6It is the nose muscle control structure diagram provided by the embodiment of the present invention.

[0057] Figure 7 It is the eyebrow control structure diagram provided by the embodiment of the present invention.

[0058] Figure 8 It is the eye-mouth connection structure diagram provided by the embodiment of the present invention.

[0059] Figure 9 It is the overall head connection structure diagram provided by the embodiment of the present invention. Detailed implementation manners

[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0061] Embodiment 1

[0062] A method for anthropomorphic expression of a robot's facial expression based on non-linear mapping and multi-server collaborative control, as Figure 1 shown, includes:

[0063] Obtaining multiple facial key points of a real human face image corresponding to each preset key part of a facial robot, calculating the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part, and then subtracting the preset original state value of the key part from the pixel deviation to obtain a pixel deviation representing the relative preset original state of the key part; through non-linear mapping, mapping the pixel deviation relative to its preset original state to an angular deviation representing the relative original position of each servo for driving the key part; generating drive data by synthesizing the angular deviations corresponding to each servo to achieve anthropomorphic expression of the robot's facial expression through multi-servers; wherein, the non-linear mapping method corresponding to each key part depends on the relative position of each servo for driving the key part to the key part; and when performing non-linear mapping, first normalize the pixel deviation relative to the preset original state based on the size of the human face.

[0064] Select multiple facial key points of a real human face corresponding to each preset key part of the facial robot, and calculate the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part as the amplitude of the expression of the key part. Among them, the reference key points are selected and normalized based on the face size during non-linear mapping, so the simulation deviation caused by different human face differences can be reduced, adapted to different human faces, and ensure that the accuracy of simulation driving for different human faces is relatively high and the generalization is high. In addition, non-linear mapping is used to convert the pixel information of multiple key points into the angular deviation of the servo relative to its original position, and a servo is configured for each different key part, and one or more servos are used to drive one key part, which can avoid errors introduced by the size of the face and the relative positions between different facial AUs. Therefore, generally speaking, the method of this embodiment can improve the accuracy of imitation driving.

[0065] As Figure 2 shown, from left to right are the dlib-detected key points, the key parts of the facial expression robot, and the overall structure of the facial robot. Among them, the blue part in the overall structure diagram of the facial robot represents the servo. In addition, referring to Figure 2 the display of the dlib-detected key points and the key parts of the facial expression robot, linear mapping cannot be directly performed during the process of mapping the real human face and the key points of the facial expression robot. The main reasons are as follows: 1) The movement direction of the facial key points is different from that of the servo; 2) One facial key point may be driven by multiple servos; 3) Multiple facial key points may be driven by one servo. At the same time, in the real human face, by calculating the error between the pixel positions of the facial key points when there is no expression and the same key points when there is an expression, although the error can be calculated by setting the initial value, there will be a large error due to different human faces. Therefore, in this embodiment, a separate and preferably mapping design and error calculation method are carried out for each key part of the robot to be applicable to any human face.

[0066] For any human face, first determine its face size for subsequent normalization calculations of each non-linear mapping. The calculation formula is as follows:

[0067]

[0068] where, H represents the human face, represents the width of the human face, represents the length of the human face, represents the AU (facial action unit) numbered 17 on the human face in dlib key point detection, and the subscripts x and y respectively represent the horizontal and vertical pixel values of the AU collected by dlib.

[0069] As a preferred embodiment, the preset key parts of the facial robot may include: the center of the eyebrows, the ends of the eyebrows, the left pupil, the right pupil, the upper eyelids, the lower eyelids, the nose, the upper lip, the lower lip, the left corner of the mouth, the right corner of the mouth, and the mouth.

[0070] When the preset key parts of the facial robot are the center of the eyebrows and the eyebrows, since the eyebrows mainly move in the vertical direction of the human face, and the corners of the eyes move less in the vertical direction when other facial muscles move, therefore, as a further preference, by taking the key point of the corner of the eye as the above reference key point, calculate the pixel deviation in the vertical direction between the key point of the center of the eyebrows and the key point of the corner of the eye below it (which can be regarded as the new state of the key part), subtract the preset original state value of the eyebrows (which can be regarded as the reference state) as the pixel deviation representing the center of the eyebrows relative to its original state, calculate the pixel deviation in the vertical direction between the key point of the end of the eyebrows and the key point of the corner of the eye below it, as the pixel deviation representing the end of the eyebrows relative to its original state; One servo for driving the center of the eyebrows and one servo for driving the end of the eyebrows are respectively arranged at the rear positions corresponding to the two eyebrows inside the facial robot. Then the non-linear mapping methods corresponding to the center of the eyebrows and the end of the eyebrows are:

[0071]

[0072] In the formula, represents the angular deviation corresponding to the pixel deviation in the vertical direction between the center of the eyebrows and the corner of the eye; represents the angular deviation corresponding to the pixel deviation in the vertical direction between the end of the eyebrows and the corner of the eye; represents the preset original state value of the eyebrows; 、 、 、 、 、 、 and are respectively the vertical pixel coordinates of the dlib detected key points 22, 23, 40, 43, 18, 27, 37 and 46 in the real face image; represents the length of the human face in the real face image. In the facial robot, AU2 and AU3 are controlled by the same servo, and AU1 and AU4 are controlled by the same servo.

[0073] Since the dlib face detection package does not directly detect the key points of the pupils, in this embodiment, the AUs around the eyes can be connected to extract the eye mask. Subsequently, binarization is performed on the mask, the contour of the lens is extracted, and the center of the circle that can contain the largest circle of the contour is found as the pupil. Based on this, as a preference, the preset key parts of the facial robot include the left pupil and the right pupil;

[0074] By taking the average of the two canthus positions of each eye as the orbital center position of that eye, and using the orbital center positions of the two eyes as reference points, calculate the deviation of each pupil relative to its respective orbital center position. The sum of the two deviations is used as the pixel deviation shared by the two pupils, representing the deviation of each pupil relative to its original state; there are two servo motors for driving each pupil, arranged side by side, one above the other, at the corresponding rear position inside the facial robot. Then, the non-linear mapping method corresponding to each pupil shared by the two pupils is as follows:

[0075]

[0076] In the formula, is the angular deviation corresponding to the pixel deviation of each pupil in the horizontal direction; is the angular deviation corresponding to the pixel deviation of each pupil in the vertical direction; is the horizontal pixel value of the left pupil in the real face image determined by calculation; is the horizontal pixel value of the right pupil in the real face image determined by calculation; 、 、 、 are the horizontal pixel coordinates of the dlib detected key points 40, 37, 46, and 43 in the real face image respectively; represents the width of the face in the real face image; 、 、 、 are the vertical pixel coordinates of the dlib detected key points 40, 37, 46, and 43 in the real face image respectively; represents the length of the face in the real face image; is the vertical pixel value of the left pupil in the real face image determined by calculation, is the vertical pixel value of the right pupil in the real face image determined by calculation.

[0077] When the preset key parts of the facial robot are the upper eyelid and the lower eyelid, since the end point of the eyelid movement is the closure of the upper and lower eyelids, that is, the vertical pixel difference of the upper and lower eyelids AU is zero, the difference between the key points of the upper and lower eyelids of the left and right eyes in the vertical direction is used as the eyelid error. Therefore, preferably, the sum of the pixel differences between the two upper eyelids and the two lower eyelids in the vertical direction is used as the pixel deviation shared by the upper eyelids and the upper eyelids, representing the deviation of the eyelids relative to their original state; there is one servo motor for driving the upper eyelid and the lower eyelid respectively, which are arranged at the corresponding rear positions of the two eyes inside the facial robot. Then, the non-linear mapping method corresponding to each eyelid shared by the upper eyelids and the upper eyelids is as follows:

[0078]

[0079] Among them, is the angular deviation corresponding to the pixel deviation of the eyelid in the vertical direction, , , , are respectively the vertical pixel coordinates of dlib detected key points 39, 44, 41, and 48 in the real face image; is the length of the face in the real face image.

[0080] When the preset key part of the facial robot is the nose, two inner corner of the eye key points are also selected as reference points. Calculate the pixel deviation in the vertical direction between the two inner corner of the eye key points and the nose tip key point, and subtract the preset original nose state value as the pixel deviation representing the nose relative to its original state; The servo for driving the nose is one, and is set at the rear position corresponding to the nose inside the facial robot. Then the non - linear mapping method corresponding to the nose is:

[0081]

[0082] In the formula, is the angular deviation corresponding to the pixel deviation of the nose in the vertical direction, is the preset original nose state value; , , are the vertical pixel coordinates of dlib detected key points 40, 43, and 31 in the real face image; is the length of the face in the real face image.

[0083] When the preset key part of the facial robot is the upper lip, since the movement of the lips and the opening and closing of the mouth will both cause the movement of the key points around the mouth, in order to distinguish the upper lip from the opening and closing of the mouth, the key points around the nose are used as references. Therefore, preferably, the nose tip key point is used as the reference point. Calculate the pixel deviation in the vertical direction between the nose tip key point and the upper lip key point, and subtract the upper lip original state value as the pixel deviation representing the upper lip relative to its original state; The servo for driving the upper lip is one, and is set at the upper rear position (avoiding the oral cavity) corresponding to the upper lip inside the facial robot. Then the non - linear mapping method corresponding to the upper lip is:

[0084]

[0085] In the formula, is the angular deviation corresponding to the pixel deviation of the upper lip in the vertical direction; , , , They are the vertical pixel coordinates of the dlib-detected key points 33, 35, 62, and 64 in the real face image respectively; is the length of the face in the real face image; is the preset original state value of the upper lip, which is determined according to the actual situation.

[0086] When the preset key part of the facial robot is the lower lip, the principle is the same as that of the upper lip movement. The chin is used as a reference. That is, preferably, the chin key point is used as a reference point, and the pixel deviation between the chin key point and the lower lip key point in the vertical direction is calculated, and the original state value of the lower lip is subtracted from it as the pixel deviation representing the lower lip relative to its original state; There is one servo for driving the lower lip, which is set at the lower rear position (avoiding the oral cavity) corresponding to the lower lip inside the facial robot. Then the non-linear mapping method corresponding to the lower lip is:

[0087]

[0088] In the formula, is the angular deviation corresponding to the pixel deviation of the lower lip in the vertical direction; , , , , , They are the vertical pixel coordinates of the dlib-detected key points 8, 9, 10, 59, 58, and 57 in the real face image respectively; is the preset original state value of the lower lip; is the length of the face in the real face image.

[0089] When the preset key parts of the facial robot are the left corner of the mouth and the right corner of the mouth, the corners of the mouth not only move horizontally but also vertically, and the movement of the corners of the mouth is often accompanied by the movement of the nose. Therefore, the inner corners of the eyes are included as a reference. Therefore, preferably, the pixel difference between the left corner key point and the right corner key point is calculated as the horizontal pixel deviation representing the corners of the mouth relative to their original state; The two inner corner key points of the eyes are used as reference points, and the pixel difference between the left corner key point and the inner corner key point above it and the pixel difference between the right corner key point and the inner corner key point above it are calculated. The sum of the two pixel differences is used as the pixel deviation representing the corners of the mouth relative to their original state; There are two servos for driving the left corner of the mouth and the right corner of the mouth respectively. The two servos corresponding to each corner of the mouth are set at the rear position corresponding to the corner of the mouth inside the facial robot. Then the non-linear mapping method shared by the left corner of the mouth and the right corner of the mouth corresponding to each corner of the mouth is:

[0090]

[0091] In the formula, is the angular deviation corresponding to the pixel deviation of the mouth corner in the horizontal direction; is the angular deviation corresponding to the pixel deviation of the mouth corner in the vertical direction; , , , are the horizontal pixel coordinates of dlib detected key points 55, 65, 49, and 61 in the real face image respectively; , , , are the vertical pixel coordinates of dlib detected key points 40, 43, 49, and 55 in the real face image respectively; is the width of the face in the real face image.

[0092] When the preset key part of the facial robot is the mouth, since the lip movement will also cause the mouth to open and close, and the chin movement will affect the face length, the face width is used as a reference, and the nose and chin are also introduced. Therefore, preferably, the pixel difference between the key point at the lower edge of the nose and the key point of the chin in the vertical direction is calculated as the pixel deviation representing the mouth relative to its original state; there is one servo for driving the mouth to open and close, which is set at the rear position corresponding to the mouth inside the facial robot. Then the non-linear mapping method corresponding to the mouth is:

[0093]

[0094] In the formula, is the angular deviation corresponding to the pixel deviation of the mouth in the horizontal direction; , , , , , are the vertical pixel coordinates of dlib detected key points 33, 34, 35, 8, 9, and 10 in the real face image respectively; is the width of the face in the real face image.

[0095] In one implementation, if all the above non - linear mapping methods are used for mapping, that is, in the facial robot, two drivers are used to drive the eyebrows, one controls the center of the two eyebrows, and the other controls the ends of the eyebrows; two drivers are used to drive the left pupil, and two drivers are used to drive the right pupil; two drivers are used to drive the eyelids, one controls the two upper eyelids, and the other controls the two lower eyelids; one driver is used to drive the nasal muscles, one driver is used to drive the upper lip, one driver is used to drive the lower lip, two drivers are used to drive the left corner of the mouth, two drivers are used to drive the right corner of the mouth, and one driver is used to drive the mouth. A total of 16 drivers, that is, 16 servos, are configured in the facial robot. Thus, 16 servo motors are used to control the movement of facial key points, enabling it to perform anthropomorphic expressions. It should be noted that the approximate positions of the 16 drivers have been described above. Considering that each driver has sufficient force and response speed for the corresponding key parts, the facial key parts can be efficiently controlled.

[0096] In addition, regarding the multi - servo collaborative control algorithm based on the facial robot, considering the single - core execution characteristic of the single - chip microcomputer, in this embodiment, the servo execution sub - function is continuously executed in the main function loop to simulate a multi - core execution operating system.

[0097] (1) Serial port interrupt: When receiving the hexadecimal data transmitted by the Bluetooth serial port, in this embodiment, the frame header of the communication protocol is set to 0xED, the frame tail is set to 0xEF, and the 16 servo units are represented by 0xF0, 0XF1…0xFF respectively. The data received thereafter are the angle deviations of each servo.

[0098] (2) Servo update sub - function: It is used to limit the maximum and minimum values that the servo can reach to protect the face of the facial robot from damage. Use if to judge the size of the current value and the target value, and make the current value continuously accumulate at the minimum speed to avoid the stiff change effect of facial expressions caused by directly setting the current value. Among them, the servo speed is controlled by the delay_ms(xxx) function.

[0099] (3) Main function: Continuously update the current value of the servo and execute the servo movement, using the continuously refreshed main program to simulate multi - core operation.

[0100] Embodiment 2

[0101] A robot facial expression anthropomorphic expression system based on non - linear mapping and multi - servo collaborative control is used to execute a robot facial expression anthropomorphic expression method based on non - linear mapping and multi - servo collaborative control as described above, including a PC - side and a facial robot - side; wherein, the PC - side is used to obtain and generate drive data based on real human face images, and the facial robot - side is used to realize the anthropomorphic expression of the robot's facial expression based on the drive data through multi - servos.

[0102] The software implementation process framework diagram is as Figure 3 shown. On the PC side, first, the original image is collected through a camera. Subsequently, the dlib module is used to detect human faces in the original image and obtain facial key points. Select the real facial key points corresponding to the key points of the expression robot to calculate the key point deviation. Generate the pixel deviation of the key parts relative to their preset original state through non-linear mapping ratio. Synthesize all pixel errors to generate drive data (such as 20-byte hexadecimal), and finally send the drive data to the robot side through the Bluetooth serial port; on the robot side, after receiving the signal through the Bluetooth serial port, convert it into a servo drive signal and perform multi-servo coordinated control through a single-chip microcomputer to finally generate the robot expression.

[0103] In addition, in combination with the positions of the various servos described in Embodiment 1, a structural design of each key part is now given as follows:

[0104] As Figure 4 shown, it is the shell of the human face robot, which is used to support the facial shape and enable the facial skin to adhere. There are gaps left in its eyes, eyebrows, and nose parts, so that the moving units can move in the gaps. The holes on both sides can connect the shell to both ends of the eye structure. Compared with other solutions, using servos to control these four parts can make the servos make different expressions such as inward-slanting eyebrows, outward-slanting eyebrows, raised eyebrows, and lowered eyebrows, and can also control the fineness of these four different expressions.

[0105] Two control servos for each eyeball form a group and are located in the inward direction of the center of the eye socket of the eyeball. The two servos in this group are placed horizontally and have a height difference. The two servos are respectively connected to the eyeball through connecting rods and the connection positions have a height difference. The two servos alternately stretch the corresponding pull rods back and forth relative to the face to achieve the control of the movement of the eyeball in any direction. The two groups of control servos corresponding to the two eyeballs are symmetric about the vertical plane of the midline of the two eyes.

[0106] The two servos for controlling the eyelids are respectively located in the inward direction of the two groups of eyeball servos and are placed vertically. The servos move in the vertical direction. One is used to control the two upper eyelids and the other is used to control the two lower eyelids. The rotation plane of the servo is the vertical plane of space.

[0107] As Figure 5 shown, it is the cheek shell of the facial robot, which is used to support the facial shape and enable the facial skin to adhere. There are gaps left in its nose and mouth parts, so that the moving units can move in the gaps. The four connecting rods on both sides and in the middle can be connected to the lower end of the eye structure. The nasal muscle control servo is located inside the center of the cheek nose shell structure and is placed vertically. The moving direction is vertical. The rotation plane of the servo is the vertical plane of space. Figure 5 The red frame area in

[0108] As shown Figure 6 , it is the nose muscle control structure. Its upper end is used to connect the eye control structure, its lower end is used to connect the mouth control structure, and the bottom screw hole is used to fix the entire nose muscle control structure. The two screw holes at its front end are used to connect the nose muscles; the servo is horizontally placed directly behind the two forward protruding linkages. The control structure part on the left side of the robot is used to connect the servo rotation linkage on one side of the servo; the control structure part on the right side of the robot is movably connected to the other side of the servo through a circular hole structure, which is used to reinforce the rotation shaft of the servo rotation linkage, and the rotation shaft fits with the circular hole structure. The rotation plane of the servo is a vertical plane in space. Different from others, fixing the servo here can make it directly and closer to connect the middle muscles of the nose, avoiding insufficient servo torque due to too long torque.

[0109] As shown Figure 7 , it is the eyebrow control structure, which is connected to the mask through the triple screw holes at both left and right ends. The eyebrow control structure includes a servo bearing structure, a center of the eyebrows movement linkage structure, and a tail of the eyebrows movement linkage structure. The servo bearing structure is used to fix the two servos that are horizontally placed behind the eyebrows and used to control the movement of the eyebrows. Among them, one side of the center of the eyebrows movement linkage structure is connected to the servo rotation linkage on one side of one of the servos, and the other side of the center of the eyebrows movement linkage structure is movably connected to the other side of the other servo through a circular hole structure, which is used to reinforce the rotation shaft of the servo rotation linkage; one side of the tail of the eyebrows movement linkage structure is connected to the servo rotation linkage on one side of the other servo, and the other side of the tail of the eyebrows movement linkage structure is movably connected to the other side of the above one of the servos through a circular hole structure, which is used to reinforce the rotation shaft of the servo rotation linkage. Different from other structures, placing the servo here can extend the length of the servo control linkage, so that when the servo rotates by the same angle, the distance of the eyebrows rising or falling becomes larger, which is beneficial to enrich the degree of expression changes. The rotation plane of the servo is a vertical plane in space.

[0110] As shown Figure 8 , it is the eye-mouth connection structure. The upper part is used to place the eye structure, the lower part is used to connect a single servo, and the other side is used to fix the rotation shaft so that it can use one servo to control the opening and closing of the chin. The position of the servo is inside the circular hole, the placement direction is horizontal, and the rotation plane of the servo is a vertical plane in space. Figure 8 The red frame area in

[0111] As shown Figure 9 , it is the overall head connection structure. Its upper end is used to connect the eye-mouth connection structure, its lower end is used to connect the mouth structure, and the bottom screw hole is used to fix the entire head. The servos are closely attached to the side of the overall head connection structure, with a total of four, all vertically placed relative to the side, and the rotation plane of the servos is a vertical plane in space. Figure 9The red frame area in it is the corresponding servo setting position.

[0112] The hardware structure design of the facial expression robot in this embodiment is based on the Solidworks 2023 platform. The robot body is built by drawing 3D facial expression robot parts and printed using PLA material through the Bambu Lab A1mini 3D printer. The control unit of the facial expression robot in this embodiment consists of 1 STM32F103C8T6 minimum system board, 15 SG90 servos, 1 MG996R servo, and 1 PCA9685 servo driver module. On the premise that the facial robot designed in this embodiment can generate realistic expressions, as few servos as possible are used, and the proposed non-linear mapping method is used to achieve realistic movement of key facial parts.

[0113] For the structure designed in the present invention, a method for imitating a real human face to drive the facial robot is proposed. Multiple key points of the real human face are mapped to the joint parts of the facial robot through the proposed mapping function. Finally, the single-chip microcomputer is used for cooperative control to improve the authenticity of the realistic expression of the facial robot. The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.

[0114] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for anthropomorphic expression of robot facial expressions based on non - linear mapping and multi - server collaborative control, characterized in that, Including: Obtaining a plurality of facial key points of a real human face corresponding to each preset key part of a facial robot, calculating the pixel deviation between the plurality of facial key points and the reference key points preset for quantifying the expression amplitude of the key part, and subtracting the preset original state value of the key part from the pixel deviation to obtain a pixel deviation representing the relative preset original state of the key part; Through non-linear mapping, mapping the pixel deviation relative to its preset original state to an angular deviation representing the relative original position of each servo for driving the key part, and generating drive data by synthesizing the angular deviations corresponding to each servo, so as to realize anthropomorphic expression of the facial expression of the robot through multiple servos; wherein, the non-linear mapping method corresponding to each key part depends on the relative setting position of each servo for driving the key part to the key part; and when performing non-linear mapping, first normalize the pixel deviation relative to the preset original state based on the size of the human face; Wherein, the preset key parts of the facial robot include the upper lip; Taking the nose tip key point as the reference point, calculating the pixel deviation between the nose tip key point and the upper lip key point in the vertical direction, and subtracting the upper lip original state value as the pixel deviation representing the upper lip relative to its original state; there is one servo for driving the upper lip, which is arranged at the upper rear position corresponding to the upper lip inside the facial robot, then the non-linear mapping method corresponding to the upper lip is: Where E uplip is the angular deviation corresponding to the pixel deviation of the upper lip in the vertical direction; H33 y , H35 y , H62 y , H64 y are the vertical pixel coordinates of the dlib-detected key points (33), (35), (62), and (64) in the real face image, respectively; H h is the length of the face in the real face image; is the preset original state value of the upper lip; The preset key parts of the facial robot further include the lower lip; Taking the chin key point as the reference point, calculating the pixel deviation between the chin key point and the lower lip key point in the vertical direction, and subtracting the lower lip original state value as the pixel deviation representing the lower lip relative to its original state; there is one servo for driving the lower lip, which is arranged at the lower rear position corresponding to the lower lip inside the facial robot, then the non-linear mapping method corresponding to the lower lip is: Where, E downlip is the angular deviation corresponding to the pixel deviation of the lower lip in the vertical direction; H8 y , H9 y , H10 y , H59 y , H58 y , H57 y are the vertical pixel coordinates of the dlib detected key points (8), (9), (10), (59), (58) and (57) in the real face image respectively; is the preset original state value of the lower lip; H h is the length of the face in the real face image.

2. The method for anthropomorphic expression of robot facial expressions according to claim 1, characterized in that, The preset key parts of the facial robot include the center of the eyebrows and the eyebrows; By taking the eye corner key point as the reference key point, calculating the pixel deviation between the center of the eyebrows key point and the eye corner key point below it in the vertical direction, subtracting the preset original state value of the eyebrows as the pixel deviation representing the center of the eyebrows relative to its original state, and calculating the pixel deviation between the end of the eyebrows key point and the eye corner key point below it in the vertical direction as the pixel deviation representing the end of the eyebrows relative to its original state; There is one servo for driving the center of the eyebrows and the end of the eyebrows respectively, which are arranged at the rear positions corresponding to the two eyebrows inside the facial robot, then the non-linear mapping methods corresponding to the center of the eyebrows and the end of the eyebrows are: In the formula, represents the angular deviation corresponding to the pixel deviation between the center of the eyebrows and the corners of the eyes in the vertical direction; represents the angular deviation corresponding to the pixel deviation between the ends of the eyebrows and the corners of the eyes in the vertical direction; represents the preset original state value of the eyebrows; H22 y 、H23 y 、H40 y 、H43 y 、H18 y 、H27 y 、H37 y and H46 y are respectively the vertical pixel coordinates of the dlib detected key points (22), (23), (40), (43), (18), (27), (37) and (46) in the real face image; H h represents the length of the face in the real face image.

3. The method for anthropomorphic expression of robot facial expressions according to claim 1, characterized in that The preset key parts of the facial robot include the nose; Taking the two inner eye corner key points as the reference points, calculating the pixel deviation between the two inner eye corner key points and the nose tip key point in the vertical direction, and subtracting the preset original state value of the nose as the pixel deviation representing the nose relative to its original state; there is one servo for driving the nose, which is arranged at the rear position corresponding to the nose inside the facial robot, then the non-linear mapping method corresponding to the nose is: Where E nose is the angular deviation corresponding to the pixel deviation of the nose in the vertical direction, is the preset original nose state value; H40 y and H43 y and H31 y are the vertical pixel coordinates of the dlib detected key points (40), (43) and (31) in the real face image; H h is the length of the face in the real face image.

4. A robot facial expression anthropomorphic expression system based on non-linear mapping and multi-server collaborative control, characterized in that, A method for anthropomorphic expression of a robot's facial expression based on non - linear mapping and multi - server collaborative control as described in any one of claims 1 to 3, including a PC - side and a facial robot; wherein, the PC - side is used to obtain and generate drive data based on real human face images, and the facial robot is used to achieve anthropomorphic expression of the robot's facial expression through multi - servers based on the drive data.

5. The robot facial expression anthropomorphic expression system according to claim 4, wherein, The facial robot includes a nose muscle control structure; Among them, a screw hole is provided at the bottom of the nose muscle control structure for fixing the entire nose muscle control structure; a forward - protruding link is provided at the front end of the nose muscle control structure for connecting the nose muscle through the screw hole; the servo for controlling the nose is horizontally placed directly behind the forward - protruding link, and the part of the nose muscle control structure on one side of the robot is used to connect the servo - rotating link on one side of the servo; the part of the nose muscle control structure on the other side of the robot is movably connected to the other side of the servo through a circular - hole structure for strengthening the rotating shaft of the servo - rotating link; the servo - rotating plane is a vertical plane in space.

6. The robot facial expression anthropomorphic expression system according to claim 4, wherein The facial robot includes an eyebrow control structure; the eyebrow control structure is composed of a servo - bearing structure, a center - of - eyebrow movement link structure, and an end - of - eyebrow movement link structure; Among them, the servo - bearing structure is used to fix two horizontally - placed servos for controlling the movement of the eyebrows and located behind the eyebrows; one side of the center - of - eyebrow movement link structure is connected to the servo - rotating link on one side of one of the servos, and the other side of the center - of - eyebrow movement link structure is movably connected to the other side of the other servo through a circular - hole structure for strengthening the rotating shaft of the servo - rotating link; one side of the end - of - eyebrow movement link structure is connected to the servo - rotating link on one side of the other servo, and the other side of the end - of - eyebrow movement link structure is movably connected to the other side of one of the above - mentioned servos through a circular - hole structure for strengthening the rotating shaft of the servo - rotating link; the rotating plane of each servo is a vertical plane in space.

Citation Information

Patent Citations

  • Head mechanism of robot, robot and control method of robot

    CN112775991A