Facial robot expression cooperative control method and system based on nonlinear mapping
Through the method based on nonlinear mapping and multi-server collaborative control, the low accuracy of anthropomorphic expression of robot facial expressions and stiff facial movements in the prior art is solved, and a higher imitation-driven accuracy and authenticity of facial expressions are achieved.
Patent Information
- Application Number
- CN202510599607.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
When the prior art realizes the anthropomorphic expression of robot facial expressions, there are problems such as low accuracy of imitation drive, stiff facial movements, and high errors of different face drives.
Using a method based on nonlinear mapping and multi-server collaborative control, by obtaining multiple facial key points of real face images, calculating pixel deviations and performing nonlinear mappings, driving data is generated to realize the anthropomorphic expression of robot facial expressions.
It improves the accuracy of imitation drive, reduces the stiffness of facial movements, adapts to the differences between different faces, and enhances the authenticity and generalization ability of facial expressions.
Smart Images

Figure CN120126201A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to robot facial expressions, and more specifically, relates to a method for anthropomorphic expression of robot facial expressions based on non-linear mapping and multi-server collaborative control. Background Art
[0002] In humanoid embodied robots, emotional expression is the core. As the carrier of emotional output, the human face, a facial expression robot can simulate emotions such as smiling and frowning, improving the naturalness of human-robot interaction.
[0003] One of the reasons for designing humanoid robots to be humanoid is to make them easier to be received and understood. The current design of humanoid robots mainly focuses on the hands and legs, while the face, as a key part of emotional expression, has relatively less research. Currently, one of the most important ways to achieve facial expression is through an electronic screen display. The electronic screen can directly generate various expressions and can express various emotions for the robot. However, the electronic screen does not have the three-dimensional sense of the human face, so it will create a sense of distance.
[0004] Therefore, designing a realistic three-dimensional facial expression robot can effectively solve the problem of insufficient robot realism. By using a mechanical structure design to simulate the movement of facial muscles and inputting control signals into the control unit of facial muscles at the same time, the movement of facial muscles can be achieved, thereby realizing the generation of facial robot expressions. If a fixed linear signal is output to the facial robot, although the target expression can be achieved, the expression change of the facial robot will be very rigid, mainly reflected in the fixed speed of the facial movement unit (AU) controlled by the control unit.
[0005] In an actual scenario, for a facial robot driven by imitating a real human face, the determination of the error of its joint points is very important. Although the given value can be used to achieve the follow-up movement of the facial robot's AU, the tiny jitter error of the human face detecting the facial AU will also affect the movement of the AU. And robots driven by different human faces cannot be achieved through a single linear mapping. The size of the face and the position between different facial AUs will introduce different errors, resulting in a relatively high error rate of imitation driving. Summary of the Invention
[0006] In view of the above defects or improvement requirements of the prior art, the present invention provides a method and system for anthropomorphic expression of robot facial expressions based on non-linear mapping and multi-server collaborative control, aiming to improve the accuracy of imitation driving.
[0007] To achieve the above object, according to one aspect of the present invention, there is provided a method for anthropomorphic expression of robot facial expressions based on non-linear mapping and multi-server collaborative control, including: Obtain multiple facial key points corresponding to each preset key part of the facial robot, calculate the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part, subtract the preset original state value of the key part from the pixel deviation to obtain a pixel deviation representing the relative preset original state of the key part; through non-linear mapping, map the pixel deviation relative to its preset original state to an angular deviation representing the relative original position of each servo for driving the key part, and generate drive data by synthesizing the angular deviations corresponding to each servo, so as to achieve anthropomorphic expression of the robot's facial expression through multiple servos; wherein, the non-linear mapping method corresponding to each key part depends on the relative position of each servo for driving the key part to the key part; and during non-linear mapping, first normalize the pixel deviation relative to the preset original state based on the face size.
[0008] Further, the preset key parts of the facial robot include the left pupil and the right pupil; By taking the average of the two eye corner positions of each eye as the orbital center position of the eye, and using the orbital center positions of the two eyes as reference points, calculate the deviation of the pupil of each eye relative to its own orbital center position, and the sum of the two deviations is used as the pixel deviation shared by the two pupils representing the relative original state of each pupil; there are two servos for driving each pupil, which are arranged side by side in a high-low manner at the rear position corresponding to the pupil inside the facial robot, then the non-linear mapping method corresponding to each pupil shared by the two pupils is:
[0009] In the formula, is the angular deviation corresponding to the pixel deviation of each pupil in the horizontal direction; is the angular deviation corresponding to the pixel deviation of each pupil in the vertical direction; is the horizontal pixel value of the left eye pupil in the real face image determined by calculation; is the horizontal pixel value of the right eye pupil in the real face image determined by calculation; 、 、 、 are the horizontal pixel coordinates of dlib detection key points 40, 37, 46 and 43 in the real face image respectively; represents the width of the face in the real face image; 、 、 、 are the vertical pixel coordinates of dlib detection key points 40, 37, 46 and 43 in the real face image respectively; Represents the length of the face in the real face image; is the vertical pixel value of the left eye pupil determined by calculation in the real face image, is the vertical pixel value of the right eye pupil determined by calculation in the real face image; The preset key parts of the facial robot include the upper eyelids and the lower eyelids; The sum of the pixel differences between the key points of the two upper eyelids and the two lower eyelids in the vertical direction is used as the pixel deviation representing the relative state of the eyelids from their original state; One servo for driving the upper eyelids and one for driving the lower eyelids are respectively arranged at the rear positions corresponding to the two eyes inside the facial robot. Then, the non - linear mapping method corresponding to each eyelid shared by the upper eyelids and the lower eyelids is:
[0010] Among them, is the angular deviation corresponding to the pixel deviation of the eyelid in the vertical direction, 、 、 、 are respectively the vertical pixel coordinates of the dlib - detected key points 39, 44, 41, and 48 in the real face image; is the length of the face in the real face image.
[0011] Furthermore, the preset key parts of the facial robot include the upper lip; Taking the key point of the nose tip as a reference point, calculate the pixel deviation between the key point of the nose tip and the key point of the upper lip in the vertical direction, and subtract the original state value of the upper lip as the pixel deviation representing the relative state of the upper lip from its original state; The servo for driving the upper lip is one, which is arranged at the upper rear position corresponding to the upper lip inside the facial robot. Then, the non - linear mapping method corresponding to the upper lip is:
[0012] In the formula, is the angular deviation corresponding to the pixel deviation of the upper lip in the vertical direction; 、 、 、 are respectively the vertical pixel coordinates of the dlib - detected key points 33, 35, 62, and 64 in the real face image; is the length of the face in the real face image; is the preset original state value of the upper lip; The preset key parts of the facial robot also include the lower lip; Taking the chin key point as a reference point, calculate the pixel deviation between the chin key point and the lower lip key point in the vertical direction, and subtract the original state value of the lower lip from it as the pixel deviation representing the relative state of the lower lip from its original state; there is one servo for driving the lower lip, which is set at the lower rear position corresponding to the lower lip inside the facial robot, then the non-linear mapping method corresponding to the lower lip is:
[0013] In the formula, is the angular deviation corresponding to the pixel deviation of the lower lip in the vertical direction; , , , , , are respectively the vertical pixel coordinates of dlib detected key points 8, 9, 10, 59, 58 and 57 in the real face image; is the preset original state value of the lower lip; is the length of the face in the real face image.
[0014] Furthermore, the preset key parts of the facial robot include the left corner of the mouth and the right corner of the mouth; Calculate the pixel difference between the left corner of the mouth key point and the right corner of the mouth key point as the horizontal pixel deviation representing the relative state of the corners of the mouth from their original state; taking the two inner corner key points as reference points, calculate the pixel difference between the left corner of the mouth key point and the inner corner key point above it and the pixel difference between the right corner of the mouth key point and the inner corner key point above it, and take the sum of the two pixel differences as the pixel deviation representing the relative state of the corners of the mouth from their original state; there are two servos for driving the left corner of the mouth and the right corner of the mouth respectively, and the two servos corresponding to each corner of the mouth are set at the rear position corresponding to this corner of the mouth inside the facial robot, then the non-linear mapping method shared by the left corner of the mouth and the right corner of the mouth corresponding to each corner of the mouth is:
[0015] In the formula, is the angular deviation corresponding to the pixel deviation of the corners of the mouth in the horizontal direction; is the angular deviation corresponding to the pixel deviation of the corners of the mouth in the vertical direction; , , , are respectively the horizontal pixel coordinates of dlib detected key points 55, 65, 49 and 61 in the real face image; , , , are respectively the vertical pixel coordinates of dlib detected key points 40, 43, 49 and 55 in the real face image; is the width of the face in the real face image.
[0016] Furthermore, the preset key parts of the facial robot include the mouth; Calculate the pixel difference between the key point at the lower edge of the nose and the key point of the chin in the vertical direction as the pixel deviation representing the relative state of the mouth from its original state; there is one servo for driving the opening and closing of the mouth, which is set at the corresponding rear position inside the facial robot where the mouth is located. Then the non-linear mapping method corresponding to the mouth is:
[0017] In the formula, is the angular deviation corresponding to the pixel deviation of the mouth in the horizontal direction; , , , , , are respectively the vertical pixel coordinates of the dlib detected key points 33, 34, 35, 8, 9, and 10 in the real face image; is the width of the face in the real face image.
[0018] Furthermore, the preset key parts of the facial robot include the center of the eyebrows and the eyebrows; By taking the key point of the corner of the eye as the reference key point, calculate the pixel deviation between the key point of the center of the eyebrows and the key point of the corner of the eye below it in the vertical direction, subtract the preset original state value of the eyebrows as the pixel deviation representing the relative state of the center of the eyebrows from its original state, and calculate the pixel deviation between the key point of the end of the eyebrows and the key point of the corner of the eye below it in the vertical direction as the pixel deviation representing the relative state of the end of the eyebrows from its original state; there is one servo for driving the center of the eyebrows and the end of the eyebrows respectively, which are set at the corresponding rear positions inside the facial robot where the two eyebrows are located. Then the non-linear mapping methods corresponding to the center of the eyebrows and the end of the eyebrows are:
[0019] In the formula, represents the angular deviation corresponding to the pixel deviation between the center of the eyebrows and the corner of the eye in the vertical direction; represents the angular deviation corresponding to the pixel deviation between the end of the eyebrows and the corner of the eye in the vertical direction; represents the preset original state value of the eyebrows; , , , , , , and They are the vertical pixel coordinates of dlib-detected key points 22, 23, 40, 43, 18, 27, 37, and 46 in a real face image, respectively. It represents the length of the face in a real face image.
[0020] Furthermore, the preset key parts of the facial robot include the nose. Taking the two inner canthus key points as reference points, calculate the pixel deviation between the two inner canthus key points and the tip of the nose key point in the vertical direction, and subtract the preset original nose state value from it as the pixel deviation representing the nose relative to its original state. There is one servo for driving the nose, which is set at the rear position corresponding to the nose inside the facial robot. Then the non-linear mapping method corresponding to the nose is:
[0021] In the formula, is the angular deviation corresponding to the pixel deviation of the nose in the vertical direction, is the preset original nose state value; , , They are the vertical pixel coordinates of dlib-detected key points 40, 43, and 31 in a real face image. is the length of the face in a real face image.
[0022] According to another aspect of the present invention, there is provided a robot facial expression anthropomorphic expression system based on non-linear mapping and multi-servo collaborative control, which is characterized in that it is used to execute a robot facial expression anthropomorphic expression method based on non-linear mapping and multi-servo collaborative control as described above, including a PC side and a facial robot. Among them, the PC side is used to obtain and generate driving data based on a real face image, and the facial robot is used to realize the anthropomorphic expression of the robot facial expression based on the driving data through multi-servos.
[0023] Furthermore, the facial robot includes a nose muscle control structure. Among them, the bottom end of the nose muscle control structure is provided with a screw hole for fixing the entire nose muscle control structure. The front end of the nose muscle control structure is provided with a forward protruding link for connecting the nose muscle through the screw hole. The servo for controlling the nose is horizontally placed directly behind the forward protruding link. The part of the nose muscle control structure on one side of the robot is used to connect the servo rotation link on one side of the servo. The part of the nose muscle control structure on the other side of the robot is movably connected to the other side of the servo through a circular hole structure for strengthening the rotation axis of the servo rotation link. The servo rotation plane is a space vertical plane.
[0024] Further, the facial robot includes an eyebrow control structure, which is composed of a servo carrier structure, a medial eyebrow movement link structure, and a lateral eyebrow movement link structure. Among them, the servo carrier structure is used to fix two servo motors that are horizontally placed behind the eyebrows and used to control the movement of the eyebrows. One side of the medial eyebrow movement link structure is connected to the servo rotation link on one side of one of the servo motors, and the other side of the medial eyebrow movement link structure is movably connected to the rotation axis of the servo rotation link on the other side of the other servo motor through a circular hole structure for strengthening the rotation axis of the servo rotation link. One side of the lateral eyebrow movement link structure is connected to the servo rotation link on one side of the other servo motor, and the other side of the lateral eyebrow movement link structure is movably connected to the rotation axis of the servo rotation link on the other side of one of the above servo motors through a circular hole structure for strengthening the rotation axis of the servo rotation link. The rotation plane of each servo motor is a vertical plane in space.
[0025] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the technical solution provided by the present invention mainly has the following beneficial effects: 1. The present invention proposes a method for anthropomorphic expression of robot facial expressions based on non-linear mapping and multi-servo cooperation control. Multiple facial key points of a real human face image corresponding to each preset key part of the facial robot are selected, and the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part is calculated as the amplitude of the expression of the key part. Among them, the reference key points are selected and normalized based on the size of the human face during non-linear mapping, so that the simulation deviation caused by different human face differences can be reduced, adapted to different human faces, and the accuracy of simulation driving for different human faces can be ensured to be relatively high, with high generalization. In addition, non-linear mapping is used to convert the pixel information of multiple key points into the angular deviation of the servo motor relative to its original position, and its own servo motor is configured for different key parts, and one or more servo motors are used to drive one key part, which can avoid errors introduced by the size of the face and the relative positions between different facial AUs. Therefore, generally speaking, the method of this embodiment can improve the accuracy of imitation driving. 2. The present invention further preferably designs the number, position, reference key points, and non-linear mapping formula of the servo motors corresponding to different key parts, and more specifically realizes the anthropomorphic expression of expressions applicable to any human face.
[0026] 3. The present invention innovatively designs a facial robot structure, and sets the control structures of the key parts of the facial robot by simulating the way of muscle pulling, including the nose muscle control structure and the eyebrow control structure, and combines it with the driving data of the eyebrow control structure obtained based on the mapping function on the PC side, so that more realistic expression can be completed with fewer servo motors. Description of the Drawings
[0027] Figure 1 It is a flowchart of a method for anthropomorphic expression of a robot's facial expressions based on non - linear mapping and multi - server collaborative control provided by an embodiment of the present invention.
[0028] Figure 2 It is a comparison diagram of key points detected by dlib and key parts of the facial robot provided by an embodiment of the present invention.
[0029] Figure 3 It is a software implementation flowchart of the facial robot provided by an embodiment of the present invention.
[0030] Figure 4 It is a structure diagram of the facial shell provided by an embodiment of the present invention.
[0031] Figure 5 It is a structure diagram of the cheek and nose shell provided by an embodiment of the present invention.
[0032] Figure 6 It is a structure diagram of nose muscle control provided by an embodiment of the present invention.
[0033] Figure 7 It is a structure diagram of eyebrow control provided by an embodiment of the present invention.
[0034] Figure 8 It is a structure diagram of the eye - mouth connection provided by an embodiment of the present invention.
[0035] Figure 9 It is a structure diagram of the overall head connection provided by an embodiment of the present invention. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0037] Embodiment 1 A method for anthropomorphic expression of a robot's facial expressions based on non - linear mapping and multi - server collaborative control, as Figure 1 shown, includes: Obtain multiple facial key points corresponding to each preset key part of the facial robot, calculate the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part, and subtract the preset original state value of the key part from the pixel deviation to obtain a pixel deviation representing the relative preset original state of the key part; through non-linear mapping, map the pixel deviation relative to its preset original state to an angular deviation representing the relative original position of each servo for driving the key part; generate drive data by synthesizing the angular deviations corresponding to each servo, so as to achieve anthropomorphic expression of the robot's facial expression through multiple servos; where the non-linear mapping method corresponding to each key part depends on the relative position of each servo for driving the key part to the key part; and during non-linear mapping, first normalize the pixel deviation relative to the preset original state based on the face size.
[0038] Select multiple facial key points corresponding to each preset key part of the facial robot, calculate the pixel deviation between the multiple facial key points and the reference key points preset for quantifying the expression amplitude of the key part as the amplitude of the expression of the key part. Among them, the reference key points are selected and normalized based on the face size during non-linear mapping. Therefore, the simulation deviation caused by differences in different faces can be reduced, adapted to different faces, and ensure that the accuracy of simulation driving for different faces is relatively high, with high generalization. In addition, non-linear mapping is used to convert the pixel information of multiple key points into the angular deviation of the servo relative to its original position, and its own servo is configured for different key parts. One or more servos are used to drive one key part, which can avoid errors introduced by the size of the face and the relative position between different facial AUs. Therefore, generally speaking, the method of this embodiment can improve the accuracy of imitation driving.
[0039] As Figure 2 shown, from left to right are the key points detected by dlib, the key parts of the facial expression robot, and the overall structure of the facial robot. Among them, the blue part in the overall structure diagram of the facial robot represents the servo. In addition, see Figure 2The display of key points detected by dlib and the key parts of the facial expression robot cannot directly perform linear mapping during the process of mapping real human faces to the key points of the facial expression robot. The main reasons are as follows: 1) The movement direction of facial key points is different from that of the servo motors; 2) One facial key point may be driven by multiple servo motors; 3) Multiple facial key points may be driven by one servo motor. At the same time, in real human faces, by calculating the error between the pixel positions of facial key points when the face has no expression and the same key points when the face has an expression, although the error can be calculated by setting an initial value, there will be a large error due to different human faces. Therefore, in this embodiment, a separate and preferably mapping design and error calculation method are carried out for each key part of the robot to be applicable to any human face.
[0040] For any human face, first determine its facial size for subsequent normalization calculations of various non-linear mappings. The calculation formula is as follows:
[0041] Where H represents the human face, represents the width of the human face, represents the length of the human face, represents AU (Facial Action Unit) numbered 17 on the human face in dlib key point detection. The subscripts x and y respectively represent the horizontal and vertical pixel values of the AU collected by dlib.
[0042] As a preferably implementation method, the preset key parts of the facial robot may include: the center of the eyebrows, the ends of the eyebrows, the left pupil, the right pupil, the upper eyelids, the lower eyelids, the nose, the upper lip, the lower lip, the left corner of the mouth, the right corner of the mouth, and the mouth.
[0043] When the preset key parts of the facial robot are the center of the eyebrows and the eyebrows, since the eyebrows mainly move in the vertical direction of the human face, and the corners of the eyes move less in the vertical direction under the movement of other facial muscles, therefore, as a further preference, by using the corner key points of the eyes as the above reference key points, calculate the pixel deviation in the vertical direction between the key point of the center of the eyebrows and the corner key point of the eyes below it (which can be regarded as the new state of the key part), subtract the preset original state value of the eyebrows (which can be regarded as the reference state) as the pixel deviation representing the center of the eyebrows relative to its original state, calculate the pixel deviation in the vertical direction between the key point of the end of the eyebrows and the corner key point of the eyes below it as the pixel deviation representing the end of the eyebrows relative to its original state; One servo motor is used to drive the center of the eyebrows and the end of the eyebrows respectively, and they are respectively set at the rear positions corresponding to the two eyebrows inside the facial robot. Then the non-linear mapping methods corresponding to the center of the eyebrows and the end of the eyebrows are:
[0044] In the formula, Indicates the angular deviation corresponding to the pixel deviation between the eyebrow center and the eye corner in the vertical direction; Indicates the angular deviation corresponding to the pixel deviation between the eyebrow tail and the eye corner in the vertical direction; Indicates the preset original state value of the eyebrows; 、 、 、 、 、 、 and Are respectively the vertical pixel coordinates of dlib detected key points 22, 23, 40, 43, 18, 27, 37, and 46 in the real face image; Indicates the length of the face in the real face image. Facial robot AU2 and AU3 are controlled by the same servo, and AU1 and AU4 are controlled by the same servo.
[0045] Since the dlib face detection package does not directly detect the key points of the pupils, in this embodiment, the AUs around the eyes can be connected to extract the eye mask. Subsequently, binarization is performed on the mask to extract the contour of the lens, and the center of the largest circle that can contain the contour is found as the pupil. Based on this, preferably, the preset key parts of the facial robot include the left pupil and the right pupil; By taking the average of the two eye corner positions of each eye as the orbital center position of the eye, and using the orbital center positions of the two eyes as reference points, calculate the deviation of each pupil relative to its own orbital center position. The sum of the two deviations is used as the pixel deviation shared by the two pupils to represent the deviation of each pupil from its original state; There are two servos for driving each pupil, arranged side by side, high and low, at the corresponding rear position inside the facial robot. Then, the non - linear mapping method corresponding to each pupil shared by the two pupils is:
[0046] In the formula, Is the angular deviation corresponding to the pixel deviation of each pupil in the horizontal direction; Is the angular deviation corresponding to the pixel deviation of each pupil in the vertical direction; Is the horizontal pixel value of the left eye pupil determined by calculation in the real face image; Is the horizontal pixel value of the right eye pupil determined by calculation in the real face image; 、 、 、 Are respectively the horizontal pixel coordinates of dlib detected key points 40, 37, 46, and 43 in the real face image; represents the width of the face in the real face image; , , , are respectively the vertical pixel coordinates of dlib detected key points 40, 37, 46, and 43 in the real face image; represents the length of the face in the real face image; is the vertical pixel value of the left eye pupil determined by calculation in the real face image, is the vertical pixel value of the right eye pupil determined by calculation in the real face image.
[0047] When the preset key parts of the facial robot are the upper eyelid and the lower eyelid, since the movement end point of the eyelid is the closure of the upper and lower eyelids, that is, the vertical pixel difference between the upper and lower eyelids AU is zero, the vertical difference between the key points of the upper and lower eyelids of the left and right eyes is used as the eyelid error. Therefore, preferably, the sum of the pixel differences between the two upper eyelids and the two lower eyelids in the vertical direction is used as the pixel deviation representing the relative state of the eyelid from its original state; there is one servo for driving the upper eyelid and one servo for driving the lower eyelid, which are respectively set at the rear positions corresponding to the two eyes inside the facial robot. Then, the non - linear mapping method corresponding to each eyelid shared by the upper eyelid and the upper eyelid is:
[0048] Among them, is the angular deviation corresponding to the pixel deviation of the eyelid in the vertical direction, , , , are respectively the vertical pixel coordinates of dlib detected key points 39, 44, 41, and 48 in the real face image; is the length of the face in the real face image.
[0049] When the preset key part of the facial robot is the nose, two inner corner key points are also selected as reference points. Calculate the pixel deviation between the two inner corner key points and the tip of the nose key point in the vertical direction, and subtract the preset original state value of the nose as the pixel deviation representing the relative state of the nose from its original state; there is one servo for driving the nose, which is set at the rear position corresponding to the nose inside the facial robot. Then, the non - linear mapping method corresponding to the nose is:
[0050] In the formula, is the angular deviation corresponding to the pixel deviation of the nose in the vertical direction, is the preset original state value of the nose; , , are the vertical pixel coordinates of the dlib-detected key points 40, 43, and 31 in the real face image; is the length of the face in the real face image.
[0051] When the preset key part of the facial robot is the upper lip, since both lip movement and mouth opening and closing will cause the key points around the mouth to move, in order to distinguish the upper lip from mouth opening and closing, the key points around the nose are used as a reference. Therefore, preferably, the tip-of-nose key point is used as the reference point, and the pixel deviation between the tip-of-nose key point and the upper-lip key point in the vertical direction is calculated, and its subtraction from the original state value of the upper lip is used as the pixel deviation representing the upper lip relative to its original state; the servo for driving the upper lip is one, and is set at the position corresponding to the upper lip inside the facial robot at the upper rear (avoiding the oral cavity), then the non-linear mapping method corresponding to the upper lip is:
[0052] In the formula, is the angular deviation corresponding to the pixel deviation of the upper lip in the vertical direction; , , , are the vertical pixel coordinates of the dlib-detected key points 33, 35, 62, and 64 in the real face image respectively; is the length of the face in the real face image; is the preset original state value of the upper lip, which is determined according to the actual situation.
[0053] When the preset key part of the facial robot is the lower lip, the principle is the same as that of the upper lip movement. The chin is used as a reference, that is, preferably, the chin key point is used as the reference point, and the pixel deviation between the chin key point and the lower-lip key point in the vertical direction is calculated, and its subtraction from the original state value of the lower lip is used as the pixel deviation representing the lower lip relative to its original state; the servo for driving the lower lip is one, and is set at the position corresponding to the lower lip inside the facial robot at the lower rear (avoiding the oral cavity), then the non-linear mapping method corresponding to the lower lip is:
[0054] In the formula, is the angular deviation corresponding to the pixel deviation of the lower lip in the vertical direction; , , , , , are the vertical pixel coordinates of the dlib-detected key points 8, 9, 10, 59, 58, and 57 in the real face image respectively; is the preset original state value of the lower lip; is the length of the face in the real face image.
[0055] When the preset key parts of the facial robot are the left and right corners of the mouth, the corners of the mouth not only move horizontally but also vertically, and the movement of the corners of the mouth is often accompanied by the movement of the nose. Therefore, the corners of the eyes are included as a reference. Therefore, preferably, the pixel difference between the key points of the left and right corners of the mouth is calculated as the horizontal pixel deviation representing the relative original state of the corners of the mouth; the two inner corner key points are used as reference points, and the pixel difference between the key point of the left corner of the mouth and the inner corner key point above it and the pixel difference between the key point of the right corner of the mouth and the inner corner key point above it are calculated, and the sum of the two pixel differences is used as the pixel deviation representing the relative original state of the corners of the mouth; there are two servo motors for driving the left and right corners of the mouth respectively, and the two servo motors corresponding to each corner of the mouth are arranged at the rear position corresponding to the corner of the mouth inside the facial robot. Then, the non-linear mapping method corresponding to each corner of the mouth shared by the left and right corners of the mouth is:
[0056] In the formula, is the angular deviation corresponding to the pixel deviation of the corner of the mouth in the horizontal direction; is the angular deviation corresponding to the pixel deviation of the corner of the mouth in the vertical direction; , , , are the horizontal pixel coordinates of the dlib detection key points 55, 65, 49, and 61 in the real face image respectively; , , , are the vertical pixel coordinates of the dlib detection key points 40, 43, 49, and 55 in the real face image respectively; is the width of the face in the real face image.
[0057] When the preset key part of the facial robot is the mouth, since the movement of the lips will also cause the mouth to open and close, and the movement of the chin will affect the length of the face, the width of the face is used as a reference, and the nose and chin are also introduced. Therefore, preferably, the pixel difference between the key point at the lower edge of the nose and the key point of the chin in the vertical direction is calculated as the pixel deviation representing the relative original state of the mouth; the servo motor for driving the opening and closing of the mouth is one, and it is arranged at the rear position corresponding to the mouth inside the facial robot. Then, the non-linear mapping method corresponding to the mouth is:
[0058] In the formula, is the angular deviation corresponding to the pixel deviation of the mouth in the horizontal direction; , , , , , are respectively the vertical pixel coordinates of the dlib detected key points 33, 34, 35, 8, 9, and 10 in the real face image; is the width of the face in the real face image.
[0059] In one implementation, if all the above non - linear mapping methods are used for mapping, that is, in the facial robot, the eyebrows are driven by 2 drivers, one controls the center of the two eyebrows, and the other controls the ends of the eyebrows; the left pupil is driven by 2 drivers, and the right pupil is driven by 2 drivers; the eyelids are driven by 2 drivers, one controls the two upper eyelids, and the other controls the two lower eyelids; the nasal muscles are driven by 1 driver, the upper lip is driven by 1 driver, the lower lip is driven by 1 driver, the left corner of the mouth is driven by 2 drivers, the right corner of the mouth is driven by 2 drivers, and the mouth is driven by 1 driver. A total of 16 drivers, that is, 16 servos, are configured in the facial robot. Thus, 16 servo motors are used to control the movement of the facial key points, enabling it to perform anthropomorphic expressions. It should be noted that the approximate positions of the 16 drivers have been described above. Considering that each driver has sufficient force and response speed for the corresponding key parts, the facial key parts can be efficiently controlled.
[0060] In addition, regarding the multi - servo collaborative control algorithm based on the facial robot, considering the single - core execution characteristic of the single - chip microcomputer, in this embodiment, the servo execution sub - function is continuously executed in the main function loop to simulate a multi - core execution operating system.
[0061] (1) Serial port interrupt: When receiving the hexadecimal data transmitted by the Bluetooth serial port, in this embodiment, the frame header of the communication protocol is set to 0xED, the frame tail is set to 0xEF, and the 16 servo units are respectively represented by 0xF0, 0XF1…0xFF. The data received thereafter are the angular deviations of each servo.
[0062] (2) Servo update sub - function: Used to limit the maximum and minimum values that the servo can reach to protect the face of the facial robot from damage. Use if to judge the size of the current value and the target value, and make the current value continuously accumulate at the minimum speed to avoid the stiff change effect of facial expressions caused by directly setting the current value. Among them, the servo speed is controlled by the delay_ms(xxx) function.
[0063] (3) Main function: Continuously update the current value of the servo and execute the servo movement, using the continuously refreshed main program to simulate multi - core operation.
[0064] Example 2 A robot facial expression anthropomorphic expression system based on non - linear mapping and multi - server collaborative control is used to execute a robot facial expression anthropomorphic expression method based on non - linear mapping and multi - server collaborative control as described above, including a PC side and a facial robot side; wherein, the PC side is used to acquire and generate drive data based on real human face images, and the facial robot side is used to realize robot facial expression anthropomorphic expression through multi - servers based on the drive data.
[0065] The software implementation process framework diagram is as Figure 3 shown. On the PC side, first, the original picture is collected through a camera, and then the dlib module is used to detect human faces in the original image and obtain facial key points. Select the real facial key points corresponding to the key points of the expression robot to calculate the key point deviation, generate the pixel deviation representing the relative preset original state of the key parts through non - linear mapping ratio, synthesize all pixel errors, generate (for example, 20 - byte hexadecimal) drive data, and finally send the drive data to the robot side through the Bluetooth serial port; on the robot side, after receiving the signal through the Bluetooth serial port, it is converted into a servo drive signal, and multi - server collaborative control is performed through a single - chip microcomputer, and finally the robot expression is generated.
[0066] In addition, combining the positions of each server described in Example 1, a structural design of each key part is given as follows: As Figure 4 shown, it is the shell of the human - face robot, which is used to support the facial shape and enable the facial skin to adhere. There are gaps left in its eyes, eyebrows, and nose parts, so that the support of the moving unit can move in these gaps. The holes on both sides can connect the shell to both ends of the eye structure. Compared with other solutions, using servo motors to control these four parts can make the servo motors make different expressions such as inward - slanting eyebrows, outward - slanting eyebrows, raising eyebrows, and pressing eyebrows, and can also control the fineness of these four different expressions.
[0067] Two control servo motors for each eyeball form a group, located in the inward direction of the center of the eye socket of the eyeball. The two servo motors in this group are placed horizontally and have a height difference. The two servo motors are respectively connected to the eyeball through connecting rods and the connection positions have a height difference. The two servo motors alternately stretch the corresponding pull rods forward and backward relative to the face to realize the control of the movement of the eyeball in any direction. The two groups of control servo motors corresponding to the two eyeballs are symmetric about the mid - vertical plane of the line connecting the two eyes.
[0068] The two servo motors for controlling the eyelids are respectively located in the inward direction of the two groups of eyeball servo motors, placed vertically, and the servo motors move in the vertical direction. One is used to control the two upper eyelids, and the other is used to control the two lower eyelids. The rotation plane of the servo motor is a spatial vertical plane.
[0069] AsFigure 5 As shown in the figure, it is the cheek shell of the facial robot, which is used to support the facial shape, enable the facial skin to adhere, and leave gaps in the nose and mouth parts so that the moving unit can move in the gaps. The four connecting rods on both sides and in the middle can be connected to the lower end of the eye structure. The nasal muscle control servo is located inside the center of the cheek nose shell structure, placed vertically, and the moving direction is vertical. The servo rotation plane is the space vertical plane. Figure 5 The red frame area in it is the corresponding servo setting position.
[0070] As Figure 6 shown in the figure, it is the nose muscle control structure. Its upper end is used to connect the eye control structure, its lower end is used to connect the mouth control structure, and the bottom screw hole is used to fix the entire nose muscle control structure. The two front screw holes at its front end are used to connect the nose muscles; the servo is placed horizontally directly behind the two forward protruding connecting rods. The control structure part on the left side of the robot is used to connect the servo rotation connecting rod on one side of the servo; the control structure part on the right side of the robot is movably connected to the other side of the servo through a circular hole structure, which is used to reinforce the rotation shaft of the servo rotation connecting rod, and the rotation shaft fits with the circular hole structure. The servo rotation plane is the space vertical plane. Different from others, fixing the servo here can make it directly and closer to connect the middle nose muscles, avoiding insufficient servo torque due to too long torque.
[0071] As Figure 7 shown in the figure, it is the eyebrow control structure, which is connected to the mask through the triple screw holes at both ends. The eyebrow control structure includes a servo bearing structure, a center of the eyebrows movement connecting rod structure and an end of the eyebrows movement connecting rod structure. The servo bearing structure is used to fix the two servos that are used to control the eyebrow movement and are placed horizontally behind the eyebrows. Among them, one side of the center of the eyebrows movement connecting rod structure is connected to the servo rotation connecting rod on one side of one of the servos, and the other side of the center of the eyebrows movement connecting rod structure is movably connected to the other side of the other servo through a circular hole structure, which is used to reinforce the rotation shaft of the servo rotation connecting rod; one side of the end of the eyebrows movement connecting rod structure is connected to the servo rotation connecting rod on one side of the other servo, and the other side of the end of the eyebrows movement connecting rod structure is movably connected to the other side of one of the above servos through a circular hole structure, which is used to reinforce the rotation shaft of the servo rotation connecting rod. Different from other structures, placing the servo here can extend the length of the servo control connecting rod, so that when the servo rotates by the same angle, the distance of the eyebrows rising or falling becomes larger, which is beneficial to enrich the degree of expression changes. The servo rotation plane is the space vertical plane.
[0072] As Figure 8 shown in the figure, it is the eye-mouth connection structure. The upper part is used to place the eye structure, the lower part is used to connect a single servo, and the other side is used to fix the rotation shaft so that it can use one servo to control the opening and closing of the chin. The servo position is inside the circular hole, placed horizontally, and the servo rotation plane is the space vertical plane. Figure 8The red frame area in it is the corresponding servo setting position.
[0073] As Figure 9 shown, it is the overall connection structure of the head. Its upper end is used to connect the eye-mouth connection structure, its lower end is used to connect the mouth structure, and the bottom screw hole is used to fix the entire head. The servos are closely attached to the side of the overall head connection structure, a total of four, all vertically placed relative to the side. The servo rotation plane is a vertical plane in space. Figure 9 The red frame area in it is the corresponding servo setting position.
[0074] The hardware structure design of the facial expression robot in this embodiment is based on the Solidworks2023 platform. The robot body is built by drawing 3D facial expression robot parts and printed using PLA material through the Bambu Lab A1mini 3D printer. The control unit of the facial expression robot in this embodiment consists of 1 STM32F103C8T6 minimum system board, 15 SG90 servos, 1 MG996R servo, and 1 PCA9685 servo driver module. On the premise that the facial robot designed in this embodiment can generate realistic expressions, as few servos as possible are used, and the proposed non-linear mapping method is used to achieve realistic movement of key facial parts.
[0075] For the structure designed in the present invention, a method for driving the facial robot by imitating a real human face is proposed. Multiple key points of the real human face are mapped to the joint parts of the facial robot through the proposed mapping function. Finally, the single-chip microcomputer is used for cooperative control to improve the authenticity of the realistic expression of the facial robot. The relevant technical solutions are the same as those in Embodiment 1 and will not be elaborated here.
[0076] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for anthropomorphic expression of robot facial expression based on nonlinear mapping and multi-server collaborative control, characterized in that: include: Acquire multiple facial key points of a real human face image corresponding to each preset key part of the facial robot, calculate the pixel deviation between the multiple facial key points and the preset reference key points used to quantify the expression amplitude of the key part, and subtract the preset original state value of the key part from the pixel deviation to obtain the pixel deviation representing the key part relative to its preset original state; Through nonlinear mapping, the pixel deviation relative to its preset original state is mapped to represent the angular deviation of each servo used to drive the key part relative to its original position, and the driving data is generated by integrating the angular deviations corresponding to each servo, so as to realize the anthropomorphic expression of the robot's facial expression through multiple servos; wherein, the nonlinear mapping method corresponding to each key part depends on the setting position of each servo used to drive the key part relative to the key part; and during the nonlinear mapping, the pixel deviation relative to the preset original state is first normalized based on the size of the human face.
2. The method for expressing robot facial expressions anthropomorphically as claimed in claim 1, characterized in that: The preset key parts of the facial robot include the left pupil and the right pupil; The average of the two eye corner positions of each eye is taken as the orbital center position of the eye, and the orbital center positions of the two eyes are taken as reference points to calculate the deviation of the pupils of the two eyes relative to the respective orbital center positions. The sum of the two deviations is taken as the pixel deviation shared by the two pupils, which represents the pixel deviation of each pupil relative to its original state. Two servers are used to drive each pupil, and they are arranged in parallel at the rear position corresponding to the pupil inside the facial robot. Then the nonlinear mapping method corresponding to each pupil shared by the two pupils is: In the formula, is the angle deviation corresponding to the pixel deviation of each pupil in the horizontal direction; is the angle deviation corresponding to the pixel deviation of each pupil in the vertical direction; is the horizontal pixel value of the left pupil in the real face image determined by calculation; is the horizontal pixel value of the right pupil in the real face image determined by calculation; , , , These are the horizontal pixel coordinates of dlib detection key points 40, 37, 46 and 43 in the real face image; Indicates the width of the face in a real face image; , , , They are the vertical pixel coordinates of dlib detection key points 40, 37, 46 and 43 in real face images; Represents the length of the face in a real face image; is the vertical pixel value of the left pupil in the real face image determined by calculation, is the vertical pixel value of the right pupil in the real face image determined by calculation; The preset key parts of the facial robot include the upper eyelid and the lower eyelid; The sum of the pixel differences between the two upper eyelid key points and the two lower eyelid key points in the vertical direction is taken as the pixel deviation shared by the upper eyelid and the lower eyelid representing the eyelid relative to its original state; One server is used to drive the upper eyelid and one server is used to drive the lower eyelid. They are respectively set at the rear positions corresponding to the two eyes inside the facial robot. Then the nonlinear mapping method corresponding to the upper eyelid and the eyelids shared by the upper eyelids is: in, is the angle deviation corresponding to the pixel deviation of the eyelid in the vertical direction, , , , They are the vertical pixel coordinates of dlib detection key points 39, 44, 41, and 48 in real face images; is the length of the face in the real face image.
3. The method for anthropomorphic expression of robot facial expression as claimed in claim 1, characterized in that: The preset key parts of the facial robot include the upper lip; The nose tip key point is used as a reference point, and the pixel deviation between the nose tip key point and the upper lip key point in the vertical direction is calculated, and the pixel deviation representing the upper lip relative to its original state is subtracted from the original state value of the upper lip; there is one servo for driving the upper lip, which is set at the upper back position corresponding to the upper lip inside the facial robot, and the nonlinear mapping method corresponding to the upper lip is: In the formula, is the angle deviation corresponding to the pixel deviation of the upper lip in the vertical direction; , , , They are the vertical pixel coordinates of dlib detection key points 33, 35, 62 and 64 in the real face image; is the length of the face in the real face image; is the preset original state value of the upper lip; The preset key parts of the facial robot also include the lower lip; The chin key point is used as a reference point, the pixel deviation between the chin key point and the lower lip key point in the vertical direction is calculated, and the original state value of the lower lip is subtracted from the pixel deviation representing the lower lip relative to its original state; There is one servo for driving the lower lip, which is set at the lower back position corresponding to the lower lip inside the facial robot. The nonlinear mapping method corresponding to the lower lip is: In the formula, is the angle deviation corresponding to the pixel deviation of the lower lip in the vertical direction; , , , , , They are the vertical pixel coordinates of dlib detection key points 8, 9, 10, 59, 58 and 57 in real face images; is the preset original state value of the lower lip; is the length of the face in the real face image.
4. The method for anthropomorphic expression of robot facial expression as claimed in claim 1, characterized in that: The preset key parts of the facial robot include the left corner of the mouth and the right corner of the mouth; Calculate the pixel difference between the left mouth corner key point and the right mouth corner key point as the lateral pixel deviation representing the mouth corner relative to its original state; Taking the two inner eye corner key points as reference points, calculating the pixel difference between the left mouth corner key point and the inner eye corner key point above it, and the pixel difference between the right mouth corner key point and the inner eye corner key point above it, and taking the sum of the two pixel differences as the pixel deviation representing the mouth corner relative to its original state; There are two servers for driving the left and right corners of the mouth respectively. The two servers corresponding to each corner of the mouth are set at the rear position corresponding to the corner of the mouth inside the facial robot. Then the nonlinear mapping method corresponding to each corner of the mouth shared by the left and right corners of the mouth is: In the formula, is the angle deviation corresponding to the pixel deviation of the mouth corner in the horizontal direction; is the angle deviation corresponding to the pixel deviation of the mouth corner in the vertical direction; , , , These are the horizontal pixel coordinates of dlib detection key points 55, 65, 49 and 61 in real face images; , , , They are the vertical pixel coordinates of dlib detection key points 40, 43, 49 and 55 in real face images; is the width of the face in the real face image.
5. The method for anthropomorphic expression of robot facial expression as claimed in claim 1, characterized in that: The preset key parts of the facial robot include a mouth; Calculate the pixel difference between the key point of the lower edge of the nose and the key point of the chin in the vertical direction as the pixel deviation representing the mouth relative to its original state; There is one servo for driving the mouth to open and close, which is set at the rear position corresponding to the mouth inside the facial robot. The nonlinear mapping method corresponding to the mouth is: In the formula, is the angle deviation corresponding to the pixel deviation of the mouth in the horizontal direction; , , , , , They are the vertical pixel coordinates of dlib detection key points 33, 34, 35, 8, 9 and 10 in the real face image; is the width of the face in the real face image.
6. The method for anthropomorphic expression of robot facial expression according to any one of claims 1 to 5, characterized in that: The preset key parts of the facial robot include the center of the eyebrows and the eyebrows; By taking the eye corner key point as the reference key point, calculating the pixel deviation of the eyebrow center key point and the eye corner key point below it in the vertical direction, subtracting the preset eyebrow original state value from the value as the pixel deviation representing the eyebrow center relative to its original state, calculating the pixel deviation of the eyebrow tail key point and the eye corner key point below it in the vertical direction as the pixel deviation representing the eyebrow tail relative to its original state; One servo is used to drive the center of the eyebrows and one servo is used to drive the tail of the eyebrows. They are respectively set at the rear positions corresponding to the two eyebrows inside the facial robot. The nonlinear mapping method corresponding to the center of the eyebrows and the tail of the eyebrows is: In the formula, Indicates the angle deviation corresponding to the pixel deviation between the center of the eyebrow and the corner of the eye in the vertical direction; Indicates the angle deviation corresponding to the pixel deviation between the eyebrow tail and the eye corner in the vertical direction; Indicates the preset original state value of eyebrows; , , , , , , and They are the vertical pixel coordinates of dlib detection key points 22, 23, 40, 43, 18, 27, 37 and 46 in real face images; Represents the length of the face in a real face image.
7. The method for anthropomorphic expression of robot facial expression according to any one of claims 1 to 5, characterized in that: The preset key parts of the facial robot include a nose; The two inner eye corner key points are used as reference points, and the pixel deviation between the two inner eye corner key points and the nose tip key point in the vertical direction is calculated, and the pixel deviation representing the nose relative to its original state is subtracted from the preset nose original state value; there is one servo for driving the nose, which is set at the rear position corresponding to the nose inside the facial robot, and the nonlinear mapping method corresponding to the nose is: In the formula, is the angle deviation corresponding to the pixel deviation of the nose in the vertical direction, is the preset original state value of the nose; , , The vertical pixel coordinates of key points 40, 43 and 31 detected by dlib in real face images; is the length of the face in the real face image.
8. A robot facial expression anthropomorphic expression system based on nonlinear mapping and multi-servo collaborative control, which is characterized by: A method for anthropomorphic expression of robot facial expressions based on nonlinear mapping and multi-servo collaborative control as described in any one of claims 1 to 7, comprising a PC and a facial robot; wherein the PC is used to obtain and generate drive data based on real human face images, and the facial robot is used to achieve anthropomorphic expression of robot facial expressions through multiple servos based on the drive data.
9. The robot facial expression anthropomorphic expression system as claimed in claim 8, characterized in that: The facial robot includes a nose muscle control structure; Among them, a screw hole is provided at the bottom end of the nose muscle control structure for fixing the entire nose muscle control structure; a forward protruding connecting rod is provided at the front end of the nose muscle control structure for connecting the nose muscles through the screw hole; the servo for controlling the nose is horizontally placed directly behind the forward protruding connecting rod, and the nose muscle control structure part located on one side of the robot is used to connect the servo rotating connecting rod located on one side of the servo; the nose muscle control structure part located on the other side of the robot is movably connected to the other side of the servo through a circular hole structure, which is used to reinforce the rotating axis of the servo rotating connecting rod; the servo rotation plane is a vertical plane in space.
10. The robot facial expression anthropomorphic expression system as claimed in claim 8, characterized in that: The facial robot includes an eyebrow control structure; the eyebrow control structure is composed of a server bearing structure, an eyebrow center motion link structure and an eyebrow tail motion link structure; Among them, the server supporting structure is used to fix two servers that are used to control the movement of eyebrows and are placed horizontally behind the eyebrows; one side of the eyebrow center motion link structure is connected to the server rotating link located on one side of one of the servers, and the other side of the eyebrow center motion link structure is movably connected to the other side of the other server through a circular hole structure, which is used to reinforce the rotating axis of the server rotating link; one side of the eyebrow tail motion link structure is connected to the server rotating link located on one side of the other server, and the other side of the eyebrow tail motion link structure is movably connected to the other side of one of the above-mentioned servers through a circular hole structure, which is used to reinforce the rotating axis of the server rotating link; the rotation plane of each server is a vertical plane in space.
Citation Information
Patent Citations
Active-shape-model-algorithm-based method for analyzing face expression
CN104951743A
Human face dynamic expression detection method and device, equipment and storage medium
CN111382648A
Head mechanism of robot, robot and control method of robot
CN112775991A
Robot eye-head collaborative gazing behavior control method based on bionic principle
CN114872036A
Robot real-time expression simulation method and device
CN116597484A