Skin type robot facial expression control method and system based on vision and language model

Through the combination of visual and language models, facial key points characteristics are obtained and emotional analysis is performed, and the fluency and anthropomorphism of facial expression control of skin-type robots is solved, real-time, realistic driving and subtle dynamic expression of robot facial expressions are realized.

CN120347759APending Publication Date: 2025-07-22JIANGSU YUNMU ZHIZAO TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510734967.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing skin-type robot facial expression control methods rely on pre-programmed mechanical movements, and expression changes lack fluency and subtle dynamics, making it difficult to achieve emotional intelligent response.

Method used

Multimodal data fusion, large language model and facial feature detection technology are used to obtain facial key points characteristics through visual and language models, calculate servo parameter instructions, and use large language models for emotion analysis and fine-tuning to drive robot facial expressions.

Benefits of technology

Real-time driving of robot facial expressions and realistic subtle expressions are realized, improving the interactive ability and anthropomorphic effect of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347759A_ABST
    Figure CN120347759A_ABST
Patent Text Reader

Abstract

The invention discloses a skin type robot facial expression control method and system based on a vision and language model. The method comprises the following steps of S1, obtaining an input image; capturing an image corresponding to the target expression through a camera installed at the eyes of the robot; S2, extracting facial key point features of the input image; s3, preliminarily acquiring a parameter instruction of a robot face steering engine; s4, a robot face steering engine is driven; according to the skin type robot facial expression driving method developed by the invention, visual and language information is analyzed in a staged manner, the real-time driving effect of the robot facial expression can be realized, and the driving of the robot facial expression is finely adjusted by utilizing a large language model; the robot can show good distinction degree in some subtle expressions, and the developed skin type robot has vivid facial expressions and can be applied to multiple interaction fields such as text travel, medical treatment and education.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of robot control, and particularly relates to a skin-type robot facial expression control method and system based on vision and language models. Background Art

[0002] Skin-type robots are an important research direction in the field of robotics in recent years. Due to their closer resemblance to humans in terms of appearance, interaction capabilities, and tactile experience, they are widely used in fields such as culture, tourism, and entertainment, public services and government affairs, and medical care and education. The research and development of skin-type robots involve multidisciplinary intersections such as materials science, artificial intelligence, bionics, and psychology. Its technological breakthroughs play a leading role in the robot industry. The artificial intelligence involved includes a series of technologies such as gait control, facial expression control, emotion recognition and generation of robots. The bionic structure of the face of skin-type robots makes their expressions anthropomorphic. At the same time, due to the large number of detailed areas required on the face, it poses challenges to driving.

[0003] In the research and development of most existing skin-type humanoid robots, facial expression control mainly relies on pre-programmed mechanical movements and limited expression templates. For example, servo motors or linear drives are used to control the mechanical structure under the artificial skin, enabling the robot to execute several predefined basic expressions such as smiling and frowning. However, due to the limitations of the mechanical structure, the expression changes lack smoothness and subtle dynamic changes. At the same time, since the pre-programmed expressions cannot be dynamically adjusted according to the interaction scenario, it is difficult to achieve emotional intelligent responses, giving a sense of "mechanicalness".

[0004] Therefore, the present invention comprehensively applies multi-modal data (such as voice and vision), large language models, and facial feature detection technologies to improve the accuracy and anthropomorphism of facial expression driving of skin-type robots, thereby enhancing the interaction capabilities and attractiveness of robots in interactive scenarios. The proposed robot expression driving method will be widely applied to the facial driving of various bionic robots, providing a more convenient and reliable solution for the field of human-computer interaction. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: for the ultra-realistic real-time expression driving of skin-type robots, combined with multi-modal information fusion, large language models, and face detection technologies to control the facial expressions of robots, so as to generate ultra-realistic real-time facial expressions close to humans.

[0006] To solve the above technical problem, the technical solution adopted by the present invention is: a skin-type robot facial expression control method based on vision and language models, comprising the following steps:

[0007] Step S1: Obtain an input image;

[0008] Capture the image corresponding to the target expression through the camera installed in the robot's eyes, or read the image corresponding to the target expression through the reading device inside the robot.

[0009] Step S2: Extract the facial key point features of the input image;

[0010] First, use a pre-trained vision model (such as DINO) to perform preliminary feature extraction on the input image, then apply a face detection model to detect whether there is a face in the image and locate the key facial regions, and then apply a face mesh model to generate a face mapping mesh containing a predetermined number (478) of three-dimensional facial key points.

[0011] Step S3: Initially obtain the parameter instructions for the robot's facial servos;

[0012] First, select several predefined groups of three-dimensional facial key points, and the key points correspond to different regions of the face, such as eyes, eyebrows, lips, corners of the mouth, etc. Then calculate the absolute distance or relative distance between the selected key points, and use the Euclidean distance to calculate the distance between two facial key points. The distance is expressed as:

[0013]

[0014] where (x1, y1) represents the coordinates of one of the selected key points on the x-axis and y-axis, and (x2, y2) represents the coordinates of the other point on the x-axis and y-axis.

[0015] At the same time, calculate the positional relationship of the key points relative to the reference points (such as the bottom of the nose point, the reference point between the eyebrows), which is represented by the coordinate offset of the key points relative to the reference points (such as the bottom of the nose, the reference point between the eyebrows). For example, calculate the horizontal offset of the center point of the lips relative to the bottom of the nose to control the left and right movement of the mouth. Calculate the vertical offset of the eyebrow key points relative to the reference point between the eyebrows to control the up and down movement of the eyebrows. The coordinate offset is expressed as:

[0016] offset_x = point_x - reference_x;

[0017] offset_y = point_y - reference_y;

[0018] where offset_x and offset_y respectively represent the amounts of coordinate offset on the x-axis and y-axis, point_x and point_y represent the coordinates of the key points, and reference_x and reference_y represent the coordinates of the reference points.

[0019] Subsequently, normalize the calculated distance or positional relationship according to the inherent facial dimensions (such as the inter-pupillary distance or the distance between the inner corners of the eyes). Use the inherent facial dimension (usually the inter-pupillary distance or the distance between the inner corners of the eyes) as a reference, and divide other distances and positional relationships by this reference distance to obtain normalized eigenvalue. For example, divide the distance between the opened eyes by the inter-pupillary distance to obtain the normalized degree of eye opening; divide the distance of the opened mouth by the inter-pupillary distance to obtain the normalized degree of mouth opening.

[0020] Finally, convert the normalized eigenvalue into the preliminary target control parameters (pulse width values) corresponding to each servo (26 servos) through a preset basic mapping function (such as a linear mapping or a piecewise linear mapping function) or rule. Specifically, divide the range of the normalized eigenvalue into multiple sub-intervals, and use different linear mapping functions within each sub-interval.

[0021] Step S4: Drive the facial servos of the robot;

[0022] Perform sentiment analysis on the input image and / or relevant context information using a large language model (such as Llama38B), and quantify the sentiment category or sentiment intensity index output by the large language model into an adjustment factor or adjustment mode. According to the adjustment factor or adjustment mode, modify the basic mapping function or rule in step S3, or directly perform offset or scaling adjustment on the calculated preliminary target control parameters to enhance or weaken the amplitude or intensity of specific facial actions, so that the final expression output is more consistent with the analyzed emotional state. For example, when analyzing the emotion of "joy", increase the movement amplitude of the servos related to the upward curl of the corners of the mouth.

[0023] Each servo has a specific number and control range (minimum and maximum pulse widths). Define these parameters using a dictionary (SERVO_SPECS):

[0024] SERVO_SPECS = {

[0025] 1: [1360, 1570, "RightEyeLR", "1570towardsnose"], 17: [1400, 1600, "LeftEyeLR", "1400towardsnose"],

[0026] 2: [1330, 1550, "RightEyeUD", "1550up"], 18: [1420, 1640, "LeftEyeUD", "1420up"],

[0027] 3: [1400, 1650, "RightUpperEyelid", "1400closed"], 19: [1340, 1650, "LeftUpperEyelid", "1340open"],

[0028] 4: [1750, 1900, "RightLowerEyelid", "1900open"], 20: [1100, 1300, "LeftLowerEyelid", "1300closed"],

[0029] 5: [1100, 1800, "RightBrowInner", "1100down"], 6: [1300, 1700, "RightBrowOuter", "1700up"],

[0030] 7: [1500, 1670, "RightEyeCorner", "1670up"], 8: [1480, 1850, "RightUpperLip", "1850up"],

[0031] 9: [1500, 1850, "MidUpperLip", "1850up"], 10: [1150, 1500, "LeftUpperLip", "1150up"],

[0032] 11: [1400, 1700, "RightUpperMouthCorner", "1700up"], 12: [1000, 1700, "RightLowerMouthCorner", "1700up"],

[0033] 13: [1480, 1600, "MouthOpen / Close", "1480closed"], 14: [1400, 1500, "MouthLR", "1400left"],

[0034] 21: [1100, 1800, "LeftBrowInner", "1800down"], 22: [1100, 1700, "LeftBrowOuter", "1700down"],

[0035] 23: [1350, 1500, "LeftEyeCorner", "1350up"], 24: [1400, 1550, "RightLowerLip", "1400down"],

[0036] 25: [1400, 1650, "MidLowerLip", "1650in"], 26: [1200, 1700, "LeftLowerLip", "1700down"],

[0037] 27: [1360, 1600, "LeftUpperMouthCorner", "1360up"], 28: [1410, 1650, "LeftLowerMouthCorner", "1650down"],

[0038] }

[0039] Among them, the information represented by each piece of data is as follows:

[0040] · Servo number: uniquely identifies a servo, such as 1, 2, 3, etc.

[0041] · Control range: represents the minimum and maximum pulse width values that the servo can accept, such as [1360, 1570].

[0042] · Function description: describes the facial movement controlled by the servo, such as "RightEyeLR" (right eye left - right movement).

[0043] · Movement direction description: further explains the relationship between the pulse width value and the movement direction, such as "1570towardsnose" (when the pulse width is 1570, the eye turns towards the nose).

[0044] Based on the calculated normalized eigenvalue, through the mapping function, a target pulse width value is generated for each servo. These values are organized into a dictionary (servo_commands), and then formatted into a specific string for sending to the servo controller.

[0045] A skin - type robot facial expression control system based on vision and language models includes the following:

[0046] · Image acquisition unit, used to acquire an input image containing face information;

[0047] · Facial feature extraction unit, used to process the input image using a vision processing model and extract facial feature information including three - dimensional facial key point coordinates;

[0048] · Preliminary parameter calculation unit, used to calculate preliminary parameter instructions for controlling the robot's facial servos according to the three - dimensional facial key point coordinates;

[0049] · An emotion analysis and fine-tuning unit for performing emotion analysis on the input image and / or context information using a large language model to obtain an emotion state index, and fine-tuning the preliminary parameter instructions according to the index to obtain the final parameter instructions;

[0050] A robot control unit for driving the robot's facial servo in real time according to the final parameter instructions to simulate corresponding facial expressions.

[0051] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:

[0052] 1. The skin-type robot facial expression driving method developed by the present invention analyzes visual and language information in a phased manner, and can achieve the real-time driving effect of the robot's facial expressions.

[0053] 2. The present invention uses a large language model to fine-tune the driving of the robot's facial expressions, so that the robot can also show good discrimination in some subtle expressions.

[0054] 3. The skin-type robot developed by the present invention has realistic facial expressions and can be applied to multiple interactive fields such as culture and tourism, medical care, and education. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.

[0056] Figure 1 It is the overall block diagram of the skin-type robot facial expression control system of the present invention.

[0057] Figure 2 It is the flowchart of the skin-type robot facial feature extraction unit of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts in the embodiments of the present invention belong to the scope of protection of the present invention.

[0059] Embodiment 1

[0060] The following specifically introduces that the present invention provides a skin-type robot facial expression control method and system based on visual and language models, including the following steps:

[0061] Step S1: Capture the image corresponding to the target expression through the camera installed in the robot's eyes, or read the image corresponding to the target expression through the reading device inside the robot.

[0062] Step S2: First, use a pre-trained vision model (such as DINO) to perform preliminary feature extraction on the input image. Then, apply a face detection model to detect whether there is a face in the image and locate the key facial regions. Next, apply a face mesh model to generate a face mapping mesh containing a predetermined number (478) of 3D facial key points.

[0063] Step S3: First, select several pre-defined groups of 3D facial key points, where the key points correspond to different regions of the face, such as eyes, eyebrows, lips, and corners of the mouth. Then, calculate the absolute distance or relative distance between the selected key points. The Euclidean distance is used to calculate the distance between two facial key points, and the distance is expressed as:

[0064]

[0065] where (x1, y1) represents the coordinates of one of the selected key points on the x-axis and y-axis, and (x2, y2) represents the coordinates of another point on the x-axis and y-axis.

[0066] At the same time, calculate the positional relationship of the key points relative to the reference points (such as the nasal base point, the glabella reference point), which is represented by the coordinate offset of the key points relative to the reference points (such as the nasal base, the glabella reference point). For example, calculate the horizontal offset of the center point of the lips relative to the nasal base to control the left and right movement of the mouth. Calculate the vertical offset of the eyebrow key points relative to the glabella reference point to control the up and down movement of the eyebrows. The coordinate offset is expressed as:

[0067] offset_x = point_x - reference_x;

[0068] offset_y = point_y - reference_y;

[0069] where offset_x and offset_y respectively represent the amounts of coordinate offset on the x-axis and y-axis, point_x and point_y represent the coordinates of the key points, and reference_x and reference_y represent the coordinates of the reference points.

[0070] Immediately afterwards, normalize the calculated distance or positional relationship according to the inherent facial dimensions (such as the inter-pupillary distance or the inter-canthal distance). Take the inherent facial dimension (usually the inter-pupillary distance or the inter-canthal distance) as the benchmark, divide other distances and positional relationships by this benchmark distance to obtain normalized feature values. For example, divide the distance of the eyes being open by the inter-pupillary distance to obtain the normalized degree of eye opening; divide the distance of the mouth being open by the inter-pupillary distance to obtain the normalized degree of mouth opening.

[0071] Finally, the normalized eigenvalue is converted into the preliminary target control parameters (pulse width values) corresponding to each servo (26 servos) through a preset basic mapping function (such as a linear mapping or piecewise linear mapping function) or rule. Specifically, the range of the normalized eigenvalue is divided into multiple sub-intervals, and different linear mapping functions are used within each sub-interval.

[0072] Step S4: Use a large language model (such as Llama 38B) to perform sentiment analysis on the input image and / or relevant context information, quantify the sentiment category or sentiment intensity index output by the large language model into an adjustment factor or adjustment mode, and modify the basic mapping function or rule in Step S3 according to the adjustment factor or adjustment mode, or directly perform an offset or scaling adjustment on the calculated preliminary target control parameters to enhance or weaken the amplitude or intensity of specific facial actions, so that the final expression output more conforms to the analyzed sentiment state. For example, when the "joy" sentiment is analyzed, increase the movement amplitude of the servo related to the upward curl of the corners of the mouth.

[0073] Each servo has a specific number and control range (minimum and maximum pulse widths). These parameters are defined by using a dictionary (SERVO_SPECS):

[0074] SERVO_SPECS = {

[0075] 1: [1360, 1570, "RightEyeLR", "1570towardsnose"], 17: [1400, 1600, "LeftEyeLR", "1400towardsnose"],

[0076] 2: [1330, 1550, "RightEyeUD", "1550up"], 18: [1420, 1640, "LeftEyeUD", "1420up"],

[0077] 3: [1400, 1650, "RightUpperEyelid", "1400closed"], 19: [1340, 1650, "LeftUpperEyelid", "1340open"],

[0078] 4: [1750, 1900, "RightLowerEyelid", "1900open"], 20: [1100, 1300, "LeftLowerEyelid", "1300closed"],

[0079] 5: [1100, 1800, "RightBrowInner", "1100down"], 6: [1300, 1700, "RightBrowOuter", "1700up"],

[0080] 7: [1500, 1670, "RightEyeCorner", "1670up"], 8: [1480, 1850, "RightUpperLip", "1850up"],

[0081] 9: [1500, 1850, "MidUpperLip", "1850up"], 10: [1150, 1500, "LeftUpperLip", "1150up"],

[0082] 11: [1400, 1700, "RightUpperMouthCorner", "1700up"], 12: [1000, 1700, "RightLowerMouthCorner", "1700up"],

[0083] 13: [1480, 1600, "MouthOpen / Close", "1480closed"], 14: [1400, 1500, "MouthLR", "1400left"],

[0084] 21: [1100, 1800, "LeftBrowInner", "1800down"], 22: [1100, 1700, "LeftBrowOuter", "1700down"],

[0085] 23: [1350, 1500, "LeftEyeCorner", "1350up"], 24: [1400, 1550, "RightLowerLip", "1400down"],

[0086] 25: [1400, 1650, "MidLowerLip", "1650in"], 26: [1200, 1700, "LeftLowerLip", "1700down"],

[0087] 27: [1360, 1600, "LeftUpperMouthCorner", "1360up"], 28: [1410, 1650, "LeftLowerMouthCorner", "1650down"],

[0088] }

[0089] Among them, the information represented by each piece of data is as follows:

[0090] · Servo number: uniquely identifies a servo, such as 1, 2, 3, etc.

[0091] · Control range: represents the minimum and maximum pulse width values that the servo can accept, such as [1360, 1570].

[0092] · Function description: describes the facial movement controlled by the servo, such as "RightEyeLR" (right eye left - right movement).

[0093] · Movement direction description: further explains the relationship between the pulse width value and the movement direction, such as "1570towardsnose" (when the pulse width is 1570, the eye turns towards the nose).

[0094] Based on the calculated normalized eigenvalue, a target pulse width value is generated for each servo through a mapping function. These values are organized into a dictionary (servo_commands), and then formatted into a specific string for sending to the servo controller.

[0095] A skin - type robot facial expression control system based on vision and language models includes the following:

[0096] · Image acquisition unit, used to acquire an input image containing face information;

[0097] · Facial feature extraction unit, used to process the input image using a vision processing model and extract facial feature information including three - dimensional facial key point coordinates;

[0098] · Preliminary parameter calculation unit, used to calculate preliminary parameter instructions for controlling the robot's facial servos according to the three - dimensional facial key point coordinates;

[0099] · Emotion analysis and fine - tuning unit, used to perform emotion analysis on the input image and / or context information using a large - language model to obtain an emotion state index, and fine - tune the preliminary parameter instructions according to this index to obtain the final parameter instructions;

[0100] Robot control unit, used to drive the robot's facial servos in real - time according to the final parameter instructions to simulate corresponding facial expressions.

[0101] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A method for controlling the facial expressions of a skin-type robot based on a vision and language model, characterized in that, It includes the following steps: Step S1: Obtain the input image; Capture the image corresponding to the target expression through the camera installed in the robot's eyes, or read the image corresponding to the target expression through the reading device inside the robot. Step S2: Extract the facial key point features of the input image; First, use a pre-trained vision model (such as DINO) to perform preliminary feature extraction on the input image, then apply a face detection model to detect whether there is a face in the image and locate the key facial regions, and then apply a face mesh model to generate a face mapping mesh containing a predetermined number (478) of three-dimensional facial key points. Step S3: Initially obtain the parameter instructions of the robot's facial servos; First, select several predefined groups of three-dimensional facial key points, and the key points correspond to different regions of the face, such as eyes, eyebrows, lips, corners of the mouth, etc. Then calculate the absolute distance or relative distance between the selected key points. The Euclidean distance is used to calculate the distance between two facial key points, and the distance is expressed as: where (x1, y1) represents the coordinates of one of the selected key points on the x-axis and y-axis, and (x2, y2) represents the coordinates of the other point on the x-axis and y-axis. At the same time, calculate the positional relationship of the key points relative to the reference points (such as the nasal base point, the glabella reference point), which is represented by the coordinate offset of the key points relative to the reference points (such as the nasal base, the glabella reference point). For example, calculate the horizontal offset of the center point of the lips relative to the nasal base to control the left-right movement of the mouth. Calculate the vertical offset of the eyebrow key point relative to the glabella reference point to control the up-down movement of the eyebrows. The coordinate offset is expressed as: offset_x = point_x - reference_x; offset_y = point_y - reference_y; where offset_x and offset_y respectively represent the amounts of coordinate offset on the x-axis and y-axis, point_x and point_y represent the coordinates of the key points, and reference_x and reference_y represent the coordinates of the reference points. Immediately normalize the calculated distance or positional relationship according to the inherent facial dimensions (such as the inter-pupillary distance or the inter-canthal distance). Take the inherent facial dimensions (usually the inter-pupillary distance or the inter-canthal distance) as the benchmark, and divide other distances and positional relationships by this benchmark distance to obtain the normalized eigenvalue. For example, divide the distance of the eyes being open by the inter-pupillary distance to obtain the normalized eye-opening degree; divide the distance of the mouth being open by the inter-pupillary distance to obtain the normalized mouth-opening degree. Finally, convert the normalized eigenvalue into the preliminary target control parameters (pulse width values) corresponding to each servo (26 servos) through a preset basic mapping function (such as a linear mapping or a piecewise linear mapping function) or rule. Specifically, divide the range of the normalized eigenvalue into multiple sub-intervals, and use different linear mapping functions within each sub-interval. Step S4: Drive the robot's facial servos; Perform sentiment analysis on the input image and / or relevant context information using a large language model (such as Llama 38B), and quantify the sentiment category or sentiment intensity index output by the large language model into an adjustment factor or adjustment mode. Modify the basic mapping function or rule in step S3 according to the adjustment factor or adjustment mode, or directly perform offset or scaling adjustment on the calculated preliminary target control parameter to enhance or weaken the amplitude or intensity of a specific facial action, so that the final expression output is more in line with the analyzed emotional state. For example, when the "joy" emotion is analyzed, increase the movement amplitude of the servo related to the upward curl of the corners of the mouth. Each servo has a specific number and control range (minimum and maximum pulse widths). These parameters are defined by using a dictionary (SERVO_SPECS): SERVO_SPECS = { 1: [1360, 1570, "RightEyeLR", "1570towardsnose"], 17: [1400, 1600, "LeftEyeLR", "1400towardsnose"], 2: [1330, 1550, "RightEyeUD", "1550up"], 18: [1420, 1640, "LeftEyeUD", "1420up"], 3: [1400, 1650, "RightUpperEyelid", "1400closed"], 19: [1340, 1650, "LeftUpperEyelid", "1340open"], 4: [1750, 1900, "RightLowerEyelid", "1900open"], 20: [1100, 1300, "LeftLowerEyelid", "1300closed"], 5: [1100, 1800, "RightBrowInner", "1100down"], 6: [1300, 1700, "RightBrowOuter", "1700up"], 7: [1500, 1670, "RightEyeCorner", "1670up"], 8: [1480, 1850, "RightUpperLip", "1850up"], 9: [1500, 1850, "MidUpperLip", "1850up"], 10: [1150, 1500, "LeftUpperLip", "1150up"], 11: [1400, 1700, "RightUpperMouthCorner", "1700up"], 12: [1000, 1700, "RightLowerMouthCorner", "1700up"], 13: [1480, 1600, "MouthOpen / Close", "1480closed"], 14: [1400, 1500, "MouthLR", "1400left"], 21: [1100, 1800, "LeftBrowInner", "1800down"], 22: [1100, 1700, "LeftBrowOuter", "1700down"], 23: [1350, 1500, "LeftEyeCorner", "1350up"], 24: [1400, 1550, "RightLowerLip", "1400down"], 25: [1400, 1650, "MidLowerLip", "1650in"], 26: [1200, 1700, "LeftLowerLip", "1700down"], 27: [1360, 1600, "LeftUpperMouthCorner", "1360up"], 28: [1410, 1650, "LeftLowerMouthCorner", "1650down"], } Among them, the information represented by each piece of data is as follows: · Servo number: uniquely identifies a servo, such as 1, 2, 3, etc. · Control range: represents the minimum and maximum pulse width values that the servo can accept, such as [1360, 1570]. · Function description: describes the facial movement controlled by the servo, such as "RightEyeLR" (right eye left - right movement). · Movement direction description: further explains the relationship between the pulse width value and the movement direction, such as "1570towardsnose" (when the pulse width is 1570, the eye rotates towards the nose). Based on the calculated normalized eigenvalue, a target pulse width value is generated for each servo through a mapping function. These values are organized into a dictionary (servo_commands) and then formatted into a specific string for sending to the servo controller. A skin-type robot facial expression control system based on a vision and language model, characterized in that, It includes the following: · An image acquisition unit for acquiring an input image containing face information; · A facial feature extraction unit for processing the input image using a vision processing model to extract facial feature information including three - dimensional facial key point coordinates; · A preliminary parameter calculation unit for calculating preliminary parameter instructions for controlling the robot's facial servos according to the three - dimensional facial key point coordinates; · An emotion analysis and fine - tuning unit for performing emotion analysis on the input image and / or context information using a large - language model to obtain an emotion state index, and fine - tuning the preliminary parameter instructions according to this index to obtain the final parameter instructions; A robot control unit for driving the robot's facial servos in real - time according to the final parameter instructions to simulate corresponding facial expressions.

Citation Information

Cited By

  • Voice-driven facial expression control method and system for humanoid robot

    CN121083626A

  • Humanoid robot voice-driven facial expression control method and system

    CN121083626B

  • Face repair system and face repair method

    CN121883487A