A lip motion control method, device and storage medium for a humanoid robot

By combining a dynamic pixel library and multiple control motors, high-fidelity synchronization of lip movements in humanoid robots was achieved, solving the problem of lip-syncing in Mandarin Chinese and improving human-computer interaction experience and speech intelligibility.

CN122392561APending Publication Date: 2026-07-14HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-04-01
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing humanoid robot lip movement models struggle to achieve natural lip-sync with Mandarin Chinese, resulting in poor mechanical abruptness and interactive experience, especially in noisy environments where speech intelligibility is insufficient.

Method used

By employing a dynamic visual library and capturing facial motion features, a spatiotemporal trajectory mapping between phonemes and lip movements is established. Multiple control motors are used to achieve a smooth expression of lip movements. Combined with dimensionality reduction mapping and nonlinear cosine function processing, lip movements that conform to the rules of Chinese pronunciation are generated.

Benefits of technology

It achieves high-fidelity synchronization of humanoid robot lip movements, improves human-computer interaction experience and speech intelligibility, solves the problem of lip shape and pronunciation coordination, and adapts to information transmission in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392561A_ABST
    Figure CN122392561A_ABST
Patent Text Reader

Abstract

The application discloses a kind of humanoid robot lip action control method, device and storage medium, belong to humanoid robot control technical field, method includes: setting dynamic visual library, the dynamic visual library includes multiple phoneme corresponding visual mapping table, the visual is the space-time trajectory based on time sequence change of lip part;Obtain to be pronounced text, and to be pronounced text is disassembled into basic phoneme, and the visual corresponding to each basic phoneme is obtained based on dynamic visual library;Visual is mapped to multiple control motor, and control motor drives humanoid robot to generate corresponding lip action.The application discards the traditional single frame static mapping scheme, by corresponding to the space-time trajectory based on time sequence change of corresponding lip part, realizes pronunciation and dynamic three-dimensional mouth shape corresponding;Then by mapping lip space-time trajectory to multiple control motor, realize the corresponding matching of sound and lip action, solve the problem of existing humanoid robot facial expression action distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method, device and storage medium for controlling lip movements of a humanoid robot. Background Art

[0002] With the gradual integration of humanoid robots into real scenarios such as public services and medical escort, multimodal human-computer interaction has become a current research hotspot. To establish long-term trust and a natural interaction experience, robots not only need to have the ability of intelligent dialogue, but also need to show non-verbal behaviors that conform to human intuition, such as facial expressions and head postures. Among them, lip synchronization closely related to speech plays a crucial role in eliminating the sense of incongruity in interaction. Research shows that human speech perception is essentially a process of audiovisual multimodal fusion, and the pronunciation dynamics at the visual level will directly reshape the brain's decoding of auditory signals. Therefore, when the lip movements of a robot do not match the speech in terms of time and space, it will not only destroy the user's immersion, but also trigger a strong cognitive conflict, thereby exacerbating the "uncanny valley" effect. In addition, under strong noise interference, the mouth shape dynamics in the visual modality can significantly improve speech intelligibility, and its gain effect is equivalent to increasing the signal-to-noise ratio by about 15 dB, which enables the robot to more efficiently transmit information to users in noisy scenarios such as shopping malls and stations.

[0003] Based on the requirement of consistency matching between the speech and lip movements of a bionic robot, there are mainly two forms of existing mainstream lip expressions. One is to construct a "phoneme-viseme" mapping dictionary based on visemes, convert the recognized speech units into preset mouth shape targets, and then generate a continuous facial action sequence through linear or polynomial interpolation algorithms. However, the current viseme modeling is mainly static. The existing viseme modeling reduces visemes to two-dimensional single static target states, ignoring the continuous dynamics of speech generation. This way of replacing complete actions with a single mouth shape will cause the lip shape of the robot to show a mechanical step feeling and lack natural transitions. In addition, the existing lip movement expression models are mainly based on the English language rule system, and they face many challenges when migrating to Mandarin Chinese. Complex syllables containing medial vowels in Chinese (such as: gua), whose movements are accompanied by multi-stage continuous sliding from slightly closed, round lips to wide open, are difficult to adapt to by the existing lip expression models. Summary of the Invention

[0004] In view of one or more of the above defects or improvement requirements of the prior art, the present invention provides a method for controlling lip movements of a humanoid robot to solve the problem of distorted facial expression actions of existing humanoid robots.

[0005] To achieve the above object, the present invention provides a method for controlling lip movements of a humanoid robot, which includes the following steps: S1. Set up a dynamic visual pixel library, which contains multiple phoneme-corresponding visual pixel mapping tables, where the visual pixel is the spatiotemporal trajectory of the lips based on temporal changes. S2. Obtain the text to be pronounced, break it down into basic phonemes, and obtain the visuals corresponding to each basic phoneme based on the dynamic visual library. S3. Map the visual pixels to multiple control motors, and control the motors to drive the humanoid robot to generate corresponding lip movements.

[0006] As a further improvement of the present invention, the setting of the dynamic pixel library includes: Capture facial movement features corresponding to continuous pronunciation; The audio is segmented at the peak of the continuous signal stream based on the pronunciation motion cycle, and the motion trajectory of a single phoneme is extracted. The motion trajectory of a single phoneme is normalized by the standard trajectory along the time axis, resulting in a dynamic positional feature space with spatiotemporal consistency for a single phoneme.

[0007] As a further improvement of the present invention, the setting of the dynamic pixel library also includes: Multiple phonemes are dimensionality-reduced and mapped to multiple visual elements; wherein each visual element contains at least one of the phonemes.

[0008] As a further improvement of the present invention, the facial motion feature is a facial blending shape coefficient based on the ARKit standard related to lip movement.

[0009] As a further improvement of the present invention, the mapping of visual pixels to multiple control motors in S3 includes: Obtain the basic phonemes corresponding to the visual pixels, map the basic phonemes to a sequence of 1 to 3 bottom visual pixels of length, fuse the two- and three-visual pixels in the bottom visual pixel sequence, and express the fused visual pixels by controlling the motor.

[0010] As a further improvement of the present invention, the dual-visual-pixel fusion is a fusion of initial consonant and final vowel, in any... At any given moment, the facial motion vector resulting from the fusion of initials and finals for: (Formula 1) in, and They represent the initials and finals respectively. The time corresponds to the 27-dimensional lip hybrid deformation state vector. It is a nonlinear cosine function with an a-th power bias.

[0011] As a further improvement of the present invention, the three-pixel fusion is the fusion of initial consonants and diphthongs. At any given time, the facial motion vector after the fusion of initial consonants and diphthongs... For the following piecewise functions: (Formula 3) In this function, the first segment represents the transition from initial consonant to final vowel 1, the second segment represents the transition from final vowel 1 to final vowel 2, V1 is the continuous visual pixel state vector of initial consonant, V2 is the continuous visual pixel state vector of final vowel 1, and V3 is the continuous visual pixel state vector of final vowel 2. The time segmentation ratio, This is the interpolation weight function.

[0012] As a further improvement of the present invention, mapping the visual element to multiple control motors in step S3 is equivalent to mapping the visual element to multiple control motors in a dimension-reduced manner. The dimension-reduced mapping is achieved through the following formula: (Formula 4) in, This represents a smoothed 27-dimensional digital hybrid shape vector sequence at time t. This represents the 14-dimensional physical motor control command sequence after dimensionality reduction mapping. Assign weights to the 27-dimensional digital hybrid shape vector sequence to the 14-dimensional physical motor control command sequence.

[0013] The present invention also includes a humanoid robot lip movement control device, comprising: The data acquisition module is used to acquire the visual elements corresponding to the phonemes, wherein the visual elements are the spatiotemporal trajectories of the lips based on temporal changes; The mapping module is used to reduce the dimensionality of phonemes and map them to visual pixels; The control module includes multiple control motors mounted on the humanoid robot. The control module receives text to be pronounced, decomposes the text into basic phonemes, maps the basic phonemes into a sequence of 1-3 low-level visual pixels, fuses the two- and three-viewpoint visual pixels in the low-level visual pixel sequence, and expresses the fused visual pixels through the control motors. The present invention also includes a computer storage medium storing instructions for executing the humanoid robot lip movement control method.

[0014] The aforementioned improved technical features can be combined with each other as long as they do not conflict with each other.

[0015] In summary, the beneficial effects of the above-described technical solutions conceived by this invention compared with the prior art include: (1) The humanoid robot lip movement control method of the present invention abandons the traditional single-frame static mapping scheme. It realizes the correspondence between pronunciation and dynamic three-dimensional mouth shape by mapping phonemes to the corresponding spatiotemporal trajectory of the lips based on temporal changes. Then, by mapping the spatiotemporal trajectory of the lips to multiple control motors, the lip movement expression is realized by using multiple control motors, thereby achieving the correspondence and matching between sound and lip movement, and solving the problem of distortion in the facial expression of existing humanoid robots. The present invention achieves a smooth display of the facial expression of humanoid robots by matching pronunciation with lip movement, enabling humanoid robots to achieve high-fidelity and physically consistent lip-shape synchronization, solving the problem of coordination between pronunciation and lip shape, and improving the human-computer interaction experience. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the humanoid robot lip movement control method in an embodiment of the present invention. Figure 2 This is a system framework diagram of the humanoid robot lip movement control method in an embodiment of the present invention; Figure 3 This is a flowchart illustrating the creation of the dynamic pixel library in this embodiment of the invention; Figure 4 These are lip shape variation diagrams of two typical visual pixels in embodiments of the present invention; Figure 5 This is a schematic diagram of the humanoid robot device in an embodiment of the present invention. Figure 5 (a) represents a humanoid robot platform. Figure 5 (b) represents the degrees of freedom markers for the lips and jaw; Figure 5 (c) represents a demonstration of humanoid robot lip movements; Figure 6 These are curves showing the changes in the average acceleration across nine key dimensions based on various methods within the test corpus in this embodiment of the invention. Figure 7 This is a comparison chart of Jawopen values ​​and real-person values ​​for trajectories generated by various methods based on the test corpus in this embodiment of the invention. Figure 8 This is a diagram showing the facial expression changes of a humanoid robot performing a real-time test in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0018] Furthermore, unless otherwise stated, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0019] Example: Please see Figure 1 , Figure 2 As shown, the humanoid robot lip movement control method in a preferred embodiment of the present invention includes the following steps: A dynamic visual pixel library is set up, which contains multiple phoneme-to-visual pixel mapping tables. The visual pixels are the spatiotemporal trajectories of the lips based on temporal changes. The text to be pronounced is obtained, and it is broken down into basic phonemes. Based on the dynamic visual library, the visuals corresponding to each basic phoneme are obtained. The visual pixels are mapped to multiple control motors, which then drive the humanoid robot to generate corresponding lip movements.

[0020] The humanoid robot lip movement control method of this invention abandons the traditional single-frame static mapping scheme. Instead, it maps phonemes to the corresponding spatiotemporal trajectories of the lips based on temporal changes, achieving a correspondence between pronunciation and dynamic three-dimensional mouth shape. Then, by mapping the spatiotemporal trajectories of the lips to multiple control motors, it utilizes these motors to express lip movements, achieving a matching between sound and lip action, thus solving the problem of distorted facial expressions in existing humanoid robots. This invention achieves smooth display of humanoid robot facial expressions by matching pronunciation with lip movements, enabling high-fidelity lip-shape synchronization that conforms to the laws of physical motion, solving the coordination problem between pronunciation and lip shape, and improving the human-computer interaction experience.

[0021] Furthermore, as an optional embodiment of the present invention, the setting of the dynamic pixel library in the present invention includes: Capture facial movement features corresponding to continuous stiffening; The audio is segmented at the peak of the continuous signal stream based on the pronunciation motion cycle, and the motion trajectory of a single phoneme is extracted. The motion trajectory of a single phoneme is normalized by the standard trajectory along the time axis, resulting in a dynamic positional feature space with spatiotemporal consistency for a single phoneme.

[0022] This invention primarily captures facial vocal expressions and then segments continuous vocalizations into individual phonemes, thereby establishing a correspondence between individual phonemes and facial expressions (motion trajectories). Building upon this, the invention further normalizes the motion trajectories of individual phonemes, mapping them one-to-one with standard trajectories on the time axis to obtain motion trajectories within a standard time period, resulting in a dynamic visual position feature space with spatiotemporal consistency for each individual phoneme. Specifically, this invention records the pronunciation process of Mandarin initials and finals using a depth camera at a 60Hz sampling rate, acquiring a continuous stream of facial hybrid shape coefficient data. To address minor differences in the duration of multiple vocalizations, a dynamic time warping mechanism is introduced to resample similar pixel sequences to a standard length, facilitating strict temporal alignment during subsequent interpolation and co-pronunciation fusion. Optionally, after obtaining the dynamic pixel pool, this invention performs smoothing and cleaning using a moving average filter, precisely removing high-frequency noise from the visual sensor while retaining the low-frequency principal components of vocalization deformation, ensuring high smoothness and physiological rationality of the motion trajectory.

[0023] As an optional embodiment of the present invention, the facial motion features in the present invention are facial blending shape coefficients related to lip movements under the ARKit standard. Specifically, the present invention uses depth camera technology to track facial features of lip muscle movements during continuous pronunciation in a standard data format; then, it selects the ARKit standard containing 52 facial blending shape coefficients to construct a parameterized space, and then extracts 27 core bases strongly correlated with lip movements from it; the facial features corresponding to each phoneme acquired by the depth camera are expressed through the 27 core bases, and the motion trajectory of a single phoneme includes the spatial trajectory of the 27 core bases based on time changes. After acquiring the motion trajectories corresponding to all phonemes, the standard trajectories of each motion trajectory along the time axis are normalized, and finally a dynamic view position feature space with high spatiotemporal consistency is obtained. It can be understood that the dynamic view position feature space here refers to the continuous dynamic change map composed of 27 core bases within a single standard time.

[0024] In one specific embodiment of the present invention, a device equipped with a TrueDepth camera system is used as a high-precision data acquisition terminal. This acquisition terminal integrates an infrared projector, an infrared camera, and a topological dot matrix, constructing a 3D topological structure by projecting over 30,000 invisible infrared dots onto the face. Simultaneously, the device coordinates with the Live Link Face data transmission protocol to achieve data synchronization. During data acquisition, the subject repeatedly pronounces specific target phonemes, and the subtle movements of facial muscles are analyzed in real time using the ARKit framework, sampling at a constant sampling rate of 60 frames per second, outputting a real-time data stream containing 52 sets of standard mixed shape coefficients. Then, 27 motion parameters directly related to lip shape are extracted and recorded from the data stream, and... Figure 3The process involves building a 3D dynamic pixel library.

[0025] Furthermore, as an optional embodiment of the present invention, the setting of the dynamic pixel library in the present invention further includes: Multiple phonemes are downgraded to mappings, and multiple phoneme combinations are mapped to multiple visual elements; where each visual element contains at least one phoneme.

[0026] Specifically, the Mandarin Chinese phonological system contains 21 initials and 39 finals, resulting in hundreds of pronunciation combinations. If each pronunciation combination were represented by one or more control motors, the lips of a humanoid robot simply could not accommodate so many motors, leading to abnormal lip movements. However, in facial animation, the richness of acoustic and visual spaces differs; auditory dissimilarities often exhibit high visual similarity, i.e., homographs. Therefore, to eliminate redundant pronunciation expressions, this invention performs a many-to-one dimensionality reduction mapping of phonetic combinations based on kinematic features. By merging phonemes with similar morphology and dynamic expression, a dictionary containing 14 core dynamic visual elements is constructed to achieve the mapping from phonemes to lip movements. The specific visual element mapping table is as follows:

[0027] This invention uses 14 dynamic visual pixels to represent phonemes in the Mandarin Chinese phonological system, thereby significantly reducing the difficulty of controlling motors to express lip movements. Specifically, the dictionary of the 14 dynamic visual pixels in this invention follows the visual clustering principle of the MPEG-4 FBA standard and is adapted to the characteristics of Mandarin pronunciation, solving the problem of complex medial and phonological coordination in Mandarin pronunciation, and accurately expressing Mandarin pronunciation movements, as detailed below. Figure 4 As shown, it intuitively demonstrates the spatiotemporal evolution of the visual pixel V1_BPM pronunciation cycle with the ARKit format JawOpen coefficient, and V2_F with the MouthUpperUp coefficient (i.e., the average of the left and right MouthUpperUp coefficients), highlighting the correspondence between lip movements and related values.

[0028] Furthermore, as an optional embodiment of the present invention, step S3 of the present invention, mapping visual pixels to multiple control motors, includes: Obtain the basic phonemes corresponding to the visual pixels, map the basic phonemes to a sequence of 1 to 3 bottom visual pixels of length, fuse the two- and three-visual pixels in the bottom visual pixel sequence, and express the fused visual pixels by controlling the motor.

[0029] Specifically, based on the dynamic visual element library, each visual element can be limited to a standard duration to obtain the basic mouth shape sequence of each phoneme. However, in the expression of continuous phonemes, although simple linear splicing can produce continuous actions, the actions will feel disjointed and cannot reproduce the complex coordinating effects in Chinese. For example, when pronouncing the syllable "tu", the lips already show the rounded tendency of the vowel / u / the moment they pronounce the consonant / t / . Therefore, it is necessary to adjust the transition of the phoneme to the corresponding visual element for specific pronunciations. This invention considers the underlying visual element sequence in the basic phonemes and integrates the expression of different types of visual elements based on the rules of Chinese pronunciation, so that the generated lip shape is closer to the actual expression. Specifically, this invention first deconstructs Chinese characters into underlying basic phonemes, and then maps the basic phonemes to underlying visual element sequences of length 1 to 3; for pure simple vowels (such as "a") or compound vowels with pre-recorded complete dynamic features (such as "ei"), their pronunciation is not affected by other factors, and they are expressed as single visual elements. For Chinese characters whose pronunciations contain 2-3 underlying visual elements, a visual element fusion mechanism is designed to approximate the actual expression in Chinese.

[0030] Specifically, bipixel fusion includes the fusion of "initial consonant and final vowel": In any At any given moment, the facial motion vector resulting from the fusion of initials and finals for: (Formula 1) in, and They represent the initials and finals respectively. The time corresponds to the 27-dimensional lip hybrid deformation state vector. For a nonlinear cosine function with a power bias, The dynamic fusion weights vary with t, where This is a normalized timeline. Specifically, when... When the initial consonant is close to 0, the action is dominated by the initial consonant; when it is close to 1, the final vowel is dominant. However, in actual pronunciation, the initial consonant often accounts for a very short percentage of the actual acoustic duration (usually less than 20%). If the actual acoustic time is strictly linearly fused at a uniform speed in the control of a physical robot, the underlying control motors will lack the response time to complete the full stroke of key actions such as lip closing and lip biting, resulting in severe "visual swallowing" and high-frequency vibration. To balance "shortness of pronunciation" and "physical integrity of action," a weighting is used... A nonlinear cosine function with a power of a bias was designed. Specifically: (Formula 2) By intervening in this cosine function curve, the speed of the lip movements at the start and end positions can be ensured to be smooth.

[0031] Optionally, typically when simulating a case where the initial consonant accounts for a small proportion of the acoustic time, the initial retention time of the initial consonant control weight will be artificially lengthened (e.g., setting a=0.7).

[0032] Alternatively, the fusion of three visual elements may include the fusion of "initial consonant and diphthong": At any given moment, the facial motion vector resulting from the fusion of the initial consonant and the diphthong. For piecewise functions: (Formula 3) The first function segment represents the transition from initial consonant to final vowel 1, and the second function segment represents the transition from final vowel 1 to final vowel 2. This invention divides the complex long-cycle pronunciation into two continuous sub-transition states on the time axis, setting the syllable to deconstruct three continuous visual state vectors as V1 (initial consonant), V2 (final vowel 1), and V3 (final vowel 2). This is the interpolation weighting function. In the normalized pronunciation period... Within, by time segmentation ratio The visual pixel action is divided into two expressive phases. Optionally, the temporal segmentation ratio... This is an empirical coefficient.

[0033] Specifically, in the first function segment, the sub-stage normalized time is: ,along with arrive As x increases from 0 to 1, the weight gradually shifts from biased towards V1 to biased towards V1.

[0034] In the second function segment, the sub-stage normalized time is: As τ increases from λ to 1 and x increases from 0 to 1, the weights gradually shift from biased towards V2 to biased towards V3.

[0035] This invention mathematically decouples complex high-dimensional spatiotemporal trajectories into two steps through a cascading mechanism, thereby reducing the risk of swallowing intermediate sounds during the expression of three visual elements and endowing the humanoid robot's lip movement control device with delicate Chinese rhythmic expression capabilities.

[0036] Furthermore, as an optional embodiment of the present invention, step S3 of mapping visual pixels to multiple control motors includes: dimensionality-reduced mapping of visual pixels to multiple control motors. Dimensionality reduction mapping is achieved through the following formula: (Formula 4) in, This represents a smoothed 27-dimensional digital hybrid shape vector sequence at time t. This represents the sequence of 14-dimensional physical motor control commands after dimensionality reduction mapping. This represents the weight allocation from the 27-dimensional digital hybrid shape vector sequence to the 14-dimensional physical motor control command sequence. Specifically, here... This invention establishes a mapping matrix between the baseline data of human facial vocalization acquired by a visual facial capture device and the basic physical response curve during robot execution, relative to a 27-dimensional digital hybrid deformation. Due to the highly nonlinear deformation characteristics of the facial silicone skin stretching and the physical tolerances caused by mechanical linkages, purely mathematical linear mapping often fails to achieve ideal human-like realism. Therefore, this invention utilizes high-frequency dynamic lip-shape testing captured by a depth camera to manually compensate and fine-tune the lip-shape parameters under motor control. This compensates for the nonlinear physical errors of the flexible facial material of the humanoid robot, endowing the robot with bionic lip dynamics and realizing a closed-loop link from digital generation algorithm to physical entity execution.

[0037] Furthermore, regarding the humanoid robot lip movement control method of this invention, a high-fidelity humanoid robot head is used as a verification platform. A hybrid transmission method combining linkage mechanisms and cable drives is employed to realize the humanoid robot's movement. The internal control motor of the humanoid robot simulates the contraction and relaxation of human facial muscles by traction of anchor points on the inner side of flexible silicone skin. For the lip shape generation task, this invention features a high-density configuration of 14 degrees of freedom in the lip area, resulting in a humanoid robot with flexible movement that can reproduce the full spectrum of mouth shape dynamics, including lip closure, lip rounding, lip abduction, and jaw opening and closing. Specifically, as shown below... Figure 5 As shown.

[0038] To comprehensively evaluate the performance of the algorithm, this invention constructed a visual pixel-covered reading test corpus containing 20 characters. This dataset consists of four representative Chinese short sentences (each with five syllables, see table below). Sample selection followed the principles of "maximizing phoneme diversity" and "high kinematic difficulty," aiming to cover the pronunciation and mouth shape of all key initials and finals in Mandarin. This test set covers highly challenging pronunciation scenarios: for example, S1 focuses on the coarse articulation of diphthongs, requiring smooth processing of the three-stage continuous deformation from slightly open to rounded lips to fully open lips when pronouncing / gua / ; S3 contains a large number of bilabial consonants ( / b / , / p / , / m / ) to verify the ability of the upper and lower lips to completely close, thus addressing the visual swallowing phenomenon of the baseline method; S4 tests the system's high-frequency dynamic response limit during rapid switching between flat and rounded consonants.

[0039]

[0040] To comprehensively evaluate the effectiveness of the proposed audio driver framework and the independent contributions of each core module, this invention designs four alignment and ablation driving strategies, specifically defined as follows: Method A: Traditional static baseline method. It extracts only static visual target points and relies on simple linear interpolation to generate transition trajectories, representing the most basic engineering implementation paradigm and serving as a lower bound for performance evaluation.

[0041] Method B: Dynamic Visual Pixel Direct Approach. This method introduces the high-fidelity dynamic curve library constructed in this paper, but employs a hard time-axis segmentation strategy at phoneme boundaries. The aim is to verify the quality of native dynamic visual pixel data and, conversely, demonstrate the necessity of introducing a syllable fusion algorithm.

[0042] Method C: Basic Co-pronunciation Method. This method introduces a co-pronunciation fusion algorithm based on a perceptual dictionary, building upon Method B. A backend kinematic smoothing module is not yet configured; it is specifically used to isolate and verify the effect of the co-pronunciation algorithm on improving facial motion coherence.

[0043] Method D: Complete Physical Perception Method. Building upon Method C, this method further integrates audio RMS energy dynamic amplitude modulation and moving average filtering mechanisms. The aim is to verify the core value of physical energy perception and kinematic filtering in filtering out mechanical jitter, ensuring hardware safety, and improving anthropomorphism.

[0044] In the experimental phase, this invention used facial motion capture equipment to collect visual motion capture data and synchronized video recordings of real subjects reading the four sets of test corpora. The clean audio extracted from the video was used as the driving input for the four algorithms to obtain the average acceleration change curves for nine key dimensions based on the test corpus under each method. This curve served as the basis for subsequent quantitative evaluation and anthropomorphism analysis, specifically as follows: Figure 6 As shown.

[0045] Furthermore, to assess the friendliness of the generated trajectory to the robot's physical actuators, this invention introduces average acceleration as a metric for motion smoothness analysis, as shown in the table below. Jerk, as the derivative of acceleration with respect to time, directly reflects the jerking and impact intensity of the system's motion. Excessive Jerk can lead to overheating, unnatural vibrations, or even damage to the mechanical structure.

[0046]

[0047] Although all 27 mixed shape coefficients are related to the lips and palate, some degrees of freedom (left and right chin offset and lip crease) are often dormant or slightly activated in natural pronunciation. Including such features in the global evaluation would dilute the actual error and jitter. Therefore, this invention extracts the nine most active key dimensions of the lip region (data with significantly higher variance than other dimensions), and calculates their average acceleration using real facial data as a benchmark to quantify the smoothing cost, such as... Figure 6As shown, Method B, lacking a phoneme transition algorithm, exhibits a high jump value during pronunciation switching, resulting in an average Jerk peak value as high as 20.84. While Method C employs co-pronunciation between adjacent visual pixels, its dynamic data dictionary interpolation leads to more numerical variations compared to Method A, resulting in a slightly higher Jerk value. Both methods also exhibit significant numerical abrupt changes. In contrast, Method D demonstrates superior kinematic smoothness, with a Jerk value (1.01) even slightly lower than that of human motion capture data (1.78). This is because real optical motion capture inevitably includes facial muscle tremors and high-frequency capture noise from marker points. It can be seen that the humanoid robot lip movement control method in this invention effectively filters out micro-vibrations while restoring the macroscopic pronunciation amplitude, providing a safe and reliable data trajectory for the robot actuator from a fundamental kinematic perspective.

[0048] Furthermore, to objectively evaluate the accuracy and naturalness of the generated lip movements, this invention introduces Pearson correlation coefficient (PCC) and root mean square error (RMSE) as evaluation metrics. PCC assesses the morphological similarity between the generated trajectory and the motion capture trajectory of a real person, reflecting the synchronization capability in the algorithm's temporal phase, pronunciation rhythm, and fluctuation trend. RMSE assesses the deviation of the generated trajectory in absolute spatial amplitude, reflecting the accuracy of the physical opening and closing scale. In feature selection, this invention isolates local lip deformations and designates mandibular opening as the core feature for human-likeness evaluation. This is because the temporomandibular joint is the fundamental node for lip and jaw movements; subtle deformations such as lip stretching, closing, and protrusion are all superimposed on the physical spatial framework determined by mandibular opening and closing. Mandibular opening directly determines the basic tension of the lip muscles and the global spatial coordinates; if the JawOpen trajectory is distorted, the local movements attached to it will produce severe global visual misalignment. Simultaneously, there is a direct physical mapping between the periodic opening and closing of the mandible and the short-term vocal capacity envelope and the core vowel formants. Furthermore, as the articulatory site with the largest displacement amplitude and the most visually significant impact, even a slight temporal misalignment or amplitude overshoot in the mandibular trajectory can instantly disrupt the natural flow of lip-sound synchronization. Therefore, the trajectory quality of JawOpen constitutes an important criterion for anthropomorphism assessment. The table below shows the objective anthropomorphism assessment structure for four methods under two evaluation metrics.

[0049]

[0050]

[0051] It is evident that traditional baseline methods have significant limitations: Method A, lacking dynamic transitions, results in a rigid trajectory with an average PCC of only 0.155, failing to fit the rhythm of human speech. Methods B and C, by introducing co-articulation or dynamic transition mechanisms, improve phase alignment (average PCCs of 0.443 and 0.489, respectively), but based on... Figure 7 The JawOpen numerical comparison chart shows that, due to the lack of fine amplitude modulation capabilities in audio processing, both methods often exhibit excessive mouth opening, resulting in RMSE values ​​as high as 0.305 and 0.280, respectively. In contrast, Method D in this invention demonstrates significant advantages. Thanks to the combined effect of energy mapping and EMA filtering mechanisms, Method D's average PCC jumps to 0.595, achieving the best rhythmic fit among the four test sentences; furthermore, its average RMSE is reduced to 0.177 (a 36.8% reduction compared to Method C).

[0052] Furthermore, to verify the physical conversion effect of objective indicators, this paper sends the generated high-frequency trajectory data to the actuator of a highly realistic humanoid robot for real-time testing, such as... Figure 8 As shown in the figure, it can be seen that the humanoid robot lip movement control method in this invention basically achieves a similarity to the mouth shape of a real person.

[0053] The present invention also includes a humanoid robot lip movement control device for executing a humanoid robot lip movement control method, comprising: The data acquisition module is used to acquire the visual image corresponding to the phoneme, which is the spatiotemporal trajectory of the lips based on temporal changes; The mapping module is used to reduce the dimensionality of phonemes and map them to visual pixels; The control module contains multiple control motors mounted on the humanoid robot. The control module is used to receive the text to be pronounced, decompose the text into basic phonemes, map the basic phonemes into a sequence of 1 to 3 low-level visual pixels, fuse the two- and three-viewpoint visual pixels in the low-level visual pixel sequence, and express the fused visual pixels through the control motors.

[0054] The present invention also includes a computer storage medium storing instructions for performing the steps of the humanoid robot lip movement control method.

[0055] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for controlling the lip movements of a humanoid robot, characterized in that, Includes the following steps: S1. Set up a dynamic visual pixel library, which contains multiple phoneme-corresponding visual pixel mapping tables, where the visual pixel is the spatiotemporal trajectory of the lips based on temporal changes. S2. Obtain the text to be pronounced, break it down into basic phonemes, and obtain the visuals corresponding to each basic phoneme based on the dynamic visual library. S3. Map the visual pixels to multiple control motors, and control the motors to drive the humanoid robot to generate corresponding lip movements.

2. The humanoid robot lip movement control method according to claim 1, characterized in that, The settings for the dynamic pixel library include: Capture facial movement features corresponding to continuous pronunciation; The audio is segmented at the peak of the continuous signal stream based on the pronunciation motion cycle, and the motion trajectory of a single phoneme is extracted. The motion trajectory of a single phoneme is normalized by the standard trajectory along the time axis, resulting in a dynamic positional feature space with spatiotemporal consistency for a single phoneme.

3. The humanoid robot lip movement control method according to claim 2, characterized in that, The setting of the dynamic pixel library also includes: Multiple phonemes are dimensionality-reduced and mapped to multiple visual elements; wherein each visual element contains at least one of the phonemes.

4. The humanoid robot lip movement control method according to claim 3, characterized in that, The facial motion features are facial blending shape coefficients related to lip movement under the ARKit standard.

5. The humanoid robot lip movement control method according to any one of claims 1 to 5, characterized in that, The mapping of visual pixels to multiple control motors in S3 includes: Obtain the basic phonemes corresponding to the visual pixels, map the basic phonemes to a sequence of 1 to 3 bottom visual pixels of length, fuse the two- and three-visual pixels in the bottom visual pixel sequence, and express the fused visual pixels by controlling the motor.

6. The humanoid robot lip movement control method according to claim 5, characterized in that, The fusion of the two visual pixels is the fusion of the initial consonant and the final vowel, in any At any given moment, the facial motion vector resulting from the fusion of initials and finals for: in, and They represent the initials and finals respectively. The time corresponds to the 27-dimensional lip hybrid deformation state vector. It is a nonlinear cosine function with an a-th power bias.

7. The humanoid robot lip movement control method according to claim 5, characterized in that, The three-pixel fusion refers to the fusion of initial consonants and diphthongs. At any given time, the facial motion vector resulting from the fusion of initial consonants and diphthongs... For the following piecewise functions: In this function, the first segment represents the transition from initial consonant to final vowel 1, the second segment represents the transition from final vowel 1 to final vowel 2, V1 is the continuous visual pixel state vector of initial consonant, V2 is the continuous visual pixel state vector of final vowel 1, and V3 is the continuous visual pixel state vector of final vowel 2. The time segmentation ratio, This is the interpolation weight function.

8. The humanoid robot lip movement control method according to any one of claims 1 to 5, characterized in that, In step S3, mapping visual elements to multiple control motors is equivalent to dimensionality-reduced mapping of visual elements to multiple control motors. This dimensionality-reduced mapping is achieved through the following formula: in, This represents a smoothed 27-dimensional digital hybrid shape vector sequence at time t. This represents the 14-dimensional physical motor control command sequence after dimensionality reduction mapping. Assign weights to the 27-dimensional digital hybrid shape vector sequence to the 14-dimensional physical motor control command sequence.

9. A humanoid robot lip movement control device, characterized in that, include: The data acquisition module is used to acquire the visual elements corresponding to the phonemes, wherein the visual elements are the spatiotemporal trajectories of the lips based on temporal changes; The mapping module is used to reduce the dimensionality of phonemes and map them to visual pixels; The control module includes multiple control motors mounted on the humanoid robot. The control module is used to receive the text to be pronounced, decompose the text into basic phonemes, map the basic phonemes into a sequence of 1 to 3 low-level visual pixels, fuse the two- and three-viewpoint visual pixels in the low-level visual pixel sequence, and express the fused visual pixels through the control motors.

10. A computer storage medium, characterized in that, The computer storage medium stores instructions for each step of the method for controlling the lip movements of a humanoid robot.