Phoneme time axis driven high-definition video mouth shape automatic synthesis method
Through the phoneme timeline driven method, combined with the frame index mapping table and three-dimensional facial feature points, the synchronization and posture consistency problems of lip synthesis in high-resolution videos are solved, and high-precision generation of lip synthesis in high-definition videos is achieved.
Patent Information
- Application Number
- CN202511121728.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing video lip-syncing technology has problems such as complex speech-to-video mapping in high-resolution videos, insufficient synchronization accuracy in the time dimension, and misalignment between lip movements and speech rhythm, resulting in inaccurate lip-syncing generation and abrupt transitions.
Using a phoneme timeline-driven approach, the correspondence between phonemes and lip shapes is accurately modeled by constructing a phoneme timeline, a frame index mapping table, and three complementary derivation chains. Combined with three-dimensional facial feature points and head posture trajectories, the time synchronization and posture consistency of lip-shape synthesis are achieved.
It achieves high-precision alignment of audio and video in the temporal dimension, and consistency of lip shape and head posture in the spatial dimension, improving the accuracy and realism of lip-sync video generation. It is suitable for application scenarios such as digital humans and voice-driven animation.
Smart Images

Figure CN120640052A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a phoneme timeline-driven high-definition video lip-sync automatic synthesis method. Background Art
[0002] In the current field of video generation and speech synthesis, lip-syncing technology for realistic speech and video synchronization has become a key support link in digital humans, virtual anchors, human-computer interaction, and film and television post-production. Traditional methods generally rely on deep learning models to extract features from the input speech, map them to lip movements or lip rendering parameters through learning, and generate or modify facial image sequences based on the prediction results in the rendering module to achieve face-driven audio and video synchronization. However, although current speech-driven video lip-syncing technology has achieved widespread application, it still faces multiple technical bottlenecks, especially in the areas of high-resolution video, accurate speech synchronization, smooth lip transitions, and interpretable modeling.
[0003] Existing lip synthesis technologies can be broadly categorized into data-driven end-to-end models and parameter-controlled image manipulation methods. Typical examples of the former include speech-to-video synthesis frameworks based on LSTM, Transformer, VAE, or diffusion models. These methods typically use speech features such as Mel spectrum, MFCC, or wav2vec as input, and combine historical frame lip shapes or images as conditions to generate a single frame or segment of facial image output. Their advantage lies in their strong overall modeling capabilities and the lack of need to explicitly model the regular associations between speech and lip shapes. However, these methods suffer from the following issues:
[0004] First, the mapping of speech to video is a complex high-dimensional to high-dimensional conversion process. Although the end-to-end model can learn the association between features through a large amount of training data, it lacks the inherent constraints of language structure and pronunciation principles. As a result, the lip movements driven by speech may appear to be "wrong sound but right shape" or "right shape but wrong sound". That is, the lip shape generated by the model is visually reasonable but inconsistent with the actual speech, or misaligned with the speech sequence.
[0005] Second, existing models lack high-precision synchronization mechanisms in the temporal dimension. Most methods perform frame-by-frame predictions at a fixed frame rate rather than being controlled by external speech structure. This makes it prone to cumulative delays or synchronization drift in fast speech, pauses, or nonlinear speech segments. This is particularly true in the task of generating high-definition, long-term video sequences. Because phoneme boundaries are not explicitly modeled, the boundaries of lip movements cannot be accurately located, resulting in abrupt, disjointed, or blurred transitions between adjacent phonemes. Summary of the Invention
[0006] To address the aforementioned technical issues, a new method for automatic lip-sync synthesis in high-definition video driven by a phoneme timeline is provided. By constructing a phoneme timeline, a frame index mapping table, and three complementary derivation chains, the correspondence between phonemes and lip shapes is accurately modeled. Phase continuity and traceable calibration of articulation points are achieved at the keyframe level. By combining three-dimensional facial feature points with head posture trajectories, posture-consistent lip-sync synthesis is achieved. This method offers advantages such as structural interpretability, high temporal synchronization accuracy, strong visual naturalness, and high-resolution adaptability. It significantly improves the accuracy and realism of lip-sync video generation and is suitable for a variety of application scenarios, including digital humans, voice-driven animation, and virtual broadcasting.
[0007] In order to achieve the above objects, the technical solution adopted by the present invention is:
[0008] A phoneme timeline-driven high-definition video lip-sync automatic synthesis method, the method comprising:
[0009] Step 1: Obtain input text or speech, use a fixed phoneme set and deterministic pronunciation rules to generate a phoneme sequence and corresponding start and end time points, and establish a phoneme timeline. Using the timestamp of the original video as a slave clock, construct a frame index mapping table that maps the target frames of the original video to the phoneme timeline one by one. Output the phoneme timeline, frame index mapping table, and target frame set.
[0010] Step 2: Under the constraints of the phoneme timeline and frame index mapping table, phoneme slicing and phase calibration are performed sequentially. Based on a fixed and limited lip shape dictionary and articulation position rules, three complementary derivation chains are established simultaneously. A single forward decision without feedback and the principle of minimum interval preservation are used to complete conflict resolution and key frame determination, resulting in a lip shape key frame sequence that meets phase continuity and boundary traceability.
[0011] Step 3: Extract 3D facial feature points and head posture trajectories based on the original video, align the lip-sync keyframe sequence with the head posture trajectory according to the frame index mapping table to obtain a posture-consistent lip-sync sequence; output the video lip-sync synthesis result based on the posture-consistent lip-sync sequence with the original video resolution and synchronized with the phoneme timeline.
[0012] Furthermore, in step 2, the process of sequentially executing phoneme slicing and phase calibration includes: dividing the duration of each phoneme into three types of phoneme slices: starting phase, transition phase and steady-state phase; according to the frame index mapping table, each target frame is uniquely assigned to its corresponding phoneme slice, and the phase type label is recorded to form a phase-labeled frame column.
[0013] Furthermore, in step 2, based on a fixed and limited dictionary of mouth shape configurations and rules of articulation positions, three complementary deduction chains are established simultaneously, including: a main chain, a first side chain, and a second side chain; the main chain is a one-to-one mapping from phonemes to mouth shape categories, used to generate a basic mouth shape sequence without transitions; the first side chain is used for rhyme sequence connection, sequentially connecting the starting phase, transition phase, and steady-state phase belonging to the same syllable to generate segment labels of monotonic opening and closing trajectories and lip shape change trajectories; the second side chain corresponds to the articulation position contacts, and is used to mark discrete event sites of lip closure, tongue tip contact, and soft palate opening for phonemes that require closure, friction, or explosion based on the deterministic contact rules of the lips, gums, hard palate, soft palate, and glottis; the mouth shape categories of the main chain, the segment labels of the first side chain, and the event sites of the second side chain are merged on the same timeline according to the phoneme time axis to obtain a three-way chain merge table.
[0014] Furthermore, in step 2, the phoneme timeline is first read, which consists of phoneme records arranged in chronological order, and each phoneme record contains a phoneme identifier and a start and end time point; the frame index mapping table is read, and the frame index mapping table gives the correspondence between the target frame and the phoneme timeline, and each record contains the target frame number, the corresponding phoneme identifier and the frame timestamp; the mouth shape configuration dictionary and the pronunciation position contact rules are loaded; the mouth shape configuration dictionary is a fixed and finite mapping set, which gives a one-to-one mapping relationship between phoneme categories and mouth shape categories; the pronunciation position contact rules are a fixed rule set, which gives the contact conditions and corresponding event types of the lips, gums, hard palate, soft palate and glottis.
[0015] Furthermore, in step 2, the process of generating the main chain includes: traversing each phoneme record in the order of the phoneme time axis, searching the mouth shape category corresponding to the phoneme category in the mouth shape dictionary, and generating a basic mouth shape entry; adding time boundary information to each basic mouth shape entry, and the time boundary is consistent with the start time point and end time point of the phoneme record; writing all basic mouth shape entries into the main chain in time sequence to obtain a basic mouth shape sequence without transition; any entry in the main chain only covers the time range of the phoneme record to which it belongs, and does not cross the time boundary of adjacent phonemes.
[0016] Furthermore, in step 2, the generation process of the first side chain includes: dividing the phoneme time axis into syllables, and the time range of the syllable is determined by the minimum starting time point and the maximum ending time point of the continuous phoneme records constituting the syllable; within each syllable, according to the order of the starting phase, transition phase and steady-state phase, the corresponding target frame numbers are connected in series to generate a segment label of the monotonic opening and closing trajectory; the segment label records the segment type, the starting target frame number, and the ending target frame number; within each syllable, according to the deterministic order of the lip shape from contraction to relaxation or from relaxation to contraction, a segment label of the lip shape change trajectory is generated. The segment label records the lip shape change type, the starting target frame number, and the ending target frame number; the segment label of the monotonic opening and closing trajectory and the segment label of the lip shape change trajectory are written into the first side chain in chronological order; any segment label shall not exceed the time range of the syllable to which it belongs; the process of generating the second side chain includes: according to the pronunciation position contact rule, the phoneme time axis is checked one by one to determine the phoneme set that needs to be closed, rubbed or exploded; for the phonemes belonging to double lip closure, the discrete event site of lip closure is marked within the time range of the phoneme record; the time position of the event site is preferentially taken as one of the starting time point and the ending time point of the phoneme record or two; for phonemes that belong to tongue tip contact, the discrete event sites of tongue tip contact are marked within the time range of the phoneme recording; the time position of the event site is preferentially the target frame number close to the middle of the phoneme; for phonemes that belong to soft palate opening, the discrete event sites of soft palate opening are marked within the time range of the phoneme recording; the time position of the event site is preferentially the target frame number near the starting time point of the phoneme recording; the above discrete event sites are written into the second side chain in chronological order; each event site is bound to a target frame number. If the event time point is not equal to any target frame timestamp, it is assigned to the target frame number closest in time.
[0017] Furthermore, in step 2, the lip-sync keyframe sequence is obtained through the following process: searching the overlapping regions of records at the time boundaries of adjacent phonemes in the three-pronged chain merge table, and generating seam segments based on a fixed seam strategy; seam segments exist only within the common boundaries of adjacent phonemes and do not cross common boundaries; based on the segment labels of the first side chain and the discrete event sites of the second side chain, keyframes are marked at phase boundaries, seam boundaries, and event sites; keyframes are discrete values at the target frame number. For short-duration segments that appear in the three-pronged chain merge table and are below the minimum duration standard, they are merged into segments that are adjacent in time and of the same category; the merging operation does not change the position of the marked keyframes; all keyframes are combined into a lip-sync keyframe sequence in chronological order.
[0018] Furthermore, step 3 specifically includes: in each target frame of the original video, calibrating three-dimensional facial feature points based on a fixed facial geometry template; obtaining a head posture trajectory according to the time sequence of the target frames, wherein the head posture trajectory includes the head rotation and head translation of the target frame; assigning each key frame in the lip shape key frame sequence to the corresponding target frame number according to the frame index mapping table; reading the head rotation and head translation of the target frame from the head posture trajectory and attaching them to the corresponding lip shape key frame; between adjacent lip shape key frames, using spherical linear interpolation for head rotation and linear interpolation for head translation to obtain a posture-consistent lip shape sequence; generating a mouth area grid on the target frame with the three-dimensional facial feature points as anchor points; based on According to the lip shape configuration of the frame in the pose-consistent lip sequence, the mesh is subjected to analytical geometric deformation and written into the mouth area; when the required lip shape combination does not exist in the original video, a fixed oral geometry template and lip-tooth texture patch that match the current posture are selected and embedded into the mouth area after posture transformation; a transition zone of fixed width is set at the boundary of the mouth area, and boundary feathering is performed; based on the facial reference area of the same frame, color transfer is performed to make the mouth area consistent with the reference area in brightness and chroma; using the phoneme timeline as the only master clock, the replaced and processed mouth area is written back to the corresponding target frame, keeping the frame timestamp and resolution unchanged; the output video lip-synthesis result is consistent with the resolution of the original video and synchronized with the phoneme timeline.
[0019] Furthermore, when the time position of a lip-sync keyframe is not equal to the timestamp of any target frame, it is assigned to the target frame number with the smallest time difference; if there are two equidistant target frames, the one with the smaller time is taken; gesture-consistent interpolation is implemented between adjacent lip-sync keyframes without crossing the common boundaries of adjacent phonemes.
[0020] Compared with the existing technology, the present invention has the following advantages: it can achieve high-precision alignment of audio and video in the temporal dimension and maintain the consistency of lip shape configuration and head posture in the spatial dimension. By using the phoneme timeline as the only master clock, the present invention avoids the problem of misalignment between lip movements and speech rhythm in traditional data-driven methods, significantly improving the timing accuracy of lip synthesis results. On this basis, each phoneme is divided into three phases: onset, transition, and steady state through phoneme slicing and phase calibration. Combining a fixed lip shape configuration dictionary with the articulation contact point rules, three complementary derivation chains are established to achieve structured modeling of the pronunciation process, effectively solving the problem of the lack of physical explanation of speech configuration in existing methods. The present invention also uses a key frame arbitration mechanism to calibrate discrete key frames at seam boundaries, event sites, and phase transitions, constructing a traceable and interpolable lip shape key frame sequence. Furthermore, combining three-dimensional facial feature point extraction and head posture estimation, the present invention synchronizes the lip shape sequence with head rotation and translation parameters in each frame, so that the generated lip shape maintains posture consistency during the dynamic process, avoiding visual jumps and configuration drift. Furthermore, the present invention utilizes high-precision geometric mesh deformation and texture fusion technology to achieve a natural transition between the mouth region and the original video in terms of boundaries, color, and resolution, ensuring the generated video possesses high definition and a high sense of realism. Overall, the present invention balances language structure drive, timing accuracy, visual continuity, and voice matching, making it suitable for digital human broadcasting, voice-to-video conversion, virtual character driving, and other applications requiring high lip-sync accuracy and image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of a phoneme timeline-driven high-definition video lip-syncing automatic synthesis method proposed by the present invention;
[0022] Figure 2 Schematic diagram of a phoneme time axis and frame index mapping table in an embodiment of the present invention;
[0023] Figure 3 Schematic diagram of the working principles of three complementary deduction chains in an embodiment of the present invention;
[0024] Figure 4 This is a graph of a phoneme phase slicing experiment in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0026] Reference Figure 1 As shown, a phoneme timeline driven high-definition video lip-sync automatic synthesis method, the method comprising:
[0027] Step 1: Obtain input text or speech, use a fixed phoneme set and deterministic pronunciation rules to generate a phoneme sequence and corresponding start and end time points, and establish a phoneme timeline. Using the timestamp of the original video as a slave clock, construct a frame index mapping table that maps the target frames of the original video to the phoneme timeline one by one. Output the phoneme timeline, frame index mapping table, and target frame set.
[0028] Step 1 first establishes a unified data entry for input text or speech. When the input type is text, a fixed set of phonemes and deterministic pronunciation rules are used by a preset word-to-phoneme mapping module to generate a phoneme sequence and corresponding start and end time points. When the input type is speech, an end-to-end forced alignment model is first used to segment the continuous speech stream. Then, acoustic analysis is performed on each segment using the same fixed set of phonemes and deterministic pronunciation rules to obtain phoneme identifiers and synchronously output start and end time points. Finally, all phoneme records are written in chronological order on a linear timeline to establish a phoneme timeline. Because the phoneme timeline is the sole master clock source for all subsequent timing operations, time accuracy and stability must be guaranteed during the generation process. To this end, this embodiment adopts a dual-buffered writing method: a primary buffer is used to record the start and end time points of the current phoneme in real time, and a secondary buffer is used to temporarily store the gap information between the end time of the previous phoneme and the start time of the current phoneme. If the gap length is less than half the duration of a single frame, it is directly incorporated into the current phoneme, thus avoiding the occurrence of unmapped zero-length gaps in the phoneme timeline. After completing the establishment of the phoneme timeline, the system starts the coupling process with the original video, sets the original video decoder to frame-by-frame pull mode and reads the native timestamp of each frame. This timestamp is regarded as the reference node of the slave clock with the timestamp of the original video as the slave clock.
[0029] On this basis, the system extracts the target frames from the original video in frame order. The target frame selection strategy can be set to full-frame acquisition or acquisition at a fixed downsampling rate as needed, as long as it ensures that the target frame with a time difference of no more than half a frame can be found at each time point within the coverage range of the phoneme timeline. Subsequently, the system constructs a frame index mapping table that corresponds the target frames of the original video to the phoneme timeline one by one. The specific approach is to use binary search to locate the phoneme record to which the timestamp belongs in the phoneme timeline after parsing the timestamp of the target frame. If the target frame timestamp falls between the start and end time points of a phoneme record, the mapping relationship between the target frame number and the phoneme identifier is recorded; if the target frame timestamp falls exactly at the junction of two adjacent phoneme records, the target frame number is mapped to the phoneme identifier with a smaller time difference based on the comparison result with the time difference between the two phoneme start time points; at the same time, the target frame timestamp field is retained in the mapping table entry for subsequent phase operation calls. In order to improve the processing throughput in long-duration high-definition video scenarios, the mapping table construction process adopts a partition parallel mechanism. First, the target frame set is divided into equal-time segments according to the decoding order of the original video. Each partition is completed by an independent thread to complete the target frame timestamp reading, phoneme timeline query and mapping entry generation, and then the partition results are merged according to the global time order. After all the mapping entries are written, a complete frame index mapping table can be obtained. At this point, the embodiment outputs the phoneme timeline, frame index mapping table and target frame set through the above process, laying a unified and precise timing foundation for the subsequent execution of phoneme slicing and phase calibration under the constraints driven by the phoneme timeline, ensuring that the subsequent steps can still maintain frame-level synchronization accuracy when processing large-scale high-definition content and avoid the impact of cumulative errors on the quality of lip synthesis.
[0030] Step 2: Under the constraints of the phoneme timeline and frame index mapping table, phoneme slicing and phase calibration are performed sequentially. Based on a fixed and limited lip shape dictionary and articulation position rules, three complementary derivation chains are established simultaneously. A single forward decision without feedback and the principle of minimum interval preservation are used to complete conflict resolution and key frame determination, resulting in a lip shape key frame sequence that meets phase continuity and boundary traceability.
[0031] Step 2 uses the phoneme time axis and frame index mapping table as strict timing constraints. First, it enters the phoneme slicing stage. The system reads the phoneme records arranged continuously on the phoneme time axis, and instantly divides the duration of each phoneme into three types of phoneme slices in the memory: starting phase, transition phase, and steady-state phase. The phase type label is written to the slice by looking up the table; then the frame index mapping table is called to match the target frames within the time range of the phoneme record one by one, uniquely assign the target frame number to the phoneme slice to which it belongs, and synchronously record the phase type, thereby completing the phase calibration and generating a phase-labeled frame column. In this process, all write operations follow the real-time writing and delay verification double buffering strategy to ensure phase consistency. Then the system loads a fixed and limited dictionary of mouth shape configurations and pronunciation position rules and initializes three complementary deduction chains. The main chain writes basic mouth shape entries without transitions and adds time boundary information when traversing the phoneme time axis in sequence according to the one-to-one mapping of phonemes to mouth shape categories, ensuring that any entry in the main chain only covers the time range of the phoneme record to which it belongs and does not cross the boundary of adjacent phonemes; the first side chain sequentially connects the starting phase, transition phase and steady-state phase of the same syllable within the syllable time range output by the syllable division module to generate a monotonic opening and closing trajectory segment label, and within the same syllable, the lip shape changes from contraction to relaxation or from contraction to relaxation. Segment labels for lip shape trajectory are generated sequentially from relaxation to contraction. Both types of segment labels are strictly written into the first side chain and cannot exceed the time range of the syllable to which they belong. The second side chain, based on the articulation position rule, determines the set of phonemes that require closure, friction, or explosion while traversing the phoneme timeline. It also writes discrete event locations for events such as lip closure, tongue tip contact, and soft palate opening. The target frame number selection strategy for event locations prioritizes the start and end time points of the phoneme record, or the target frame number near the middle of the phoneme. If the event time point cannot be strictly aligned with any target frame timestamp, it is assigned to the target frame with the smallest time difference to ensure event traceability. After the three chains are independently generated, the system performs a three-chain merge operation in a unified temporal coordinate system. The main chain's lip shape category, the first side chain's segment labels, and the second side chain's event locations are merged into a three-chain merge table along the phoneme timeline. The merge table is stored using a time index plus category segment double-key indexing scheme to facilitate subsequent retrieval. Next, the system enters a single forward decision phase without feedback. When traversing the three-way chain merge table, if it finds that there is an overlapping area of records at the time boundary of adjacent phonemes, the conflict detection module is triggered. The module first generates patchwork segments based on a fixed patchwork strategy and ensures that the patchwork segments only exist within the common boundary of adjacent phonemes and do not cross the common boundary. Then, the patchwork segments and phase segments are compared by weight based on the shortest interval preservation principle. The shorter segments are selected for retention, and the longer segments are truncated at the boundary without affecting the rest. The entire process does not require backtracking or secondary adjustment of the already decided results, ensuring that the computational complexity increases linearly.The merge table after the adjudication is sent to the key frame calibration module, which combines the segment labels of the first side chain and the discrete event sites of the second side chain to calibrate key frames at each phase boundary, seam boundary and event site. The position of the key frame is a discrete value on the target frame number and must be consistent with the timestamp to which it is bound. For short-term segments that appear in the merge table and are lower than the shortest duration standard, the system immediately calls the shortest interval merging subroutine to merge them into segments that are adjacent in time and of the same category. During the merging process, the calibrated key frame positions are not moved, so as to avoid short-term segments destroying the continuity of lip movements. All key frames are automatically combined in memory in chronological order to form a lip key frame sequence. The system also performs an integrity check on the key frame sequence to ensure that the sequence covers the entire phoneme timeline and there are no duplicate or missing timestamps. At this point, step 2 outputs a lip-sync keyframe sequence that meets phase continuity and boundary traceability through a continuous process of phoneme slicing, phase calibration, construction of three complementary derivation chains, single forward arbitration without feedback, conflict resolution based on the shortest interval maintenance principle, and keyframe positioning. This lays an accurate and stable keyframe foundation for subsequent posture-consistent interpolation and high-definition lip-sync synthesis.
[0032] Step 3: Extract 3D facial feature points and head posture trajectories based on the original video, align the lip-sync keyframe sequence with the head posture trajectory according to the frame index mapping table to obtain a posture-consistent lip-sync sequence; output the video lip-sync synthesis result based on the posture-consistent lip-sync sequence with the original video resolution and synchronized with the phoneme timeline.
[0033] In the implementation process of the present invention, step 3 uses the lip shape key frame sequence output from step 2 as the only lip shape reference. First, the target frame is extracted frame by frame in the original video decoding thread according to the time sequence of the frame index mapping table. A fixed facial geometry template is called in each target frame to perform multi-scale feature fitting. The three-dimensional facial feature point set is obtained by minimum residual iterative calculation using the predefined contour, facial features and perioral control points on the template. At the same time, the head posture parameters are extracted in real time accompanied by the projection matrix inversion process. The head posture parameters are composed of a rotation vector and a translation vector and have a one-to-one correspondence with the target frame timestamp, thereby forming a complete head posture trajectory. Next, the system calls the frame index mapping table to accurately bind each key frame in the lip key frame sequence to the original video timeline according to the target frame number, synchronously reads the rotation vector and translation vector of the corresponding target frame and writes them into the key frame record, thereby generating a key frame queue with the initial posture; to ensure the temporal continuity between lip movement and head movement, the system uses spherical linear interpolation for the rotation vector and linear interpolation for the translation vector in the adjacent interval of the key frame, and the interpolation results are written in real time to the newly generated posture-consistent lip sequence, which is the same length as the phoneme timeline on the timeline and the interpolation value at the key frame is completely consistent with the original parameter.
[0034] Subsequently, the system is driven by the posture-consistent lip sequence in the rendering thread, and constructs a mouth area mesh on each target frame with three-dimensional facial feature points as anchors. The mesh vertices correspond to the template control points one by one, and the analytic geometric deformation parameters corresponding to the lip category of the frame are found in the lip shape mapping module. According to the analytic geometric deformation parameters, vertex-by-vertex displacement and scaling operations are performed on the mesh vertex coordinates. After the deformation is completed in the local coordinate system, it is mapped back to the target frame image coordinate system through the head posture rotation matrix and translation vector and written into the mouth area; when there is a native mouth texture matching the current lip category in the original video frame, the system gives priority to the native texture to avoid unnecessary pixel replacement, otherwise the template model and its lip and tooth texture slice closest to the current head posture are selected from the fixed oral geometry template library, and embedded into the mouth area after posture transformation and resolution consistency processing; in order to reduce replacement Traces, the system automatically generates a fixed-width transition zone at the boundary of the mouth area, realizes boundary feathering through Gaussian weighted fusion, and performs color transfer with the brightness distribution and chromaticity distribution of the facial reference area in the same frame as the target, so that the mouth area and facial skin are consistent in brightness and chromaticity channels; the above operations are immediately written back to the target frame after completion in the image buffer, keeping the timestamp and resolution of the target frame unchanged, and sending the result frame to the output buffer; when all target frames are processed, the output buffer writes out the final frame stream in the original video time sequence, forming a video lip-synthesis result that is consistent with the original video resolution and synchronized with the phoneme time axis. The lip shape of the result at any moment is directly driven by the phoneme time axis and is completely consistent with the head posture in space, thereby ensuring the stable unity of lip movement logic, visual coherence and high-definition rendering accuracy in long-term and complex corpus scenarios.
[0035] Furthermore, in the implementation framework of the present invention, the core motivation for phoneme slicing and phase calibration is to accurately describe the evolution of the morphological changes of the vocal organs with the smallest observable time domain unit. Therefore, the system actively divides the duration of each phoneme into three types of phoneme slices: the starting phase, the transition phase, and the steady-state phase. The starting phase corresponds to the instantaneous convergence segment of the vocal organs from the previous target configuration to the current phoneme specific configuration. At this time, the lips, tongue tip or soft palate are in the initial stage of rapid closing, lifting or opening. The transition phase describes the dynamic segment of the vocal organs crossing between two stable configurations. The rate of change of the vocal tract morphology is the largest and it mixes the synergistic characteristics of the previous and next phonemes. The steady-state phase covers the segment where the vocal organs have reached the current phoneme configuration and remain relatively still in this configuration. The airflow pattern and resonance peak position tend to be stable, and the acoustic characteristics are quasi-constant output. After calculating the total duration of each phoneme by reading the phoneme timeline, the system assigns three boundary segments in memory using a proportional partitioning strategy. It then retrieves all target frames within these three segments using a frame index mapping table, writes the target frame number to the corresponding phoneme slice, and records the phase type label in an additional field. To avoid cross-phoneme confusion, target frames are assigned based on a unique attribution principle: once a target frame is assigned to a phoneme slice, it will not appear in other phoneme slices, thus forming a temporally non-overlapping sequence of phase-labeled frames. Specifically, the bilabial plosive / p / visually exhibits typical strong onset characteristics: the onset phase shows rapid lip closure and a rapid increase in muscle tension, the transition phase is characterized by a sudden opening of the lips with a slight lip rebound, and the steady-state phase is almost negligible because the lips immediately transition to the next phoneme configuration after the airflow is released. Therefore, when processing / p / , the onset phase occupies the majority of the phoneme duration and corresponds to the largest number of target frames. Taking the nasal sound / m / as an example, its onset phase includes the subtle movement of the lips from half-closed to fully closed, the transition phase is the process of the soft palate descending to open the nasal passage, and the steady-state phase shows a stable shape in which the lips remain closed while the cheeks slightly bulge. Therefore, the number of steady-state phase frames in the phase-annotated frame series is significantly greater than the onset and transition phases. For the vowel / a / , the onset phase is characterized by an increase in the speed of mandibular sinking and a widening of the lip shape. The transition phase shows a continued expansion of the oral cavity volume and a rapid approach to the resonant position. The steady-state phase is characterized by the vertical opening of the oral cavity remaining unchanged and the tongue surface fixed in the low area. When selecting frames, the system detects the two local extremes of the tongue back at its lowest point and the lip aperture at its largest point and adds the adjacent target frames to the steady-state slice accordingly, so that the phase-annotated frame series stably covers the entire vowel. Through the above-mentioned refinement, the three types of phoneme slices are not only precisely separated in time, but also capture the phased laws of the deformation of the vocal organs at the visual dynamics level. This provides a high-resolution timing benchmark for the subsequent mapping, conflict resolution, and keyframe positioning of the three complementary deduction chains, ensuring that the final synthesized lip movement maintains physical consistency and visual coherence at both the shot scale and the speech scale.
[0036] Furthermore, when the system enters the stage of constructing three complementary derivation chains, the primary prerequisite is that a fixed and limited lip shape dictionary and articulation rules have been loaded, and that the phoneme timeline and phase annotation frame columns remain in a random access state in memory. The generation logic of the main chain is based on a one-to-one mapping of phonemes to lip shape categories. It focuses on the central configuration presented by each phoneme in the steady-state phase. Therefore, when traversing the phoneme timeline, it directly retrieves the lip shape category paired with the current phoneme identifier in the lip shape dictionary, and then writes the lip shape category along with the start and end time points of the phoneme record into the main chain entry. The basic lip shape sequence obtained in this way without transitions is arranged in time sequence by phonemes and does not overlap. Taking the syllable "ma" as an example, the phonemes / m / and / a / are mapped to the closed and open lip shape categories respectively. Since the steady-state duration of / m / is extremely short and the steady-state duration of / a / is relatively long, it can be seen on the main chain that the duration segment of the closed lip shape category is extremely narrow while the open lip shape category occupies a significantly longer time range, which visually foreshadows the necessity of subsequent transition.
[0037] The first side chain performs rhyme sequence connection, taking as input the syllable segmentation results and a phase-labeled frame sequence. Within a syllable, the system concatenates the target frame numbers of the onset, transition, and steady-state phases in chronological order, assigning a segment label to each phase continuum and indicating the segment type. If multiple opening and closing cycles occur within a syllable, such as the diphthong "ai," which undergoes a distinct contraction-extension process from / a / to / i / , the first side chain generates two monotonic opening and closing trajectory labels: the first marking rapid jaw lowering until the oral cavity reaches maximum volume, and the second marking tongue tip elevation and lip retraction until closure stabilizes. For example, in the bilabial plosive sequence "ba," the transition phase of / b / overlaps with the onset phase of / a / , resulting in a short-duration but large-amplitude opening and closing segment label in the first side chain. The subsequent steady-state phase of / a / is labeled as a monotonic opening and closing segment with slow lip relaxation. Each segment label records the starting target frame number, the ending target frame number and their order within the syllable, thereby providing sparse control points for subsequent curve fitting.
[0038] The second side chain focuses on discrete articulation event sites. It scans the phoneme timeline one by one according to the articulation rules to determine whether to trigger lip closure, tongue tip contact, or soft palate opening. Taking / p / as an example, it satisfies both lip closure and explosion conditions. Therefore, the system marks a lip closure event site at the start time point of / p / and an explosion release event site near its end time point. Looking at the fricative / s / , which corresponds to alveolar friction, the system inserts a tongue tip contact event site in the middle of the phoneme; the nasal / n / requires soft palate opening, and the system places the soft palate opening event site at the target frame number near the start of / n / . In the second side chain, each event site is bound to a precise target frame number and contains an event type identifier. These discrete points will be directly converted into keyframes later.
[0039] Once the three chains have been written, the system initiates the three-chain merging process. This process uses the phoneme timeline as a unified coordinate, aligning the main chain entries, the first side chain segment labels, and the second side chain event sites to the same timeline. These are then sorted by timestamps and written into a three-chain merge table. The time index in the merge table refers to the target frame number rather than the absolute time, facilitating subsequent direct retrieval in the video frame domain. For the "ma" example, the merged table contains overlapping entries for lip closure, soft palate opening event sites, and lip opening and closing segment labels at the common boundary between the end of / m / and the beginning of / a / . The system then determines the keyframes and adjudicates the spliced segments based on the shortest interval during subsequent non-feedback adjudication. For the diphthong "ai," the merged table shows a continuous transition from mouth opening to lip closing, interspersed with multiple tongue tip contact event sites. These sites are prioritized during adjudication to ensure traceability of tongue tip movements. Through this merge table, the system is able to integrate steady-state spatial information, dynamic timing information, and articulation contact information into the same data structure, which not only provides full-dimensional context for subsequent key frame positioning, but also avoids the retrieval overhead caused by random crossing of multi-source linked lists, fundamentally improving the computational efficiency and timing consistency of high-definition lip-synthesis automatic synthesis on long speech sequences.
[0040] Furthermore, before the system processing flow enters the main chain generation phase, the runtime environment first synchronously loads the phoneme timeline and the frame index mapping table into the cache through memory mapping. Each phoneme record in the phoneme timeline stores its phoneme identifier, start time point, and end time point in the form of a structure, while each record in the frame index mapping table provides three fields: target frame number, corresponding phoneme identifier, and frame timestamp. This parallel design of the two tables allows both time sequence retrieval and random positioning to be completed within constant time complexity. Subsequently, the system calls the persistent resource manager to load the mouth shape dictionary and pronunciation position contact rules. The mouth shape dictionary is a fixed and finite mapping set that merges all occurring phoneme categories into several distinguishable mouth shape categories, such as mapping / p / , / b / , / m / to "closed lip", mapping / f / to "lower lip teeth", mapping / t / , / d / , / n / , / l / to "pre-alveolar", and mapping / t / , / d / , / n / , / l / to "pre-alveolar". / 、 / / is mapped to "hard palate fricative", / k / , / g / , / h / are mapped to "soft palate opening", / a / is mapped to "largest opening", / i / is mapped to "flat lip grin", / u / is mapped to "rounded lip convergence", / e / is mapped to "medium opening", / / is mapped to "rounded lips half open", so that even when facing hundreds of phonemes, only about a dozen mouth shape categories need to be maintained, thus controlling the upper limit of the rendering template. The articulation contact rules also exist in the form of a fixed rule set, which formulates Boolean conditions and event type mappings for the five major parts of the lips, gums, hard palate, soft palate and glottis. For example, the rule table stipulates that "if the phoneme belongs to / p / , / b / , / m / , / / then trigger the lip closure event" "If the phoneme belongs to / t / , / d / , / n / , / l / then trigger the tongue tip alveolar contact event" "If the phoneme belongs to / k / , / g / , / / then trigger the soft palate lifting and nasal cavity closing events" "If the phoneme belongs to / / triggers the glottal closure event", and for / s / , / z / , / / 、 / / and other fricatives uniformly mark alveolar friction events, for / x / , / / etc. mark the hard palate friction event. After loading, the system hashes the two tables into hash tables for constant time query.
[0041] The main chain generation process begins with the first phone record on the phoneme timeline. The system reads the phoneme identifier and performs a lookup in the lip shape dictionary to determine the phoneme's lip shape category. A basic lip shape entry is then created, combining this category with the start and end time points of the phoneme record and writing it into a pre-allocated sequential array. For phonemes that are temporally adjacent but have distinct articulation locations, such as the segments " / b / 0-100ms, / a / 100-300ms, / t / 300-350ms" on the timeline, the main chain will consecutively write three basic lip shape entries: "Closed Lip 0-100ms," "Maximum Opening 100-300ms," and "Pre-Alveolar 300-350ms." Each entry strictly covers the time range of its respective phoneme and does not cross adjacent phoneme boundaries. Because the main chain completely ignores transition phases, the above sequence does not include any information about the transition from closed lip to open lip or from open lip to pre-alveolar lip. This dynamic is then supplemented by the first and second side chains. For example, the vowel string " / i / 0-120ms, / u / 120-240ms" is returned by dictionary lookup as "flat open lips 0-120ms" and "rounded lips 120-240ms." The main chain maintains the hard boundary between the two without providing a connecting path, thus providing a clear time switching point for subsequent feedback-free judgment. The main chain uses a sequential write cache and a lockstep counter during writes to ensure that the temporal order of entries is not disrupted even in a multi-threaded environment. An atomic bit flag is also appended to the end of the entry to indicate that the entry has not yet been conflict-checked by the first and second side chains. This flag is reset to zero during the subsequent three-chain merge phase to indicate that the entry has been safely integrated into the global time sequence. Through this design, the main chain not only quickly provides the steady-state spatial reference of each phoneme in the entire corpus, but also uses a fixed and limited lip shape dictionary to compress high-dimensional pronunciation classifications into a limited lip shape template, so that the storage and computing scale of high-definition lip rendering can be controlled within a predictable range; at the same time, the instantiation of the articulation contact point rules ensures that subsequent discrete event sites can be accurately aligned with the main chain time boundary, reserving reliable anchor points for key frame determination and high-precision transition interpolation.
[0042] Furthermore, during actual operation, the generation of the first side chain is organized around the syllable, a higher-level phonological unit. Its fundamental purpose is to abstract the opening and closing rhythms and lip shape gradients scattered among continuous phonemes into two clear segment sequences, thereby providing constraints for the subsequent interpolation trajectory. The system first calls the existing pinyin word segmenter or international phonetic symbol segmentation algorithm to divide the phoneme timeline into syllables. The division results are written to the buffer in chronological order. The time range of each syllable is directly determined by the minimum start time point and maximum end time point of the internal continuous phoneme records, forming a non-overlapping closed interval in time sequence. The system then reads the phase-annotated frame sequence within the same syllable and concatenates these target frame numbers in the natural order of the starting phase, transition phase, and steady-state phase to generate a monotonic opening and closing trajectory segment label. The segment type field "opening and closing trajectory" and the starting and ending target frame numbers are written into the label structure. For example, the syllable "ba" contains the transition of / p / and the steady state of / a / . The system identifies a monotonically increasing opening and closing trajectory from lip closure to full mouth opening, with its starting frame typically occurring in the frame before the end of / p / and its ending frame falling midway through the steady state of / a / . The system then retrieves the extreme points of lip curvature within the same syllable based on the deterministic order of lip shape transitions from contraction to relaxation or relaxation to contraction. Using interpolation, it performs a one-time global minimum and maximum search on the curvature sequence. It then generates segment labels for the lip shape change trajectory, using the target frame numbers near the extreme points as boundaries. The label structure contains the segment type field "lip shape change" as well as the starting and ending target frame numbers. For the diphthong "ai," for example, the initial jaw lowering and lip widening produce a lip shape change segment from contraction to relaxation, followed by a relaxation-to-contraction segment caused by tongue tip elevation and reconvergence of the mouth. Both labels are written into the first side chain, maintaining their temporal order. When inserting segment labels, the system also performs boundary validation checks to ensure that the start and end target frame numbers of any label do not exceed the time range of the syllable to which it belongs. If a boundary violation is detected, the label is automatically truncated to the syllable boundary. After the first side chain is generated, each segment label has a defined time period and clear functional classification, providing a sparse control segment for continuous curve fitting for subsequent feedback-free adjudication.
[0043] Relative to the continuous trajectory abstraction of the first side chain, the second side chain focuses on capturing discrete events of articulatory contact points. The system traverses the phoneme timeline item by item according to the articulatory contact point rules. First, it checks in the rule table whether the current phoneme meets conditions such as bilabial closure, tongue tip contact, or velum opening. If it meets the conditions, it enters the event site annotation process. For phonemes with bilabial closure, such as / p / , / b / , / m / , the system will first generate a lip closure event site at the starting time point of the phoneme record and generate a burst or release site near the ending time point. If the duration of the phoneme is short and there are less than two frames between the start and the end, only the starting site is retained. For phonemes that require the tongue tip to contact the alveolar ridge, such as / t / , / d / , / n / , / l / , the system calculates the midpoint of the phoneme time range and rounds it to the nearest target frame number, and marks this target frame number as the tongue tip contact event site. The midpoint position is also applicable in the case of fricatives / s / , / z / because the tongue tip has reached the position closest to the alveolar ridge during the frication process. For nasal sounds with velum opening, such as / m / , / n / , / / , the system tends to select the target frame number closest to the reference frame near the starting point of the phoneme record to annotate the velum opening event site, so as to ensure that the opening of the nasal passage can occur before the start of acoustic nasalization. If the event time point is not equal to any target frame timestamp, the target frame number with the smallest time difference is found through binary search and this number is used to avoid displacement errors caused by subsequent interpolation. All event sites are sorted by timestamp and written into the second side chain. The linked list node contains three fields: event type, target frame number, and original timestamp. Taking the phrase "关" / g u a n / as an example, / g / triggers a velum elevation and closure event site, / u / triggers a lip constriction event site, and / n / triggers a double event site of tongue tip-alveolar contact and velum opening. These discrete nodes are non-overlappingly distributed on the timeline, meeting both the requirement of traceability of contact points and providing clear anchor points for key frame positioning.
[0044] When the first side chain and the second side chain each complete writing and pass the boundary check, the system can merge them with the main chain on the comprehensive time coordinate. The time segments or discrete points carried by each of the three are uniformly mapped to the dimension of the target frame number, realizing multi-channel information fusion of steady-state configuration, dynamic transition, and contact events. The whole process maintains a one-way scan without backtracking, ensuring that the computational complexity is linearly related to the number of phonemes and still enabling real-time output of high-precision mouth shape key frames in the case of high-definition long video segments.
[0045] Furthermore, at runtime, once the three-way chain merge table is written and indexed in ascending order by target frame number, the system immediately initiates a conflict detection scanner. The scanner uses phoneme time boundaries as segment indexes for a sliding window. Whenever the window crosses the common time interval of adjacent phonemes, it checks whether there is any overlap of lip category entries, opening and closing segment labels, or contact event sites within that interval. If overlap is detected, the seam generator is triggered. The seam generator follows a fixed seam strategy: it first calculates the duration of each entry in the overlapping segment, retains the entry with the shortest duration that is semantically compatible with the adjacent segment, and truncates the remaining entries at the boundary to form a seam segment of limited length. The seam segment is marked as a "transition" category and exists only within the common boundary. After generation, it is immediately written back to the merge table without spreading to either side. The system then reads the monotonic opening and closing trajectory segment labels and lip shape change trajectory segment labels provided by the first side chain, and combines them with the discrete event sites given by the second side chain to perform a single forward traversal of the merge table: when the cursor encounters a starting phase boundary, a seam segment boundary, or an event site, a key frame is marked at the target frame number. In order to avoid interpolation redundancy caused by excessive key frame density, the system performs distance threshold filtering on continuous calibration points. If the interval between two adjacent points on the time axis is less than half a frame, only the first point is retained. Next, the segment optimizer checks the duration of all segments in the merge table, and merges short-term segments that are lower than the shortest duration standard into the long segment that is closest in time and has the same lip shape category. The merging operation strictly keeps the target frame number of the calibrated key frame unchanged. If the categories on both sides of the short-term segment are different, the front segment is merged first to ensure semantic continuity. Once the incorporation is complete, the system again collects all keyframe numbers in chronological order. The resulting lip-keyframe sequence contains both the steady-state spatial anchor point for each phoneme and the rapid deformation inflection points caused by splicing segments and articulation events. The time interval within the sequence is variable but never less than half a frame, ensuring that spherical linear interpolation and analytic geometric deformation do not introduce visible jitter or redundant calculations due to over-sparse or over-dense sampling. Through this process, centered on splicing segment control, event location assurance, and shortest interval incorporation optimization, the system compresses multi-source timing information into a series of discrete, traceable, and fully comprehensive keyframes, providing both sufficient and streamlined frame-level support for subsequent posture-consistent interpolation.
[0046] Furthermore, in the implementation architecture of the present invention, the core goal of step 3 is to seamlessly integrate the lip-sync keyframe sequence containing only static spatial information with the actual head motion trajectory of the original video, ensuring that the generated lip-sync is both strictly compliant with the phoneme timeline in the temporal dimension and physically consistent with the facial posture in the spatial dimension. To this end, the system first extracts the target frame one by one in the original video decoding thread according to the frame index mapping table sequence and calls a fixed facial geometry template to perform multi-scale feature alignment. This template has been determined offline through air blowing experiments and laser scanning to have a set of 3D vertex coordinates that match the average head model height. At runtime, the template is projected onto the grayscale gradient extreme value region of the current frame using a photometric consistency metric and an iterative closest point algorithm. An affine transformation is then calculated with the goal of minimizing pixel residuals. The final output is a set of 3D facial feature points covering the brow peak, orbital bones, nasal bridge, lip beads, and chin, totaling approximately 60 control points. The system then uses the PnP solver in the perspective projection model to convert the template coordinate system back to the camera coordinate system, obtaining the head rotation and translation vectors for the target frame and writing them into the head pose trajectory. Because each pose entry carries a native timestamp, subsequent interpolation can be performed directly on the real sampling interval without resampling. Next, the system uses a hash map to assign each keyframe in the lip keyframe sequence to a target frame number in the frame index map. If the keyframe timestamp is not exactly equal to any target frame, a binary search is performed to select the frame number with the smallest time difference to bind. If the absolute differences are the same, the smaller time difference is used to ensure that the discrete frame numbers of the keyframes do not jump back or forth. After binding, the pipeline enters the pose consistency phase: For the interval between two consecutive keyframes, the system calls the quaternion spherical linear interpolation function slerp() in rotation space to connect the two points along the unit 4D sphere at a constant speed. Simultaneously, a first-order linear interpolation function is called in Euclidean space to directly connect the translation vectors. The interpolation interval strictly stops at the common boundary of adjacent phonemes to prevent cross-phoneme pose mixing. The resulting posture-consistent lip sequence corresponds one-to-one with the original video target frame on the time axis, and each sequence unit carries three types of information: head rotation, head translation, and lip type category.
[0047] The rendering thread then generates a local mouth mesh based on the pose-consistent lip sequence, using 3D facial landmarks as anchor points on the corresponding target frame. The mesh's topology directly inherits the template's mouth partition, and its vertex quality is controlled between 2,000 and 4,000 through density-adaptive sampling to balance deformation accuracy and gridding efficiency. The system reads the current frame's mouth shape category from a lookup table and retrieves the analytic geometry deformation parameter set, which describes the displacement directions and scaling factors of the lip margin, cleft, cheek muscles, and jawbone. After performing a per-vertex coordinate transformation on the mesh, the system multiplies the deformed coordinates by the head rotation matrix and then adds the translation vector, ultimately writing the mouth region pixel blocks in the target frame's image coordinate system. If the original video frame already contains a native mouth texture matching the target mouth shape category, the native texture is directly mapped onto the deformed mesh surface using a differential motion field algorithm. Otherwise, the system retrieves the mouth geometry model and high-definition lip and tooth texture patch that best matches the current head pitch and yaw angles from a fixed mouth geometry template library. After performing pose transformation and resolution consistency, the mesh is embedded into the mouth mesh. To eliminate replacement artifacts, the system extends a fixed-width transition band outside the grid boundary. Within this band, the original and new pixels are weighted and fused using Gaussian weights. The width of the transition band is linearly proportional to the target frame resolution, ensuring that high-definition scenes are not overly blurred while also avoiding hard edges in low-resolution scenes. Bilateral texture statistical matching is then performed using the nose-to-cheek bone region in the same frame as the reference domain. Separate quantile mapping is performed on the luminance and chrominance channels to ensure that the average brightness and hue distribution of the mouth region is consistent with that of the surrounding skin, addressing color temperature drift caused by uneven lighting or multiple cameras.
[0048] After all pixel operations are completed in the off-screen buffer, the buffer is written back to the corresponding target frame memory page with the timestamp of the original video as the primary key to ensure that the frame order is consistent with the original encoding. Since the phoneme timeline is used as the only master clock from beginning to end, the entire rendering pipeline does not need to make additional corrections to the frame rate or sampling interval, thereby avoiding the lip-syncing audio phenomenon caused by dual clock drift. When the last frame is written out, the encoder performs lossless repackaging according to the original video container format and resolution. The output video file visually retains the original fineness and the lip shape at any point in time strictly corresponds to the phoneme at that moment, meeting the dual demands of synchronization and clarity for application scenarios such as long dialogues, emotional interpretations, and low-latency live broadcasts.
[0049] Figure 2The construction principle and corresponding relationship of the phoneme time axis and frame index mapping table in the present invention are demonstrated in detail. The figure consists of three parts: the upper, middle and lower parts, which are the phoneme time axis, the original video target frame sequence and the frame index mapping table respectively. In the phoneme time axis part, five phonemes are arranged in chronological order: phoneme A, phoneme B, phoneme C, phoneme D and phoneme E, and their corresponding time intervals are t0-t1, t1-t2, t2-t3, t3-t4 and t4-t5 respectively. Each phoneme has a certain start and end time point, forming a continuous phoneme sequence. Phoneme A occupies the time interval t0-t1 and has a relatively short duration; phoneme B occupies the time interval t1-t2 and has a longer duration; phoneme C occupies the time interval t2-t3 and has a short duration; phoneme D occupies the time interval t3-t4 and has a medium duration; phoneme E occupies the time interval t4-t5 and has a medium duration. The original video target frame sequence shows 11 target frames, from F1 to F11, arranged using the original video's timestamps as a slave clock. Each target frame has a fixed time position and frame number, forming a discrete frame sequence. The distribution density of target frames is related to the original video's frame rate, and frame density may vary within different time periods. A frame index mapping table, a core technical component, establishes a one-to-one correspondence between the phoneme timeline and the original video's target frames. This mapping table contains three key fields: target frame number, corresponding phoneme identifier, and frame timestamp. Specifically, the mapping relationship is as follows: target frames F1-F3 correspond to phoneme A, with timestamps t0-t1; target frames F4-F5 correspond to phoneme B, with timestamps t1-t2; and target frames F6-F9 correspond to the combination of phonemes C, D, E, and C, with timestamps t2-t5. The red arrows clearly indicate the mapping relationship between the target frames and the corresponding phoneme intervals, achieving a precise correspondence between the continuous phoneme timeline and the discrete video frame sequence. This mapping mechanism ensures the accuracy of time synchronization during the subsequent lip synthesis process, and provides time reference and frame-level index support for phoneme-driven video lip synthesis.
[0050] Figure 3This paper systematically demonstrates the working principles of three complementary derivation chains, one of the core technologies of this invention. This mechanism combines a main chain with two side chains to achieve a multi-dimensional mapping conversion from phonemes to mouth shapes. The phoneme timeline, located at the top of the diagram, contains five consecutive phonemes: / a / , / i / , / p / , / l / , and / e / . These phonemes are arranged in pronunciation order and temporal sequence, forming the basis of a complete phoneme sequence. The main chain, as the first derivation chain, implements a one-to-one mapping function between phonemes and mouth shape categories. Based on a fixed and limited dictionary of mouth shape configurations, the main chain converts the input phoneme sequence into a corresponding mouth shape sequence: phoneme / a / is mapped to mouth shape A, phoneme / i / is mapped to mouth shape I, phoneme / p / is mapped to mouth shape P, phoneme / l / is mapped to mouth shape L, and phoneme / e / is mapped to mouth shape E. This mapping process strictly adheres to deterministic pronunciation rules, ensuring that each phoneme has a unique and accurate correspondence with the mouth shape category. The basic mouth shape sequence generated by the main chain does not include transition effects, providing a stable baseline for subsequent processing. The first side chain is specifically responsible for rhyme sequence connection, performing phase analysis and trajectory generation for phonemes belonging to the same syllable. This side chain divides the phoneme sequence into syllables and generates two types of segment labels for each syllable: monotonic opening and closing trajectories and lip shape change trajectories. Syllable 1 includes monotonic opening and closing trajectories for the onset phase, transition phase, and steady-state phase; syllable 2 includes the corresponding lip shape change trajectories. The processing of the first side chain ensures smooth transitions and continuity between phonemes within the same syllable. The second side chain processes articulation contact events based on deterministic contact rules for the lips, alveoli, hard palate, soft palate, and glottis. This side chain identifies phonemes that require special articulatory movements, such as lip closure for / p / and tongue tip contact for / l / . The second side chain annotates discrete articulatory events at corresponding time points, including event site 1 and event site 2. These event sites precisely indicate the occurrence of specific articulatory movements. The final output is a three-way chain merge table, which combines the main chain's lip-shape categories, the first side chain's segment labels, and the second side chain's event locations on the same timeline along the phoneme timeline. This merging process employs a single-pass forward decision principle with no feedback, resolving potential conflicts between different derivation chains and generating a unified "main chain + first side chain + second side chain merged result." This merged result provides a comprehensive technical foundation for subsequent keyframe determination and lip-shape sequence generation.
[0051] Figure 4The experimental data curves detail the dynamic changes in mouth opening and closing during phoneme phase slicing and the technical implementation mechanism for phase calibration. The figure uses a two-dimensional coordinate system, with the horizontal axis representing time and the vertical axis representing the quantified value of mouth opening and closing. The time axis extends from t0 to t5, covering the complete articulatory cycle of five consecutive phonemes. The duration of each phoneme is scientifically divided into three distinct phase types: onset phase, transition phase, and steady-state phase. The phoneme / a / corresponds to the time interval t0-t1, primarily encompassing the onset phase; the phoneme / i / corresponds to the time interval t1-t2, encompassing the transition phase; the phoneme / u / corresponds to the time interval t2-t3, encompassing the steady-state phase; the phoneme / o / corresponds to the time interval t3-t4, again encompassing the transition phase; and the phoneme / e / corresponds to the time interval t4-t5, encompassing the onset phase. This phase division fully accounts for the physiological characteristics of human pronunciation and the dynamics of mouth shape changes. The mouth-opening degree curve is represented by a smooth, continuous curve with values ranging from 0 to 1.0, objectively reflecting the quantitative changes in the degree of mouth opening and closing during pronunciation. The curve starts at an opening degree of 0.2 and gradually rises to 0.6 with the pronunciation of the phoneme / a / , reflecting the pronunciation characteristics of an open phoneme. During the / i / phase, the opening degree further increases to a peak value close to 0.8, corresponding to the high opening degree requirement of this phoneme. Subsequently, during the steady-state phase of the / u / phoneme, the opening degree remains relatively stable at around 0.8, reflecting the characteristics of the steady-state phase. During the / o / phase, the opening degree begins to decrease to around 0.4, finally returning to a level of 0.6 during the / e / phase. Five keyframes are marked on the experimental curve, corresponding to the characteristic moments of each phase. Keyframe 1 is located at the characteristic point of the starting phase, with an opening degree of approximately 0.2; keyframe 2 is located at the peak point of the transition phase, with an opening degree reaching 0.8; keyframe 3 is located in the plateau region of the steady-state phase, with an opening degree maintained at 0.8; keyframe 4 is located in the descending transition phase, with an opening degree decreasing to 0.4; and keyframe 5 is located at the final starting phase, with an opening degree stabilized at 0.6. These keyframes were determined in accordance with the principle of maintaining the shortest interval and the requirement of phase continuity, providing precise time anchors and numerical references for subsequent lip interpolation and video synthesis. Phase boundaries are clearly marked by vertical dashed lines, ensuring boundary traceability and phase calibration accuracy. This experimental data validates the effectiveness of the phoneme phase slicing algorithm and provides important theoretical support and experimental evidence for phoneme timeline-driven lip synthesis technology.
[0052] The following example illustrates the complete implementation of a phoneme timeline driven high-definition video lip-syncing automatic synthesis method in a real engineering scenario. Assume that the input is a period of Original video, resolution , frame rate , and gives the Chinese test sentence "Mom holds baby" synchronized with the video. The system uses the International Phonetic Alphabet, and the corresponding phoneme sequence is recorded as This paper will demonstrate the calculation process using a fixed phoneme duration model, where each phoneme duration is set to ,total phonemes, satisfying .
[0053] First, create a phoneme timeline. Define the phoneme record Start and end time of the strip Phoneme Timeline Master Clock Then construct the frame index mapping table. Assume that the video single frame time interval , No. Frame timestamp .
[0054] in accordance with , assign a unique phoneme identifier to each frame and get the mapping Next, perform phoneme slicing and phase calibration. For phonemes : .exist Mark the starting phase, Mark the transition phase, Mark the steady-state phase; each frame determines the phase label by the timestamp landing point to form a phase-labeled frame sequence .
[0055] Loading the lip shape dictionary : , , , Loading place of articulation contact rules :Labioli trigger the lip closure event, plosive triggers the plosive release event, and rounded vowels trigger the lip contraction event. Main chain traversal Write the basic lip sequence in sequence The first side chain gives the start and end frame numbers of the opening and closing segments and the lip shape change segments in the syllables "ma", "ma", "bao", and "bao"; the second side chain gives the start and end frame numbers of / m / , / b / , and / / Generate event site, such as / m / start frame Mark the lip closure, / b / end frame Record the blast release. After the three chains are merged, perform conflict detection. Assume that / m / and / / The main chain closed lip segment and the first side chain open segment overlapped at the common boundary The stitching strategy takes shorter segments, retains the open segments, and cuts off the closed lip segments to the boundaries. Then, key frames are marked at the phase boundaries, stitching boundaries, and event locations. The minimum key frame interval threshold is Short-term fragments are not persistent enough are merged into adjacent similar segments. Finally, the key frame number sequence is obtained .
[0056] Build head pose trajectory. Execute for each target frame Solve and get the rotation quaternion With translation vector If the interval between two key frames is , interpolation time ,but , , where the subscript Indicates the keyframe at the endpoint of the interval. All non-keyframe poses are calculated by this formula, limited to the common boundary of the same phoneme and not crossing the boundary. In the rendering stage, a four-layer subdivision mouth area mesh is generated in each frame with the three-dimensional facial feature points as anchor points. . Opposite vertex Perform analytical deformation ,in is the scaling matrix obtained by looking up the table according to the lip shape category, is the rotation matrix, is the translation vector, all three are predefined in the template coordinate system. If the current frame requires "rounded lip convergence", take 、 、 After deformation, the vertex is then transformed into the head posture Transform and write to the image plane. If the original frame has no matching texture, the lip template texture in the library is called, and the texture resolution is , embedded in the mouth grid after perspective mapping. The boundary feathering is done with Gaussian kernel half-height width The color transfer is done by piecewise linear mapping to make the brightness of the mouth area mean Aligned with the average brightness of the cheek reference area, mapping formula .
[0057] The output stage keeps the original frame timestamp The processed frame stream is passed to the H.265 encoder with the same bit rate as the resolution. , packaged as MP4. The final composite video is A continuous movement from lip closure to maximum opening was observed, which was precisely synchronized with the phoneme / mɑ / . , meeting the live broadcast synchronization threshold Require.
[0058] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A phoneme timeline driven high-definition video lip-sync automatic synthesis method, characterized in that: The method comprises: Step 1: Obtain input text or speech, use a fixed phoneme set and deterministic pronunciation rules to generate a phoneme sequence and corresponding start and end time points, and establish a phoneme timeline. Using the timestamp of the original video as a slave clock, construct a frame index mapping table that maps the target frames of the original video to the phoneme timeline one by one. Output the phoneme timeline, frame index mapping table, and target frame set. Step 2: Under the constraints of the phoneme timeline and frame index mapping table, phoneme slicing and phase calibration are performed sequentially. Based on a fixed and limited lip shape dictionary and articulation position rules, three complementary derivation chains are established simultaneously. A single forward decision without feedback and the principle of minimum interval preservation are used to complete conflict resolution and key frame determination, resulting in a lip shape key frame sequence that meets phase continuity and boundary traceability. Step 3: Extract 3D facial feature points and head posture trajectories based on the original video, align the lip-sync keyframe sequence with the head posture trajectory according to the frame index mapping table to obtain a posture-consistent lip-sync sequence; output the video lip-sync synthesis result based on the posture-consistent lip-sync sequence with the original video resolution and synchronized with the phoneme timeline.
2. The phoneme timeline driven high-definition video lip-sync automatic synthesis method according to claim 1, wherein: In step 2, the process of performing phoneme slicing and phase calibration in sequence includes: dividing the duration of each phoneme into three types of phoneme slices: starting phase, transition phase and steady-state phase; according to the frame index mapping table, each target frame is uniquely assigned to its corresponding phoneme slice, and the phase type label is recorded to form a phase-labeled frame column.
3. The phoneme timeline driven high-definition video lip-sync automatic synthesis method according to claim 2, wherein: In step 2, based on a fixed and limited dictionary of mouth shape configurations and articulation position rules, three complementary derivation chains are established simultaneously, including: a main chain, a first side chain, and a second side chain. The main chain is a one-to-one mapping from phonemes to mouth shape categories, used to generate a basic mouth shape sequence without transitions. The first side chain is used for rhyme sequence connection, sequentially connecting the onset phase, transition phase, and steady-state phase belonging to the same syllable to generate segment labels for monotonic opening and closing trajectories and lip shape change trajectories. The second side chain corresponds to the articulation position contacts, and is used to annotate discrete event sites of lip closure, tongue tip contact, and soft palate opening for phonemes that require closure, friction, or explosion based on the deterministic contact rules of the lips, alveoli, hard palate, soft palate, and glottis. The mouth shape categories of the main chain, the segment labels of the first side chain, and the event sites of the second side chain are merged on the same timeline according to the phoneme timeline to obtain a three-way chain merge table.
4. The phoneme timeline driven high-definition video lip-sync automatic synthesis method according to claim 3, wherein: In step 2, the phoneme timeline is first read. The phoneme timeline consists of phoneme records arranged in chronological order, and each phoneme record contains a phoneme identifier and a start and end time point. The frame index mapping table is read. The frame index mapping table gives the correspondence between the target frame and the phoneme timeline. Each record contains the target frame number, the corresponding phoneme identifier, and the frame timestamp. The mouth shape configuration dictionary and the articulation contact rules are loaded. The mouth shape configuration dictionary is a fixed and finite mapping set that provides a one-to-one mapping relationship between phoneme categories and mouth shape categories. The articulation contact rules are a set of fixed rules that provide the contact conditions and corresponding event types of the lips, gums, hard palate, soft palate and glottis.
5. The phoneme timeline driven high-definition video lip-syncing automatic synthesis method according to claim 4, wherein: In step 2, the process of generating the main chain includes: traversing each phoneme record in the order of the phoneme time axis, searching the lip shape category corresponding to the phoneme category in the lip shape dictionary, and generating a basic lip shape entry; adding time boundary information to each basic lip shape entry, where the time boundary is consistent with the start time point and end time point of the phoneme record; writing all basic lip shape entries into the main chain in chronological order to obtain a basic lip shape sequence without transitions; any entry in the main chain only covers the time range of the phoneme record to which it belongs, and does not cross the time boundary of adjacent phonemes.
6. The phoneme timeline driven high-definition video lip-sync automatic synthesis method according to claim 5, wherein: In step 2, the generation process of the first side chain includes: dividing the phoneme time axis into syllables, and the time range of the syllable is determined by the minimum starting time point and the maximum ending time point of the continuous phoneme records constituting the syllable; within each syllable, according to the order of the starting phase, transition phase and steady-state phase, the corresponding target frame numbers are connected in series to generate a segment label of the monotonic opening and closing trajectory; the segment label records the segment type, the starting target frame number, and the ending target frame number; within each syllable, according to the deterministic order of the lip shape from contraction to relaxation or from relaxation to contraction, the segment label of the lip shape change trajectory is generated; the segment label records the lip shape change type, the starting target frame number, and the ending target frame number; the segment label of the monotonic opening and closing trajectory and the segment label of the lip shape change trajectory are written into the first side chain in time sequence; any segment label shall not exceed the time range of the syllable to which it belongs; the generation process of the second side chain includes: according to the contact rule of the pronunciation position Then, the phoneme time axis is checked one by one to determine the set of phonemes that need to be closed, rubbed or exploded; for phonemes that belong to lip closure, the discrete event sites of lip closure are marked within the time range of the phoneme record; the time position of the event site is preferentially one or two of the start time point and the end time point of the phoneme record; for phonemes that belong to tongue tip contact, the discrete event sites of tongue tip contact are marked within the time range of the phoneme record; the time position of the event site is preferentially the target frame number close to the middle of the phoneme; for phonemes that belong to soft palate opening, the discrete event sites of soft palate opening are marked within the time range of the phoneme record; the time position of the event site is preferentially the target frame number near the start time point of the phoneme record; the above discrete event sites are written into the second side chain in chronological order; each event site is bound to a target frame number, and if the event time point is not equal to any target frame timestamp, it is assigned to the target frame number closest in time.
7. The phoneme timeline driven high-definition video lip-sync automatic synthesis method according to claim 6, wherein: In step 2, the lip-sync key frame sequence is obtained through the following process: searching for overlapping areas of records at the time boundaries of adjacent phonemes in the three-way chain merge table, and generating seam segments according to a fixed seam strategy; the seam segments only exist within the common boundaries of adjacent phonemes and do not cross the common boundaries; based on the segment labels of the first side chain and the discrete event sites of the second side chain, key frames are marked at the phase boundaries, seam boundaries and event sites; the key frames are discrete values on the target frame numbers; for short-term segments that appear in the three-way chain merge table and are lower than the shortest duration standard, they are merged into segments that are adjacent in time and have the same category; the merging operation does not change the position of the marked key frames; all key frames are combined into a lip-sync key frame sequence in chronological order.
8. The phoneme timeline driven high-definition video lip-sync automatic synthesis method according to claim 7, wherein: Step 3 specifically includes: in each target frame of the original video, calibrating the three-dimensional facial feature points according to the fixed facial geometry template; obtaining the head posture trajectory according to the time sequence of the target frames, wherein the head posture trajectory includes the head rotation and head translation of the target frame; assigning each key frame in the lip shape key frame sequence to the corresponding target frame number according to the frame index mapping table; reading the head rotation and head translation of the target frame from the head posture trajectory and attaching them to the corresponding lip shape key frame; between adjacent lip shape key frames, using spherical linear interpolation for head rotation and linear interpolation for head translation to obtain a posture-consistent lip shape sequence; generating a mouth area grid on the target frame with the three-dimensional facial feature points as anchor points; and The lip shape configuration of the frame in the lip sequence is consistent, and the mesh is subjected to analytic geometric deformation and written into the mouth area. When the required lip shape combination does not exist in the original video, a fixed oral geometry template and lip-tooth texture patch that match the current posture are selected and embedded into the mouth area after posture transformation. A transition zone of fixed width is set at the boundary of the mouth area to perform boundary feathering. The color is transferred based on the facial reference area of the same frame to make the mouth area consistent with the reference area in brightness and chroma. The phoneme timeline is used as the only master clock to write the replaced and processed mouth area back to the corresponding target frame, keeping the frame timestamp and resolution unchanged. The video lip synthesis result is output with the same resolution as the original video and synchronized with the phoneme timeline.
9. The phoneme timeline driven high-definition video lip-syncing automatic synthesis method according to claim 8, wherein: When the time position of a lip keyframe is not equal to the timestamp of any target frame, it is assigned to the target frame number with the smallest time difference; if there are two equidistant target frames, the smaller one is taken; pose-consistent interpolation is performed between adjacent lip keyframes without crossing the common boundaries of adjacent phonemes.
Citation Information
Patent Citations
System and method for animated lip synchronization
CA2959862A1
Three-dimensional vocal organ animation method combining physiological model and data driving model
CN103218841A
Chinese speech synthesis method based on phonemes and rhythm structures
CN110534089A
Virtual face generation method
CN113781610A
Mouth shape animation generation method and device, electronic equipment and storage medium
CN115423904A
Cited By
Speech recognition and speech synthesis optimization method and system based on large model
CN121034309A
Digital population type synchronization method and device based on phonon driving, equipment and medium
CN121545542A
Digital human lip sound synchronous driving method and system
CN122067554A
Digital human lip sync driving method and system
CN122067554B