A large model-based speech recognition and speech synthesis optimization method and system

By extracting lip movement features and audio feature timestamps to generate a dynamic offset compensation parameter sequence, and using a large model to reconstruct missing or delayed speech segments, the problem of misalignment between speech and facial movements in existing technologies is solved, achieving high consistency and natural output between speech and lip movements.

CN121034309BActive Publication Date: 2026-01-23LUSTER LIGHTWAVE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511477160.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-23
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing technologies rely on fixed alignment strategies and static time offset estimation mechanisms, which are difficult to adapt to the dynamic differences in different speech performances and lack effective compensation mechanisms for extreme cases, resulting in misalignment between speech and facial movements, affecting the consistency and naturalness of the interactive experience.

Method used

An initial spatiotemporal offset sequence is generated by extracting lip movement features and audio feature timestamps. A dynamic offset compensation parameter sequence is generated by combining it with a pre-constructed benchmark alignment template. A large model is used for joint processing to reconstruct missing or delayed target speech segments and generate speech waveforms that are phase-synchronized with lip movements.

Benefits of technology

It achieves high-precision synchronous acquisition of voice and visual signals, solves the benchmark deviation problem caused by asynchronous signal sources in traditional solutions, adapts to sudden changes in speech rate and accent differences, generates coherent voice output, and improves the realism and naturalness of the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034309B_ABST
    Figure CN121034309B_ABST
Patent Text Reader

Abstract

The application provides a speech recognition and speech synthesis optimization method and system based on a large model, which extracts lip movement features, lip feature time stamps and audio feature time stamps by acquiring real-time speech input signals and image frame sequences of a user's face area to generate an initial space-time offset sequence; generates a dynamic offset compensation parameter sequence based on the initial space-time offset sequence and in combination with a reference alignment template in a speech visual synchronization data set; jointly processes the dynamic offset compensation parameter sequence, the lip movement features and the real-time speech input signals by using a large model to reconstruct a target speech segment, generates a corrected phoneme sequence in combination with the lip movement features, retrieves a mouth shape parameter group corresponding to the corrected phoneme sequence from a phoneme mouth shape mapping rule library to generate a speech waveform synchronized with the lip movement phase; and the application improves the naturalness, immersion and robustness in a speech missing or delayed scenario of human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a speech recognition and speech synthesis optimization method and system based on a large model. BACKGROUND

[0002] In the current application scenarios of increasingly popular human-computer interaction, speech recognition and speech synthesis technology is widely used in intelligent assistants, remote conferences, and barrier-free systems in many fields. With the increasing demand of users for naturalness and real-time of interaction, the system not only needs to accurately recognize the speech content, but also needs to maintain synchronization with the speaker's facial movements when outputting speech, in order to improve the realism and immersion of user experience. Especially in video conferencing and other scenarios with high visual and auditory coordination requirements, the timing consistency between speech and lip movements becomes an important factor affecting the quality of interaction.

[0003] To address the above technical needs, a representative existing solution adopts a multi-modal fusion strategy, extracts speech features from audio signals and facial image features from videos, and respectively performs timestamp labeling. Then, a timing alignment algorithm is used to estimate the offset between the two. On this basis, the system introduces a neural network model to jointly model speech and visual information, thereby optimizing speech recognition results and generating more natural speech output. However, the existing solution still has limitations in handling complex and variable speech performances. Because it relies on fixed alignment strategies and static time offset estimation mechanisms, it is difficult to adapt to the dynamic differences of different speech performances; in the face of extreme cases of speech loss or severe delay, there is a lack of effective compensation mechanism, resulting in obvious misalignment between the final output speech and facial movements, affecting the consistency and naturalness of the overall interaction experience. SUMMARY

[0004] The present application provides a speech recognition and speech synthesis optimization method and system based on a large model, to solve the problems in the prior art that rely on fixed alignment strategies and static time offset estimation mechanisms, making it difficult to adapt to the dynamic differences of different speech performances; lack of effective compensation mechanism for extreme cases, resulting in insufficient consistency and naturalness of the overall interaction experience.

[0005] In a first aspect, the present application provides a speech recognition and speech synthesis optimization method based on a large model, comprising:

[0006] Obtaining a real-time speech input signal and a sequence of image frames of a user's facial region, and extracting lip movement features from the sequence of image frames;

[0007] Identifying the lip feature timestamps of the sequence of image frames and the audio feature timestamps of the real-time speech input signal to generate an initial spatio-temporal offset sequence;

[0008] generate a dynamic offset compensation parameter sequence based on the initial spatio-temporal offset sequence and a reference alignment template in a pre-constructed speech-visual synchronization dataset;

[0009] reconstruct the missing or delayed target speech segment by jointly processing the dynamic offset compensation parameter sequence, the lip motion feature, and the real-time speech input signal using a large model;

[0010] generate a corrected phoneme sequence from the target speech segment and the lip motion feature, and retrieve a set of mouth shape parameters corresponding to the corrected phoneme sequence from a pre-set phoneme-mouth shape mapping rule library to generate a speech waveform synchronized with the lip motion phase.

[0011] Optionally, lip feature timestamps of the image frame sequence and audio feature timestamps of the real-time speech input signal are identified to generate an initial spatio-temporal offset sequence, including:

[0012] detecting the lip region boundary in the image frame sequence and setting an array of reference points with equal spacing within the lip region boundary;

[0013] generating a displacement dataset for each reference point according to the position changes of the reference point array in different image frames;

[0014] defining reference points located on the lip contour line and having a displacement exceeding a pre-set distance in the displacement dataset as dynamic key points;

[0015] identifying the timestamp of the image frame corresponding to the peak displacement in the motion trajectory of the dynamic key point as the lip feature timestamp;

[0016] dividing the real-time speech input signal into multiple audio segments and detecting the time points of energy amplitude jumps in each audio segment as audio feature timestamps;

[0017] pairing the lip feature timestamps and the audio feature timestamps of image frames in the same time period to obtain multiple paired timestamps, and calculating the difference between each paired timestamp to combine all the differences into an initial spatio-temporal offset sequence.

[0018] Optionally, a dynamic offset compensation parameter sequence is generated based on the initial spatio-temporal offset sequence and a reference alignment template in a pre-constructed speech-visual synchronization dataset, including:

[0019] selecting a historical reference template consistent with the user's lip opening and closing frequency from the pre-constructed speech-visual synchronization dataset, the historical reference template containing a fixed correspondence between lip feature timestamps and audio feature timestamps in a standard scenario;

[0020] connecting each reference offset in the historical reference template to form a reference curve;

[0021] extracting each offset in the initial spatio-temporal offset sequence and the corresponding audio feature timestamp, and locating the reference offset corresponding to each audio feature timestamp in the reference curve;

[0022] obtaining the corresponding offset compensation value based on the offset corresponding to each audio feature timestamp and the reference offset;

[0023] constructing a compensation index table according to the group identifier assigned to each offset compensation value and the preset basic compensation coefficient;

[0024] extracting the corresponding preset compensation coefficient from the compensation index table according to the audio feature timestamp corresponding to each offset in the initial spatio-temporal offset sequence, and integrating all preset compensation coefficients into a dynamic offset compensation parameter sequence.

[0025] Optionally, constructing a compensation index table according to the group identifier assigned to each offset compensation value and the preset basic compensation coefficient, comprises:

[0026] setting a plurality of compensation value ranges, and assigning a group identifier to each compensation value range;

[0027] determining the compensation value range to which each offset compensation value belongs, so as to mark the corresponding group identifier for each offset compensation value;

[0028] counting the total number of offset compensation values with the same group identifier, and combining the corresponding offset compensation value into an active group when the total number exceeds a preset activation threshold;

[0029] assigning a preset basic compensation coefficient to each active group, and calculating the distribution density of the offset compensation values of each active group;

[0030] adjusting the basic compensation coefficient of each active group according to the set density threshold and the distribution density of each active group to obtain the corresponding adjusted compensation coefficient;

[0031] storing the group identifier and the adjusted compensation coefficient of each active group as a key-value pair, and integrating all key-value pairs to construct a compensation index table.

[0032] Optionally, the dynamic offset compensation parameter sequence, the lip movement feature and the real-time speech input signal are jointly processed by using a large model to reconstruct the missing or delayed target speech segment, comprising:

[0033] According to the dynamic offset compensation parameter sequence, the lip movement feature is time axis stretching and compression processed to generate a time-synchronized lip movement trajectory sequence;

[0034] The lip movement trajectory sequence is fused with the real-time speech input signal to generate a multi-modal input stream;

[0035] A large model is used to detect a speech interruption area in the multi-modal input stream, and extract a lip movement trajectory segment corresponding to the speech interruption area;

[0036] Based on the lip movement trajectory segment, phoneme identifiers, duration and energy levels are generated for phoneme period analysis, contour arrangement and energy scaling processing of the speech interruption area to obtain an adjusted interruption area, and combined with anchor point frames of a normal speech area, an adjusted transition area is generated, and combined with the adjusted interruption area, a connected spectral contour is obtained.

[0037] According to the connected spectral contour and the energy level, a missing or delayed target speech segment is reconstructed.

[0038] Optionally, based on the lip movement trajectory segment, phoneme identifiers, duration and energy levels are generated for phoneme period analysis, contour arrangement and energy scaling processing of the speech interruption area to obtain an adjusted interruption area, and combined with anchor point frames of a normal speech area, an adjusted transition area is generated, and combined with the adjusted interruption area, a connected spectral contour is obtained, including:

[0039] According to the lip movement trajectory segment, a lip contour height change rate is calculated, according to the lip contour height change rate, a corresponding phoneme identifier is matched, and a phoneme boundary point in the lip movement trajectory segment is marked;

[0040] The time between adjacent phoneme boundary points is determined as a phoneme period, and according to the phoneme period, a duration and an average curvature value of the lip movement trajectory are calculated to determine an energy level;

[0041] According to the phoneme identifier and the duration, the standard spectral contour of each phoneme identifier in the speech interruption area is arranged to obtain an arranged contour, and according to the energy level, the arranged contour is scaled to obtain a scaled contour, and a formant gradual change sequence is inserted between the scaled contours of adjacent phoneme identifiers to obtain an adjusted interruption area;

[0042] An anchor point frame adjacent to the start end of the speech interruption area in the normal speech area is located, and the spectral contour of the anchor point frame is extracted as an anchor point spectral contour;

[0043] Starting from the anchor point spectrum profile, the spectrum profile of a preset number of frames is copied frame by frame towards the speech interruption region to form a gradual transition area;

[0044] The position of the resonant peak of the spectral profile in the gradual transition region is adjusted to the position of the resonant peak of the standard spectral profile corresponding to the target phoneme identifier to obtain the adjusted transition region.

[0045] The adjusted transition region and the adjusted interruption region are connected to obtain the connected spectrum profile.

[0046] Optionally, based on the target speech segment and phoneme time period, a modified phoneme sequence is generated, and a set of lip shape parameters corresponding to the modified phoneme sequence is retrieved from a preset phoneme lip shape mapping rule base to generate a speech waveform synchronized with the lip movement phase, including:

[0047] The target speech segment is processed by the speech recognition branch of the large model to generate an initial phoneme sequence;

[0048] The displacement rate of the lip movement trajectory is extracted from multiple phoneme time periods, and the duration of the corresponding phoneme in the initial phoneme sequence is scaled using the displacement rate to generate a modified phoneme sequence.

[0049] The lip shape parameter set corresponding to the modified phoneme sequence is retrieved from the preset phoneme lip shape mapping rule library, and the lip shape pattern code of the lip shape parameter set is compared with the lip movement feature to calculate the shape difference value of each phoneme time period. The lip shape parameter set includes lip shape pattern code, tongue position coordinates and vocal tract expansion.

[0050] The shape difference value is used to correct the tract expansion of the lip shape parameter set in order to generate an optimized lip shape parameter sequence;

[0051] Based on the phoneme identifiers and durations of the corresponding phonemes in the modified phoneme sequence, the tongue position coordinates and vocal tract expansion of the optimized lip shape parameter sequence, a speech waveform synchronized with the phase of lip movement is generated.

[0052] Secondly, the present invention provides a speech recognition and speech synthesis optimization system based on a large model, comprising:

[0053] The acquisition module is used to acquire real-time voice input signals and image frame sequences of the user's facial region, and extract lip movement features from the image frame sequences;

[0054] The recognition module is used to identify the lip feature timestamps of the image frame sequence and the audio feature timestamps of the real-time speech input signal to generate an initial spatiotemporal offset sequence.

[0055] The generation module is used to generate a dynamic offset compensation parameter sequence based on the initial spatiotemporal offset sequence and the reference alignment template in the pre-built speech-visual synchronization dataset.

[0056] The reconstruction module is used to utilize a large model to jointly process the dynamic offset compensation parameter sequence, the lip movement features, and the real-time speech input signal to reconstruct the missing or delayed target speech segment.

[0057] The retrieval module is used to generate a modified phoneme sequence based on the target speech segment and the lip movement features, and retrieve the lip shape parameter group corresponding to the modified phoneme sequence from a preset phoneme lip shape mapping rule library to generate a speech waveform that is phase-synchronized with the lip movement.

[0058] Thirdly, the present invention provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a large-model-based speech recognition and speech synthesis optimization method as described in the first aspect above.

[0059] Fourthly, the present invention provides a computer storage medium storing a computer program, which, when executed by a computer, implements a large-model-based speech recognition and speech synthesis optimization method as described in the first aspect.

[0060] In this invention, a sequence of image frames of a real-time voice input signal and a user's facial region is acquired, and lip movement features are extracted from the image frame sequence. The timestamps of the lip features in the image frame sequence and the timestamps of the audio features in the real-time voice input signal are identified to generate an initial spatiotemporal offset sequence. Based on the initial spatiotemporal offset sequence, a dynamic offset compensation parameter sequence is generated by combining a reference alignment template from a pre-constructed speech-visual synchronization dataset. Using a large model, the dynamic offset compensation parameter sequence, the lip movement features, and the real-time voice input signal are jointly processed to reconstruct a missing or delayed target speech segment. Based on the target speech segment and the lip movement features, a corrected phoneme sequence is generated, and a set of lip shape parameters corresponding to the corrected phoneme sequence is retrieved from a preset phoneme lip shape mapping rule library to generate a speech waveform synchronized with the lip movement phase. The technical solution provided by this invention achieves synchronous acquisition of speech and visual signals, extracts high-precision lip dynamic features, provides basic data support for cross-modal temporal alignment, and solves the benchmark deviation problem caused by asynchronous signal sources in traditional solutions. By matching the spatiotemporal changes in lip movement trajectory and speech energy, the degree of audio-visual asynchrony is quantified, breaking through the limitations of traditional fixed time window alignment and providing accurate input for dynamic compensation. Personalized lip movement patterns are used to match the benchmark alignment template to generate adaptive compensation parameters, overcoming the shortcomings of static alignment strategies in adapting to changes in speech rate and accent differences. By driving speech spectrum reconstruction through lip movement, coherent audio is generated in scenarios with speech interruption or high latency, solving the audio-visual fragmentation problem caused by the lack of compensation mechanisms in existing technologies. By fusing lip movement features to correct phoneme duration and tract parameters, phoneme-level lip-syncing is achieved, improving the temporal consistency between synthesized speech and lip movement. Furthermore, based on the dynamic offset compensation parameter sequence, the lip movement features are subjected to time-axis scaling to generate a time-synchronized lip movement trajectory sequence. The lip movement trajectory is fused with the real-time speech generation multimodal input stream. A large model is used to detect speech interruption regions and extract corresponding lip movement trajectory segments. Phoneme parameters are generated based on the lip movement trajectory segments, and the spectral contour of the interruption region is reconstructed through energy scaling and formant gradation mechanisms to obtain the adjusted interruption region. Combined with the adjusted transition region generated by anchor frames in the normal speech region, a coherent target speech segment is finally synthesized. Specifically, the temporal scaling of the lip movement trajectory and the generation of personalized phoneme parameters can adapt to complex scenarios such as sudden changes in speech rate and accent differences, overcoming the rigidity of traditional fixed alignment strategies. The lip movement trajectory-driven spectral reconstruction and gradation transition mechanism generates natural and coherent audio even when speech is completely missing or has high latency, eliminating audio-visual misalignment and improving the interactive experience in high-noise and weak-network environments.

[0061] These or other aspects of the invention will become more apparent from the following description of the embodiments. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 The flowchart of a speech recognition and speech synthesis optimization method based on a large model provided by the present invention is shown.

[0064] Figure 2 A schematic diagram of the structure of a speech recognition and speech synthesis optimization system based on a large model provided by the present invention is shown;

[0065] Figure 3 A schematic diagram of the structure of a computing device provided by the present invention is shown. Detailed Implementation

[0066] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0067] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0069] To address the challenges of dynamic temporal alignment in complex contexts and the decline in audiovisual synchronization caused by missing or delayed speech in existing speech recognition and synthesis technologies, current solutions rely on fixed alignment strategies and static temporal offset estimation mechanisms. These solutions struggle to achieve precise matching between speech and lip movements when faced with complex speech characteristics such as variations in speaking style, accent differences, or sudden changes in speech rate. Particularly when speech input is missing or delayed, the system lacks effective compensation mechanisms, leading to significant misalignment between the output speech and facial movements, severely impacting the realism and naturalness of the interactive experience. To solve these problems, this invention extracts lip movement features from image frames of the user's facial region and combines this with audio feature timestamps to generate an initial spatiotemporal offset sequence, constructing dynamic offset information reflecting the actual speaking rhythm. Based on this, a dynamic offset compensation parameter sequence adapted to the current context is generated using a pre-constructed speech-visual synchronization dataset with a benchmark alignment template. Leveraging the powerful multimodal modeling capabilities of a large model, the dynamic offset compensation parameter sequence, lip movement features, and real-time speech input signal are jointly encoded to achieve high-quality reconstruction of missing or delayed speech segments. Ultimately, by matching the reconstructed speech content with lip movement features and combining the mapping relationship between phonemes and lip shapes, a speech waveform highly synchronized with facial movements is generated, improving the stability and immersion of the speech recognition and synthesis system in complex contexts. Figure 1 A flowchart of a speech recognition and speech synthesis optimization method based on a large model is provided as an embodiment of the present invention, such as... Figure 1 As shown, the method includes:

[0070] Step 101: Acquire a sequence of image frames of the real-time voice input signal and the user's facial region, and extract lip movement features from the image frame sequence;

[0071] In this step, the real-time voice input signal refers to the continuous audio stream captured by the microphone when the user speaks, containing sound wave amplitude, frequency components, and timing information, used to analyze the speech content and temporal characteristics. The image frame sequence refers to the set of video frames of the user's face continuously captured by the camera at a fixed frame rate (e.g., 30 frames per second), each frame containing a red, green, and blue pixel matrix, used to extract the lip movement trajectory. Lip movement features refer to dynamic parameters quantified from the image frame sequence, including the rate of change of lip contour height (the speed of change of the vertical distance from the upper lip to the lower lip), opening and closing angle (the angle between the line connecting the corners of the mouth and the horizontal plane), and displacement trajectory (the path of movement of the lip center point).

[0072] In this embodiment of the invention, a microphone array is used to acquire the user's real-time voice input signal (including raw audio data containing time-domain waveforms and frequency-domain spectra), while a camera captures a sequence of image frames (a set of facial video frames arranged in chronological order) of the user's facial region. Based on a facial key point detection algorithm, the lip region is located, and lip motion features (including three-dimensional motion parameters such as changes in lip contour height, opening and closing speed, and displacement trajectories of the left and right corners of the mouth) are extracted to provide visual dynamic features for subsequent spatiotemporal alignment.

[0073] Step 102: Identify the lip feature timestamps of the image frame sequence and the audio feature timestamps of the real-time speech input signal to generate an initial spatiotemporal offset sequence;

[0074] In this embodiment of the invention, the boundary of the lip region in the image frame sequence is detected, and an array of reference points with equal spacing is set within the boundary; a displacement dataset for each reference point is generated based on the positional change of the reference point array in different image frames; reference points located on the lip contour line in the dataset and whose displacement exceeds a preset distance are defined as dynamic key points to identify the lip feature timestamps of the dynamic key points; the moment of energy amplitude jump in the real-time speech input signal is detected as the audio feature timestamp; the lip feature timestamps and audio feature timestamps corresponding to the image frames in the same time period are paired to obtain multiple paired timestamps, and the difference between each paired timestamp is calculated, and all differences are combined into an initial spatiotemporal offset sequence.

[0075] Step 103: Based on the initial spatiotemporal offset sequence, and combined with the reference alignment template in the pre-constructed speech-visual synchronization dataset, generate a dynamic offset compensation parameter sequence;

[0076] In this embodiment of the invention, a historical reference template consistent with the user's lip opening and closing frequency is selected from a pre-constructed speech-visual synchronization dataset; the reference offsets in the historical reference template are connected to form a reference curve; each offset and its corresponding time stamp in the initial spatiotemporal offset sequence are extracted, and the reference offset corresponding to each time stamp is located in the reference curve; based on the offset corresponding to each audio feature timestamp and the reference offset, the corresponding offset compensation value is obtained; a compensation index table is constructed according to the grouping identifier assigned to each offset compensation value and the preset basic compensation coefficient; according to the time stamp corresponding to each offset in the initial spatiotemporal offset sequence, the corresponding preset compensation coefficient is extracted from the compensation index table, and all preset compensation coefficients are integrated to obtain a dynamic offset compensation parameter sequence.

[0077] Step 104: Using a large model, the dynamic offset compensation parameter sequence, the lip movement features, and the real-time speech input signal are jointly processed to reconstruct the missing or delayed target speech segment.

[0078] In this embodiment of the invention, the lip movement features are subjected to time-axis scaling processing based on the dynamic offset compensation parameter sequence to generate a time-synchronized lip movement trajectory sequence; this sequence is then fused with the real-time speech input signal to generate a multimodal input stream; a large model is used to detect speech interruption regions in the multimodal input stream, and the corresponding lip movement trajectory segments are extracted to generate phoneme identifiers, durations, and energy levels. Phoneme time-segment analysis, contour arrangement, and energy scaling processing are then performed on the speech interruption regions to obtain adjusted interruption regions. Combined with anchor frames of normal speech regions, adjusted transition regions are generated. Combined with the adjusted interruption regions, a connected spectral contour is obtained. Combined with the energy level, the missing or delayed target speech segments are reconstructed.

[0079] Step 105: Based on the target speech segment and the lip movement features, generate a modified phoneme sequence, and retrieve the lip shape parameter group corresponding to the modified phoneme sequence from the preset phoneme lip shape mapping rule library to generate a speech waveform that is synchronized with the lip movement phase.

[0080] In this embodiment of the invention, the target speech segment is processed by the speech recognition branch of a large model to generate an initial phoneme sequence; the displacement rate of the lip movement trajectory is extracted from multiple phoneme time periods to scale the duration of the corresponding phonemes in the initial phoneme sequence to generate a corrected phoneme sequence; the lip shape parameter group corresponding to the corrected phoneme sequence is retrieved from a preset phoneme lip shape mapping rule library, and the lip shape pattern encoding therein is compared with the lip movement features to calculate the shape difference value of each phoneme time period; the vocal tract expansion of the lip shape parameter group is corrected using the shape difference value in the lip shape parameter group to generate an optimized lip shape parameter sequence; based on the phoneme identifier and duration of the corresponding phoneme in the corrected phoneme sequence, the tongue position coordinates and vocal tract expansion of the optimized lip shape parameter sequence, a speech waveform synchronized with the lip movement phase is generated.

[0081] This invention generates dynamic compensation parameters by matching the lip opening and closing frequency with the benchmark alignment template, and combines this with displacement rate-driven phoneme duration scaling to solve the problem of audio-visual asynchrony caused by sudden changes in speech rate and accent differences. Based on lip movement features, phoneme parameters are generated and spectrum reconstruction is driven. Combined with anchor point gradual fusion, natural and coherent audio is output when speech is missing or there is high latency, eliminating lip-reading mismatch and ultimately improving the accuracy of speech recognition and the synchronization of synthesis in complex scenarios.

[0082] This invention provides a specific embodiment. Step 102 involves identifying the lip feature timestamps of the image frame sequence and the audio feature timestamps of the real-time speech input signal to generate an initial spatiotemporal offset sequence. This specifically includes the following steps:

[0083] Step 201: Detect the boundary of the lip region in the image frame sequence, and set an array of reference points with equal spacing within the boundary of the lip region;

[0084] In this step, the lip region boundary refers to the closed polygonal region formed by connecting the outer contours of the upper and lower lips as determined by the facial key point detection algorithm, used to define the range of lip motion analysis. The reference point array refers to the grid of coordinate points set at fixed intervals (e.g., 5 pixels) inside the lip region boundary, used to quantify the microscopic movements of the lips.

[0085] In this embodiment of the invention, an image frame sequence is processed by a facial key point detection algorithm based on a convolutional neural network to locate the boundary points of the upper and lower lips and connect them to form the boundary of the lip region. Inside this boundary, a reference point array is deployed according to grid coordinates (one reference point is set every 5 pixels) to form a reference point array covering the lip movement area.

[0086] Step 202: Generate a displacement dataset for each reference point based on the positional changes of the reference point array in different image frames;

[0087] In this step, the position change refers to the coordinate difference (Δx, Δy) of the same reference point between adjacent image frames, reflecting the local displacement vector of the lip. The displacement dataset refers to a structured collection of data that stores the position changes of all reference points in chronological order, in the format {point ID:[(Δx, Δy)-t1,(Δx, Δy)-t2,...]}.

[0088] In this embodiment of the invention, optical flow is used to track the movement trajectory of each reference point between adjacent frames, and the position change (Δx=x{t}-x{t-1}, Δy=y{t}-y{t-1}) is calculated by the coordinate difference between the current image frame and the previous image frame. The position change of each reference point is recorded in time series for all image frames to form a displacement dataset (structure: {reference point ID: [(Δx1,Δy1),(Δx2,Δy2),...]}).

[0089] Step 203: Define the reference points in the displacement data that are concentrated on the lip contour line and whose displacement exceeds a preset distance as dynamic key points;

[0090] In this step, the displacement of the lip contour line refers to the displacement vector magnitude (√(Δx²+Δy²)) of a reference point located on the lip contour line. The preset distance refers to the displacement threshold for determining significant lip movement (e.g., 0.5 mm), set based on the physiological characteristics of human lip movement. The dynamic key point refers to a reference point that simultaneously meets the conditions of being located on the lip contour line and having a single-frame displacement exceeding the preset distance, representing the effective lip movement position.

[0091] In this embodiment of the invention, reference points that meet two conditions in the displacement dataset are selected. Condition 1 is that the reference point must be located on the lip contour line, that is, the shortest distance between the reference point coordinates and the boundary of the lip region is less than 2 pixels. Condition 2 is that the displacement exceeds a preset distance, that is, the magnitude of the single-frame displacement vector (√(Δx²+Δy²)) is greater than the threshold of 0.5mm. Reference points that meet both conditions are marked as dynamic key points.

[0092] Step 204: Identify the timestamp of the image frame corresponding to the peak displacement in the motion trajectory of the dynamic key point as the lip feature timestamp;

[0093] In this step, the motion trajectory refers to the path of positional changes of dynamic keypoints across consecutive frames, formed by connecting coordinate sequences from the displacement dataset. The displacement magnitude refers to the magnitude of the displacement vector in each frame of the motion trajectory, used to quantify the intensity of lip movements. The lip feature timestamp is the precise time marker (unit: milliseconds) corresponding to when the displacement magnitude in the motion trajectory reaches its peak.

[0094] In this embodiment of the invention, extreme value analysis is performed on the motion trajectory of each dynamic key point to locate the image frame with the largest displacement vector magnitude (i.e., the image frame where the displacement reaches its peak value), and the timestamp of the image frame in the image sequence is extracted (accurate to the millisecond level) as the lip feature timestamp characterizing the significant movement of the lips.

[0095] Step 205: Divide the real-time voice input signal into multiple audio segments, and detect the moment when the energy amplitude jumps in each audio segment as the audio feature timestamp;

[0096] In this step, an audio segment refers to a real-time speech input signal segment divided into fixed durations (e.g., 20ms). An energy amplitude jump refers to a sudden event where the short-term energy ratio between adjacent audio segments exceeds a set threshold (e.g., 3 times). An audio feature timestamp refers to the precise time stamp of the energy amplitude jump.

[0097] In this embodiment of the invention, the real-time voice input signal is processed in frames with a window of 20ms to obtain multiple audio segments. The short-time energy of each audio segment is calculated (i.e., the sum of the squares of the amplitudes of all sampling points divided by the image frame length). When the energy ratio of adjacent image frames exceeds 3 times, it is determined that an energy amplitude jump has occurred, and the precise time of the jump is recorded as the audio feature timestamp.

[0098] Step 206: Pair the lip feature timestamps and audio feature timestamps corresponding to the image frames in the same time period to obtain multiple paired timestamps, calculate the difference between each paired timestamp, and combine all the differences into an initial spatiotemporal offset sequence.

[0099] In this step, the paired timestamp refers to the combination of the closest lip feature timestamp and audio feature timestamp within the same time period (t_lip, t_sound). The initial spatiotemporal offset sequence refers to the set of paired timestamp differences [Δt1, Δt2, ..., Δtn] arranged in chronological order, quantifying the degree of audio-visual asynchrony.

[0100] In this embodiment of the invention, within a 500ms time window, the lip feature timestamp is paired with the nearest audio feature timestamp to generate a paired timestamp. The difference between each pair of timestamps (Δt = tlip - tcon) is calculated, and all Δt values ​​are arranged in chronological order to form an initial spatiotemporal offset sequence (e.g., [Δt1, Δt2, ..., Δtn]).

[0101] The embodiments of the present invention can adaptively capture motion events of different speaking styles (such as rapid lip movements or weak pronunciation) through dynamic key point screening, overcoming the dependence of traditional schemes on predefined feature points; through millisecond-level pairing of displacement peak timestamps and energy jump timestamps, a reliable spatiotemporal offset sequence can still be generated when network packet loss or environmental noise causes partial signal loss, providing accurate input for subsequent compensation.

[0102] This invention provides a specific embodiment. Step 103 involves generating a dynamic offset compensation parameter sequence based on the initial spatiotemporal offset sequence and a reference alignment template from a pre-constructed speech-visual synchronization dataset. This specifically includes the following steps:

[0103] Step 301: Select a historical reference template that matches the user's lip opening and closing frequency from the pre-built speech-visual synchronization dataset. The historical reference template contains a fixed correspondence between lip feature time markers and audio feature time markers in a standard scenario.

[0104] In this step, the pre-constructed speech-visual synchronization dataset refers to a pre-collected set of data on the strict synchronization of lip movements and audio between a standard speaker and the audio in a quiet environment. It includes time-stamped pairs and a baseline offset to provide a personalized calibration benchmark. The user's lip opening and closing frequency refers to the number of times the lips complete an opening-closing cycle per unit time (per second), calculated by dividing the number of times the lip contour height change rate exceeds a threshold by time, reflecting the speech rhythm characteristics. The historical baseline template refers to a subset of standard speaker data in the pre-constructed speech-visual synchronization dataset that is closest to the current user's opening and closing frequency, containing a fixed temporal mapping relationship for that speaker. The standard scenario refers to an ideal acquisition environment with no environmental noise, facing the camera directly, and speaking at a constant speed, used to ensure the reliability of the baseline data. The lip feature time stamp refers to the precise moment (milliseconds) when the displacement in the lip movement trajectory reaches its peak, characterizing significant lip movement events. The audio feature time stamp refers to the precise moment when the energy amplitude jumps in the audio signal, characterizing significant articulation events. The fixed correspondence refers to the fixed time difference between the lip movement feature time stamp and the audio feature time stamp in the historical template, reflecting the physiological articulation delay of the standard speaker.

[0105] In this embodiment of the invention, the user's lip opening and closing frequency is obtained by calculating the number of times the user's lips open and close per unit time (i.e., the number of times the lip opening and closing frequency changes by the number of times the lip contour height change rate exceeds the threshold ÷ time). The user then retrieves historical benchmark templates with a frequency difference of less than 10% from a pre-built speech-visual synchronization dataset. These templates contain a fixed correspondence between lip feature time markers (i.e., peak lip movement moments) and audio feature time markers (i.e., moments of energy surges) in standard scenarios (such as quiet environments and front-facing cameras). For example, the lip movement feature time markers and audio feature time markers for the plosive / p / are strictly 50ms apart.

[0106] Step 302: Connect the various reference offsets in the historical reference template to form a reference curve;

[0107] In this step, the reference offset refers to the fixed time difference between the lip movement feature time marker and the audio feature time marker in a standard scene, serving as the calibration reference value. The reference curve refers to a continuous curve formed by linearly connecting the reference offsets of historical templates in chronological order, used to provide a reference value at any given time.

[0108] In this embodiment of the invention, the reference offsets arranged in chronological order in the historical reference template are extracted, and adjacent reference offsets are connected by line segments using linear interpolation to generate a continuous and smooth reference curve. The horizontal axis represents the lip movement feature time marker and the audio feature time marker, and the vertical axis represents the reference offset.

[0109] Step 303: Extract each offset and its corresponding audio feature timestamp from the initial spatiotemporal offset sequence, and locate the reference offset corresponding to each audio feature timestamp in the reference curve;

[0110] In this embodiment of the invention, each offset and its corresponding audio feature timestamp in the initial spatiotemporal offset sequence are read. Using the audio feature timestamp as the abscissa value, the ordinate value is found on the reference curve to obtain the reference offset of the corresponding audio feature timestamp. For example, when the audio feature timestamp = 3.2s, the corresponding reference offset on the reference curve is 0.04s.

[0111] Step 304: Based on the offset and baseline offset corresponding to each audio feature timestamp, obtain the corresponding offset compensation value;

[0112] In this step, the offset compensation value refers to the difference between the offset and the reference offset, reflecting the asynchronous deviation of the current scene relative to the standard scene.

[0113] In this embodiment of the invention, the offset compensation value is calculated by subtracting the reference offset corresponding to the same audio feature timestamp from the initial offset. For example, if the initial offset is 0.05s and the reference offset is 0.04s, then the offset compensation value of the audio feature timestamp is 0.01s.

[0114] Step 305: Construct a compensation index table based on the grouping identifier assigned to each offset compensation value and the preset basic compensation coefficient;

[0115] In this step, the group identifier refers to the interval number (e.g., groups 1-4) divided according to the offset compensation value, used for classifying and managing compensation parameters. The preset base compensation coefficient refers to the initial scaling factor preset for each group, used for dynamically adjusting the time axis. The compensation index table refers to the set of key-value pairs storing group identifiers and adjusted compensation coefficients, supporting fast table lookup response.

[0116] In this embodiment of the invention, multiple compensation value ranges are set, and a group identifier is assigned to each compensation value range to label each offset compensation value with the corresponding group identifier; the total number of offset compensation values ​​with the same group identifier is counted, and when the total number exceeds a preset activation threshold, the corresponding offset compensation values ​​are combined into an active group; a preset basic compensation coefficient is assigned to each active group, and the distribution density of the offset compensation values ​​of each active group is calculated; based on the set density threshold and the distribution density of each active group, the basic compensation coefficient is adjusted to obtain the corresponding adjusted compensation coefficient; the group identifier and the adjusted compensation coefficient of each active group are associated and stored as key-value pairs, and all key-value pairs are integrated to construct a compensation index table.

[0117] Step 306: Based on the audio feature timestamp corresponding to each offset of the initial spatiotemporal offset sequence, extract the corresponding preset compensation coefficient from the compensation index table, and integrate all preset compensation coefficients into a dynamic offset compensation parameter sequence.

[0118] In this step, the preset compensation coefficient refers to the final scaling factor retrieved from the compensation index table, used to correct the time offset. The dynamic offset compensation parameter sequence refers to all preset compensation coefficients arranged in chronological order, driving subsequent time-domain scaling processing.

[0119] In this embodiment of the invention, based on the audio feature timestamp of the initial spatiotemporal offset sequence, the group ID to which the corresponding offset compensation value belongs is found. For example, when the audio feature timestamp is t=3.2s, the offset compensation value is 0.01s, and the corresponding group ID is 3. The coefficient 1.32 of group ID=3 is retrieved from the compensation index table as the preset compensation coefficient, and the dynamic offset compensation parameter sequence [1.32, 1.28, ...] is generated in time order.

[0120] This invention addresses the reliance of traditional solutions on fixed speaking patterns by filtering historical benchmark templates based on lip opening and closing frequencies; it achieves real-time adaptation to sudden changes in speech rate through offset compensation value grouping and density driving coefficient adjustment; and the compensation index table mechanism can still output stable compensation parameters when network latency fluctuates, avoiding audio-visual misalignment.

[0121] This invention provides a specific embodiment. Step 305 involves constructing a compensation index table based on the grouping identifier assigned to each offset compensation value and the preset basic compensation coefficient. This specifically includes the following steps:

[0122] Step 311: Set multiple compensation value ranges and assign a grouping identifier to each compensation value range;

[0123] In this step, the compensation value range refers to the continuous interval in which the offset compensation value is divided according to its numerical value, such as [-0.1s, -0.05s], which is used to classify and manage compensation requirements of different magnitudes.

[0124] In this embodiment of the invention, the offset compensation value is divided into several continuous intervals according to the distribution of historical data, and each interval is used as the compensation value range, such as [-0.1s, -0.05s), [-0.05s, 0s), [0s, 0.05s), [0.05s, 0.1s). Each interval is assigned a unique number as a group identifier (such as group ID=1,2,3,4).

[0125] Step 312: Determine the compensation value range to which each offset compensation value belongs, and mark each offset compensation value with a corresponding grouping identifier;

[0126] In this embodiment of the invention, for each offset compensation value (e.g., 0.01s), the range of compensation values ​​to which its value belongs is determined, such as (0.01s∈[0s,0.05s)), and the compensation value is marked with the corresponding group identifier (group ID=3).

[0127] Step 313: Count the total number of offset compensation values ​​with the same group identifier. When the total number exceeds the preset activation threshold, combine the corresponding offset compensation values ​​into an active group.

[0128] In this step, the preset activation threshold refers to the minimum number of samples (e.g., 5) required to determine the validity of a group. Groups with fewer than this value are considered noise data and are filtered out. Active groups refer to valid groups whose number of compensation values ​​exceeds the preset activation threshold, and are used for subsequent coefficient calculations.

[0129] In this embodiment of the invention, offset compensation values ​​are aggregated according to group identifiers (e.g., group ID=3 contains 12 offset compensation values). If the number of offset compensation values ​​with the same group identifier in a group exceeds a preset activation threshold, such as a preset activation threshold of 5, then the group is marked as an active group, i.e. an effective compensation group, in order to filter out noise groups with insufficient quantity.

[0130] Step 314: Assign a preset base compensation coefficient to each active group, and calculate the distribution density of the offset compensation value for each active group.

[0131] In this step, distribution density refers to the number of offset compensation values ​​per unit interval width within the active group, reflecting the degree of data concentration.

[0132] In this embodiment of the invention, a preset basic compensation coefficient is assigned to each active group, such as group ID=3 and basic compensation coefficient=1.2. The distribution density of the offset compensation value for each active group is calculated, i.e., distribution density = number of compensation values ​​within the group ÷ group width.

[0133] Step 315: Based on the set density threshold and the distribution density of each active group, adjust the basic compensation coefficient of each active group to obtain the corresponding adjusted compensation coefficient;

[0134] In this step, the density threshold is set to the critical distribution density value that triggers coefficient adjustment (e.g., 200 samples / second), determined based on historical data statistics. The adjusted compensation coefficient is the final scaling factor after dynamically adjusting the basic compensation coefficient according to the distribution density, used for time axis correction.

[0135] In this embodiment of the invention, if the distribution density of an active group is greater than a set density threshold (e.g., threshold = 200 groups / second), the basic compensation coefficient is increased proportionally. The adjusted compensation coefficient is calculated as: basic compensation coefficient × (1 + excess ratio). For example, if the distribution density of an active group is 240 > the set density threshold of 200, the basic compensation coefficient is adjusted to 1.2 × (10.2) = 1.44, and the adjusted compensation coefficient is generated.

[0136] Step 316: Associate and store the group identifier and adjusted compensation coefficient of each active group as key-value pairs, and integrate all key-value pairs to build a compensation index table;

[0137] In this step, key-value pairs refer to the mapping units that store group identifiers and compensation coefficients (such as {group ID: coefficient}), which are the basic building blocks of the compensation index table.

[0138] In this embodiment of the invention, the group identifier (group ID=3) of the active group and the adjusted compensation coefficient (1.44) are combined into a key-value pair ({3:1.44}), and all key-value pairs are summarized to form a compensation index table (structured as {group ID1: coefficient 1, group ID: coefficient 2, ...}).

[0139] This invention improves the reliability of compensation parameters by filtering low-frequency noise groups (such as abnormal shifts caused by coughing) through a preset activation threshold; solves the compensation distortion problem caused by full processing of noise data in traditional solutions; automatically enhances the compensation intensity when the distribution density exceeds the threshold, breaking through the rigid limitation of fixed coefficients; and improves the adaptability to sudden changes in speech rate (such as rapid speech in debate scenarios).

[0140] This invention provides a specific embodiment. Step 104 involves jointly processing the dynamic offset compensation parameter sequence, the lip movement features, and the real-time speech input signal to reconstruct the missing or delayed target speech segment. This specifically includes the following steps:

[0141] Step 401: Perform time-axis scaling on the lip movement features according to the dynamic offset compensation parameter sequence to generate a time-synchronized lip movement trajectory sequence;

[0142] In this step, the lip motion trajectory sequence refers to the lip motion data sequence after scaling the time axis by the dynamic offset compensation parameter. It includes contour height and curvature values ​​sorted by frame, reflecting the lip movement state after time synchronization.

[0143] In this embodiment of the invention, the time axis of the original lip movement features is scaled based on a dynamic offset compensation parameter sequence (e.g., [1.3, 1.28, ...]), including: positive compensation (dynamic offset compensation parameter > 1): expanding the time axis to a parameter multiple of the original length (e.g., 1.3 times) to reduce lip movement speed to match delayed audio; negative compensation (dynamic offset compensation parameter < 1): compressing the time axis to a parameter multiple of the original length (e.g., 0.8 times) to accelerate lip movement speed to match ahead audio; after the above positive and negative compensation, a lip movement trajectory sequence synchronized with real-time speech is generated, including a time-corrected sequence of lip contour height and curvature values.

[0144] Step 402: Fuse the lip movement trajectory sequence with the real-time speech input signal to generate a multimodal input stream;

[0145] In this step, the multimodal input stream refers to the matrix data that is stitched together with the lip movement trajectory sequence and the real-time speech input signal according to the timestamp, and is used for cross-modal feature fusion.

[0146] In this embodiment of the invention, the lip movement trajectory sequence (which is the lip contour height value of each frame) and the real-time speech input signal (which is the spectrum of each 20ms audio segment) are aligned according to the timestamp and spliced ​​into a matrix-form multimodal input stream with the structure of [timestamp, lip height, spectral features] for processing by a large model.

[0147] Step 403: Use a large model to detect speech interruption regions in the multimodal input stream and extract the lip movement trajectory segments corresponding to the speech interruption regions;

[0148] In this step, the speech interruption region refers to a continuous time period in the multimodal input stream where the audio energy is continuously zero but the lips are still moving, representing a segment where speech is lost or severely delayed. The lip movement trajectory segment refers to a subset of lip contour height and curvature data extracted from the speech interruption region, used to drive speech reconstruction.

[0149] In this embodiment of the invention, a large model is used to analyze the multimodal input stream and detect speech interruption regions, namely, segments where the audio energy is 0 for more than 5 consecutive frames and the lips are continuously moving; the corresponding lip movement trajectory segments are extracted from these segments to extract all lip contour height and curvature data within the speech interruption region.

[0150] Step 404: Based on the lip movement trajectory segment, generate phoneme identifiers, durations and energy levels to perform phoneme time segment analysis, contour arrangement and energy scaling on the speech interruption region to obtain the adjusted interruption region. Combined with the anchor frame of the normal speech region, an adjusted transition region is generated. Combined with the adjusted interruption region, a connected spectral contour is obtained.

[0151] In this step, the phoneme identifier refers to the phoneme category label (such as / ɑ / , / i / ) obtained through lip movement trajectory classification, identifying the pronunciation content. Duration refers to the duration of a single phoneme within the speech interruption region, calculated by multiplying the number of phoneme boundary frames by the single frame duration. Energy level refers to the speech energy intensity value mapped based on the average curvature of lip movements, used to control the synthesized volume. The adjusted interruption region refers to the reconstructed spectral region within the speech interruption region after phoneme segmentation, spectral arrangement, and energy scaling. The normal speech region refers to the spoken speech region directly adjacent to the speech interruption region, providing an anchor point for spectral reconstruction. The anchor frame refers to the last frame of the normal speech region, serving as the starting point for spectral reconstruction of the interruption region. The adjusted transition region refers to the smoothly transitioning spectral region gradually generated towards the interruption region, starting from the anchor frame. The connected spectral profile refers to the complete spectral sequence of the adjusted transition region and the adjusted interruption region spliced ​​together in chronological order.

[0152] Step 405: Reconstruct the missing or delayed target speech segment based on the connected spectral profile and the energy level;

[0153] In this step, the target speech segment refers to the time-domain audio waveform synthesized from the concatenated spectral profile, which is used to replace missing or delayed parts in the original signal.

[0154] In this embodiment of the invention, the connected spectral profile is input to the vocoder, and an amplitude gain controlled by the energy level is superimposed to synthesize a time-domain waveform as the target speech segment, filling in the missing parts of the original signal.

[0155] This invention generates phoneme parameters and spectral contours based on lip movement trajectory segments, and synthesizes content that conforms to lip shape when speech is completely interrupted, overcoming the limitation of traditional solutions that cannot handle lip movements during silence. By generating a smooth transition area through anchor frames, it solves the problem of spectral jumps between interrupted and normal areas, significantly improving the naturalness and continuity of synthesized speech.

[0156] This invention provides a specific embodiment, step 404, which involves generating phoneme identifiers, durations, and energy levels based on the lip movement trajectory segment, to perform phoneme time-segment analysis, contour arrangement, and energy scaling on the speech interruption region to obtain an adjusted interruption region. Combined with anchor frames from the normal speech region, an adjusted transition region is generated. Finally, combined with the adjusted interruption region, a connected spectral contour is obtained. Specifically, this includes the following steps:

[0157] Step 411: Calculate the lip contour height change rate based on the lip movement trajectory segment, match the corresponding phoneme identifier based on the lip contour height change rate, and mark the phoneme boundary points in the lip movement trajectory segment.

[0158] In this step, the lip contour height change rate refers to the rate of change of the vertical distance from the upper lip to the lower lip between adjacent video frames (unit: pixels / second), reflecting the lip opening and closing rate, and is used to detect phoneme boundaries. The phoneme boundary point refers to the frame moment when the lip contour height change rate exceeds a set threshold (e.g., 50 pixels / second), marking the start or end position of phoneme switching.

[0159] In this embodiment of the invention, inter-frame difference calculation is performed on the lip movement trajectory segment: lip contour height change rate = (lip contour height of the current frame - lip contour height of the previous frame) ÷ interval between adjacent frames (0.02 seconds); when the lip contour height change rate exceeds a threshold (e.g., 50 pixels / second), the frame is marked as a phoneme boundary point (e.g., the closing moment of the plosive / p / ); each lip contour height change rate is classified into a corresponding phoneme identifier (e.g., / ɑ / , / i / ) by a pre-trained long short-term memory network model.

[0160] Step 412: Determine the time between adjacent phoneme boundary points as phoneme time intervals, and calculate the duration and average curvature value of the lip movement trajectory based on the phoneme time intervals to determine the energy level;

[0161] In this step, the phoneme time interval refers to the time period between adjacent phoneme boundary points (e.g., 0.3 seconds), used to divide the duration of a single phoneme. The lip movement trajectory refers to the sequence of lip contour height and curvature values ​​recorded chronologically within the speech interruption area. The average curvature value refers to the average lip contour curvature (curvature = 1 / radius) across all frames within the phoneme time interval, reflecting the degree of lip curvature.

[0162] In this embodiment of the invention, the time period between adjacent phoneme boundary points is defined as a phoneme period. The duration and average curvature value of the lip movement trajectory within the period are calculated. The duration is defined as: duration = end boundary point time - start boundary point time; average curvature value = sum of lip curvature of all frames within the phoneme period ÷ number of frames; energy level = average curvature value × 100.

[0163] Step 413: Arrange the standard spectral contours of each phoneme identifier in the speech interruption region according to the phoneme identifier and the duration to obtain the arranged contours, and scale the arranged contours according to the energy level to obtain the scaled contours. Insert a formant gradient sequence between the scaled contours of adjacent phoneme identifiers to obtain the adjusted interruption region.

[0164] In this step, the standard spectral profile refers to an idealized spectral template pre-stored in the phoneme lip-sync mapping rule base, containing parameters such as fundamental frequency and formant positions. The arranged profile refers to a sequence of standard spectral profiles arranged in order of phoneme identifiers. The scaled profile refers to the spectral profile after adjusting the amplitude according to the energy level. The formant gradient sequence refers to transition frames with continuously changing formant positions inserted between adjacent phoneme spectra.

[0165] In this embodiment of the invention, standard spectral contours (such as the fundamental frequency of / ɑ / at 250Hz and formant F1 = 700Hz) in the phoneme lip-sync mapping rule library are called in order of phoneme identifiers. The energy scaling process is performed on the arranged contours according to the energy level: the amplitude of the arranged contours is multiplied by (energy level ÷ 100) to generate the scaled contours. Five frames of formant gradient sequences (F1 linearly transitions from 700Hz to 1500Hz) are inserted between the scaled contours of adjacent phoneme identifiers to obtain the adjusted interruption zone.

[0166] Step 414: Locate the anchor frame adjacent to the start of the speech interruption region in the normal speech region, and extract the spectral profile of the anchor frame as the anchor spectral profile.

[0167] In this step, the spectral profile refers to the frequency domain representation of a single frame of speech signal, including the fundamental frequency, formants, and amplitude envelope. The anchor point spectral profile refers to the spectrum of the last frame of the normal speech region, serving as the reference for reconstructing the interrupted region.

[0168] In this embodiment of the invention, in the normal speech region, that is, the speech area 0.5 seconds before the speech interruption area, the last frame is located as the anchor frame, and the anchor spectral profile of the frame is extracted, such as the fundamental frequency 230Hz and the formant F1=650Hz.

[0169] Step 415: Starting from the anchor point spectrum profile, copy the spectrum profile of a preset number of frames one by one towards the speech interruption area to form a gradual transition area;

[0170] In this step, the preset frame count refers to the set value of the number of frames generated in the gradient transition zone (e.g., 5 frames), which controls the smoothness of the transition. The gradient transition zone refers to the spectrum sequence copied from the anchor point spectrum, used to connect the normal zone and the interruption zone.

[0171] In this embodiment of the invention, starting from the anchor point spectrum profile, the spectrum profile of a preset number of frames (e.g., 5 frames) is copied frame by frame towards the speech interruption area to generate a gradual transition area, the structure of which is [anchor point spectrum, copied spectrum 1, ..., copied spectrum 5]).

[0172] Step 416: Adjust the position of the resonant peak of the spectral profile in the gradual transition region to the position of the resonant peak of the standard spectral profile corresponding to the target phoneme identifier to obtain the adjusted transition region;

[0173] In this step, the target phoneme identifier refers to the first phoneme category label (e.g., / ɑ / ) of the speech interruption region, guiding the spectral adjustment target. The formant position refers to the center frequency of the energy concentration region in the spectrum, determining the timbre characteristics.

[0174] In this embodiment of the invention, the formant position (e.g., F1 value) of each frame in the gradual transition region is linearly adjusted to the formant position (e.g., 700Hz) of the standard spectral profile of the target phoneme identifier (i.e., the first phoneme / ɑ / in the speech interruption region), thus obtaining the adjusted transition region. For example: frame 1 F1: 650Hz, adjusted to 660Hz; frame 2 F1: 670Hz, adjusted to 680Hz; ... frame 5 F1: 690Hz, adjusted to 700Hz, finally generating the adjusted transition region.

[0175] Step 417: Connect the adjusted transition region and the adjusted interruption region to obtain the connected spectrum profile;

[0176] In this embodiment of the invention, the adjusted transition area (e.g., 5 frames) and the adjusted interruption area (e.g., 15 frames) are spliced ​​together in chronological order to form a complete connected spectrum profile (a total of 20 frames of spectrum sequence).

[0177] This invention addresses the problem of traditional methods being unable to segment phonemes during silence by marking phoneme boundaries based on the rate of change in lip contour height; it achieves smooth transitions between phonemes through a formant gradient sequence, eliminating the mechanical abruptness of synthesized speech; and it solves the problem of spectral break between speech interruption areas and normal speech areas by using an anchor frame-guided adjusted transition area.

[0178] This invention provides a specific embodiment, step 105, which involves generating a modified phoneme sequence based on the target speech segment and the lip movement features, and retrieving the lip shape parameter group corresponding to the modified phoneme sequence from a preset phoneme lip shape mapping rule library to generate a speech waveform synchronized with the lip movement phase. The specific steps include:

[0179] Step 501: Process the target speech segment through the speech recognition branch of the large model to generate an initial phoneme sequence;

[0180] In this step, the speech recognition branch refers to the encoder-decoder structure in the large model that specifically handles audio input, used to convert speech into a phoneme sequence. The initial phoneme sequence refers to the raw phoneme recognition result output by the speech recognition branch, containing phoneme identifiers and their unadjusted durations.

[0181] In this embodiment of the invention, the target speech segment is input into the speech recognition branch based on the transducer encoder-decoder structure. Through acoustic feature extraction and context modeling, an initial phoneme sequence (such as [ / k / , / ɑ / , / r / ]) is output. Each phoneme contains an identifier and an initial duration (such as / k / with a duration of 0.15 seconds).

[0182] Step 502: Extract the displacement rate of the lip movement trajectory from multiple phoneme time periods, and use the displacement rate to scale the duration of the corresponding phoneme in the initial phoneme sequence to generate a corrected phoneme sequence.

[0183] In this step, displacement rate refers to the average speed at which the center point of the lips moves within a single phoneme segment, reflecting the speed of lip movement. The corrected phoneme sequence refers to the phoneme sequence after scaling the duration using displacement rate.

[0184] In this embodiment of the invention, each phoneme time period is located from the lip movement features, and the displacement rate is calculated (displacement rate = total distance moved by the center point of the lips ÷ duration of the time period). This rate is used to scale the duration of the initial phoneme sequence: scaling formula: corrected duration = initial duration × displacement rate ÷ reference rate, where the reference rate = 5 mm / s, and a corrected phoneme sequence is generated.

[0185] Step 503: Retrieve the mouth shape parameter group corresponding to the modified phoneme sequence from the preset phoneme mouth shape mapping rule library, and compare the lip shape pattern code of the mouth shape parameter group with the lip movement feature to calculate the shape difference value of each phoneme time period. The mouth shape parameter group includes lip shape pattern code, tongue position coordinates and vocal tract expansion degree.

[0186] In this step, the preset phoneme lip shape mapping rule library refers to the pre-stored phoneme-lip shape parameter mapping table, where the key is the phoneme ID and the value is the lip shape code, tongue coordinates, and vocal tract expansion. The lip shape parameter set refers to the set of lip shape parameters corresponding to a single phoneme, including the lip shape pattern code, tongue position coordinates, and vocal tract expansion. The lip shape pattern code refers to abstracting the lip contour into a numerical code (e.g., 1011 represents rounded lips), used to quantify lip shape. The phoneme duration refers to the time interval from the start to the end of a single phoneme, determined by the phoneme boundary points. The shape difference value refers to the Euclidean distance difference between the real-time lip contour and the lip shape pattern code, used to measure lip shape deviation. The tongue position coordinates refer to the three-dimensional coordinates (x, y, z) of the highest point of the tongue surface, controlling the formant position of the synthesized speech. The vocal tract expansion refers to the parameter (0-1) simulating the width of the vocal tract, where 0 is fully closed and 1 is fully open.

[0187] In this embodiment of the invention, a set of mouth shape parameters matching the modified phoneme sequence is retrieved from a preset phoneme mouth shape mapping rule library, such as the parameters corresponding to / k / : lip shape code 1011, tongue body coordinates (0.2, 0.3), and vocal tract expansion degree 0.7; the lip shape pattern code (1011) and the lip shape contour coordinates in the real-time lip movement features are calculated using Euclidean distance to obtain the shape difference value.

[0188] Step 504: Use the shape difference value to correct the vocal tract expansion of the lip shape parameter group to generate an optimized lip shape parameter sequence;

[0189] In this step, the optimized lip shape parameter sequence refers to the lip shape parameter sequence after shape difference value correction, arranged in phoneme order.

[0190] In this embodiment of the invention, the vocal tract expansion is adjusted according to the shape difference value: Correction rule: If the shape difference value > threshold 0.1, then the vocal tract expansion is multiplied by (1 + shape difference value) to generate an optimized mouth shape parameter sequence, the structure of which is [phoneme ID, tongue coordinates, corrected expansion].

[0191] Step 505: Based on the phoneme identifiers and durations of the corresponding phonemes in the modified phoneme sequence, the tongue position coordinates and vocal tract expansion of the optimized mouth shape parameter sequence, generate a speech waveform that is phase-synchronized with the lip movements.

[0192] In this step, the speech waveform refers to the final output time-domain audio signal, whose phoneme duration, formant features, and lip movements are synchronized.

[0193] In this embodiment of the invention, the vocoder is driven to perform the following: loading the phoneme identifier and duration of the modified phoneme sequence (e.g., / k / identifier, 0.18 seconds); loading the tongue position coordinates and vocal tract expansion (e.g., (0.2, 0.3), 0.805) of the optimized lip shape parameter sequence; and synthesizing a speech waveform that is strictly aligned with the lip movements, wherein the vocoder controls the vocal tract width through the expansion and adjusts the formant position through the tongue coordinates.

[0194] This invention addresses the issue of lip-sound asynchrony caused by fixed duration in traditional solutions by using displacement rate-driven duration correction (e.g., automatically extending phoneme duration for slow speech in the elderly); it overcomes the limitations of single audio features to achieve cross-modal adaptive adjustment; it corrects vocal tract expansion in real time based on shape difference values ​​(e.g., increasing expansion for those with unclear speech) to improve speech clarity; and it uses joint control of tongue coordinates and lip shape encoding to ensure millimeter-level synchronization between synthesized speech and lip shape.

[0195] Figure 2 This invention provides a schematic diagram of the structure of a speech recognition and speech synthesis optimization system based on a large model, as shown in the embodiment of the invention. Figure 2 As shown, the system includes:

[0196] The acquisition module 21 is used to acquire real-time voice input signals and image frame sequences of the user's facial region, and extract lip movement features from the image frame sequences;

[0197] The recognition module 22 is used to recognize the lip feature timestamp of the image frame sequence and the audio feature timestamp of the real-time voice input signal to generate an initial spatiotemporal offset sequence.

[0198] The generation module 23 is used to generate a dynamic offset compensation parameter sequence based on the initial spatiotemporal offset sequence and the reference alignment template in the pre-built speech-visual synchronization dataset.

[0199] The reconstruction module 24 is used to use a large model to jointly process the dynamic offset compensation parameter sequence, the lip movement features and the real-time speech input signal to reconstruct the missing or delayed target speech segment.

[0200] The retrieval module 25 is used to generate a modified phoneme sequence based on the target speech segment and the lip movement features, and retrieve the lip shape parameter group corresponding to the modified phoneme sequence from a preset phoneme lip shape mapping rule library to generate a speech waveform that is phase-synchronized with the lip movement.

[0201] Figure 2 The aforementioned large-model-based speech recognition and speech synthesis optimization system can perform... Figure 1 The implementation principle and technical effects of the large-model-based speech recognition and speech synthesis optimization method described in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit performs operations in the large-model-based speech recognition and speech synthesis optimization system described in the above embodiments have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0202] In one possible design, Figure 2 The speech recognition and speech synthesis optimization system based on a large model shown in the embodiment can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0203] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.

[0204] The processing component 32 is used for the above Figure 1 The embodiment describes a speech recognition and speech synthesis optimization method based on a large model.

[0205] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0206] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0207] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.

[0208] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.

[0209] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.

[0210] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.

[0211] This invention also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The embodiment shown is a speech recognition and speech synthesis optimization method based on a large model.

[0212] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0213] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0214] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing speech recognition and speech synthesis based on a large model, characterized in that, include: Acquire real-time voice input signals and image frame sequences of the user's facial region, and extract lip movement features from the image frame sequences; The lip feature timestamps of the image frame sequence and the audio feature timestamps of the real-time speech input signal are identified to generate an initial spatiotemporal offset sequence. Based on the initial spatiotemporal offset sequence, and combined with the benchmark alignment template in the pre-constructed speech-visual synchronization dataset, a dynamic offset compensation parameter sequence is generated. Using a large model, the dynamic offset compensation parameter sequence, the lip movement features, and the real-time speech input signal are jointly processed to reconstruct the missing or delayed target speech segment. Based on the target speech segment and the lip movement features, a modified phoneme sequence is generated. The lip shape parameter group corresponding to the modified phoneme sequence is retrieved from the preset phoneme lip shape mapping rule library to generate a speech waveform that is phase-synchronized with the lip movement.

2. The method according to claim 1, characterized in that, Identifying the lip feature timestamps of the image frame sequence and the audio feature timestamps of the real-time speech input signal to generate an initial spatiotemporal offset sequence includes: The boundary of the lip region in the image frame sequence is detected, and an array of reference points with equal spacing is set within the boundary of the lip region; Based on the positional changes of the reference point array in different image frames, a displacement dataset for each reference point is generated; The reference points that are centrally located on the lip contour line and whose displacement exceeds a preset distance are defined as dynamic key points. The timestamp of the image frame corresponding to the peak displacement in the motion trajectory of the dynamic key point is used as the lip feature timestamp. The real-time voice input signal is divided into multiple audio segments, and the moment point of energy amplitude jump in each audio segment is detected as the audio feature timestamp. The lip feature timestamps corresponding to image frames within the same time period are paired with audio feature timestamps to obtain multiple paired timestamps. The difference between each paired timestamp is calculated, and all differences are combined into an initial spatiotemporal offset sequence.

3. The method according to claim 1, characterized in that, Based on the initial spatiotemporal offset sequence, and combined with the benchmark alignment template in the pre-constructed speech-visual synchronization dataset, a dynamic offset compensation parameter sequence is generated, including: A historical benchmark template consistent with the user's lip opening and closing frequency is selected from a pre-constructed speech-visual synchronization dataset. The historical benchmark template contains a fixed correspondence between lip feature time markers and audio feature time markers in a standard scenario. Connect the various reference offsets in the historical reference template to form a reference curve; Extract each offset and its corresponding audio feature timestamp from the initial spatiotemporal offset sequence, and locate the reference offset corresponding to each audio feature timestamp in the reference curve; Based on the offset and baseline offset corresponding to each audio feature timestamp, the corresponding offset compensation value is obtained; A compensation index table is constructed based on the group identifier assigned to each offset compensation value and the preset basic compensation coefficient; Based on the audio feature timestamp corresponding to each offset in the initial spatiotemporal offset sequence, the corresponding preset compensation coefficients are extracted from the compensation index table, and all preset compensation coefficients are integrated into a dynamic offset compensation parameter sequence.

4. The method according to claim 3, characterized in that, Based on the group identifier assigned to each offset compensation value and the preset base compensation coefficient, a compensation index table is constructed, including: Set multiple compensation value ranges and assign a grouping identifier to each compensation value range; Determine the range of compensation values ​​to which each offset compensation value belongs, and then label each offset compensation value with a corresponding group identifier; The total number of offset compensation values ​​with the same group identifier is counted. When the total number exceeds a preset activation threshold, the corresponding offset compensation values ​​are combined into an active group. Assign a preset base compensation coefficient to each active group, and simultaneously calculate the distribution density of the offset compensation value for each active group: Based on the set density threshold and the distribution density of each active group, the basic compensation coefficient of each active group is adjusted to obtain the corresponding adjusted compensation coefficient. The group identifier and adjusted compensation coefficient of each active group are associated and stored as key-value pairs, and all key-value pairs are integrated to build a compensation index table.

5. The method according to claim 1, characterized in that, Using a large model, the dynamic offset compensation parameter sequence, the lip movement features, and the real-time speech input signal are jointly processed to reconstruct missing or delayed target speech segments, including: The lip movement features are subjected to time-axis scaling processing based on the dynamic offset compensation parameter sequence to generate a time-synchronized lip movement trajectory sequence. The lip movement trajectory sequence is fused with the real-time speech input signal to generate a multimodal input stream; The large model is used to detect speech interruption regions in the multimodal input stream, and the corresponding lip movement trajectory segments are extracted. Based on the lip movement trajectory segment, phoneme identifiers, durations, and energy levels are generated to perform phoneme time segment analysis, contour arrangement, and energy scaling on the speech interruption region to obtain the adjusted interruption region. Combined with the anchor frame of the normal speech region, an adjusted transition region is generated. Combined with the adjusted interruption region, the connected spectral contour is obtained. Based on the connected spectral profile and the energy level, the missing or delayed target speech segments are reconstructed.

6. The method according to claim 5, characterized in that, Based on the lip movement trajectory segment, phoneme identifiers, durations, and energy levels are generated to perform phoneme time-segment analysis, contour alignment, and energy scaling on the speech interruption region, resulting in an adjusted interruption region. Combined with anchor frames from the normal speech region, an adjusted transition region is generated. Finally, based on the adjusted interruption region, a connected spectral contour is obtained, including: Based on the lip movement trajectory segment, calculate the lip contour height change rate, match the corresponding phoneme identifier based on the lip contour height change rate, and mark the phoneme boundary points in the lip movement trajectory segment. The time between adjacent phoneme boundary points is defined as a phoneme time interval, and the duration and average curvature value of the lip movement trajectory are calculated based on the phoneme time interval to determine the energy level; Based on the phoneme identifier and the duration, the standard spectral contours of each phoneme identifier in the speech interruption region are arranged to obtain the arranged contours. The arranged contours are then scaled according to the energy level to obtain the scaled contours. A formant gradient sequence is inserted between the scaled contours of adjacent phoneme identifiers to obtain the adjusted interruption region. Locate the anchor frame adjacent to the start of the speech interruption region in the normal speech region, and extract the spectral profile of the anchor frame as the anchor spectral profile. Starting from the anchor point spectrum profile, the spectrum profile of a preset number of frames is copied frame by frame towards the speech interruption region to form a gradual transition area; The position of the resonant peak of the spectral profile in the gradual transition region is adjusted to the position of the resonant peak of the standard spectral profile corresponding to the target phoneme identifier to obtain the adjusted transition region. The adjusted transition region and the adjusted interruption region are connected to obtain the connected spectrum profile.

7. The method according to claim 1, characterized in that, Based on the target speech segment and phoneme time period, a modified phoneme sequence is generated. A set of lip shape parameters corresponding to the modified phoneme sequence is retrieved from a preset phoneme lip shape mapping rule base to generate a speech waveform synchronized with lip movement phase, including: The target speech segment is processed by the speech recognition branch of the large model to generate an initial phoneme sequence; The displacement rate of the lip movement trajectory is extracted from multiple phoneme time periods, and the duration of the corresponding phoneme in the initial phoneme sequence is scaled using the displacement rate to generate a modified phoneme sequence. The lip shape parameter set corresponding to the modified phoneme sequence is retrieved from the preset phoneme lip shape mapping rule library, and the lip shape pattern code of the lip shape parameter set is compared with the lip movement feature to calculate the shape difference value of each phoneme time period. The lip shape parameter set includes lip shape pattern code, tongue position coordinates and vocal tract expansion. The shape difference value is used to correct the tract expansion of the lip shape parameter set in order to generate an optimized lip shape parameter sequence; Based on the phoneme identifiers and durations of the corresponding phonemes in the modified phoneme sequence, the tongue position coordinates and vocal tract expansion of the optimized lip shape parameter sequence, a speech waveform synchronized with the phase of lip movement is generated.

8. A speech recognition and speech synthesis optimization system based on a large model, characterized in that, include: The acquisition module is used to acquire real-time voice input signals and image frame sequences of the user's facial region, and extract lip movement features from the image frame sequences; The recognition module is used to identify the lip feature timestamps of the image frame sequence and the audio feature timestamps of the real-time speech input signal to generate an initial spatiotemporal offset sequence. The generation module is used to generate a dynamic offset compensation parameter sequence based on the initial spatiotemporal offset sequence and the reference alignment template in the pre-built speech-visual synchronization dataset. The reconstruction module is used to utilize a large model to jointly process the dynamic offset compensation parameter sequence, the lip movement features, and the real-time speech input signal to reconstruct the missing or delayed target speech segment. The retrieval module is used to generate a modified phoneme sequence based on the target speech segment and the lip movement features, and retrieve the lip shape parameter group corresponding to the modified phoneme sequence from a preset phoneme lip shape mapping rule library to generate a speech waveform that is phase-synchronized with the lip movement.

9. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement the large model-based speech recognition and speech synthesis optimization method as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that, The system contains a computer program that, when executed by a computer, implements a large-model-based speech recognition and speech synthesis optimization method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Character-driven lip sound synchronous digital human generation method and device, equipment and medium

    CN119274534A

  • Voice interaction method and device based on lip language enhancement, equipment and storage medium

    CN120600019A