Digital population profile driven method, system and media based on phoneme recognition
By using phoneme-level fine-grained recognition and dynamic mapping matching, the problem of poor lip-sync between speech caused by noise and echo interference in traditional digital human voice interaction has been solved, thus improving the naturalness and realism of digital human voice interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN LUDIE SOFTWARE TECHNOLOGY CO LTD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-06-26
AI Technical Summary
In traditional digital human voice interaction, environmental noise and echo interference affect the accuracy of phoneme recognition, resulting in poor synchronization between lip movements and speech, and reducing the naturalness and realism of the interactive experience.
By employing phoneme-level refined recognition, dynamic mapping matching, and precise muscle model-driven methods, the system collects raw speech signals for noise reduction and echo cancellation, constructs a phoneme-lip shape parameter mapping library, dynamically matches lip shape driving parameters with muscle deformation parameters, and drives the deformation of the digital human mouth and surrounding muscle models to achieve synchronization between lip shape and speech pronunciation.
It enhances the naturalness and realism of digital human voice interaction, ensuring real-time synchronization between lip movements and speech.
Smart Images

Figure CN122290625A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and more specifically, to a digital lip-tracking method, system, and medium based on phoneme recognition. Background Technology
[0002] With the rapid development of digital human technology, digital humans have been gradually applied to various interactive scenarios. Voice interaction, as the core method of communication between digital humans and users, directly determines the naturalness of the interactive experience through the synchronization of lip movements and speech. In traditional driving methods, the preprocessing effect of speech signals is not good, and interference such as environmental noise and echoes will affect the accuracy of phoneme recognition, further reducing the synchronization accuracy between lip movements and speech. Summary of the Invention
[0003] The purpose of this application is to provide a digital human lip-reading driving method, system and medium based on phoneme recognition. Through phoneme-level fine recognition, dynamic mapping matching and precise driving of muscle models, it can realize real-time synchronization between lip-reading and speech pronunciation, and improve the naturalness and realism of digital human voice interaction.
[0004] This application also provides a digital lip-reading method based on phoneme recognition, including:
[0005] The raw input speech signal is collected, and noise reduction, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information.
[0006] Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time sequence information;
[0007] Construct a preset phoneme-lip shape parameter mapping library, which includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles;
[0008] The phoneme timing information is dynamically matched with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters;
[0009] The mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to drive the model deformation and obtain the driving result. Based on the driving result, the model is processed to synchronize mouth shape and speech pronunciation to obtain synchronization result information.
[0010] Optionally, in the digital lip-reading method based on phoneme recognition described in the embodiments of this application, the input raw speech signal is acquired, and the raw speech signal is subjected to noise reduction, echo cancellation, and standardization processing to obtain preprocessed data information, specifically including:
[0011] Acquire the raw speech signal input by the user, perform preliminary detection on the raw speech signal, and analyze the basic feature information of the signal;
[0012] The basic feature information of the signal is analyzed by using spectral subtraction or Wiener filtering algorithm to obtain noise feature information. The noise feature information is then removed to obtain clean speech signal information.
[0013] The clean speech signal information is adaptively filtered to generate echo information. The echo information is then removed to obtain an echo-free signal information.
[0014] The clean speech signal information and the anechoic signal information are uniformly standardized to obtain preprocessed data information.
[0015] Optionally, in the digital lip-reading method based on phoneme recognition described in the embodiments of this application, feature extraction and phoneme recognition are performed on the preprocessed data information to obtain accurate phoneme category and time sequence information, specifically including:
[0016] Set the length information for each frame, and then divide the preprocessed data into frames.
[0017] Add a Hanning or Hamming window to the preprocessed data after frame splitting to obtain the windowed preprocessed data.
[0018] Fourier transform is performed on the windowed preprocessed data to convert the time-domain speech signal in the preprocessed data into a frequency-domain signal, thereby obtaining the spectral characteristics of each frame of preprocessed data.
[0019] A phoneme recognition model is constructed by inputting the spectral features of each frame of preprocessed data into the phoneme recognition model to obtain phoneme category and temporal sequence information.
[0020] Optionally, in the phoneme-based digital lip-shape driving method described in the embodiments of this application, a preset phoneme-lip-shape parameter mapping library is constructed. The mapping library includes lip-shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles, specifically including:
[0021] Set the phoneme category, pronunciation style, mouth shape parameters, and muscle deformation parameter types;
[0022] A unified operational standard for data collection is established, which combines the collected speech samples, lip shape parameters, muscle deformation parameters, and subsequent mapping relationships to obtain a phoneme-lip shape parameter mapping library.
[0023] Optionally, in the digital lip-sync driving method based on phoneme recognition described in the embodiments of this application, the phoneme temporal information is dynamically matched with a mapping library to obtain the corresponding lip-sync driving parameters and muscle deformation parameters, specifically including:
[0024] The time sequence information is acquired, parsed, and core information is extracted. The core information includes the category, start time, end time, and duration of each phoneme, as well as the prosodic features of the speech, to obtain the parsed information.
[0025] The parsed information is standardized to obtain parsed information in a standard format;
[0026] Set a matching threshold to compare the parsed information in the standard format with the matching threshold;
[0027] If the value is greater than or equal to the matching threshold, the corresponding mouth shape driving parameters and muscle deformation parameters are obtained.
[0028] If the value is less than the matching threshold, the time sequence information will be optimized.
[0029] Optionally, in the phoneme-based digital lip-shape driving method described in this application embodiment, lip shape parameters and muscle deformation parameters are input into a digital human mouth and surrounding muscle model to drive model deformation, obtain driving results, and perform lip-shape and speech pronunciation synchronization processing on the model based on the driving results to obtain synchronization result information, specifically including:
[0030] Call the 3D model of the digital human mouth and surrounding muscles to verify the model's running status;
[0031] If the running status meets the set running conditions, the mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to obtain the driving result.
[0032] If the running status does not meet the set running conditions, correction information is generated, and the model parameters are adjusted based on the correction information.
[0033] Secondly, embodiments of this application provide a digital lip-reading driving system based on phoneme recognition. The system includes a memory and a processor. The memory includes a program for a digital lip-reading driving method based on phoneme recognition. When the program for the digital lip-reading driving method based on phoneme recognition is executed by the processor, it implements the following steps:
[0034] The raw input speech signal is collected, and noise reduction, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information.
[0035] Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time sequence information;
[0036] Construct a preset phoneme-lip shape parameter mapping library, which includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles;
[0037] The phoneme timing information is dynamically matched with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters;
[0038] The mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to drive the model deformation and obtain the driving result. Based on the driving result, the model is processed to synchronize mouth shape and speech pronunciation to obtain synchronization result information.
[0039] Optionally, in the phoneme-recognition-based digital lip-reading system described in this application embodiment, the input raw speech signal is acquired, and the raw speech signal is subjected to noise reduction, echo cancellation, and standardization processing to obtain preprocessed data information, specifically including:
[0040] Acquire the raw speech signal input by the user, perform preliminary detection on the raw speech signal, and analyze the basic feature information of the signal;
[0041] The basic feature information of the signal is analyzed by using spectral subtraction or Wiener filtering algorithm to obtain noise feature information. The noise feature information is then removed to obtain clean speech signal information.
[0042] The clean speech signal information is adaptively filtered to generate echo information. The echo information is then removed to obtain an echo-free signal information.
[0043] The clean speech signal information and the anechoic signal information are uniformly standardized to obtain preprocessed data information.
[0044] Optionally, in the digital lip-reading system based on phoneme recognition described in this application embodiment, feature extraction and phoneme recognition are performed on the preprocessed data information to obtain accurate phoneme category and time sequence information, specifically including:
[0045] Set the length information for each frame, and then divide the preprocessed data into frames.
[0046] Add a Hanning or Hamming window to the preprocessed data after frame splitting to obtain the windowed preprocessed data.
[0047] Fourier transform is performed on the windowed preprocessed data to convert the time-domain speech signal in the preprocessed data into a frequency-domain signal, thereby obtaining the spectral characteristics of each frame of preprocessed data.
[0048] A phoneme recognition model is constructed by inputting the spectral features of each frame of preprocessed data into the phoneme recognition model to obtain phoneme category and temporal sequence information.
[0049] Thirdly, embodiments of this application also provide a computer-readable storage medium, which includes a digital lip-reading driving method program based on phoneme recognition. When the digital lip-reading driving method program based on phoneme recognition is executed by a processor, it implements the steps of the digital lip-reading driving method based on phoneme recognition as described in any of the preceding claims.
[0050] As can be seen from the above, the digital human mouth shape driving method, system, and medium provided in this application embodiment, based on phoneme recognition, acquires the input raw speech signal, performs noise reduction, echo cancellation, and standardization processing on the raw speech signal to obtain preprocessed data information; performs feature extraction and phoneme recognition on the preprocessed data information to obtain accurate phoneme category and time sequence information; constructs a preset phoneme and mouth shape parameter mapping library, the mapping library including mouth shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles; dynamically matches the phoneme time sequence information with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters; inputs the mouth shape parameters and muscle deformation parameters into the digital human mouth and surrounding muscle model, drives the model deformation, obtains the driving result, and performs synchronization processing of mouth shape and speech pronunciation on the model according to the driving result to obtain synchronization result information; through phoneme-level fine recognition, dynamic mapping matching, and precise driving of muscle models, real-time synchronization of mouth shape and speech pronunciation is achieved, improving the naturalness and realism of digital human voice interaction. Attached Figure Description
[0051] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 A flowchart of a digital lip reading driving method based on phoneme recognition provided in an embodiment of this application;
[0053] Figure 2 This is a flowchart illustrating the preprocessing of raw speech information in the digital lip-reading method based on phoneme recognition provided in this application embodiment.
[0054] Figure 3 A flowchart illustrating the lip-sync processing of a phoneme-based digital lip-sync driving method provided in this application embodiment. Detailed Implementation
[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0056] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0057] Please refer to Figure 1 , Figure 1 This is a flowchart of a phoneme-based digital lip-reading driving method according to some embodiments of this application. This phoneme-based digital lip-reading driving method is used in a terminal device and includes the following steps:
[0058] S101: Acquires the input raw speech signal, performs noise reduction, echo cancellation and standardization on the raw speech signal to obtain preprocessed data information;
[0059] S102, perform feature extraction and phoneme recognition on the preprocessed data information to obtain accurate phoneme category and time sequence information;
[0060] S103, Construct a preset phoneme and mouth shape parameter mapping library. The mapping library includes mouth shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles.
[0061] S104, dynamically match the phoneme timing information with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters;
[0062] S105: Input the mouth shape parameters and muscle deformation parameters into the digital human mouth and surrounding muscle model, drive the model to deform, obtain the driving result, and perform synchronous processing of mouth shape and speech pronunciation on the model based on the driving result to obtain synchronization result information.
[0063] like Figure 2As shown, according to an embodiment of the present invention, the input raw speech signal is acquired, and the raw speech signal is subjected to noise reduction, echo cancellation, and standardization processing to obtain preprocessed data information, specifically including:
[0064] Acquire the raw speech signal input by the user, perform preliminary detection on the raw speech signal, and analyze the basic feature information of the signal;
[0065] The basic feature information of the signal is analyzed by using spectral subtraction or Wiener filtering algorithm to obtain noise feature information. The noise feature information is then removed to obtain clean speech signal information.
[0066] The clean speech signal information is adaptively filtered to generate echo information. The echo information is then removed to obtain an echo-free signal information.
[0067] The clean speech signal information and the anechoic signal information are uniformly standardized to obtain preprocessed data information.
[0068] Specifically, the core of standardization is to unify the format of speech signals from different acquisition devices and in different acquisition scenarios, eliminate the impact of differences in devices and acquisition environments, and ensure the consistency of subsequent feature extraction and phoneme recognition. The specific steps are as follows:
[0069] Sampling rate standardization: The speech signal after noise reduction and echo cancellation is standardized to a sampling rate of 16kHz. If the original signal sampling rate is higher than 16kHz, a downsampling algorithm (such as decimation) is used to preserve the core frequency characteristics of the speech signal and avoid frequency distortion. If the original signal sampling rate is lower than 16kHz, an upsampling algorithm (such as interpolation) is used to supplement frequency information and ensure a uniform sampling rate.
[0070] Bit depth standardization: The bit depth of the speech signal is standardized to 16 bits. Through quantization, speech signals with different bit depths (such as 8 bits and 24 bits) are converted into 16-bit bit depths to ensure the amplitude accuracy of the signal, while reducing the amount of data storage and improving the efficiency of subsequent processing. During the quantization process, a uniform quantization algorithm is used to avoid the loss of speech features due to quantization errors.
[0071] Amplitude standardization: The amplitude of the speech signal is normalized and adjusted to the range of [-1,1] to eliminate the influence of volume differences between different acquisition devices; by calculating the maximum amplitude of the speech signal, all signal amplitudes are divided by the maximum amplitude to achieve amplitude uniformity, while avoiding signal saturation distortion caused by excessive amplitude or feature extraction difficulties caused by excessive amplitude.
[0072] Preprocessed data verification: After standardization, output preprocessed data information and perform final verification on the data to confirm that the data format (16kHz sampling rate, 16-bit depth), amplitude range, and signal-to-noise ratio all meet the preset requirements, and that there is no noise, echo residue, or signal distortion before it can be input into the subsequent phoneme recognition and timing information extraction stages; if the verification fails, return to the corresponding preprocessing step for reprocessing.
[0073] According to embodiments of the present invention, feature extraction and phoneme identification are performed on preprocessed data information to obtain accurate phoneme categories and time sequence information, specifically including:
[0074] Set the length information for each frame, and then divide the preprocessed data into frames.
[0075] Add a Hanning or Hamming window to the preprocessed data after frame splitting to obtain the windowed preprocessed data.
[0076] Fourier transform is performed on the windowed preprocessed data to convert the time-domain speech signal in the preprocessed data into a frequency-domain signal, thereby obtaining the spectral characteristics of each frame of preprocessed data.
[0077] A phoneme recognition model is constructed by inputting the spectral features of each frame of preprocessed data into the phoneme recognition model to obtain phoneme category and temporal sequence information.
[0078] It should be noted that before feature extraction, the preprocessed data is first validated a second time to ensure that the data meets the input requirements for feature extraction and to avoid affecting the subsequent recognition accuracy due to data anomalies. Specific validation includes: confirming that the sampling rate of the preprocessed data is 16kHz, the bit depth is 16 bits, the amplitude range is within the [-1,1] interval, the signal-to-noise ratio is ≥25dB, and there is no noise, echo residue, or signal distortion; simultaneously, the preprocessed data undergoes normalization and secondary optimization to eliminate minor amplitude differences remaining from the data acquisition process and ensure data consistency. If the validation fails, the data is returned to the speech signal preprocessing stage for reprocessing; if the validation passes, the preprocessed data is adapted to a format recognizable by the feature extraction algorithm and then proceeds to the feature extraction stage.
[0079] The core of feature extraction is to extract core vectors that can accurately represent phoneme features from clean preprocessed speech data, discard irrelevant and redundant information, and provide reliable feature support for phoneme recognition. The specific steps are as follows:
[0080] Speech framing and windowing: The adapted preprocessed speech data is segmented into frames. Taking into account the short-term stationary characteristics of speech signals, the length of each frame is set to 20-40ms (preferably 30ms), with a 10ms overlap between frames to avoid loss of phoneme information at frame boundaries and ensure the temporal continuity of continuous pronunciation. At the same time, a Hanning window or Hamming window is added to each frame of speech signal to suppress spectral leakage at frame edges, reduce interference between adjacent frames, and improve the accuracy of subsequent spectral analysis.
[0081] Frequency domain transformation and spectrum extraction: For each windowed speech signal, a Fast Fourier Transform (FFT) is performed to convert the time-domain speech signal into a frequency-domain signal, obtaining the spectral characteristics of each frame (including frequency distribution, amplitude intensity, etc.); through spectrum analysis, the effective frequency range of the speech signal (200-3400Hz, corresponding to the core frequency range of human pronunciation) is selected, and invalid frequency components outside this range are filtered out, further simplifying the feature data and reducing the complexity of subsequent processing.
[0082] Mel-frequency spectrum and MFCC feature extraction: A Mel-frequency filter bank is used to filter the frequency domain signal, simulating the frequency perception characteristics of the human auditory system (sensitive to low-frequency signals, insensitive to high-frequency signals). The linear spectrum is converted into a Mel-frequency spectrum, highlighting the core frequency features related to phoneme pronunciation. The logarithm of the Mel-frequency spectrum is taken to suppress excessive differences in spectral amplitude. Then, a discrete cosine transform (DCT) is used to remove redundant correlations in the Mel-frequency spectrum, extracting 13-40 dimension (preferably 13 dimension) Mel-frequency cepstral coefficients (MFCCs) as the speech feature vector. This feature vector can accurately characterize the pronunciation differences of different phonemes, especially the distinguishing features of easily confused phonemes (such as the Chinese "n" and "l", and the English "b" and "p"), providing a core basis for phoneme recognition.
[0083] Feature optimization and supplementation: To further enhance the representational ability of the feature vector, the extracted MFCC feature vector is optimized: the first-order and second-order differences of MFCC are calculated to supplement the dynamic change features of phoneme pronunciation and capture the temporal change pattern in the phoneme pronunciation process; at the same time, energy features are introduced to incorporate the energy value of each frame of speech into the feature vector, enriching the feature dimension and ensuring that the feature vector can fully reflect the pronunciation characteristics of the phoneme; finally, the optimized feature vector is normalized and mapped to the [0,1] interval to eliminate feature scale differences and improve the training and recognition efficiency of the subsequent phoneme recognition model.
[0084] Feature validity verification: After extraction, the feature vectors are validated. By calculating the variance and correlation of the feature vectors, valid features (feature dimensions with variance ≥ 0.1) are selected and redundant and invalid features are removed. If the number of valid feature dimensions is less than 13, the feature extraction parameters (such as frame length and number of Mel filters) are readjusted and the feature extraction steps are repeated to ensure that the feature vectors can accurately represent phoneme features.
[0085] The core of phoneme recognition is to accurately identify the phoneme category corresponding to the speech signal by training a mature recognition model based on the extracted speech feature vectors, and at the same time calculate the pronunciation duration of each phoneme to form a complete phoneme time sequence. The specific steps are as follows:
[0086] Model selection and initialization: Select a model suitable for phoneme-level recognition, preferably a DNN-HMM model combining deep neural networks (DNN) and hidden Markov models (HMM), or an end-to-end recognition model based on Transformer; during model initialization, import pre-trained parameters (trained from a large number of multilingual and multi-pronunciation style speech samples, covering all phoneme categories of the target language), and set parameters such as the model's recognition threshold and number of iterations to ensure that the model can accurately recognize phonemes of different pronunciation styles and different individuals.
[0087] Feature vector input and model inference: The optimized speech feature vectors are input into the initialized recognition model in batches. The model classifies and recognizes the feature vectors of each frame through forward inference, and determines the phoneme category corresponding to each frame. At the same time, combined with the model's temporal modeling capability, the feature vectors of adjacent frames are associated to avoid single-frame recognition errors and ensure the continuity of phoneme recognition in continuous pronunciation scenarios.
[0088] Phoneme category confirmation and denoising: The initial phoneme recognition results output by the model are denoised to remove phoneme categories with a recognition probability lower than a preset threshold (preferably 90%) to avoid misidentification; the same phonemes identified in consecutive frames are merged to confirm the accuracy of the phoneme category; at the same time, for easily confused phonemes, the recognition results are further corrected by comparing the differences in their feature vectors (such as the difference in the MFCC feature dimension of "n" and "l") to improve the accuracy of phoneme recognition (ensuring recognition accuracy ≥98%).
[0089] Temporal information extraction and calculation: Based on the identified phoneme categories and combined with the temporal information of speech framing, the pronunciation duration of each phoneme is calculated. By counting the number of consecutive frames corresponding to the same phoneme, and combining the frame length and overlap length, the actual pronunciation duration of the phoneme (unit: ms) is calculated. At the same time, the start time (corresponding to the time point of the phoneme in the first frame) and end time (corresponding to the time point of the phoneme in the last frame) of each phoneme on the speech time axis are recorded to form a preliminary phoneme temporal sequence (including phoneme category, start time, end time, and pronunciation duration).
[0090] To ensure that the phoneme timing sequence matches the actual pronunciation rhythm, the initially extracted phoneme timing information is calibrated, and the final output is accurate phoneme category and timing sequence information. The specific steps are as follows:
[0091] Timing deviation detection: Combining the prosodic features of speech (speech rate, stress, pauses), the deviation of the initial phoneme timing sequence is detected. If the speech rate is fast, it is necessary to check whether there is overlap in the timing connection of adjacent phonemes; if the speech rate is slow, it is necessary to check whether there is a problem of excessively large timing intervals between phonemes. At the same time, it is checked whether the pronunciation duration of stressed phonemes conforms to the actual pronunciation rules (the pronunciation duration of stressed phonemes is usually 10%-20% longer than that of unstressed phonemes).
[0092] Timing calibration execution: Targeted calibration is performed on detected deviations: For cases of overlapping timings of adjacent phonemes, the start time of the subsequent phoneme is adjusted to ensure smooth timing transitions; for cases of excessively large timing intervals, the start and end times of phonemes are adjusted through linear interpolation to shorten the interval; for stressed phonemes, their pronunciation duration is appropriately extended, and the start time is adjusted (by 5-10ms) to ensure that the timing is consistent with the actual pronunciation rhythm; simultaneously, the start and end times of phonemes are checked in conjunction with the speech waveform of the preprocessed data to further correct timing deviations and ensure that the timing error is ≤10ms.
[0093] Final Verification and Output: After calibration, the phoneme timing sequence is finally verified to confirm that the phoneme categories are accurate, the pronunciation duration is reasonable, the timing is smooth, and there are no missing or misidentified phonemes. After the verification is passed, the accurate phoneme category and timing sequence information is output. This information will be directly input into the subsequent phoneme-lip shape parameter mapping stage to provide an accurate timing reference for the dynamic matching of lip shape parameters.
[0094] According to an embodiment of the present invention, a preset phoneme-lip shape parameter mapping library is constructed. The mapping library includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles, specifically including:
[0095] Set the phoneme category, pronunciation style, mouth shape parameters, and muscle deformation parameter types;
[0096] A unified operational standard for data collection is established, which combines the collected speech samples, lip shape parameters, muscle deformation parameters, and subsequent mapping relationships to obtain a phoneme-lip shape parameter mapping library.
[0097] It should be noted that the core content to be covered by the mapping library includes all phoneme categories of the target language (e.g., 41 phonemes in Chinese and 44 phonemes in English), common pronunciation styles (formal, friendly, lively, calm, etc.), and the types of mouth shape parameters and muscle deformation parameters to be collected. Mouth shape parameters include lip opening and closing degree (0-100%), corner of mouth deviation (-5mm to +5mm), mandibular movement angle (0° to 30°), tongue position height (0-5 levels), lip thickness variation (0-2mm), etc.; muscle deformation parameters include the contraction / relaxation degree (0-100%) and deformation speed (mm / ms) of key muscles such as the orbicularis oris, levator labii superioris, depressor labii inferioris, and buccinator, etc., to ensure that the parameters cover the core dimensions of mouth and surrounding muscle movement.
[0098] Establish data acquisition standards: unify the hardware equipment and operating procedures for data acquisition, select high-precision facial capture equipment (such as motion capture devices and 3D scanning equipment), and ensure the accuracy of lip shape and muscle movement data acquisition by sampling frequency ≥60fps; clarify the requirements for the acquisition of pronunciation samples, each phoneme should be sampled for different durations (50-200ms) and different pronunciation intensities (light and heavy), and each pronunciation style should correspond to no less than 50 complete speech-lip shape samples to avoid the deviation of the mapping relationship caused by single samples.
[0099] Establish a data storage and processing platform: Build an embedded database (adapted for subsequent system deployment) to store collected speech samples, lip shape parameters, muscle deformation parameters, and subsequent mapping relationships; build a data preprocessing platform to clean and calibrate the collected data to ensure its accuracy and consistency, providing reliable data support for the subsequent establishment of mapping relationships.
[0100] According to an embodiment of the present invention, phoneme timing information is dynamically matched with a mapping library to obtain corresponding lip-sync driving parameters and muscle deformation parameters, specifically including:
[0101] The temporal sequence information is acquired, parsed, and core information is extracted. The core information includes the category, start time, end time, and duration of each phoneme, as well as the prosodic features of the speech, to obtain the parsed information.
[0102] The parsed information is standardized to obtain parsed information in a standard format;
[0103] Set a matching threshold to compare the parsed information in the standard format with the matching threshold;
[0104] If the value is greater than or equal to the matching threshold, the corresponding mouth shape driving parameters and muscle deformation parameters are obtained.
[0105] If the value is less than the matching threshold, the time sequence information will be optimized.
[0106] It should be noted that the output precise phoneme time sequence is parsed to extract core information, including the category, start time, end time, and duration of each phoneme, as well as the prosodic features of the speech (fundamental frequency F0, energy intensity, speech rate, and stress markers). The parsed information is then organized into a standardized matching format to ensure consistency with the parameter retrieval format of the mapping library.
[0107] Mapping library status verification: Call the mapping library storage unit to verify the running status of the mapping library, confirm that the mapping library has been built and the parameters are complete and without abnormalities, the mapping relationship verification meets the standards (parameter matching accuracy ≥ 95%), and the response latency ≤ 10ms; at the same time, based on the current voice scenario (such as virtual anchor, digital human customer service), preset pronunciation style parameters (such as formal, friendly) to provide a benchmark for subsequent accurate matching.
[0108] Matching parameter initialization: Set the core parameters for dynamic matching, including the matching threshold (parameter error ≤ 5%), interpolation smoothing coefficient (0.1-0.3), and accent parameter adjustment range (10%-20%). Initialize the matching buffer to temporarily store intermediate data during the matching process, avoid matching delays or data loss, and ensure that the matching process is efficient and accurate.
[0109] Monophone matching is the foundation of dynamic matching. For each phoneme in the phoneme temporal sequence, it combines its core information with a mapping library for precise matching to obtain the corresponding mouth shape and muscle parameters. The specific steps are as follows:
[0110] Phoneme category and style matching: Based on the current phoneme category (such as Chinese "a" and "i", English "b" and "p"), and combined with the preset pronunciation style, the corresponding basic mouth shape parameters and muscle deformation parameters are retrieved from the mapping library; if the current speech does not have a preset pronunciation style, the default style (formal style) parameters in the mapping library are automatically matched, and the pronunciation duration of the phoneme is recorded to provide a basis for parameter adjustment.
[0111] Parameter adjustment based on pronunciation duration: The retrieved basic parameters are dynamically adjusted according to the actual pronunciation duration of the current phoneme. If the pronunciation duration is higher than the average duration of the phoneme in the mapping library (e.g., average duration 80ms, actual duration 100ms), the amplitude and duration of muscle deformation are appropriately increased to avoid unnatural mouth shape deformation caused by too fast deformation. If the pronunciation duration is lower than the average duration, the deformation amplitude is reduced and the deformation speed is increased to ensure that the mouth shape and pronunciation duration are accurately matched.
[0112] Prosodic feature-based parameter optimization: Combining the fundamental frequency and energy characteristics of speech, the matching parameters are further optimized. When the speech energy is high (stressed phonemes), the amplitude of mouth shape parameters (such as lip opening and corner of mouth deviation) is increased by 10%-20%, the strength of muscle deformation is enhanced, and the mouth state during stressed pronunciation is restored. When the fundamental frequency is high (pitch is high), the tongue height and lip thickness change are appropriately adjusted to match the mouth shape characteristics of high-pitched pronunciation. When the fundamental frequency is low, the jaw movement angle is adjusted to ensure that the mouth shape matches the pitch.
[0113] Single phoneme matching verification: The matching result of each phoneme is verified, and the error between the matching parameters and the basic parameters of the mapping library is calculated. If the error is ≤5%, the matching is successful and the parameters are stored in the matching cache. If the error is >5%, the mapping library is searched again, the style parameters or pronunciation duration adaptation coefficient are adjusted, and the matching is repeated until the target is met, so as to avoid the single phoneme matching error from affecting the overall mouth shape fluency.
[0114] For continuous pronunciation scenarios, it is necessary to address the abrupt changes in mouth shape between adjacent phonemes. This can be achieved through smooth transition processing to ensure consistent and natural mouth movements. The specific steps are as follows:
[0115] Adjacent phoneme parameter correlation analysis: Retrieve the mouth shape parameters and muscle deformation parameters of two adjacent phonemes in the matching buffer, analyze the magnitude of the difference between the two (such as the difference in lip opening and closing degree, the difference in muscle contraction degree), and determine whether a smooth transition is needed. If the parameter difference is ≤10%, they can be directly connected without transition. If the parameter difference is >10%, interpolation smoothing is started to avoid abrupt mouth shape changes.
[0116] Interpolation smoothing algorithm execution: Using linear interpolation or cubic spline interpolation algorithms, based on the pronunciation duration and parameter differences of two adjacent phonemes, the mouth shape parameters and muscle deformation parameters of the transition stage are calculated to generate a transition parameter sequence; the number of transition parameters is determined according to the temporal interval of adjacent phonemes to ensure that the transition process is synchronized with the speech pronunciation rhythm, and the transition duration is consistent with the overlap duration of adjacent phonemes (usually 10ms), so as to achieve a smooth transition of mouth shape from the previous phoneme to the next phoneme.
[0117] Continuous matching coherence verification: The matching results and transition parameters of continuous phonemes are verified for coherence. The mouth movement process is simulated to check for problems such as deformation, pauses, and parameter abrupt changes. If pauses are found, the interpolation smoothing coefficient is adjusted and the number of transition parameters is increased. If parameters abruptly change, the matching parameters of adjacent phonemes are re-optimized to ensure that the mouth movement is smooth and natural during continuous pronunciation, which conforms to the mouth shape change pattern of actual human pronunciation.
[0118] By combining the contextual phoneme relationships, the matching parameters are further optimized to improve the accuracy and naturalness of the matching, adapting to scenarios such as connected speech and weak forms in actual pronunciation. The specific steps are as follows:
[0119] Contextual phoneme feature analysis: Analyze the phoneme categories and pronunciation characteristics before and after the current phoneme to determine whether there are cases of liaison or weak pronunciation (such as the liaison of "l" and "i" in the Chinese word "lian" and the weak pronunciation of "at" in the English word "not at all"), extract the correlation features of contextual phonemes, and determine the direction of parameter adjustment.
[0120] Parameter optimization: For connected speech scenarios, the matching parameters of adjacent phonemes are adjusted so that the end of the mouth shape deformation of the previous phoneme partially overlaps with the beginning of the mouth shape deformation of the next phoneme, restoring the mouth shape state during connected speech; for weak pronunciation scenarios, the amplitude of mouth shape parameters and the intensity of muscle deformation of weak pronunciation phonemes are reduced, and the deformation duration corresponding to the pronunciation duration is shortened to fit the characteristics of weak pronunciation and avoid excessively exaggerated mouth shape of weak pronunciation phonemes.
[0121] Multi-scenario adaptation and adjustment: Based on the current application scenarios of voice (such as virtual anchor live broadcast, digital human dialogue), the matching parameters are further optimized. In the live broadcast scenario, the lip shape parameter amplitude is appropriately increased to improve the visual effect. In the dialogue scenario, the parameter adjustment is made to better fit the natural human pronunciation, avoid excessive exaggeration, and ensure that the matching results are adapted to the needs of different scenarios.
[0122] The results of the entire dynamic matching process are fully validated to ensure that the matching parameters are accurate and consistent, meeting the requirements of subsequent model-driven processes. The specific steps are as follows:
[0123] Overall matching accuracy verification: Calculate the average error between all phoneme matching parameters and the standard parameters of the mapping library. If the average error is ≤5% and the single phoneme matching error is ≤5%, then the overall matching meets the standard. If the average error is >5%, readjust the matching threshold and parameter optimization coefficients, and re-execute the matching process until it meets the standard.
[0124] Timing synchronization verification: The timing of the matched mouth shape driving parameters and muscle deformation parameters is compared with the start and end times of the phoneme timing sequence to ensure that the parameter output timing is completely synchronized with the phoneme pronunciation timing, with a timing error of ≤10ms, thus avoiding misalignment between mouth shape and speech pronunciation.
[0125] Parameter standardization output: The verified mouth shape driving parameters (lip opening and closing degree, mouth corner offset, etc.) and muscle deformation parameters (muscle contraction / relaxation degree, deformation speed, etc.) are standardized to unify the parameter format and ensure that they are compatible with the input format of the digital human mouth and surrounding muscle model. The standardized parameters are sorted in chronological order and output to the digital human model driving stage in step S4 to provide core input for accurate model driving.
[0126] Through the above dynamic matching analysis steps, accurate and dynamic adaptation of phoneme temporal information and mapping library is achieved. The matched mouth shape driving parameters and muscle deformation parameters can not only accurately correspond to the pronunciation characteristics of each phoneme, but also achieve smooth mouth shape transitions for continuous pronunciation. At the same time, it adapts to different pronunciation styles and application scenarios, providing solid support for the natural and synchronous driving of digital human mouth shape, and further improving the naturalness and realism of digital human voice interaction.
[0127] like Figure 3 As shown in the embodiment of the present invention, lip shape parameters and muscle deformation parameters are input into a digital human mouth and surrounding muscle model to drive model deformation, obtain driving results, and perform lip shape and speech pronunciation synchronization processing on the model based on the driving results to obtain synchronization result information, specifically including:
[0128] Call the 3D model of the digital human mouth and surrounding muscles to verify the model's running status;
[0129] If the running status meets the set running conditions, the mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to obtain the driving result.
[0130] If the running status does not meet the set running conditions, correction information is generated, and the model parameters are adjusted based on the correction information.
[0131] Secondly, embodiments of this application provide a digital lip-reading driving system based on phoneme recognition. The system includes a memory and a processor. The memory includes a program for a digital lip-reading driving method based on phoneme recognition. When the program for the digital lip-reading driving method based on phoneme recognition is executed by the processor, it implements the following steps:
[0132] The raw input speech signal is collected, and noise reduction, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information.
[0133] Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time sequence information;
[0134] Construct a pre-defined phoneme-lip shape parameter mapping library, which includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles;
[0135] The phoneme timing information is dynamically matched with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters;
[0136] The mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to drive the model deformation and obtain the driving result. Based on the driving result, the model is processed to synchronize mouth shape and speech pronunciation to obtain synchronization result information.
[0137] According to an embodiment of the present invention, the input raw speech signal is acquired, and the raw speech signal is subjected to noise reduction, echo cancellation, and standardization processing to obtain preprocessed data information, specifically including:
[0138] Acquire the raw speech signal input by the user, perform preliminary detection on the raw speech signal, and analyze the basic feature information of the signal;
[0139] The basic feature information of the signal is analyzed by using spectral subtraction or Wiener filtering algorithm to obtain noise feature information. The noise feature information is then removed to obtain clean speech signal information.
[0140] The clean speech signal information is adaptively filtered to generate echo information. The echo information is then removed to obtain an echo-free signal information.
[0141] The clean speech signal information and the anechoic signal information are uniformly standardized to obtain preprocessed data information.
[0142] According to embodiments of the present invention, feature extraction and phoneme identification are performed on preprocessed data information to obtain accurate phoneme categories and time sequence information, specifically including:
[0143] Set the length information for each frame, and then divide the preprocessed data into frames.
[0144] Add a Hanning or Hamming window to the preprocessed data after frame splitting to obtain the windowed preprocessed data.
[0145] Fourier transform is performed on the windowed preprocessed data to convert the time-domain speech signal in the preprocessed data into a frequency-domain signal, thereby obtaining the spectral characteristics of each frame of preprocessed data.
[0146] A phoneme recognition model is constructed by inputting the spectral features of each frame of preprocessed data into the phoneme recognition model to obtain phoneme category and temporal sequence information.
[0147] A third aspect of the present invention provides a computer-readable storage medium including a phoneme-based digital lip-tracking driving method program, wherein when the phoneme-based digital lip-tracking driving method program is executed by a processor, it implements the steps of the phoneme-based digital lip-tracking driving method as described in any of the above claims.
[0148] This invention discloses a digital human mouth shape driving method, system, and medium based on phoneme recognition. It acquires raw input speech signals, performs denoising, echo cancellation, and standardization to obtain preprocessed data; extracts features and identifies phonemes from the preprocessed data to obtain precise phoneme categories and time-series information; constructs a pre-defined phoneme-mouth shape parameter mapping library, including mouth shape parameters and muscle deformation parameters corresponding to different phonemes and pronunciation styles; dynamically matches the phoneme time-series information with the mapping library to obtain corresponding mouth shape driving parameters and muscle deformation parameters; inputs the mouth shape parameters and muscle deformation parameters into a digital human mouth and surrounding muscle model to drive model deformation, obtaining a driving result; and performs synchronization processing of mouth shape and speech pronunciation based on the driving result to obtain synchronization result information. Through phoneme-level refined recognition, dynamic mapping matching, and precise muscle model driving, real-time synchronization of mouth shape and speech pronunciation is achieved, improving the naturalness and realism of digital human voice interaction.
[0149] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0150] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0151] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0152] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] Alternatively, if the integrated units of the present invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. A digital lip-reading method based on phoneme recognition, characterized in that, include: The raw input speech signal is collected, and noise reduction, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information. Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time sequence information; Construct a preset phoneme-lip shape parameter mapping library, which includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles; The phoneme timing information is dynamically matched with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters; The mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to drive the model deformation and obtain the driving result. Based on the driving result, the model is processed to synchronize mouth shape and speech pronunciation to obtain synchronization result information.
2. The digital lip-reading method based on phoneme recognition according to claim 1, characterized in that, The raw input speech signal is acquired, and denoising, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information, specifically including: Acquire the raw speech signal input by the user, perform preliminary detection on the raw speech signal, and analyze the basic feature information of the signal; The basic feature information of the signal is analyzed by using spectral subtraction or Wiener filtering algorithm to obtain noise feature information. The noise feature information is then removed to obtain clean speech signal information. The clean speech signal information is adaptively filtered to generate echo information. The echo information is then removed to obtain an echo-free signal information. The clean speech signal information and the anechoic signal information are uniformly standardized to obtain preprocessed data information.
3. The digital lip-reading method based on phoneme recognition according to claim 2, characterized in that, Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time-series sequence information, specifically including: Set the length information for each frame, and then divide the preprocessed data into frames. Add a Hanning or Hamming window to the preprocessed data after frame splitting to obtain the windowed preprocessed data. Fourier transform is performed on the windowed preprocessed data to convert the time-domain speech signal in the preprocessed data into a frequency-domain signal, thereby obtaining the spectral characteristics of each frame of preprocessed data. A phoneme recognition model is constructed by inputting the spectral features of each frame of preprocessed data into the phoneme recognition model to obtain phoneme category and temporal sequence information.
4. The digital lip-reading method based on phoneme recognition according to claim 3, characterized in that, A pre-defined phoneme-lip shape parameter mapping library is constructed. This library includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and pronunciation styles, specifically including: Set the phoneme category, pronunciation style, mouth shape parameters, and muscle deformation parameter types; A unified operational standard for data collection is established, which combines the collected speech samples, lip shape parameters, muscle deformation parameters, and subsequent mapping relationships to obtain a phoneme-lip shape parameter mapping library.
5. The digital lip-reading method based on phoneme recognition according to claim 4, characterized in that, The phoneme timing information is dynamically matched with the mapping library to obtain the corresponding lip-sync driving parameters and muscle deformation parameters, specifically including: The time sequence information is acquired, parsed, and core information is extracted. The core information includes the category, start time, end time, and duration of each phoneme, as well as the prosodic features of the speech, to obtain the parsed information. The parsed information is standardized to obtain parsed information in a standard format; Set a matching threshold to compare the parsed information in the standard format with the matching threshold; If the value is greater than or equal to the matching threshold, the corresponding mouth shape driving parameters and muscle deformation parameters are obtained. If the value is less than the matching threshold, the time sequence information will be optimized.
6. The digital lip-reading driving method based on phoneme recognition according to claim 5, characterized in that, The lip shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to drive the model deformation and obtain the driving result. Based on the driving result, the model is subjected to synchronization processing of lip shape and speech pronunciation to obtain synchronization result information, specifically including: Call the 3D model of the digital human mouth and surrounding muscles to verify the model's running status; If the running status meets the set running conditions, the mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to obtain the driving result. If the running status does not meet the set running conditions, correction information is generated, and the model parameters are adjusted based on the correction information.
7. A digital lip-reading driving system based on phoneme recognition, characterized in that, The system includes: a memory and a processor, wherein the memory includes a program for a phoneme-based digital lip-reading driving method, and when the program for the phoneme-based digital lip-reading driving method is executed by the processor, it performs the following steps: The raw input speech signal is collected, and noise reduction, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information. Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time sequence information; Construct a preset phoneme-lip shape parameter mapping library, which includes lip shape parameters and muscle deformation parameters corresponding to different phonemes and different pronunciation styles; The phoneme timing information is dynamically matched with the mapping library to obtain the corresponding mouth shape driving parameters and muscle deformation parameters; The mouth shape parameters and muscle deformation parameters are input into the digital human mouth and surrounding muscle model to drive the model deformation and obtain the driving result. Based on the driving result, the model is processed to synchronize mouth shape and speech pronunciation to obtain synchronization result information.
8. The digital lip-reading driving system based on phoneme recognition according to claim 7, characterized in that, The raw input speech signal is acquired, and denoising, echo cancellation, and standardization are performed on the raw speech signal to obtain preprocessed data information, specifically including: Acquire the raw speech signal input by the user, perform preliminary detection on the raw speech signal, and analyze the basic feature information of the signal; The basic feature information of the signal is analyzed by using spectral subtraction or Wiener filtering algorithm to obtain noise feature information. The noise feature information is then removed to obtain clean speech signal information. The clean speech signal information is adaptively filtered to generate echo information. The echo information is then removed to obtain an echo-free signal information. The clean speech signal information and the anechoic signal information are uniformly standardized to obtain preprocessed data information.
9. The digital lip-reading driving system based on phoneme recognition according to claim 8, characterized in that, Feature extraction and phoneme identification are performed on the preprocessed data to obtain accurate phoneme categories and time-series sequence information, specifically including: Set the length information for each frame, and then divide the preprocessed data into frames. Add a Hanning or Hamming window to the preprocessed data after frame splitting to obtain the windowed preprocessed data. Fourier transform is performed on the windowed preprocessed data to convert the time-domain speech signal in the preprocessed data into a frequency-domain signal, thereby obtaining the spectral characteristics of each frame of preprocessed data. A phoneme recognition model is constructed by inputting the spectral features of each frame of preprocessed data into the phoneme recognition model to obtain phoneme category and temporal sequence information.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a phoneme-based digital phasing driving method program, which, when executed by a processor, implements the steps of the phoneme-based digital phasing driving method as described in any one of claims 1 to 6.