A robust optimization method, system, and device for lip-reading retrieval model based on risk perception modeling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-08-14
AI Technical Summary
第一,现有方法对所有帧采用统一处理逻辑,未对帧级输出可靠性进行区分
[0009]其有益效果在于:先逐帧提取音频梅尔频谱特征,构建含音频特征、口型参数、视位标签、牙齿状态的口型码本,检索生成候选口型集合;再基于检索置信度、转移异常度等六项指标,完成帧级风险归一化与等级划分;融合多模态检测实现音节边界感知,明确切换约束;依托风险等级与边界信息建立分级路由规则,结合稳定缓存提供回退基准;通过索引级止损、参数级缓和抑制异常,经三重校验与多级兜底修复,最终输出无突跳、无闪烁、时序连续稳定的数字人口型驱动结果。
Smart Images

Figure CN122575400A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a robust optimization method, system, and device for a lip-reading retrieval model based on risk perception modeling. Background Technology
[0002] With the widespread application of digital human technology in scenarios such as live streaming, customer service, short videos, and virtual anchors, increasingly higher demands are being placed on the naturalness, stability, and real-time performance of digital lip-reading. Existing digital lip-reading generation methods mostly employ audio feature retrieval and matching, building a lip-reading material library or codebook, searching for the optimal lip-reading frame by frame, and directly outputting it, thereby achieving lip-reading-driven processing.
[0003] Current mainstream lip-sync generation methods have the following main shortcomings: First, existing methods use a uniform processing logic for all frames without differentiating the reliability of frame-level output. In scenarios with noise, accents, and varying speech rates, the stability of search results varies significantly, with some frames exhibiting low matching accuracy and poor reliability, and prone to producing incorrect lip movements.
[0004] Second, the direct retrieval output method lacks anomaly identification and risk control mechanisms. If a frame is retrieved incorrectly, it will directly lead to problems such as abrupt changes in mouth shape, lip deformity, tooth flickering, and lip misalignment, which seriously affect the viewing experience and reduce the realism of the digital human.
[0005] Third, existing smoothing processes mostly use global filtering or uniform interpolation methods. Although they can reduce jitter, they blur the details of normal lip movements, causing lip movement delay and inaccurate tracking, and cannot achieve a balance between stability and naturalness.
[0006] Fourth, there is a lack of tiered rollback and repair mechanisms. When high-risk frames occur, the system lacks a reliable fallback strategy and cannot quickly roll back to a stable state, which can easily lead to continuous anomalies and make it difficult to guarantee the continuity of output timing and the stability of the image. Summary of the Invention
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A robust optimization method for a lip-reading retrieval model based on risk perception modeling includes: extracting Mel-spectral features of digital human-driven audio frame by frame; constructing a lip-reading codebook containing audio features, lip-reading parameters, positional labels, and tooth states; performing approximate nearest neighbor retrieval to generate a frame-by-frame candidate lip-reading set; using retrieval uncertainty, transition anomaly, mouth movement anomaly, tooth inconsistency, continuous risk accumulation, and boundary matching as inputs to a multi-factor risk perception model; calculating frame-level risk scores after quantile truncation and normalization, and classifying low / medium / high risk levels; implementing switching boundary inference through phoneme alignment, energy mutation detection, and speech activity detection, and outputting boundary constraint features; and combining risk levels with boundary constraint features. The system executes strategic reasoning, directly outputting the optimal retrieval result for low-risk scenarios, implementing conservative backoff decisions or soft-switch interpolation reasoning for medium-risk scenarios, and prioritizing index-level stop-loss constraints and parameter-level mitigation corrections for high-risk scenarios. It standardizes retrieval behavior through stable index freezing, safe neighborhood filtering, and long-distance index removal, adaptively adjusting parameter variation based on risk values and fulfilling lip and tooth state consistency constraints. An output verification reasoning model is constructed to perform compliance reasoning on lip shape parameter ranges, tooth state matching, and inter-frame variation amplitudes, triggering multi-level backoff corrections in case of anomalies. Through risk-aware model training, hierarchical decision-making reasoning learning, retrieval constraint optimization, cache-assisted decision-making, and adaptive parameter correction, the system achieves robustness optimization for real-time reasoning scenarios.
[0008] A robust optimization system for a lip-reading retrieval model based on risk perception modeling is provided. The system is used to execute executable instructions to perform the aforementioned robust optimization method for a lip-reading retrieval model based on risk perception modeling.
[0009] Its beneficial effects are as follows: First, the audio Mel spectrum features are extracted frame by frame to construct a lip shape codebook containing audio features, lip shape parameters, view position labels, and tooth status, and a candidate lip shape set is generated. Then, based on six indicators such as retrieval confidence and transition anomaly, frame-level risk normalization and level classification are completed. Multimodal detection is integrated to realize syllable boundary perception and clarify switching constraints. Hierarchical routing rules are established based on risk level and boundary information, and a backoff benchmark is provided in combination with stable caching. Through index-level stop loss and parameter-level mitigation to suppress anomalies, and after triple verification and multi-level fallback repair, the final output is a digital lip shape driven result with no jumps, no flicker, and continuous temporal stability.
[0010] This application aims to achieve precise frame-level risk classification, differentiate the processing of low, medium, and high-risk frames, avoid detail blurring or abnormal leakage caused by uniform smoothing, and balance naturalness and stability. It introduces a dual control mechanism of index-level stop-loss and parameter-level mitigation to prevent cross-viewpoint jumps at the source and constrain lip-sync distortion at the end, effectively suppressing issues such as tooth flickering and excessive mouth opening. Furthermore, it includes a stable buffer and multi-level fallback strategy, allowing abnormal frames to quickly revert to a stable state, curbing the spread of anomalies and ensuring consistent output timing. It adapts to various scenarios such as broadcasting and dialogue, and can still output high-quality lip movements in complex environments such as low signal-to-noise ratio and varying speech rates, enhancing the realism of the digital human and the user experience. Attached Figure Description
[0011] Figure 1 A flowchart illustrating a robustness optimization method for a lip-reading retrieval model based on risk perception modeling, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a robust optimization system for a lip-reading retrieval model based on risk perception modeling, provided in an embodiment of the present invention. Detailed Implementation
[0012] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Figure 1 This paper describes a robust optimization method for a lip-reading retrieval model based on risk perception modeling, according to an exemplary embodiment of this application.
[0013] In this application embodiment, a robustness optimization method for a lip-reading retrieval model based on risk perception modeling is provided, such as... Figure 1 As shown: S101: Extract the Mel spectrum features of the digital human-driven audio frame by frame, construct a lip shape codebook containing audio features, lip shape parameters, view position labels, and tooth status, and perform approximate nearest neighbor retrieval to generate a frame-by-frame candidate lip shape set.
[0014] In one implementation, frame-by-frame acoustic feature extraction is performed on the audio driven by the digital human for subsequent lip-syncing. The audio is uniformly sampled at 16kHz and processed in mono, using a 25ms Hanning window and a 10ms frame length to divide the continuous audio into temporally continuous audio frames. Each frame is first pre-emphasized, then subjected to a Fast Fourier Transform to obtain frequency domain data, filtered using an 80-dimensional Mel filter bank, and the energy is calculated and normalized to the range of 0 to 1, ultimately obtaining a fixed-length Mel spectrum feature for each frame. The feature sequence is output frame-by-frame according to the audio temporal sequence to ensure time alignment and feature stability. For example, a 3-second audio clip of a digital human speaking normally is divided into 300 frames at 10ms intervals. 80-dimensional Mel spectrum features are extracted from each frame, outputting 30 sets of continuous features, each set corresponding to one frame of lip-syncing data.
[0015] A lip shape codebook is a standard database that maps audio features to lip shapes, used for retrieval and matching. First, a large amount of digital human video footage with natural pronunciation, normal facial expressions, and standard lip shapes is collected. Audio features, lip height, lip width, and mouth opening are extracted from each frame. Audio features are then bound to corresponding lip shapes, filtering out invalid frames, silent frames, and distorted lip shapes. K-Means clustering is used to group the audio features, with 512 cluster centers, each representing a typical lip shape. Each cluster center is labeled with its view position (closed lips, rounded lips, abducted lips, open mouth) and tooth state (closed, exposed). Audio features, lip shape parameters, view position labels, and tooth states are integrated into fixed entries, forming a lip shape codebook that can be quickly queried. For example, collecting 8 hours of standard digital human pronunciation video yields approximately 240,000 valid frames. Clustering generates 512 typical lip shape centers, each storing corresponding audio features, lip shape parameters, view position type, and tooth state, forming a complete lip shape codebook.
[0016] This algorithm uses real-time audio features to search for matching lip shapes in a lip shape codebook and outputs candidate results. The input consists of the audio features of the current frame and the complete lip shape codebook. The similarity between the current feature and all entries in the codebook is calculated, using cosine distance to measure the degree of matching. Similarity scores are sorted from high to low, and the top 5 to 20 highly matched entries are selected as candidates, while low-matching entries with similarity scores below 0.5 are filtered out. The output includes candidate indexes, matching scores, lip shape parameters, and tooth status, forming a set of candidate lip shapes for the current frame. For example, after inputting real-time audio features for a certain frame, the similarity is compared with each of the 512 data points in the codebook. The top 10 highly matched entries are selected, and two low-scoring entries are removed, resulting in 8 candidate lip shapes, each accompanied by a matching score, lip shape parameters, and tooth status.
[0017] S102 uses retrieval uncertainty, transfer anomaly degree, mouth movement anomaly degree, tooth inconsistency, continuous risk accumulation degree, and boundary matching degree as inputs to the multi-factor risk perception model. After quantile truncation and normalization, the frame-level risk score is calculated by weighted fusion to complete the classification of low / medium / high risk levels.
[0018] In one implementation, frame-by-frame acoustic feature extraction is performed on the audio driven by the digital human. Standardized preprocessing is completed based on the pre-training specifications of the risk perception model, providing standardized audio feature input for subsequent lip-sync codebook construction and candidate retrieval. The input audio is uniformly standardized at a 16kHz sampling rate and mono format, and is processed in frames using a 25ms Hanning window and a 10ms frame length. The continuous speech signal is divided into time-aligned, non-overlapping audio frame sequences to ensure temporal continuity between frames and accurate timestamp correspondence. Each frame of audio is pre-emphasized sequentially, with a fixed pre-emphasis coefficient of 0.95 to compensate for high-frequency energy attenuation and suppress low-frequency noise interference. Then, the time-domain audio signal is converted into frequency-domain spectral data through Fast Fourier Transform. Subsequently, an 80-dimensional Membrane filter bank is used for filtering to simulate the human ear's auditory sensitivity to different frequencies. Finally, the output energy of each filter is calculated and global normalization is performed, with the normalization interval fixed between 0 and 1. The final output is an 80-dimensional Membrane spectral feature vector, ensuring that all feature dimensions are uniform and the numerical range is consistent. For example, a 3-second audio clip of a digital human broadcasting is divided into 300 frames at 10ms frame lengths; 80-dimensional Mel-spectrum features are extracted from each frame, and the final output is a 300×80 temporal feature matrix with fixed feature dimensions and perfect alignment between the temporal sequence and the audio.
[0019] The lip-shape codebook is a standard database that maps audio features to lip shapes, establishing a one-to-one mapping between audio signals and lip forms to support rapid retrieval and matching. First, a massive amount of standard digital human pronunciation video material was collected, covering different speech rates, intonations, expressions, and pronunciation scenarios, with a total duration of no less than 8 hours, ensuring comprehensive lip shape coverage and strong sample representativeness. 80-dimensional Mel-spectral features were extracted frame-by-frame from the audio material, simultaneously extracting three core lip shape parameters: lip height, lip width, and mouth opening. Audio features were then frame-by-frame bound to corresponding lip shape data, filtering out invalid samples such as silent frames, noise interference frames, and frames with distorted lip shapes. Approximately 240,000 valid frames were ultimately selected, ensuring sample quality compliance. The K-Means unsupervised clustering algorithm was employed, using audio features as the core clustering basis, with a fixed number of cluster centers of 512. Each cluster center corresponds to a typical pronunciation lip shape, containing a unique audio feature distribution and a stable lip shape. The clustering convergence threshold was set to 1e-5 to ensure stable and unbiased clustering results.
[0020] Two types of standardized labels were assigned to each cluster center: positional labels were divided into four categories: closed lips, rounded lips, abducted lips, and open mouth; tooth status labels were divided into two categories: closed and exposed. The labels were manually verified to ensure accurate and unambiguous classification. The 80-dimensional audio features, lip height, lip width, mouth opening parameters, positional labels, and tooth status of each cluster center were integrated into a standard codebook entry, ultimately constructing a 512-entry, uniformly structured, and quickly searchable mouth shape codebook. For example, 8 hours of standard digital human pronunciation video were collected, and 240,000 valid frames were selected; K-Means clustering was used to generate 512 typical mouth shape centers, each storing 80-dimensional audio features, lip height, lip width, mouth opening, positional labels, and tooth status, forming a complete standard mouth shape codebook.
[0021] Based on a pre-constructed set of 512 standard lip-sync codes, and utilizing real-time extracted audio features, a frame-by-frame candidate lip-sync set is generated through near-nearest neighbor retrieval, providing alternative results for subsequent risk assessment. The retrieval input includes two types of data: 80-dimensional real-time audio features of the current frame and the complete lip-sync code. A cosine similarity algorithm is used to calculate the degree of matching between the current frame's features and the features of all entries in the code. The cosine similarity calculation formula is as follows: The similarity value ranges from 0 to 1, with the value closer to 1 indicating a higher degree of matching.
[0022] All codebook entries are sorted by similarity from highest to lowest, and the top 5 to 20 highly matched entries are selected as the initial candidate pool. A similarity filtering threshold of 0.5 is set to remove invalid entries with similarity below 0.5 to avoid interfering with subsequent risk assessment. The final output candidate set includes the index number, matching score, lip height, lip width, mouth opening, and tooth status of each candidate, ensuring that the candidate information is complete and can be directly used for subsequent risk assessment and routing decisions. For example, after inputting real-time 80-dimensional audio features of a frame, the similarity is calculated with each of the 512 codebook entries, and the top 10 highly matched entries are selected after sorting. Two invalid entries with similarity below 0.5 are removed, resulting in 8 candidate mouth shapes, each accompanied by an index, matching score, lip shape parameters, and tooth status. In addition, the candidate score is obtained by weighting the audio matching score and the transition smoothing score. , (Preferred 0.8), according to Sort the candidates from highest to lowest to obtain the Top-k candidates.
[0023] Based on six indicators—retrieval confidence, transfer anomaly, mouth movement anomaly, tooth inconsistency, continuous risk accumulation, and boundary matching—a unified approach of 1%–99% quantile truncation and Min-Max normalization was used to map them to the [0,1] interval. Weights were assigned according to the degree of impact on mouth tremors. Retrieval confidence, transfer anomaly, movement anomaly, and boundary matching were the core high-weight items, with fixed weights set as a1=0.20, a2=0.25, a3=0.20, a4=0.10, a5=0.10, and a6=0.15. Among them, a1 is the weight of retrieval confidence (retrieval credibility), which is a core high-weight item; a2 is the weight of transfer anomaly (transfer anomaly), which has the highest weight among the six indicators; a3 is the weight of mouth movement anomaly (movement anomaly), which is a core high-weight item; a4 is the weight of tooth inconsistency; a5 is the weight of continuous risk accumulation; and a6 is the weight of boundary matching (boundary matching), which is a core high-weight item. The core risk assessment indicators are clearly defined as retrieval confidence, transfer anomaly, mouth movement anomaly, tooth inconsistency, continuous risk accumulation, and boundary matching, which are used to quantify the instability of the current frame's lip shape output. To eliminate differences in indicator dimensions and avoid interference from extreme outliers, a unified method combining 1% to 99th percentile truncation and Min-Max normalization is adopted to map all indicator values to the 0-1 range, ensuring that the indicators can be weighted and integrated and that the risk judgment criteria are consistent.
[0024] A multi-factor risk perception model was built based on six indicators. The indicators were truncated to the 1%~99th percentile and normalized to [0,1] using Min-Max. The numerical distribution of all samples for each indicator was statistically analyzed, and extreme outliers below the 1% quantile and above the 99th percentile were removed. Only the middle 98% of valid data were retained for subsequent normalization to avoid outliers distorting the overall risk distribution and causing misjudgments. For example, the transfer anomaly rate data of 100,000 frames of normal lip-sync samples were statistically analyzed. The 1% quantile was 0.03 and the 99th percentile was 0.97. Outliers below 0.03 and above 0.97 were removed, and the remaining data were considered valid samples. Fixed weights a1=0.20, a2=0.25, a3=0.20, a4=0.10, a5=0.10, a6=0.15 were used to obtain the risk score for each frame. A 7-frame causal sliding window was used to calculate continuous risk, and the 85th and 95th percentiles of the samples were used for calibration with 3-fold cross-validation. =0.35、 =0.70, classifying risks into low, medium, and high levels. The weights of each indicator and the threshold for each level are iteratively adjusted in reverse based on the output verification results, and the model is trained and optimized using a large number of abnormal frames and repaired samples. For example, the weighted average risk of the six normalized indicators is 0.147, which is judged as low risk.
[0025] For the truncated valid indicator data, the Min-Max normalization formula is used to complete the numerical mapping, as follows: Normalized value = (Original value) / (Original value) (Minimum valid value) / (Maximum valid value) (Effective Minimum) Through this calculation, the value of each indicator is linearly scaled to the range of 0 to 1. The closer the value is to 1, the higher the risk. This achieves uniformity of the units of measurement for different indicators and allows for quantifiable comparison of risk levels. For example, the original effective data range for the abnormality of mouth movement in a certain frame is 0.05 to 0.95, with an original value of 0.5. The normalized calculation is (0.5) 0.05) / (0.95) 0.05) = 0.5, and the risk value after mapping is 0.5.
[0026] Based on the impact of six indicators on abnormal phenomena such as lip trembling, flickering, and abrupt jumps, fixed weights are assigned, with higher weights given to core, high-impact indicators. The total weight is 1, ensuring that the weighted total risk value is between 0 and 1. Retrieval Confidence (a1): Weight 0.20, measures the certainty of audio-lip matching; the lower the matching degree, the higher the risk. Transition Anomaly (a2): Weight 0.25, measures the abruptness of lip switching between frames; it is the core factor causing lip jumps and has the highest weight. Mouth Movement Anomaly (a3): Weight 0.20, measures whether lip movements exceed the normal speech rate range. Tooth Inconsistency (a4): Weight 0.10, measures the probability of abrupt changes in tooth state between frames, affecting tooth flickering. Continuous Risk Accumulation (a5): Weight 0.10, measures the probability of high-risk frames appearing consecutively. Boundary Matching (a6): Weight 0.15, measures whether lip abrupt changes occur at reasonable syllable boundaries. For example, the normalized values of the six indicators are 0.1, 0.2, 0.15, 0.05, 0.08, and 0.12, respectively. Substituting these values into the weighted formula: Total risk value = 0.20×0.1 + 0.25×0.2 + 0.20×0.15 + 0.10×0.05 + 0.10×0.08 + 0.15×0.12 = 0.147. Among these, the migration anomaly degree has the highest weight and the greatest impact on the total risk; the tooth inconsistency has the lowest weight and a relatively smaller impact.
[0027] To achieve an objective and stable risk level classification, a quantile method was used to set the classification thresholds based on the statistical results of the risk value distribution of a large number of normal lip morphology samples. Specifically, the 85th percentile of the risk distribution of normal samples was set as the low-risk threshold. The 95th percentile was set as the high-risk threshold. The threshold is not determined by a single empirical value; it requires accuracy calibration through three-fold cross-validation. Specifically, normal sample data is randomly divided into three groups. Two groups are used alternately as the training set and one as the validation set. The accuracy of the threshold under different data distributions is repeatedly tested to determine the optimal threshold and avoid misjudgments due to sample bias. After extensive testing and calibration, the following initial empirical value was determined: =0.35、 =0.70. This initial value is suitable for most digital human broadcasting and dialogue scenarios and can be directly used for system initialization. For example, when calculating the total risk score of 100,000 frames of normal digital human mouth shape samples, the 85th percentile corresponds to a value of 0.35, and the 95th percentile corresponds to a value of 0.70; after randomly dividing the samples into 3 groups and cross-validating, the threshold remains stable, and the final value is determined. =0.35、 =0.70.
[0028] The continuous risk accumulation degree is used to distinguish between single-frame random noise and continuous unstable segments, avoiding misjudgment of persistent high risk due to anomalies in individual frames. The calculation uses a causal sliding window of length W=7 frames, covering the current frame and the previous 6 frames. It relies only on historical time-series data and does not involve future frames, adapting to real-time streaming processing. Specifically, the calculation method is as follows: the arithmetic mean of the normalized total risk scores of the 7 frames within the window is taken as the continuous risk accumulation degree of the current frame. The formula is as follows: =a1×R1+a2×R2+a3×R3+a4×R4+a5×R5+a6×R6. Where: a1=0.20 (retrieval confidence weight), R1 is the normalized value of retrieval confidence; a2=0.25 (transition anomaly weight), R2 is the normalized value of transition anomaly; a3=0.20 (mouth movement anomaly weight), R3 is the normalized value of mouth movement anomaly; a4=0.10 (dental inconsistency weight), R4 is the normalized value of dental inconsistency; a5=0.10 (continuous risk accumulation weight), R5 is the normalized value of continuous risk accumulation; a6=0.15 (boundary matching weight), R6 is the normalized value of boundary matching.
[0029] For example, if the normalized values of the six risk indicators in a certain frame are 0.1, 0.2, 0.15, 0.05, 0.08, and 0.12 respectively, and these values are substituted into the formula for calculation: =0.20×0.1+0.25×0.2+0.20×0.15+0.10×0.05+0.10×0.08+0.15×0.12=0.02+0.05+0.03+0.005+0.008+0.018=0.147.
[0030] The frame-level total risk score is obtained by weighted summation of the six indicators. < For low risk, ≤ < For medium risk, ≥ For high-risk scenarios, low, medium, and high-risk levels are categorized, generating frame-level risk assessment results containing risk sub-items, weight parameters, and grading thresholds. After normalizing the six risk indicators, a weighted summation method is used to calculate the total risk score for a single frame, achieving multi-dimensional risk quantification and fusion. The weighted calculation strictly adheres to preset fixed weights, with the sum of all weights equal to 1, ensuring that the total risk score remains stable between 0 and 1. The closer the value is to 1, the higher the risk of the current frame's lip output.
[0031] Based on the total risk score, combined with preset grading thresholds , Each frame is divided into three risk levels: low, medium, and high. The grading standards are clear, and the thresholds are fixed, making it suitable for most digital human broadcasting and dialogue scenarios. For low-risk scenarios: < (Initial empirical value) =0.35), indicating that the candidate lip shape for the current frame is reliable and there is no obvious risk of anomalies; for medium risk: ≤ < (Initial empirical value) =0.70), indicating that the current frame has some unstable factors and needs to be handled with caution; for high-risk situations: ≥ This indicates a high risk of anomalies in the current frame, making it prone to issues such as abrupt lip movements and flickering teeth, requiring focused repair. For example, based on the above calculations, the total risk score for this frame is 0.147, which is less than... A frame with a total risk score of 0.35 is classified as a low-risk frame; a frame with a total risk score of 0.5, which is between 0.35 and 0.70, is classified as a medium-risk frame; and a frame with a total risk score of 0.8, which is greater than or equal to 0.70, is classified as a high-risk frame.
[0032] For medium-risk frames, the preferred procedure is to execute the following programmatic logic: first calculate... , , and If satisfied < , < 、( =0 and >= )or >= If so, a conservative rollback will be executed. >= , >= , =1 and <= < If so, a soft handover will be performed. If the conditions are met... >= , >= and < If the current candidate output is not found, then the current candidate output will be used directly.
[0033] After calculating risk scores and classifying risk levels, complete risk information is integrated to generate frame-level risk assessment results. These results include four core categories, with complete and traceable information to support subsequent routing decisions. The six risk sub-items are the normalized values of each indicator; the weight parameters are the fixed weight values corresponding to the six indicators; and the classification thresholds are... =0.35、 =0.70; the risk level is the low / medium / high risk judgment result corresponding to a single frame. For example, the evaluation results of the low-risk frame above are clearly stated as follows: retrieval confidence 0.1, transfer anomaly 0.2, mouth movement anomaly 0.15, tooth inconsistency 0.05, continuous risk accumulation 0.08, boundary matching degree 0.12; the weights are 0.20, 0.25, 0.20, 0.10, 0.10, 0.15 respectively; the classification threshold is... =0.35、 =0.70; the risk level is low.
[0034] S103 achieves switching boundary inference through phoneme alignment, energy mutation detection, and speech activity detection, and outputs boundary constraint features.
[0035] In one implementation, a priority-driven multimodal boundary perception fusion mechanism is built around three core technologies: phoneme-level forced alignment, short-time energy and zero-crossing rate analysis, and VAD speech activity detection. This mechanism is used to accurately identify syllable boundaries, pause boundaries, and key locations where lip-syncing is permitted in the speech of a digital human. The mechanism employs a hierarchical decision logic, prioritizing high-precision phoneme alignment results, with energy, zero-crossing rate, and VAD detection serving as supplementary fallbacks to ensure comprehensive boundary recognition, stable and reliable results, and adaptability to real-time digital human-driven scenarios.
[0036] Using the phoneme-level forced alignment algorithm, timestamp annotation for each phoneme is performed on the digital human-driven audio to accurately locate the start frame, end frame, and duration of each phoneme. The start and end frames of phonemes directly correspond to syllable boundaries and phoneme transition boundaries, which are the core basis for boundary determination, with the highest priority and the best recognition accuracy. For example, for a 3-second digital human broadcast audio containing 4 syllables, namely "ni, hao, shi, jie", after forced alignment, the annotation is: "ni" (frame 10 - frame 45), "hao" (frame 46 - frame 80), "shi" (frame 81 - frame 115), "jie" (frame 116 - frame 150), and the start and end frames of each phoneme are marked as boundaries.
[0037] Short-time energy reflects the strength of the speech signal, and the zero-crossing rate reflects the frequency mutation of the signal. The combination of the two can identify implicit syllable boundaries such as the transition between initial consonants / finals and the alternation between vowels / consonants. Calculation rule: Calculate the short-time energy and zero-crossing rate frame by frame with a frame length of 10 ms. When the mutation amplitude of the energy between adjacent frames ≥ 30% and the difference in the zero-crossing rate ≥ 20, it is determined as an energy mutation boundary. For example, for a certain segment of speech, the energy of frame 50 is 0.2, the energy of frame 51 suddenly increases to 0.38 (an increase of 90%), and the zero-crossing rate suddenly increases from 15 to 38, which is determined as a syllable transition boundary.
[0038] VAD is used to distinguish speech segments from silent segments and accurately identify the boundaries of pauses between words and pauses between sentences. Judgment rule: Output a binary result, 1 for the speech segment and 0 for the silent segment. The 3 frames before and after the silent segment are forced to be marked as pause boundaries. For example, if frame 60 - frame 68 is a silent segment between words (VAD outputs 0), then frames 57, 58, 59, 69, 70, and 71 are all marked as pause boundaries.
[0039] The fusion mechanism follows the principle of priority exclusivity + fallback complementarity as follows. First, the results of phoneme alignment are adopted. The start and end frames of phonemes / syllables are directly marked as boundaries without relying on other detections. When the phoneme alignment data is missing (such as in low signal-to-noise ratio or dialect speech scenarios), it automatically switches to the combined determination of energy / zero-crossing rate and VAD. The frames with energy mutations and the 3 frames before and after the VAD silent segments are uniformly marked as fallback boundaries to ensure that there are no blanks in boundary recognition. For example, for a low signal-to-noise ratio broadcast audio, phoneme alignment fails; the 40th frame is identified as the boundary between the initial consonant / final through energy mutation, and the 80th - 85th frames are detected as an inter-sentence silent segment through VAD, and the 3 frames before and after are marked as pause boundaries to complete the fallback of boundary recognition.
[0040] Binary boundary markers are generated frame by frame. Phoneme / syllable start and end frames, word pauses, the three frames before and after a VAD silence segment, and frames with energy abrupt changes are marked as allowed switching boundaries. The remaining frames are determined as non-boundary stable constraint regions. Based on a multimodal boundary-aware fusion mechanism, speech features are standardized and analyzed frame by frame to generate binary boundary markers. This clarifies whether each frame represents a boundary location where lip-syncing is allowed and whether it belongs to a constraint region that needs to remain stable, providing a precise basis for subsequent lip-syncing permission allocation.
[0041] Boundary determination is performed frame by frame, generating 0 / 1 binary flag bits. Specifically, flag bit = 1: it is determined to be a boundary where switching is allowed, and lip-sync restrictions can be relaxed; flag bit = 0: it is determined to be a non-boundary stable constraint region, and inter-frame continuity needs to be strengthened and large changes in lip-sync need to be restricted.
[0042] A frame is marked as an allowed switching boundary if any of the following conditions are met: phoneme start and end frames: the start or end frame of a phoneme in the phoneme-level forced alignment result; syllable start and end frames: the start or end frame of a syllable in the syllable segmentation result; word pause position: the frame corresponding to the pause between adjacent words; 3 frames before and after a VAD silence segment: the frame that VAD detection determines is a silence segment, and the 3 frames before and after the silence segment; energy mutation frame: the short-term energy difference between adjacent frames is ≥0.2 (after normalization), and is determined to be an energy mutation frame.
[0043] Frames that do not meet any of the above boundary conditions are uniformly judged as non-boundary stable constraint regions, prohibiting large-scale lip-syncing and only allowing small-scale smooth adjustments. For example, a 1-second, 100-frame digital human voice broadcasts "Hello". Phoneme alignment results show that "you" corresponds to frames 10-40 and "hello" corresponds to frames 41-70; VAD detection shows that frames 71-80 are word-interval silence segments; energy mutation shows that the energy difference between frame 40 (ending with "you") and frame 41 (starting with "hello") is 0.3 (≥0.2).
[0044] The boundary determination results are as follows: Flag = 1 (switching allowed): frames 10, 40, 41, 70, 71, 72, 73, 78, 79, and 80; Flag = 0 (stable constraint): frames 11-39, 42-69, and 81-100. Among them, frame 10 is the starting frame of the "ni" syllable, frame 40 is the ending frame of the "ni" phoneme, frame 41 is the starting frame of the "hao" phoneme, frame 70 is the ending frame of the "hao" syllable, and frames 71-80 are the silent segment and the frames before and after it, all of which are marked as boundaries; the continuous pronunciation frames in the middle are marked as constraint areas, restricting large-scale lip-syncing.
[0045] Establish linkage rules between boundary markers and lip-sync switching. Relax lip-sync switching restrictions at boundary positions and strengthen inter-frame continuity constraints at non-boundary positions to complete the division of switching permissions and solidify stable constraints. Based on generating binary boundary markers frame by frame, establish linkage rules between boundary states and lip-sync switching. The core is to formulate differentiated constraints according to boundary / non-boundary positions, clarify the switching permissions of different frames, solidify stable constraints, and ensure that lip-sync switching conforms to speech timing logic and avoids abnormal jumps.
[0046] Boundary positions include phoneme start / end frames, syllable start / end frames, word pause frames, three frames before and after a VAD silence segment, and energy mutation frames. These frames relax lip-shape switching restrictions, allowing reasonable and obvious changes in lip shape to adapt to the natural lip-shape transitions required for syllable switching and pause connections. Lip shape switching across viewpoints is allowed, such as changing from closed lips to rounded lips, from rounded lips to open lips, and from open lips to closed lips. The range of lip shape parameter changes between frames can be relaxed to 0.25 (after normalization). It only needs to match the pronunciation requirements of the current syllable, without being forcibly bound to the lip shape of the previous frame, ensuring natural and smooth lip-shape switching. For example, when a digital human pronounces "hello," the syllable boundary frame (frame 40, ending with "you") allows the lip shape to switch directly from a closed lip state (lip opening degree 0.05) to an open mouth state (lip opening degree 0.60), with an inter-frame parameter change of 0.55, conforming to the boundary switching relaxation rules.
[0047] Non-boundary positions are continuous pronunciation frames within syllables or phonemes. These frames emphasize inter-frame continuity constraints, strictly limiting abrupt changes in lip shape, allowing only small, smooth adjustments to avoid lip shape jumps and flickering. Large-scale switching across viewpoints is prohibited; only small adjustments within the same viewpoint are permitted. The variation in lip shape parameters between frames is strictly limited to within 0.06 (after normalization). A temporal continuity with the previous frame's lip shape is mandatory; changes in lip height, lip width, and mouth opening parameters must be smooth transitions to prevent abrupt distortions. For example, in the intermediate frame (frame 20, non-boundary) of the digital human pronouncing "you," if the previous frame's lip opening was 0.20, the current frame is only allowed to adjust to the range of 0.18-0.26. If the candidate lip opening is 0.50, it is considered an abnormal abrupt change and is prohibited.
[0048] By using differentiated rules, the precise division of lip-syncing permissions and the solidification of stable constraints are achieved. Specifically, for boundary positions: switching permissions are opened to adapt to the natural lip-syncing changes of speech syllables and pause transitions; for non-boundary positions: stable permissions are locked to force temporal continuity and prevent meaningless lip-syncing abrupt changes; the linkage logic is as follows: the boundary flag directly triggers the corresponding constraint rules without additional calculation, and adapts to the needs of streaming lip-syncing generation in real time.
[0049] After generating the binarized boundary markers, it is necessary to perform consistency checks on the boundary marker timing, the rationality of inter-frame switching, and the energy mutation threshold to correct detection errors and avoid misjudgments and omissions. At the same time, it is necessary to complete the accurate mapping between the voice timestamp and the video frame to ensure accurate boundary determination timing, reasonable switching logic, and audio-visual synchronization.
[0050] For boundary marker timing verification: The continuity of boundary marker timing is checked frame by frame. Frequent alternations of continuous boundaries and non-boundaries within short frames (e.g., alternating 0 / 1 markers for three consecutive frames) are prohibited. Frames with timing errors are corrected to ensure that the boundary distribution conforms to the syllable patterns of speech. For inter-frame switching rationality verification: The logic of lip movements before and after boundary frames is verified. Reasonable switching is allowed in boundary frames, while large abrupt changes are prohibited in non-boundary frames. Unreasonable boundary markers (e.g., isolated boundaries in a single frame, or marking non-boundary frames as switching boundaries) are removed. For energy mutation threshold verification: A unified energy mutation judgment threshold is established. A short-term energy difference ≥0.2 (after normalization) between adjacent frames is fixed as the judgment standard. Abnormal threshold deviations are calibrated to avoid misjudgments caused by threshold fluctuations. For example, in a speech segment, frame 40 is marked as an energy mutation boundary, frame 41 as a non-boundary, and frame 42 as a boundary again, indicating timing errors. After verification, frame 42 is corrected to be a non-boundary to ensure the continuity of boundary timing.
[0051] The timestamps of phonemes and syllables are precisely mapped to their corresponding video frames using the following formula: Frame Number = Timestamp (seconds) × Video Frame Rate (30fps). The timestamp is then rounded down to match the corresponding video frame, eliminating timing discrepancies between audio and video. For example, if the end timestamp of a phoneme is 1.667 seconds, at 30fps: 1.667 × 30 = 50.01. This is rounded down to match the 50th video frame, completing the timestamp mapping.
[0052] One to two frames are extended before and after each boundary frame as a switching safety zone. Within each safety zone, the boundary is marked as a permitted switching boundary to avoid abrupt transitions and audio-visual misalignment caused by minor timing deviations. The safety zone extension rules are as follows: For a single boundary frame: extend by one frame before and after, for a total of three safety zones; For consecutive boundary frames: extend by two frames before and after, forming a continuous safe transition segment. For example, when a phoneme ends and maps to frame 50, one frame is extended before and after, setting frames 49, 50, and 51 as switching safety zones, all marked as permitted switching boundaries; the lip movements of the digital mouth smoothly transition from the closed-lip frame (frame 49) to the open-mouth frame (frame 51), resulting in a natural transition and avoiding audio-visual timing misalignment and abrupt transitions.
[0053] By integrating multimodal detection results, boundary marker sequences, and handover constraint rules, syllable / pause boundary perception and handover constraint result information is generated, including boundary type, frame-level flag bits, handover permission threshold, and stability constraint parameters. After completing boundary timing verification, inter-frame rationality check, audio-visual timing alignment, and safe zone expansion, multimodal detection data, frame-level boundary marker sequences, and handover constraint rules are integrated to form structured, directly callable syllable / pause boundary perception and handover constraint result information, ensuring data integrity, logical closure, and direct support for subsequent routing decisions and lip-sync generation.
[0054] The multimodal detection results include four types of raw detection data: phoneme forced alignment boundaries, short-time energy mutation boundaries, zero-crossing rate mutation boundaries, and VAD silence segment boundaries. The generation basis and corresponding frame number for each boundary are clearly defined. The frame-level boundary marker sequence is a frame-by-frame binary marker (1 = allowed switching boundary, 0 = non-boundary), including boundary safety zone extended frame annotations. Switching constraint rules include three types of threshold parameters: maximum allowed amplitude for boundary switching, maximum allowed amplitude for non-boundary switching, and inter-frame mutation limit. Stability constraint parameters include the upper limit of lip shape parameter changes in non-boundary frames, tooth state preservation rules, and inter-frame smooth transition coefficients.
[0055] The final generated results contain four core modules: complete data items, clear numerical values, and direct parsing capability. Specifically, boundary types include the boundary category of the labeled frame, encompassing six categories: phoneme start / end boundaries, syllable start / end boundaries, word pause boundaries, energy mutation boundaries, safe zone extension boundaries, and non-boundary boundaries. Frame-level flags are used to correspond 0 / 1 flags for each frame, covering the complete audio timing sequence, and each frame uniquely identifies switching permissions. Switching permission thresholds include: boundary frames: maximum allowable variation of inter-frame mouth parameters 0.25 (normalized); non-boundary frames: maximum allowable variation of inter-frame mouth parameters 0.06 (normalized); energy mutation judgment threshold: energy difference between adjacent frames ≥ 0.2 (normalized). Stability constraint parameters include: upper limit of lip opening variation in non-boundary frames: ±0.06; tooth state in non-boundary frames: forced to remain consistent with the previous frame; smooth transition coefficient: 0.8 (high weight biased towards stable frames).
[0056] For example, a 1-second, 100-frame digital human voice broadcast (content: "Hello") can be integrated into the following output: For boundary types: Frame 10 (syllable start boundary), Frame 40 (phoneme end boundary), Frame 41 (phoneme start boundary), Frames 49-51 (safe zone extension boundary), Frames 71-80 (word pause boundary), and the remaining frames (non-boundary); For frame-level flags: Frames 10, 40, 41, 49, 50, 51, and 71-80 = 1, and the remaining frames = 0; For switching permission thresholds: boundary frames 0.25, non-boundary frames 0.06, and energy threshold 0.2; For stability constraint parameters: lip opening ±0.06, tooth state maintenance, and transition coefficient 0.8. The integrated result can be directly output as a structured file, with each frame corresponding to complete boundary information and constraint parameters, supporting subsequent risk routing and lip-sync generation.
[0057] S104 combines risk level and boundary constraint features to perform strategy reasoning. For low risk, the optimal retrieval result is directly output. For medium risk, a conservative backoff decision or soft switching interpolation reasoning is executed. For high risk, index-level stop-loss constraints are executed first and parameter-level mitigation corrections are performed.
[0058] In one implementation, if the current frame is determined to be of medium risk, the system does not directly adopt the current candidate result. Instead, it compares the current candidate result with the most recent stable result and prioritizes the result with better continuity and smaller switching amplitude as the current frame output. This mechanism is suitable for situations where "the current result is not obviously wrong, but the system is not confident enough," and its purpose is to avoid unnecessary abrupt changes when candidates are close together.
[0059] Preferably, conservative backoff and soft handover should not be triggered concurrently, but rather a mutually exclusive judgment logic should be used. Preferably, the program should first determine whether the current frame is eligible for handover. This determination can be automatically calculated by the program using the following indicators: 1. Candidate confidence difference. = - ,in and 1. The matching scores of the first and second candidates, respectively; 2. The parameter difference between the current candidate and the most recent stable result. = 3. Matching gain of the current relatively stable candidate results = - 4. Boundary markers A value of 1 indicates that the current location is a position where obvious switching is allowed, while a value of 0 indicates that the current location is not a position where obvious switching is allowed.
[0060] If any of the following conditions are met, the system is deemed "not eligible for switching," and a conservative rollback is preferred: < or < or( =0 and >= )or >= ,in, This represents the threshold for the difference in candidate credibility. This represents the minimum matched gain threshold required for the switching. This indicates the upper bound of the difference that allows direct switching. This represents the upper bound of the allowable soft handover difference. If the above conditions are met, it means that the current candidate is either not reliable enough, the handover benefit is insufficient, or it does not qualify for handover at the current moment. Therefore, it is safer to continue using the most recent stable result or the stable neighboring result.
[0061] exist Assuming that the correlation values of Match and P have been normalized, the preferred empirical starting value can be set as follows: =0.08, =0.03, =0.06, =0.18, where, The upper bound represents the area where "the differences are minimal and direct switching is possible". This represents the upper bound of "moderate difference, soft switching possible"; when >= At that time, a soft handover will no longer be performed, and a conservative rollback will be implemented instead.
[0062] When a system is determined to be ineligible for switching, or at medium or high risk, the following three-tiered rollback selection rules are executed in order of priority: First Priority: Retrieve the most recent stable frame (optimal rollback). Directly reuse the latest valid stable result from the stable buffer (index, lip shape parameters, and tooth state are fully preserved) without any modifications, ensuring absolutely no jumps. Second Priority: Weighted fusion of stable neighborhoods (optional soft rollback). If the system supports smooth transition, weighted fusion can be performed on the most recent stable frame plus 1-2 adjacent stable entries to output intermediate lip shape parameters, avoiding complete freezing. Third Priority: Directly reuse the previous physically frame (fallback). When the stable buffer is empty or invalid, directly reuse the output result of the previous frame to ensure temporal continuity and prevent crashes.
[0063] The rollback duration rule uses a risk-driven dynamic rollback mechanism, rather than a fixed rollback of 1 frame: 1. Once the dynamic rollback rule enters the rollback state, it continuously maintains the rollback output until one of the following exit conditions is met: the risk of the current frame drops to a low-risk range ( < ); The current frame enters the syllable / phoneme boundary ( =1) And the confidence level returns to normal; the maximum number of consecutive rollbacks is reached. 2. Maximum rollback frame limit setting: The maximum number of consecutive rollback frames is 3 to 5 frames (preferably 3 frames) to avoid unnatural lip movements due to excessive "freezing". 3. Exit rollback mode: When exiting rollback, do not hard switch back to the new candidate, but use a 1 to 2 frame soft transition to restore normal output.
[0064] If the current frame is determined to be of medium risk and the current candidate has some usability, but a direct switch would cause a significant abrupt change, a gradual transition is performed between the current candidate and the stable result, so that the current frame output falls between the two. Preferably, the soft switch mechanism is executed at the medium risk level, for situations where "the current candidate is not obviously wrong, but a direct switch would be too abrupt." Continuous transition frames can be obtained by interpolating the lip-sync parameters of the current candidate with the most recently stable lip-sync parameters, and then generating the current frame output from the interpolated intermediate lip-sync parameters. This mechanism is suitable for situations where "switching is possible, but a hard switch is not advisable," and its purpose is to transform single-frame jumps into continuous transitions, thereby reducing lip-sync jumps and tooth flickering.
[0065] Preferably, after determining that "the current frame is eligible for handover," the next step is to determine "whether a direct hard handover is suitable." The soft handover mechanism is only executed if the handover eligibility determination has been passed. If the following conditions are met, a soft handover is preferred over a direct hard handover: >= and >= and <= < and =1. This means that the current candidate is basically reliable and the relatively stable result has sufficient matching benefit. At the same time, the current position is in a position where switching is allowed, but there is still a moderate difference between the current candidate and the stable result. If the switch is made directly, it is easy to cause a sudden change. Therefore, it is more suitable to obtain a continuous transition frame through interpolation.
[0066] The soft-switching interpolation method uses linear interpolation, which is computationally efficient, suitable for real-time streaming digital human driving, has no latency, and is easy to implement. Interpolation formula, lip-sync parameters, and soft-switching output: =α +(1 α) Recent stable frame lip-sync parameters Current candidate lip shape parameter α: Transition coefficient (0≤α≤1), the larger α is, the more it is biased towards stable frames, the smaller α is, the more it is biased towards the current candidate. The transition coefficient α is calculated by combining the risk value, parameter distance, and boundary markers: α = w1· +w2· +w3·(1 Constraint: α is ultimately clamped in the interval [0.2, 0.8]. Each component contains the following: Total risk score for the current frame; : The parameter distance (normalized) between the current candidate and the stable frame; Boundary marker (1 if at the boundary, 0 if not); Weights: w1=0.4, w2=0.4, w3=0.2. Intuitive logic: Higher risk -- larger α -- more biased towards stable frames; Greater distance -- larger α -- more biased towards stable frames; Not at the boundary -- even larger α -- more conservative. The soft transition frame count uses a multi-frame gradual change instead of a single-frame hard cut to ensure a completely smooth visual experience: Transition frame count: 2-3 frames (preferred): 2-frame gradual change; Frame 1: α decreases by 50% from high to low; Frame 2: α continues to decrease to 0, completing the soft transition and automatically executing the full gradual change cycle without interruption. ≥ (≥0.70), prioritize index-level stop loss, then moderate the output parameters; if the requirements are still not met, revert to the most recent stable result.
[0067] In another implementation, based on frame-level risk levels and boundary marker information, a five-level routing decision rule is established: direct output, conservative backoff, soft handover, index-level stop-loss, and parameter-level mitigation. The entire process relies on a stable cache module to provide a reliable backoff benchmark, achieving precise risk classification and stable lip-sync output. The stable cache is a fixed-length causal sliding window, with a preferred cache length of N=20 frames. Only low-risk, validated, and sequentially continuous stable frames are written as the sole benchmark for backoff and handover, avoiding interference from abnormal frames. For example, one frame has a risk score of 0.28 and a boundary marker of 1 (syllable boundary); another frame has a risk score of 0.75 and a boundary marker of 0 (syllable middle). The system reads the risk and boundary information of both frames, matches them to direct output and index-level stop-loss paths respectively, and simultaneously retrieves the latest stable frame (index 12, lip opening degree 0.20) from the cache as the backoff benchmark.
[0068] When the total risk score at the frame level < ( =0.35), which is determined to be a low-risk frame. The current best candidate lip shape is directly output without backtracking or mitigation. The best candidate is defined as the item with the highest similarity and confidence (difference between the first and second candidates) in the candidate set, which is ≥0.05. Its index, lip shape parameters, and tooth status are directly reused. For example, if the six-item risk weighted score of a frame is 0.28 < 0.35, it is determined to be low-risk; the first candidate in the candidate set has a similarity of 0.82 and a confidence of 0.12, so this candidate (index 5, lip opening degree 0.30, teeth closed) is directly output.
[0069] when ≤ < (0.35≤ If the risk score is less than 0.70 and the boundary flag is 0 (not a boundary), a conservative rollback is executed, prioritizing the reuse of the latest stable frame in the stable cache and prohibiting abrupt changes. Rollback priority: ① Latest stable frame (most recent); ② Second newest stable frame; ③ Previous frame as a fallback. For example, if a frame has a risk score of 0.52 and a boundary flag of 0, it is classified as medium risk; the latest stable frame in the cache is index 12 with a lip opening degree of 0.20, and the parameters of this frame are directly reused to avoid abrupt changes.
[0070] when ≤ If the value is less than 0.70 and the boundary flag is 1 (syllable / pause boundary), a soft handover is performed, using linear interpolation between the current candidate and stable frames for a smooth transition. Interpolation formula: =α +(1 α) Transition coefficient α: Calculated from risk, parameter distance, and boundary, clamped to 0.2-0.8; the larger the α, the more it favors stable frames. Optimal number of transition frames: 2 frames for gradual transition, avoiding abrupt switching. For example, in a frame with risk 0.55, boundary flag = 1, current candidate lip opening degree 0.60, stable frame 0.20; calculate α = 0.5, interpolate to output 0.40, in the second frame α = 0, smoothly switching to candidate lip shape.
[0071] when ≥ For values ≥0.70, prioritize index-level stop-loss, freeze stable indices, remove long-distance anomalous candidates, and prevent cross-viewpoint jumps. The stop-loss steps are as follows: freeze the latest stable index in the cache (prohibit long-distance jumps); filter safe neighborhood entries in the candidate set that are ≤0.25 distance from the stable index; remove long-distance anomalous indices with a distance >0.25. For example, if a frame has a risk of 0.78 and the stable index is closed lip (distance from baseline 0); the candidates include open mouth (distance 0.5) and rounded lip (distance 0.15), remove open mouth candidates, and retain rounded lip safe candidates.
[0072] After index-level stop loss, parameter-level mitigation is implemented, using risk linear constraints and forced tooth lip shape matching to suppress deformities. Specifically, for amplitude clamping: inter-frame parameter variation ≤ 0.20 (normalized), lip opening clamping 0.02-0.90; for risk linear constraints: mitigation intensity k = The higher the risk, the stronger the constraint; the forced matching is as follows: when the lip opening degree is <0.15, the teeth must be closed, and when it is >0.40, exposure is allowed. For example, after stop loss, if the candidate lip opening degree is 0.85, it is clamped to 0.80; the tooth status is exposed, matching 0.80, and the compliance parameters are output.
[0073] S105 standardizes retrieval behavior through stable index freezing, safe neighborhood filtering, and long-distance index removal, and adaptively adjusts the parameter change range based on risk value to complete the consistency constraint of the lip and tooth state.
[0074] In one implementation, a dual-control mechanism is established to address two types of problems in high-risk frames: index anomalies and parameter distortions. This mechanism combines index-level loss mitigation and parameter-level mitigation. On one hand, it controls the source of candidate indices to prevent long-distance, cross-viewpoint jumps; on the other hand, it imposes end-point constraints on lip shape parameters to avoid lip distortion and tooth flickering. This overall dual governance, from the source of the index to the parameter form, ensures stable and compliant output of high-risk frames. For example, when a high-risk frame appears during digital human broadcasting, there may be long-distance entries in the candidate index that jump directly from a closed mouth to a wide-open mouth, and the lip shape parameters may also show distortions such as an excessively large mouth opening and mismatched tooth positions. In this case, the dual-control mechanism is activated, simultaneously regulating both the index and parameters, ultimately outputting a stable, natural, and abnormal lip shape result.
[0075] Index-level stop-loss is used to manage the rationality of candidate indexes, including three standardized operations: freezing stable indexes, filtering safe neighborhoods, and removing long-distance indexes. Specifically, for freezing stable indexes: the latest reliable index in the stable cache is locked first, and long-distance jumps across the benchmark are not allowed. The stable cache length is fixed at 20 frames, and only low-risk, valid frames that have passed verification are stored. For filtering safe neighborhoods: based on the stable index, candidate entries with a lip shape parameter distance less than or equal to 0.25 (normalized) are selected as safe candidates. For removing long-distance indexes: abnormal long-distance indexes with a lip shape parameter distance greater than 0.25 (normalized) are removed to prevent abrupt transitions such as from closed lips to open lips, or from rounded lips to open lips. For example, the stable index is closed lips with a lip opening degree of 0.10; candidates include rounded lips (distance 0.18) and open lips (distance 0.52). After filtering, rounded lip candidates are retained, and long-distance open lips indexes are removed to avoid abrupt index jumps.
[0076] Parametric mitigation is used to correct abnormal mouth shapes, including three standardized operations: amplitude clamping, risk-adaptive variation limitation, and forced matching of teeth and lips. Specifically, for amplitude clamping: the maximum allowable variation of mouth shape parameters between frames is set to 0.20 (normalized); the lip opening range is fixed between 0.02 and 0.90 to prevent excessive mouth opening or lip flattening. For risk-adaptive variation limitation: the mitigation intensity varies linearly with the risk value, and the mitigation intensity k = clip( (0.3, 1.0); the higher the risk, the smaller the allowed variation; high-risk frames only allow minor adjustments. For forced matching of teeth and lips: when the lip opening is less than 0.15, the teeth are forced closed; when the lip opening is greater than 0.40, the teeth are allowed to be exposed; the tooth state in the intermediate range remains consistent with the previous frame to avoid tooth flickering. For example, a high-risk frame candidate lip opening of 0.88 exceeds the upper limit and is clamped to 0.85; the risk value is 0.78, and the allowed variation between frames is limited to 0.04; the tooth state is matched to be exposed, and the final output is compliant lip shape parameters.
[0077] After completing index-level stoppage and parameter-level mitigation, high-risk frame restoration results are generated, containing three core categories: stoppage index, mitigation parameters, and matching rules. The data is complete, numerically clear, and directly parseable. The stoppage index records the frozen stable index and the filtered safe candidate indexes. The mitigation parameters include the lip opening degree after clamping, the inter-frame change, and the mitigation intensity. The matching rules are used to annotate the basis for matching tooth status with lip shape and the mandatory constraints. For example, the restoration results include a stable index (closed lip), a safe candidate index (rounded lip), mitigation parameters (lip opening degree 0.85, change 0.04), and matching rules (teeth exposure), which can be directly used for mouth shape generation.
[0078] For index-level stop-loss, the latest reliable indexes in the stable cache are frozen first, and then entries within the stable neighborhood are screened in the candidate set, while distant abnormal indexes are removed to prevent cross-viewpoint jumps. After index-level stop-loss is executed, the generated result information contains four core items: complete data, clear values, and direct readability. Specifically, these include: freezing stable indexes to mark the latest reliable indexes in the cache, including index number, corresponding lip opening degree, and tooth status; safe neighborhood candidates to list the screened compliant candidate indexes, including index number, distance from stable index parameters, and lip shape parameters; removing abnormal indexes to mark the removed distant abnormal indexes, including index number, parameter distance, and anomaly type; and stop-loss judgment criteria to specify the stop-loss trigger conditions, parameter distance thresholds, and screening rules.
[0079] For example, the index-level stop-loss result information for a high-risk frame is as follows: the frozen stable index is the closed lip index (number 12, lip opening degree 0.10, teeth closed); the safe neighborhood candidate is the rounded lip index (number 18, parameter distance 0.18, lip opening degree 0.30); the anomaly removal index is the wide-mouth index (number 25, parameter distance 0.52, long-distance jump); the stop-loss basis is a risk score of 0.78 ≥ 0.70 and a parameter distance threshold of 0.25. The integrated result clearly records the entire stop-loss process data and can be directly used for subsequent parameter mitigation and lip shape generation.
[0080] For parameter-level mitigation, the inter-frame variation amplitude is linearly limited according to the risk intensity, and lip opening clamping and forced matching of tooth state with lip shape are performed to constrain the deformation range of high-risk frames. Parameter-level mitigation is used for mouth shape correction in high-risk frames, controlling the inter-frame variation amplitude according to the risk level, and simultaneously standardizing lip opening and tooth state to avoid mouth deformities and tooth flickering. The mitigation rules are clear and the values are fixed, directly used for parameter calibration of high-risk frames. The mitigation intensity is adjusted linearly with the risk value; the higher the risk, the smaller the allowable change, preventing abrupt changes. The mitigation intensity formula is k=clip( ,0.3,1.0); The maximum allowable variation is Δallow=Δmax×(1 k), where Δmax = 0.20 (a fixed value after normalization). For example, high-risk frames. =0.78, k=0.78, Δallow=0.20×(1 0.78)=0.04, the inter-frame lip shape parameter variation is controlled within 0.04, with only minor adjustments.
[0081] The degree of lip opening is limited to a reasonable range to prevent excessive mouth opening or lip flattening. The clamping formula is as follows: =clip( (0.02, 0.90), lower limit 0.02, upper limit 0.90 (after normalization). For example, the original lip opening degree of a high-risk frame is 0.88, after clamping: =0.85, which falls within the compliant range.
[0082] The tooth position must correspond to the lip opening degree to avoid showing teeth when the mouth is closed or closing teeth when the mouth is open. The matching rule is as follows: lip opening degree < 0.15: teeth are forced to close; lip opening degree > 0.40: teeth are allowed to be exposed; 0.15~0.40: teeth are consistent with the previous frame. For example, if the lip opening degree is 0.85 (> 0.40) after clamping, the tooth position is forced to be exposed to match the lip shape.
[0083] Specifically, high-risk frames =0.78, the original parameters were lip opening degree 0.88 and tooth closure. The mitigation intensity was k=0.78, Δallow=0.04; lip opening degree clamping was 0.85; tooth matching was 0.85>0.40, so it was changed to exposure; the final parameters were lip opening degree 0.85, tooth exposure, no deformity, and no flickering.
[0084] Integrating index control results and parameter constraint information, high-risk frame repair results are generated, including stop-loss indexes, mitigation parameters, and matching rules. After completing index-level stop-loss and parameter-level mitigation, the entire process of index control and parameter constraint information is integrated to generate high-risk frame repair results. The results contain three core elements: complete data, clear numerical values, and direct readable information, used to support stable final lip-sync output.
[0085] Specifically, it includes the following: the stop-loss index includes the frozen stable index number, the corresponding lip opening degree, the tooth status, the safe candidate index number, and the parameter distance; the mitigation parameters include the mitigation intensity, the allowable change between frames, the lip opening degree after clamping, the lip height, and the lip width; the matching rules include the tooth matching basis, the deformation constraint conditions, and the risk judgment criteria.
[0086] For example, the repair result of a high-risk frame is as follows: the stop-loss index is frozen stability index number 12, lip opening degree 0.10, and teeth closed; the safety candidate index number is 18, and the parameter distance is 0.18; the mitigation parameters include mitigation intensity 0.78, inter-frame allowable variation 0.04, clamped lip opening degree 0.85, lip height 0.42, and lip width 0.65; the matching rules are: teeth matching is based on lip opening degree > 0.40, deformation constraint is inter-frame variation ≤ 0.04, and risk judgment criterion is total risk ≥ 0.70. The integrated result clearly records the entire repair process data and can be directly used for stable digital phreography output.
[0087] S106: Construct an output verification inference model to perform compliance inference on the range of lip shape parameters, tooth state matching, and inter-frame change amplitude, and trigger multi-level backoff correction when anomalies occur; through risk perception model training, hierarchical decision reasoning learning, retrieval constraint optimization, cache-assisted decision-making, and parameter adaptive correction, complete the model robustness optimization for real-time inference scenarios.
[0088] In one implementation, in conjunction with the quality control requirements of digital lip shape output, the output verification inference model is used to link risk perception, hierarchical decision-making, retrieval constraints, cache scheduling, and parameter adaptive full-link optimization logic. A triple verification mechanism is introduced, which includes lip opening compliance threshold, tooth state matching rules, and inter-frame jump limit. A multi-level conservative backoff strategy is also established to form a complete control link of verification identification, anomaly handling, and stable output. This proactively intercepts problems such as lip shape deformity, tooth flickering, and inter-frame abrupt changes, ensuring the final output compliance.
[0089] The acceptable threshold for lip opening / closing is represented by a normalized value of 0-1, with a reasonable range of 0.02 to 0.90. A value below 0.02 is considered lip flattening, and a value above 0.90 is considered excessive mouth opening, both of which are abnormal. For example, if a frame shows a lip opening / closing value of 0.92, exceeding the upper limit, it is considered an abnormal lip shape.
[0090] The tooth status matching rules are as follows: when the lip opening degree is <0.15, the teeth must be closed; when the lip opening degree is >0.40, the teeth are allowed to be exposed; in the range of 0.15~0.40, the teeth must be consistent with the previous frame. Exposing teeth with the mouth closed or closing teeth with the mouth open are both abnormal. For example, if the lip opening degree in a certain frame is 0.10 and the teeth are exposed, it is judged as an abnormal tooth flickering.
[0091] The inter-frame jump limits are as follows: the L2 normalized distance of the inter-frame aperture parameter ≤ 0.25. If it exceeds this limit, it is judged as an abnormal inter-frame jump. For example, if the lip opening degree of the previous frame is 0.20 and that of the current frame is 0.50, the parameter distance 0.30 > 0.25 is judged as an inter-frame jump.
[0092] If the verification fails, a three-level fallback is executed sequentially according to priority to ensure no abnormal output. Specifically, the first-level fallback is: firstly, it rolls back to the most recent stable frame in the stable buffer and reuses its complete parameters such as lip opening and closing, and tooth status. The second-level fallback is: if there is no stable frame, it directly reuses the output parameters of the previous frame that have passed verification. The third-level fallback is: if there are consecutive anomalies, it freezes the parameters of the current frame and keeps the state of the previous frame unchanged.
[0093] For example, in a digital human's broadcast frame, the current frame has a lip opening of 0.92 and teeth closed, while the previous frame had a lip opening of 0.20. The triple check shows that the lip opening exceeds the upper limit and the inter-frame distance 0.72 > 0.25, so the check fails. A first-level fallback is used, which retrieves the latest stable frame (lip opening of 0.22 and teeth closed) from the cache and reuses it. Finally, the output is a stable frame parameter with no excessive mouth opening or abnormal frame abrupt changes.
[0094] The output lip shape is standardized and verified frame by frame, focusing on three key indicators: whether the lip opening degree falls within the normal range, whether the tooth exposure state and mouth opening amplitude correspond, and whether the lip shape changes excessively between consecutive frames. Abnormal and illegal lip shape frames are accurately identified through item-by-item comparison. The lip opening degree is judged by normalized values, with the normal range set at 0.02 to 0.90. Values below 0.02 are considered lip flattening, and values above 0.90 are considered excessive mouth opening, both of which are abnormal. The tooth state must match the lip opening degree: when the lip opening degree is less than 0.15, the teeth must be closed; when it is greater than 0.40, they are allowed to be exposed. The range of 0.15 to 0.40 must be consistent with the previous frame. Closed mouth with exposed teeth and open mouth with closed teeth are both flickering abnormalities. The lip shape changes between frames are measured by normalized distance, with a maximum allowable value of 0.25; values exceeding this are considered abrupt changes. For example: the measured lip opening angle in a certain frame is 0.92, exceeding the upper limit; the teeth are closed, but the lip opening angle is greater than 0.40, which is a mismatch; and the change value from the previous frame is 0.72, exceeding 0.25. Since none of the three items meet the standard, it is directly judged as an illegal lip shape frame.
[0095] Three standardized checks are performed frame-by-frame to verify whether the lip opening degree is within a preset reasonable range, whether the tooth exposure state matches the mouth opening amplitude, and whether the mouth shape changes in adjacent frames exceed the limit, accurately identifying illegal and abnormal mouth shape frames. Specifically, each frame output undergoes three fixed checks, verified item by item according to a unified standard: first, check if the lip opening degree is within the preset normal range; second, check if the tooth exposure corresponds to the mouth opening amplitude; and finally, check if the mouth shape changes between consecutive frames exceed the range. Abnormal frames are accurately identified through a standardized process. The preset reasonable range for lip opening degree is 0.02 to 0.90. Below the lower limit indicates lip flattening, and above the upper limit indicates excessive mouth opening, both directly judged as abnormal. Tooth matching follows a fixed rule: lip opening degree < 0.15 must have teeth closed, > 0.40 can have teeth exposed, the middle range remains consistent with the previous frame, and mismatch indicates tooth flickering. The inter-frame change limit is fixed at 0.25; exceeding this limit indicates an abnormal inter-frame jump. For example, a frame with a lip opening angle of 0.92 is outside the reasonable range; the mouth is open but the teeth are closed, which is a mismatch; the change between frames is 0.72, exceeding the limit. All three checks fail, and it is judged as an illegal abnormal lip shape frame. For example, a frame with a lip opening angle exceeding the normal upper limit, the teeth being closed while the mouth is open, and the change in lip shape compared to the previous frame is too large, failing all three checks, is judged as an illegal lip shape frame. A frame with a lip opening angle of 0.92 exceeds the normal upper limit of 0.90; the mouth is clearly open, but the teeth are closed; the change in lip shape compared to the previous frame is 0.72, exceeding the limit of 0.25. Comparing to the three standards, the lip opening angle exceeds the upper limit, the teeth are mismatched, and the inter-frame change exceeds the limit; all three requirements are not met, and it is directly judged as an illegal lip shape frame, entering the subsequent fallback repair process.
[0096] When verification fails, a three-tiered fallback operation is executed sequentially: reverting to the most recent stable frame, reusing the previous frame, and freezing the frame to complete frame-level repair. All verification, fallback, and parameter correction data are aggregated and fed back to each module for iterative optimization, leveraging end-to-end iteration to achieve robust optimization of the real-time inference scenario model. When verification results fail, the three-tiered fallback operation is executed sequentially according to priority, with each tier reverting to the fallback to gradually complete frame-level repair, curbing the spread of anomalies and ensuring continuous and stable lip shape output without flickering or distortion. When verification fails, the latest stable frame is retrieved from the stable cache for direct reuse with the highest priority. The stable cache is a fixed-length temporal cache of 20 frames, storing only low-risk, valid stable frames that have passed verification. Its index, lip opening / closing, and tooth status parameters are directly reused to ensure continuous and non-jumping lip shape status. For example, if a frame fails verification (lip opening / closing 0.92, inter-frame variation 0.72), the most recent stable frame in the cache is retrieved: index 12, lip opening / closing 0.22, teeth closed, and directly reused for output.
[0097] When the stable buffer is empty or there are no available stable frames, a second fallback is executed, directly reusing the output parameters that have passed verification in the previous frame to ensure timing continuity and prevent discontinuities. For example, if there are no stable frames in the buffer, and the previous frame had a lip opening / closing of 0.20 and teeth closing, these parameters are directly reused for output.
[0098] If multiple consecutive frames fail the verification and the anomaly persists, a three-level fallback is implemented, freezing the parameters of the current frame to completely maintain the output state of the previous frame, thus curbing the spread of the anomaly and preventing continuous screen flickering. For example, if three consecutive frames fail the verification, the current frame is frozen, and the lip opening / closing value of 0.20 and the teeth closing state of the previous frame remain unchanged.
[0099] Specifically, a certain frame is an illegal lip-sync frame: lip opening / closing is 0.92, teeth closing is 0.72, and the inter-frame variation is 0.72, all three checks fail. The first-level fallback is to retrieve the most recent stable frame from the buffer (lip opening / closing 0.22, teeth closing) and reuse it; if the buffer is empty, the second-level fallback is used, reusing the previous frame (lip opening / closing 0.20, teeth closing); if there are continuous anomalies, the third-level fallback is used, freezing the current frame and maintaining the state of the previous frame; the final output is a stable lip-sync frame, without flickering or distortion.
[0100] Integrating verification results, rollback strategies, and repair information, the system outputs a stable, continuous, and seamless digital mouth shape-driven result without sudden jumps or flickering. The final output contains three core modules: complete data items, clearly defined values, and direct parsing capability. Specifically, these include: Verification Conclusions: annotating the frame-by-frame judgment results for lip opening, tooth matching, and inter-frame jumps; Fallback Execution Strategy: annotating the fallback level used, reused frame information, and frozen parameter status; and Frame-Level Repair Data: annotating the mouth shape parameters, tooth status, and inter-frame changes before and after repair. Specifically, the verification conclusions specify a lip opening compliance range of 0.02~0.90; tooth matching is graded according to mouth opening; the inter-frame jump limit is 0.25; the fallback execution strategy involves level one rollback to the most recent stable frame, level two reuse of the previous frame, and level three freezing of the current frame; and the frame-level repair data records the lip opening, lip height, lip width, and tooth status before and after repair.
[0101] For example, the output of a 1-second, 100-frame digital human broadcast shows the following: Verification results indicate that frame 45 has an excessive lip opening angle of 0.92, a mismatch in tooth closure, and a jump of 0.72 between frames; the fallback strategy is to reuse cached stable frames at level one; frame-level repair data records that the lip opening angle is 0.92 and the tooth closure is 0.22 before repair and the tooth closure is 0 after repair. After integration, the results are cleaned of abnormal frames, and the final output is a digital human mouth shape driven result with no sudden jumps in mouth shape, no flickering teeth, and a continuous and stable temporal sequence. The final output information is categorized and archived, including verification conclusions, fallback types, and repair parameters, and the verification threshold and routing rules are continuously optimized based on iterative data.
[0102] like Figure 2As shown, a robust optimization system for a lip-reading retrieval model based on risk perception modeling includes: The audio feature retrieval module 201 is used to extract the Mel spectrum features of the digital human-driven audio frame by frame, construct a mouth shape codebook containing audio features, mouth shape parameters, view position labels, and tooth status, and perform approximate nearest neighbor retrieval to generate a frame-by-frame candidate mouth shape set. The risk assessment and grading module 202 is used to take retrieval uncertainty, transfer anomaly degree, mouth movement anomaly degree, tooth inconsistency, continuous risk accumulation degree, and boundary matching degree as inputs to the multi-factor risk perception model. After quantile truncation and normalization, the module is weighted and fused to calculate the frame-level risk score and complete the classification of low / medium / high risk levels. The boundary-aware constraint module 203 is used to achieve switching boundary reasoning through phoneme alignment, energy mutation detection, and speech activity detection, and output boundary constraint features. The hierarchical routing decision module 204 is used to perform strategy reasoning by combining risk level and boundary constraint features. For low risk, the optimal retrieval result is directly output. For medium risk, conservative backoff decision or soft switching interpolation reasoning is executed. For high risk, index-level stop loss constraints are executed first and parameter-level mitigation correction is performed. The risk repair execution module 205 is used to standardize retrieval behavior through stable index freezing, safe neighborhood filtering, and long-distance index removal. It adaptively adjusts the parameter change range based on the risk value and completes the consistency constraint of the lip-tooth state. The output verification fallback module 206 is used to build an output verification inference model to perform compliance inference on the range of lip shape parameters, tooth state matching, and inter-frame change amplitude. When an anomaly occurs, it triggers multi-level backoff correction. Through risk perception model training, hierarchical decision reasoning learning, retrieval constraint optimization, cache-assisted decision-making, and parameter adaptive correction, the robustness optimization of the model for real-time inference scenarios is completed.
[0103] A computing device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any robust optimization method for a lip-reading retrieval model based on risk perception modeling.
[0104] The methods and / or embodiments in this application can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processing unit, it performs the functions defined in the methods of this application.
[0105] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0106] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.
Claims
1. A robust optimization method for a lip-reading retrieval model based on risk perception modeling, characterized in that, include: The Mel spectrum features of the digital human-driven audio are extracted frame by frame. A lip shape codebook containing audio features, lip shape parameters, view position labels, and tooth status is constructed. Approximate nearest neighbor retrieval is performed to generate a frame-by-frame candidate lip shape set. Using retrieval uncertainty, transfer anomaly, mouth movement anomaly, tooth inconsistency, continuous risk accumulation, and boundary matching as inputs to the multi-factor risk perception model, the frame-level risk score is calculated after quantile truncation and normalization, and low / medium / high risk level classification is completed. Switching boundary reasoning is achieved through phoneme alignment, energy mutation detection, and speech activity detection, and boundary constraint features are output. Combining risk level and boundary constraint characteristics, the strategy reasoning is executed. For low risk, the optimal retrieval result is directly output. For medium risk, a conservative backoff decision or soft switching interpolation reasoning is executed. For high risk, index-level stop-loss constraints are executed first and parameter-level mitigation correction is performed. The retrieval behavior is standardized by freezing stable indexes, filtering safe neighborhoods, and removing distant indexes. The parameter change range is adaptively adjusted based on the risk value, and the consistency constraint of the lip and tooth state is completed. Construct an output verification inference model to perform compliance inference on the range of lip shape parameters, tooth state matching, and inter-frame change amplitude, and trigger multi-level backoff correction when anomalies occur; By training the risk perception model, learning hierarchical decision reasoning, optimizing retrieval constraints, caching-assisted decision-making, and adaptive parameter correction, the robustness of the model for real-time reasoning scenarios is optimized.
2. The robustness optimization method for the lip-reading retrieval model based on risk perception modeling according to claim 1, characterized in that, Using retrieval uncertainty, transfer anomaly degree, mouth movement anomaly degree, tooth inconsistency, continuous risk accumulation degree, and boundary matching degree as inputs to a multi-factor risk perception model, the model performs weighted fusion calculations after quantile truncation and normalization to determine frame-level risk scores, thus classifying the risk into low / medium / high levels, including: By extracting the Mel spectrum features of the digital human-driven audio frame by frame, a lip shape codebook containing audio features, lip shape parameters, view position labels, and tooth status is constructed, and an approximate nearest neighbor search is performed to generate a frame-by-frame candidate lip shape set. The multi-factor risk perception model is based on six indicators: retrieval confidence, transfer anomaly, mouth movement anomaly, tooth inconsistency, continuous risk accumulation, and boundary matching. All indicators are uniformly mapped to the [0,1] interval using a 1%–99% quantile truncation and Min-Max normalization method. Weights are assigned according to the degree of impact on mouth tremor. Retrieval confidence, transfer anomaly, movement anomaly, and boundary matching are the core high-weight items, with fixed weights set as a1=0.20, a2=0.25, a3=0.20, a4=0.10, a5=0.10, and a6=0.15, where a1 is the retrieval confidence weight; a2 is the transfer anomaly weight; a3 is the mouth movement anomaly weight; a4 is the tooth inconsistency weight; a5 is the continuous risk accumulation weight; and a6 is the boundary matching weight. Based on the 85th percentile of the risk distribution of normal samples, set 95th percentile setting Combined with 3-fold cross-validation calibration threshold, empirical initial value =0.35、 =0.70, the continuous risk accumulation degree is calculated by using a causal sliding window of length W=7 frames to calculate the risk mean of the last W frames, and to distinguish between single frame noise and continuous unstable segments; The total frame-level risk score is obtained by weighted summation of the six indicators. < For low risk, ≤ < For medium risk, ≥ For high-risk cases, the system classifies them into low, medium, and high-risk levels and generates frame-level risk assessment results containing risk sub-items, weight parameters, and classification thresholds.
3. The robustness optimization method for the lip-reading retrieval model based on risk perception modeling according to claim 1, characterized in that, Switching boundary reasoning is achieved through phoneme alignment, energy mutation detection, and speech activity detection, outputting boundary constraint features, including: A multimodal boundary-aware fusion mechanism is constructed based on phoneme-level forced alignment, short-time energy and zero-crossing rate analysis, and VAD speech activity detection. Phoneme alignment results are given the highest priority, while energy / zero-crossing rate and VAD detection are used as fallback when there is no alignment information. Binarized boundary markers are generated frame by frame, and the starting and ending frames of phonemes / syllables, the pause positions between words, the three frames before and after the VAD silence segment, and the energy mutation frames are marked as allowed switching boundaries, while the remaining frames are determined as non-boundary stable constraint regions. Establish linkage rules between boundary markers and lip-sync switching, relax lip-sync switching restrictions at boundary positions, strengthen inter-frame continuity constraints at non-boundary positions, and complete the division of switching permissions and solidification of stable constraints. Consistency checks are performed on boundary marker timing, inter-frame switching rationality, and energy mutation threshold. Phoneme / syllable timestamps are mapped to corresponding video frames, and 1-2 frames are extended before and after the boundary to form a safe switching zone, ensuring that boundary determination is accurately aligned with speech and video timing. By integrating multimodal detection results, boundary marker sequences, and switching constraint rules, syllable / pause boundary perception and switching constraint result information is generated, which includes boundary type, frame-level flag bits, switching permission threshold, and stability constraint parameters.
4. The robustness optimization method for the lip-reading retrieval model based on risk perception modeling according to claim 1, characterized in that, Search behavior is standardized through stable index freezing, safe neighborhood filtering, and long-distance index removal. Parameter variation is adaptively adjusted based on risk values, and consistency constraints on lip-and-tooth states are implemented, including: Combining the requirements for high-risk frame index management and parameter constraints, a dual control mechanism of index-level stop loss and parameter-level mitigation is introduced to determine a standardized handling scheme for freezing stable indexes, screening safe neighborhoods, removing long-distance indexes and amplitude clamping, limiting risk adaptation changes, and forcibly matching teeth and lips. For index-level stop-loss, first freeze the latest reliable index in the stable cache, then filter entries within the stable neighborhood in the candidate set and remove abnormal indexes from far away to prevent cross-view jumps. For parameter-level mitigation, the inter-frame variation amplitude is linearly limited according to the risk intensity, and lip opening clamping and forced matching of tooth state and lip shape are performed to constrain the deformation range of high-risk frames. Integrate index control results with parameter constraint information to generate high-risk frame repair result information containing stop-loss indexes, mitigation parameters, and matching rules.
5. The robustness optimization method for the lip-reading retrieval model based on risk perception modeling according to claim 4, characterized in that, Construct an output verification inference model to perform compliance inference on the range of lip shape parameters, tooth state matching, and inter-frame change amplitude, and trigger multi-level backoff correction when anomalies occur; Through risk perception model training, hierarchical decision-making reasoning learning, retrieval constraint optimization, caching-assisted decision-making, and adaptive parameter correction, the robustness optimization of the model for real-time reasoning scenarios is completed, including: In line with the quality control requirements of digital population output, a triple compliance verification mechanism and a hierarchical backoff strategy are built based on the output verification inference model. This mechanism links risk perception, hierarchical decision-making, retrieval constraints, cache scheduling, and parameter adaptive full-link optimization logic. In line with the quality control requirements of digital population output, a triple verification mechanism is introduced, which includes lip opening compliance threshold, tooth state matching rules, and inter-frame jump limit. A multi-level conservative backoff strategy is also established. Frame by frame, check whether the lip opening is within a reasonable range, whether the tooth state matches the mouth opening, and whether the mouth shape changes between frames exceed the limit, and identify illegal mouth shape frames; If the verification fails, three fallback operations are performed sequentially: rollback to the most recent stable frame, reuse the previous frame, and freeze the frame to complete frame-level repair. By summarizing the verification conclusions, fallback execution records, and parameter correction data, the entire model inference configuration is optimized in a coordinated manner, ultimately outputting lip-sync data that is smooth, flicker-free, and sequentially coherent, thereby improving the overall robustness of real-time inference scenarios.
6. A robust optimization system for a lip-reading retrieval model based on risk perception modeling, characterized in that, The system is used to execute executable instructions to perform the robust optimization method for the lip-reading retrieval model based on risk perception modeling as described in any one of claims 1 to 5.
7. An electronic device, characterized in that, include: First processor; And a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the robust optimization method for lip-reading retrieval model based on risk perception modeling as described in any one of claims 1 to 5 by executing the executable instructions.
8. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the robust optimization method for lip-reading retrieval model based on risk perception modeling as described in any one of claims 1 to 5.