Audio-lip synchronization auditing method and device in short phrase scene, equipment and medium

CN122531418APending Publication Date: 2026-08-07PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-06-24
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]本发明提供一种短话术场景下音唇同步审核方法、装置、设备及介质,以解决现有技术中短话术音唇同步审核存在的非语音干扰表示难以排除、音唇对齐表示鲁棒性不足以及匹配方式无法动态适配导致的审核准确性低、误报率高的问题

Benefits of technology

[0008]上述短话术场景下音唇同步审核方法、装置、设备及介质所实现的方案,通过引入基于最优传输理论的Sinkhorn算法与内置垃圾桶机制,在生成音唇对齐目标表示的过程中自动识别并剔除静音片段、环境噪声及发音前后无关嘴部运动等非语音干扰表示,使模型在有效信息稀少的短话术场景下能够聚焦于真实的说话音频特征与嘴巴运动特征,显著提升了音唇对齐特征表示的鲁棒性与纯净度;同时,通过将聚类中心作为网络可训练参数的端到端联合学习机制,对音频和视频局部表示进行自适应聚类聚合,将局部特征对齐问题转化为聚类分配问题,既降低了模型对大规模训练数据的依赖,又自然实现了同类口型与发音特征的语义对齐,有效克服了短话术场景下同一口型对应多个发音导致的特征混淆与误报问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531418A_ABST
    Figure CN122531418A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, and discloses a method, device and equipment for audio-lip synchronization auditing in a short speech scene and a medium, the method comprising: obtaining audio and video data to be audited; extracting a plurality of audio local representations from an audio data segment and a plurality of video local representations from a video data segment; generating a target representation for representing an audio-lip alignment state based on the plurality of audio local representations and the plurality of video local representations; matching the target representation with a preset reference representation set to obtain a matching result; and outputting an auditing conclusion on whether the audio and video data to be audited has a third party answering or abnormal customer lip movement according to the matching result. The application can be applied to the fields of financial technology and medical health, and solves the problems of low auditing accuracy and high false positive rate caused by the difficulty in excluding non-speech interference representation, the insufficient robustness of audio-lip alignment representation and the inability of the matching method to dynamically adapt in the short speech audio-lip synchronization auditing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology and is applied in the fields of financial technology and healthcare. In particular, it relates to a method, device, equipment and medium for lip-sync verification in short speech scenarios. Background Technology

[0002] With the increasing prevalence of online video lending in the financial risk control field, the need for automated identification of fraudulent behaviors such as unauthorized responses and lip-syncing has become urgent. However, existing lip-syncing verification methods are mainly based on deep learning models, such as SyncNet and TalkNet. These methods are mostly designed for long speech scenarios and face significant challenges in short speech scenarios (such as customers answering "yes" or "agree") common in loan verification. First, short speech has an extremely short duration, usually less than 1 second, resulting in very little effective lip-sync information in the audio and video, making it difficult for the model to extract stable and discriminative features. Second, when extracting and analyzing segments, existing methods easily mix in the feature sequence with the customer's natural mouth movements before and after pronunciation, environmental noise, and a large number of silent segments, leading to a large amount of non-speech interference participating in subsequent calculations and severely interfering with the accurate generation of lip-syncing states. In addition, the same lip shape may correspond to multiple pronunciations, and the lip shape changes significantly when different people pronounce, making the features extracted by the model insufficiently separable, easily triggering false alarms and resulting in an excessively high system alarm rate. Existing technologies lack effective mechanisms to exclude these non-voice interference representations, making it difficult to generate robust lip-alignment representations in short speech scenarios with scarce information. They also cannot accurately match using dynamically adaptable references, becoming a key bottleneck limiting the accuracy of large-scale automated compliance review. Summary of the Invention

[0003] This invention provides a method, apparatus, device, and medium for lip-phone synchronization verification in short speech scenarios, in order to solve the problems of low verification accuracy and high false alarm rate in the existing technology of lip-phone synchronization verification of short speech, which are difficult to eliminate non-speech interference, lack robustness of lip-phone alignment representation, and cannot dynamically adapt the matching method.

[0004] In a first aspect, the present invention provides a method for lip-phonetic synchronization verification in short speech scenarios, including: Obtain the audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; Multiple audio local representations are extracted from the audio data segment, and multiple video local representations are extracted from the video data segment; Based on the multiple audio local representations and the multiple video local representations, a target representation for characterizing the lip alignment state is generated, wherein, in the process of generating the target representation, non-speech interference representations in the multiple audio local representations and the multiple video local representations are excluded; The target representation is matched with a preset set of reference representations to obtain a matching result; Based on the matching results, output a conclusion on whether the audio and video data to be reviewed indicates that someone else answered on behalf of the user or that the customer's lip movements were abnormal.

[0005] Secondly, the present invention provides a lip-phonetic synchronization verification device for short speech scenarios, comprising: The acquisition module is used to acquire audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; The extraction module is used to extract multiple audio local representations from the audio data segment and multiple video local representations from the video data segment; A generation module is used to generate a target representation for characterizing the lip alignment state based on the plurality of audio local representations and the plurality of video local representations, wherein, in the process of generating the target representation, non-speech interference representations in the plurality of audio local representations and the plurality of video local representations are excluded; The matching module is used to match the target representation with a preset reference representation set to obtain a matching result; The output module is used to output a review conclusion based on the matching results, indicating whether the audio and video data to be reviewed has been answered by someone else or whether the customer's lip movements are abnormal.

[0006] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned lip-sync verification method in the short speech scenario.

[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-mentioned lip-sync verification method in the context of short speech.

[0008] The proposed solution for lip-sync verification in short speech scenarios, including the method, device, equipment, and medium, introduces the Sinkhorn algorithm based on optimal transmission theory and a built-in garbage can mechanism. During the generation of the lip-sync target representation, it automatically identifies and removes non-speech interference representations such as silent segments, environmental noise, and irrelevant mouth movements before and after pronunciation. This allows the model to focus on real speech audio features and mouth movement features in short speech scenarios where effective information is scarce, significantly improving the robustness and purity of the lip-sync feature representation. Simultaneously, by using an end-to-end joint learning mechanism with cluster centers as trainable parameters, it adaptively clusters and aggregates local audio and video representations, transforming the local feature alignment problem into a clustering assignment problem. This reduces the model's dependence on large-scale training data and naturally achieves semantic alignment of similar lip shapes and pronunciation features, effectively overcoming the feature confusion and false alarm problems caused by multiple pronunciations corresponding to the same lip shape in short speech scenarios.

[0009] Furthermore, the core reason why lip-sync verification in related technologies is difficult to flexibly adapt to different business scenarios is that the model's judgment logic is rigid and lacks dynamically adjustable reference criteria. This solution, however, transforms lip-sync verification into a similarity-based feature retrieval and matching task by building normal and abnormal feature retrieval libraries offline. This allows for dynamic updates and expansion of the reference representation set based on actual business needs. This design enables the system to easily adapt to compliance verification requirements under different loan products and different dialogue protocols, significantly improving the method's scenario adaptability and scalability.

[0010] Furthermore, existing technologies that rely on fixed thresholds or single-feature comparisons for anomaly detection suffer from low accuracy and high false alarm rates in short-speech scenarios. This solution, however, calculates the bidirectional similarity between the target representation and both normal and abnormal reference representation sets, and combines this with a pre-defined dual-threshold comparison mechanism to output tiered judgment results. A review mechanism is triggered when automatic judgment fails. This tiered judgment design effectively controls the alarm rate while ensuring accurate identification of high-risk cases and fallback handling of low-confidence cases. Technically, it balances review efficiency and accuracy, making large-scale automated lip-sync fraud screening truly feasible in real-world financial risk control scenarios.

[0011] In summary, this solution can solve the problems of low accuracy and high false alarm rate in the existing technology of short speech lip-phonetic synchronization verification, which are difficult to eliminate non-speech interference, lack of robustness of lip-phonetic alignment representation, and inability to dynamically adapt the matching method. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart illustrating a method for simultaneous lip-phonetic verification in a short speech scenario according to an embodiment of the present invention.

[0014] Figure 2 yes Figure 1 A flowchart of step S130.

[0015] Figure 3 yes Figure 1 Another flowchart of step S130.

[0016] Figure 4 yes Figure 1 A flowchart of step S140.

[0017] Figure 5 yes Figure 1 A flowchart of step S150.

[0018] Figure 6 This is a schematic diagram of a lip-sync verification device for short speech scenarios in one embodiment of the present invention.

[0019] Figure 7 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention.

[0020] Figure 8 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 As shown in the figure, an embodiment of the present invention provides a flowchart of a method for simultaneous lip-sync verification in a short speech scenario, which includes the following steps.

[0023] Step S110: Obtain the audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments.

[0024] Specifically, the audio and video data to be reviewed can be any media data containing synchronized audio and video information. For example, in the fintech field, when users apply for online loans or conduct other financial transactions requiring personal confirmation via remote video, financial institutions' systems require users to record a compliance confirmation video. For instance, when confirming loan amounts or signing electronic contracts, users must clearly state standardized phrases such as "I confirm the application" or "I agree to the terms of the agreement" in front of the camera, while the system simultaneously records a video file containing the user's facial lip movements and audio. This video, as the audio and video data to be reviewed, is transmitted to the backend compliance review system to detect fraudulent activities such as someone else answering on their behalf or the customer lip-syncing. The video data segment is a continuous sequence of image frames containing the user's face and lip area, and the audio data segment is the corresponding speech signal sequence. In the healthcare field, when patients apply for medical insurance reimbursement or conduct remote consultations through telemedicine platforms, they also need to submit a verification video containing synchronized audio and lip information as proof of their informed consent. For example, when a patient undergoes a remote follow-up consultation and prescription refill, the system requires the patient to record a confirmation video. The patient must clearly state statements such as "I confirm this visit" and "I agree to use this prescription" while facing the camera. This video, obtained by the healthcare platform, serves as audio and video data for review. It verifies the consistency between lip movements and spoken content, thus determining if the confirmation was made by the patient themselves and preventing impersonation for fraudulent medical appointments or insurance fraud. The video segment records the patient's facial and lip movement sequences, while the audio segment records their confirmation speech.

[0025] It should be noted that the collection and subsequent processing of the audio and video data to be reviewed must be based on the user's knowledge and explicit authorization, and must strictly comply with applicable data security and privacy protection regulations.

[0026] Step S120: Extract multiple audio local representations from the audio data segment and extract multiple video local representations from the video data segment.

[0027] It should be noted that the audio and video local representations in step S120 are compact vector sequences extracted from the original audio and video signals based on a deep learning feature extraction network. Their extraction logic is closely integrated with the data characteristics of different industries. In the fintech field, customer statements in remote loan review videos are typically short, abrupt sounds such as "yes," "agree," and "confirm." Furthermore, the recording environment is often accompanied by non-speech interference such as office background noise and system prompts. The signal-to-noise ratio of the speech signal is low, and the effective information duration is extremely short. Therefore, the extracted audio local representation must accurately capture the acoustic features of the short statements, and the video local representation must precisely depict the correspondence between lip micro-movement trajectories and pronunciation actions. In the healthcare field, the recording scenarios for remote patient confirmation videos are more complex and diverse. For example, the ward environment may contain background noise such as the sound of medical equipment running or other people talking. Moreover, patients often have difficulty speaking clearly due to their condition, which increases the difficulty of extracting lip-phonetic features. Therefore, the feature extraction network is required to have good noise robustness and tolerance to non-standard pronunciation lip shapes, and to be able to stably extract effective local representations from low-quality audio and video data, laying the foundation for subsequent clustering and compliance review.

[0028] Step S130: Based on the plurality of audio local representations and the plurality of video local representations, a target representation for characterizing the lip alignment state is generated, wherein, in the process of generating the target representation, non-speech interference representations in the plurality of audio local representations and the plurality of video local representations are excluded.

[0029] It should be noted that the goal of step S130 is to fuse local features from audio and video to form a compact and robust target representation for subsequent matching verification. The core challenges of this process lie in two aspects: first, how to effectively semantically align and aggregate local representations from different modalities and time steps; and second, how to accurately identify and remove non-speech interference representations such as silence, background noise, and non-speech mouth movements before or during aggregation. In the fintech field, this directly relates to the ability to extract lip-sync features of a customer's true intent from short responses containing significant environmental noise; in the healthcare field, it determines whether it is possible to stably capture valid biometric features confirmed by the patient, even under conditions of significant differences in patient physical condition and complex recording environments.

[0030] In some embodiments of the present invention, such as Figure 2 As shown, step S130 includes the following steps: Step S1311: Initialize a preset number of cluster centers; Step S1312: Calculate the distance between each audio local representation and each video local representation and each cluster center; Step S1313: Assign each audio local representation and each video local representation to the nearest cluster center based on the calculated distance; Step S1314: All local representations assigned to the same cluster center are fused to obtain the target representation.

[0031] Specifically, in step S1311, the method for initializing the preset number of cluster centers can be any feasible method, such as random initialization, K-Means++ initialization, or initialization based on prior knowledge distribution. The number of cluster centers is set according to the richness of the types of dialogue and the complexity of lip movements in the business scenario. Each cluster center corresponds to a combination of lip movements and pronunciation features with similar semantics. In the fintech field, for standardized dialogues such as "I confirm," "I agree," and "Yes" in loan approval scenarios, a number of cluster centers commensurate with the number of dialogue types can be preset. In the healthcare field, to address the diverse confirmation expressions and lip movements of patients, the number of cluster centers can be appropriately increased to cover the changes in pronunciation and lip movements caused by different ages, accents, and facial structures.

[0032] Specifically, in step S1312, the distance between each local representation and each cluster center can be calculated based on any vector space distance metric, such as Euclidean distance, cosine distance, or Mahalanobis distance. In the fintech field, using cosine distance as a metric can better focus on the consistency of feature directions and effectively suppress amplitude interference introduced by differences in recording volume or facial size. In the healthcare field, using Mahalanobis distance can fully consider the differences in feature distribution among patient groups due to age, disease status, etc., making the distance calculation more realistically reflect semantic similarity.

[0033] Specifically, in step S1313, the process of assigning each local representation to the nearest cluster center based on distance can employ either hard or soft assignment. In the fintech field, for videos with clear dialogue and controllable environments, hard assignment can quickly complete cluster mapping, ensuring the timeliness of large-scale audits. In the healthcare field, for videos with varying recording quality, soft assignment allows locally represented characters with blurred boundaries to have a certain probability of being assigned to multiple nearby cluster centers, enhancing robustness against noise and unclear lip movements.

[0034] Specifically, in step S1314, the fusion of all local representations assigned to the same cluster center can be based on vector mean fusion, attention-weighted fusion, or max-pooling fusion. Taking attention-weighted fusion as an example, an attention weight network is used to calculate the importance score of each local representation under the same cluster center, and then a weighted sum is performed accordingly. In the fintech field, attention-weighted fusion enables the model to automatically focus on local representations highly related to speech pronunciation, suppressing interference from non-speech mouth movements. In the healthcare field, this method can adaptively reduce the weight of low-discriminative representations caused by unclear speech and small lip movements, improving focus on core pronunciation movements.

[0035] In some embodiments of the present invention, after step S1311, the following steps are further included: Step S13111: Obtain the training dataset, which contains multiple labeled normal audio and video samples and abnormal audio and video samples; Step S13112: The cluster centers are used as trainable parameters to construct an end-to-end joint learning model with the feature extraction network used to extract local representations; Step S13113: Simultaneously update the weight parameters of the feature extraction network and the parameters of the cluster centers using the backpropagation algorithm.

[0036] It should be noted that the end-to-end joint learning mechanism described in steps S13111 to S13113 integrates the initialization and optimization process of cluster centers into the entire training process of the lip-sync feature extraction network. This makes the cluster centers no longer a set of fixed preset vectors, but learnable parameters that can adaptively adjust according to the distribution characteristics of the training data. This design enables the feature extraction network and cluster centers to optimize collaboratively, and the local representation vectors extracted by the network and the semantic prototype vectors of the cluster centers gradually achieve the best fit during the training process. In the fintech field, compliance annotation data for loan review videos usually contains a large number of positive and negative samples marked by reviewers, such as others answering on behalf of others and customers lip-syncing. Using this annotation data to perform end-to-end joint training of the network and cluster centers can enable the cluster centers to automatically focus on key feature patterns related to various speech lip movements, while driving the feature extraction network to specifically enhance its ability to characterize the lip-sync features of short speech. In the field of healthcare, the confirmed video annotation data of telemedicine platforms involve patient groups of different ages and disease states, whose mouth movements and pronunciation features have significant individual differences. Through end-to-end joint learning, cluster centers can automatically summarize the joint prototype of mouth shape and pronunciation with generalization ability from diverse patient samples, so that the model can still maintain stable lip alignment feature extraction performance when facing patients with large differences in health status.

[0037] Specifically, in step S1211, the training dataset can be obtained in any feasible way. The training dataset consists of multiple labeled normal audio and video samples and multiple abnormal audio and video samples. Normal audio and video samples refer to compliant samples where the audio content matches the lip movements in the video and there are no instances of proxy answers or lip-syncing issues. Abnormal audio and video samples refer to non-compliant samples where there are instances of someone else answering on behalf of the customer, the customer lip-syncing, or other inconsistencies between the lip movements and speech. The labeling information can come from the judgment results of human reviewers or can be automatically generated according to business rules. In the field of fintech, the construction of the training dataset can rely on the loan review video archives accumulated by financial institutions in the past. The risk control review team labels each historical video with normal or abnormal tags and divides them into training set, validation set, and test set according to a certain ratio. The ratio of normal samples to abnormal samples in the training set can be balanced or maintain the original distribution based on the fraud occurrence rate in actual business. In the healthcare field, training datasets can be obtained from informed consent confirmation video records stored on telemedicine platforms. The videos are labeled by medical compliance personnel or review systems. The labeling information can include category tags such as normal confirmation, suspected proxy answering, and abnormal lip-syncing. To fully consider the diversity of the patient population, the dataset should cover patient confirmation video samples from different age groups, genders, and dialect backgrounds.

[0038] It should be noted that the personal audio and video samples included in the training dataset are used for model training after being anonymized and desensitized, with the explicit authorization and consent of the relevant users and in compliance with applicable laws and regulations on personal information protection, to ensure that users' privacy rights are not infringed.

[0039] Specifically, in step S1212, the cluster centers are used as trainable parameters to construct an end-to-end joint learning model with the feature extraction network. This can be achieved by defining the cluster centers as trainable variables in the deep learning network and incorporating them into the computation graph along with other weight parameters of the feature extraction network. This allows the gradient of the loss function during training to propagate back to the cluster center parameters through the clustering assignment process. This design is similar to the mechanism in the NetVLAD network where K-Means cluster centers are used as network layer parameters in end-to-end training. In the fintech field, when constructing a joint learning model, the feature extraction network used to extract local audio and video representations can be used as the front-end encoder, with the cluster centers as the intermediate clustering aggregation layer, followed by a similarity calculation module. The loss function during training can be designed to minimize the distance between the aggregated representation of normal samples and the representation in the normal reference library, minimize the distance between the aggregated representation of abnormal samples and the representation in the abnormal reference library, and maximize the difference between the aggregated representations of normal and abnormal samples. In the field of healthcare, the design of joint learning models can further enhance robustness to individual differences. By introducing a domain adaptive module into the network structure, the model can automatically eliminate the interference of lip-sync and pronunciation changes caused by individual factors such as patient age and disease type when extracting local audio and video representations. At the same time, the prototype features learned by the cluster centers during training have the ability to generalize across patient groups.

[0040] Specifically, in step S1213, the process of simultaneously updating the weight parameters of the feature extraction network and the parameters of the cluster centers through the backpropagation algorithm can be implemented using any suitable optimization algorithm, such as batch stochastic gradient descent, Adam optimizer, or AdamW optimizer. In each training iteration, a batch of training samples is input into the joint learning model. Forward propagation calculates the target representation of each sample. The loss function value is calculated based on the target representation and the corresponding sample label. Then, the gradient of the loss function with respect to the weight parameters of each layer of the feature extraction network and the cluster center parameters is calculated using the chain rule. Finally, these parameters are updated based on the gradient values. The gradient of the cluster center parameters originates from the distance calculation and allocation operation between the local representation vector and the cluster center during the clustering assignment process. Through the backpropagation algorithm, the cluster center parameters are updated in a direction that makes the clustered representations of similar samples more compact and the clustered representations of dissimilar samples more dispersed. In the fintech field, due to the large scale and number of training data samples, a batch stochastic gradient descent combined with cosine annealing learning rate adjustment strategy can be adopted to obtain stable and convergent model parameters while ensuring training efficiency. The trained feature extraction network and cluster center parameters together constitute a lip-sync feature representation model suitable for short speech loan verification scenarios. In the healthcare field, considering the high cost of patient data annotation and the relatively limited sample size, a pre-training weight initialization strategy can be introduced during joint learning. That is, the feature extraction network is pre-trained on a general audio and video dataset, and then fine-tuned on medical confirmation video data. At the same time, the cluster center parameters are updated with a small learning rate to obtain a stable and highly generalizable lip-sync verification model with limited labeled data.

[0041] Understandably, by deeply integrating the cluster center optimization process into the training process of the feature extraction network through the aforementioned end-to-end joint learning mechanism, compared to the traditional approach of executing feature extraction and clustering operations independently in stages, this mechanism enables the co-evolution of feature representations and cluster prototypes. This allows the local representation vectors learned by the network to naturally distribute in a direction conducive to cluster aggregation, and the cluster centers can accurately capture various lip-sync and pronunciation feature patterns based on the actual data distribution. In the fintech field, this mechanism provides a solid technical foundation for large-scale automated credit compliance review. The jointly trained model can accurately distinguish between normal speech videos and abnormal videos such as those where someone else answers on behalf of the customer or the customer lip-syncs, not only improving review accuracy but also significantly reducing reliance on manually labeled data and accelerating the deployment and iteration efficiency of the model in new business scenarios. In the healthcare field, end-to-end joint learning enables the lip-sync verification model to adapt to the highly heterogeneous patient population in remote medical scenarios. Through adaptive learning of cluster centers, it achieves effective summarization and stable extraction of diverse lip-sync and pronunciation features, effectively ensuring the automation level and credibility of patient identity verification and informed consent review.

[0042] In some embodiments of the present invention, such as Figure 3 As shown, step S130 includes the following steps: Step S1321: Based on optimal transmission theory, construct the transmission cost matrix from each local audio representation and each local video representation to each cluster center; Step S1322: Solve for the optimal transmission plan of the transmission cost matrix to obtain the optimal allocation probability of each local representation to each cluster center; Step S1323: Mark the local representations with an allocation probability lower than a preset threshold as the non-speech interference representations; Step S1324: Before generating the target representation, local representations marked as non-speech interference representations are removed from the dataset to be fused.

[0043] Specifically, in step S1321, the rows of the constructed transport cost matrix correspond to the local representation vectors, and the columns correspond to the cluster center vectors. Each element in the matrix represents the transport cost required to assign a local representation to a cluster center. The cost can be calculated based on metrics such as Euclidean distance and cosine distance; the larger the distance, the higher the cost. In the fintech field, for review videos with relatively controllable environments, a dedicated "trash can" row or column can be added to the cost matrix, and learnable cost parameters can be assigned to it to learn the optimal non-speech interference removal threshold. In the healthcare field, the cost can be dynamically weighted based on the energy value of the audio local representation or the motion amplitude of the video local representation, making low-energy silent segments and subtle non-speech mouth movements more easily identified as high transport costs.

[0044] Specifically, in step S1322, the optimal transmission plan can be solved using the Sinkhorn algorithm. This algorithm transforms the optimal transmission problem into a differentiable approximate solution problem by introducing an entropy regularization term, and approximates the optimal transmission plan matrix by iteratively performing row and column normalization operations. In the fintech field, to meet the real-time review requirements of large batches of videos, fewer iterations and adjustments to the entropy regularization coefficient can be used to balance accuracy and efficiency; in the healthcare field, where higher accuracy is required, more iterations can be used and the trash can capacity can be dynamically adjusted to automatically adapt to videos with different noise levels.

[0045] Specifically, in step S1323, non-speech interference representations are labeled by analyzing the optimal allocation probability matrix obtained in step S1322. If the maximum allocation probability of a local representation vector across all effective cluster centers is still lower than a preset threshold, or its allocation probability in the trash can category is higher than the preset threshold, it is labeled as a non-speech interference representation. The preset threshold can be calibrated based on validation set performance to achieve a balance between recall and effective information retention. In the fintech field, thresholds can be set based on historical data to ensure that silence, noise, and non-speech mouth movements are stably eliminated; in the healthcare field, segmented or more lenient thresholds can be set for elderly patients or patients with mild speech to retain their effective but weak features.

[0046] Specifically, in step S1324, removing local representations marked as non-speech interference from the dataset to be fused is a data cleaning operation performed before the fusion operation in step S1314. This operation is achieved by filtering local representation vector indices, retaining only valid local representations related to speech pronunciation actions. In the fintech field, after cleaning, the dataset to be fused contains only high-quality representations reflecting the lip-sync characteristics of customer speech, significantly reducing the interference of silence and noise on the review judgment. In the healthcare field, after removing interference representations caused by patient breathing sounds, swallowing movements, or ward background noise, the remaining local representations can more accurately represent the true pronunciation and lip movements of the patient's confirmation speech, providing high-quality input for subsequent matching and effectively reducing the risk of normal confirmation being misjudged as abnormal due to environmental noise.

[0047] Understandably, through the aforementioned mechanism, step S130 generates a high-quality target representation through the synergistic effect of feature aggregation and noise removal. The cluster center alignment and fusion mechanism solves the core problem of feature alignment difficulties caused by limited available information and large individual differences in short speech scenarios, transforming local feature alignment into a learnable clustering assignment problem. Meanwhile, the non-speech interference elimination mechanism based on optimal transmission theory effectively solves the problem of traditional methods being sensitive to interference such as silence and noise, achieving automated feature purification by sending interfering features to the "garbage can." The combination of these two mechanisms enables this method to extract robust and discriminative lip-sync feature representations in complex application scenarios such as fintech and healthcare, laying a solid foundation for subsequent high-precision matching and verification.

[0048] In some embodiments of the present invention, the following steps are included after step S131: Step S1311: Initialize the transmission plan matrix and set the regularization coefficient; Step S1312: Iteratively perform row normalization and column normalization operations to update the transmission plan matrix; Step S1313: When the difference between the transmission plan matrices of two adjacent iterations is less than a preset convergence threshold, the iteration is stopped; Step S1314: During the iteration process, a preset trash can category is introduced, and non-voice interference is represented as the allocation entry of the trash can category.

[0049] It should be noted that the Sinkhorn iterative solution mechanism described in steps S1311 to S1314 is a key step in transforming the local representation allocation problem into a differentiable optimization process based on optimal transport theory. Its design fully considers the actual data characteristics of the fintech and healthcare fields. In the fintech field, the short speech features of loan review videos require the solution process to balance efficiency and accuracy. By setting appropriate regularization coefficients and convergence thresholds, it is possible to ensure accurate removal of non-speech interference while meeting the timeliness requirements of large-scale real-time review. In the healthcare field, patient videos have high noise complexity and significant individual differences. By introducing garbage bin categories and an adaptive iteration mechanism, it is possible to effectively adapt to diverse recording environments and patient lip-sync features, ensuring the robustness of the target representation generation.

[0050] Specifically, in step S1311, the method for initializing the transmission plan matrix and setting the regularization coefficient can be any feasible method. The dimension of the transmission plan matrix is ​​the same as that of the transmission cost matrix. Initialization can be performed by setting each element in the matrix to a uniform distribution value, that is, assuming that each local representation is initially assigned to each cluster center and garbage bin category with equal probability, ensuring that subsequent iterations can gradually converge to the optimal allocation scheme from an unbiased state. The regularization coefficient is used to control the weight of the entropy regularization term in the Sinkhorn algorithm, and its value directly affects the smoothness and solution speed of the transmission plan matrix. In the field of fintech, for the high throughput requirements of loan review videos, a lower regularization coefficient can be set to make the transmission plan matrix closer to the accurate optimal transmission solution, thereby improving the recognition accuracy of non-speech interference representations, while appropriately increasing the matrix initialization size to adapt to the needs of batch processing videos of different lengths. In the field of healthcare, considering the noise diversity of patient confirmation videos, a higher regularization coefficient can be set to keep the transmission plan matrix smooth and avoid the model being overly sensitive to interference due to the extremely low allocation probability of individual noisy frames. At the same time, a differentiated initial matrix distribution can be preset according to the characteristics of the patient group, such as giving a more conservative allocation tendency to videos of elderly patients.

[0051] Specifically, in step S1312, the process of iteratively performing row normalization and column normalization operations to update the transmission plan matrix involves alternately scaling the matrix in the row and column directions to ensure the matrix meets row and column constraints, thereby gradually approaching the optimal transmission plan. Each row normalization operation divides the element value of each row of the matrix by the sum of the elements in that row, ensuring that the sum of the probability distributions of each local representation to all cluster centers and garbage bins is 1. Column normalization divides the element value of each column of the matrix by the sum of the elements in that column, ensuring that the total quality of the local representations received by each cluster center and garbage bin meets the preset capacity. In the fintech field, the iterative process can be executed in parallel on GPUs, fully utilizing the parallelism of matrix operations to accelerate computation. A maximum iteration limit can be set; when the limit is reached, iteration stops even if complete convergence has not been achieved, ensuring fast review response times. In the healthcare field, the iterative process can be combined with the signal-to-noise ratio of video segments to dynamically adjust the normalization operation. For example, when low audio energy and a large number of silent segments are detected, more receiving capacity can be automatically allocated to the trash can category during column normalization, so that non-voice interference representations are more effectively directed to the trash can.

[0052] Specifically, in step S1313, iteration stops when the difference between the transport plan matrices of two adjacent iterations is less than a preset convergence threshold. Convergence is determined by calculating the norm distance (e.g., Frobenius norm) between the transport plan matrices obtained from two adjacent iterations. The preset convergence threshold can be set according to the accuracy requirements of the actual application scenario. The smaller the threshold, the more accurate the transport plan, but the required number of iterations may increase. In the fintech field, the optimal convergence threshold can be determined through parameter search during the offline training phase of the model, minimizing the average number of iterations while ensuring that the non-speech detection recall rate meets business targets, thereby achieving efficient processing during online inference. In the healthcare field, a relatively lenient convergence threshold can be set to accelerate the stopping of iterations. This, combined with the limited upper limit of the number of iterations in step S1312, prevents excessive iterations caused by individual complex noise samples. After stopping iterations, the transport plan is fine-tuned through post-processing to ensure that the trash can category fully captures non-speech interference.

[0053] Specifically, in step S1314, a preset trash can category is introduced during the iteration process. Non-speech interference is represented as the allocation entry point for the trash can category. This introduces an additional virtual receiving node, i.e., the trash can, into the optimal transmission problem. Its receiving cost is set to a learnable parameter or a fixed empirical value different from the effective cluster centers. During the Sinkhorn iteration, the columns of the trash can category also participate in the normalization operation and have a specified capacity constraint. This capacity constraint controls the total quality that the trash can can receive from the local representation array. When the allocation probability of a local representation is low across all effective cluster centers and the transmission cost to the trash can is relatively small, after iteration, this local representation will be mostly or completely allocated to the trash can category. Thus, the probability of the corresponding trash can dimension in the allocation probability matrix is ​​higher than a threshold, and this representation is identified as a non-speech interference representation. In the fintech field, the capacity of the trash can can be learned end-to-end along with the cluster center parameters during training, enabling the model to automatically learn the optimal capacity boundary from labeled data to distinguish effective speech features from background silence and noise. In the healthcare field, the cost of receiving trash cans can be designed in multiple levels to address medical device noise and non-verbal patient sounds, further improving the granularity of interference identification. At the same time, the trash can capacity can be dynamically adjusted in conjunction with the video signal-to-noise ratio evaluation module to cope with the differences in noise interference intensity under different recording environments.

[0054] Step S140: Match the target representation with a preset reference representation set to obtain a matching result.

[0055] It should be noted that the matching mechanism in step S140 is a process of comparing the target representation of the audio / video to be reviewed with a pre-constructed reference representation set. Its execution logic is closely integrated with the data characteristics of different industries. In the fintech field, the types of dialogue in loan review videos are relatively standardized. A normal reference representation set can cover the compliant confirmation audio / video features of different customer groups under standard recording conditions, while an abnormal reference representation set can include various known fraud patterns such as impersonation and lip-syncing. Through two-way matching, the similarity between the video to be reviewed and the normal compliant pattern and the abnormal fraud pattern can be assessed simultaneously, providing multi-dimensional quantitative basis for comprehensive judgment. In the healthcare field, the construction of the reference representation set for remote patient confirmation videos needs to fully consider the diversity of the patient group. A normal reference representation set should cover the audio / video sample features of patients confirming themselves in different age groups and disease states, while an abnormal reference representation set can include illegal sample features such as impersonation and lip-syncing. Through the matching mechanism, accurate qualitative analysis and risk rating of the video to be reviewed can be achieved.

[0056] In some embodiments of the present invention, such as Figure 4 As shown, step S140 includes the following steps: Step S141: Obtain the offline constructed normal reference representation set, which contains target representations of multiple normal audio and video samples; Step S142: Obtain the offline constructed anomaly reference representation set, which contains target representations of multiple anomaly audio and video samples; Step S143: Calculate the first average similarity between the target representation of the audio / video to be reviewed and each target representation in the normal reference representation set, and the second average similarity between the target representation and each target representation in the abnormal reference representation set. Step S144: Combine the first average similarity value and the second average similarity value into a matching result vector.

[0057] Specifically, in step S141, the method for obtaining the offline-constructed normal reference representation set can be any feasible method. The normal reference representation set is a target representation set pre-calculated and stored before the system goes live, using audio and video samples that have been confirmed as normal by compliance reviewers. This is achieved through the same local representation extraction, clustering aggregation, and non-voice interference removal processes as in steps S120 and S130. The size and coverage of the normal reference representation set can be set according to the needs of the actual business scenario. In the fintech field, the normal reference representation set can be subdivided according to the type of dialogue, such as constructing normal reference subsets for different dialogue categories, such as confirmation videos containing the customer saying "yes," authorization videos containing the customer saying "I agree," and contract signing videos containing the customer saying "confirm," to adapt to the standardized dialogue differences corresponding to different types of loan products and improve the targeting and accuracy of matching. In the field of healthcare, the construction of normal reference representation sets requires special attention to the heterogeneity of patient groups. The database can be classified according to patient age group, gender, whether or not they wear assisted breathing devices, etc. For example, separate normal confirmation reference subsets for elderly patients and normal confirmation reference subsets for middle-aged patients can be constructed. This can reduce the impact of individual differences on similarity calculation to a certain extent during matching. At the same time, the reference representation set can be incrementally updated from newly approved compliant videos on a regular basis to maintain the timeliness and representativeness of the samples in the database.

[0058] Specifically, in step S142, the method for obtaining the offline-constructed abnormal reference representation set is consistent with the construction process of the normal reference representation set. The difference lies in that its samples come from abnormal audio and video recordings that have been confirmed by reviewers to contain instances of someone else answering on behalf of the user, customers lip-syncing, or other inconsistencies between sound and lip movements. The abnormal reference representation set can also be further refined and categorized according to fraud type. In the fintech field, the abnormal reference representation set can be subdivided into subsets such as subsets of someone else answering on behalf of the user, subsets of customers lip-syncing, and subsets of partial audio replacement, each corresponding to different types of fraud methods. When the similarity between the video to be reviewed and a certain abnormal subset is significantly high, the review system can not only determine that the video is abnormal, but also provide a preliminary inference of the abnormality type, providing reference clues for subsequent manual review or risk tracing. In the healthcare field, the abnormal reference set can include labeled abnormal samples such as confirmations made by family members on behalf of patients, patients mimicking lip movements, and video audio track splicing and tampering. These samples may come from the accumulation of historical fraud cases or test data actively constructed by the platform. Through the continuous accumulation and subdivision of the abnormal reference library, the system can gradually enhance its ability to detect new variants of proxy answers or lip-syncing fraud.

[0059] Specifically, in step S143, the process of calculating the first average similarity between the target representation of the audio / video to be reviewed and each target representation in the normal reference representation set, and the second average similarity between the target representation and each target representation in the abnormal reference representation set, is achieved through vector similarity measurement and statistical summarization. The similarity measurement method can employ any vector similarity calculation method, such as cosine similarity, the reciprocal of Euclidean distance, or the Pearson correlation coefficient. Taking cosine similarity as an example, the cosine value of the angle between the target representation to be reviewed and each normal target representation in the normal reference representation set is calculated to obtain multiple first cosine similarity values. These values ​​are then averaged arithmetically or by weighted average to obtain the first average similarity value; similarly, the second average similarity value is calculated. In the fintech field, given the potentially large size and numerous samples in the normal reference representation set, to improve the computational efficiency of online review, an approximate nearest neighbor retrieval index can be pre-built for the target representation in the normal reference representation set. This index structure can be constructed using a vector similarity retrieval engine or an approximate nearest neighbor retrieval library. During matching, only the most similar normal target representations to the target representation to be reviewed are retrieved for averaging, significantly reducing the time spent on similarity calculation while maintaining matching accuracy. In the healthcare field, a weighted averaging strategy can be used for matching calculations. The weights of each sample in the reference representation set are dynamically adjusted based on the patient's basic attribute information (such as age group and gender), giving reference samples with attributes similar to the patient to be reviewed a higher weight in the similarity mean calculation. This achieves a degree of personalized matching and improves the tolerance of the similarity calculation results to individual patient differences.

[0060] Specifically, in step S144, the combination of the first and second similarity mean values ​​into a matching result vector can be achieved through a simple vector concatenation operation. The matching result vector is a one-dimensional vector containing two elements: the first element is the first similarity mean, representing the overall similarity between the video to be reviewed and the normal compliant mode; the second element is the second similarity mean, representing the overall similarity between the video to be reviewed and the abnormal violation mode. This result vector can serve as direct input for subsequent compliance judgment logic. In the fintech field, the two-dimensional values ​​of the matching result vector can be further fed into a lightweight classification decision network or a pre-defined rule engine. The network or engine then integrates the two-dimensional information to output the final review conclusion. Simultaneously, the matching result vector itself can be recorded in the review log, providing traceable quantitative evidence for business analysis. In the healthcare field, matching result vectors can be visualized on healthcare service platforms, such as displaying the normal and abnormal matching degrees of videos to be reviewed in the form of a two-dimensional bar chart. This helps reviewers quickly understand the risk status of videos to be reviewed. Matching result vectors can also serve as a key quantitative indicator in the patient identity authenticity assessment system. They can be integrated with other identity verification factors (such as facial recognition results and voiceprint comparison results) to provide multi-dimensional data support for comprehensive identity authentication decisions in telemedicine scenarios.

[0061] Understandably, the mechanism described above, which involves bidirectional matching of the target representation with both normal and abnormal reference representation sets to generate a matching result vector, offers a more comprehensive characterization of the relative position of the video under review in both the acceptable and non-compliant dimensions, compared to traditional methods that rely solely on a single threshold or feature comparison for anomaly judgment. This effectively avoids the one-sided decision-making problems caused by judging based solely on unidirectional similarity. In the fintech field, this mechanism provides financial institutions with a more accurate and interpretable quantitative assessment basis for large-scale automated compliance review. The system can not only determine whether a video is abnormal but also provide reviewers with an intuitive reference to the degree of risk through the magnitude and absolute value of the two similarity mean values ​​in the matching result vector. This effectively reduces the false positive rate and ensures the accurate identification of high-risk cases. In the healthcare field, the bidirectional matching mechanism significantly enhances the robustness and applicability of patient identity verification, enabling the review system to take into account individual patient differences and the diversity of fraud patterns. Through the multi-dimensional evaluation index of the matching result vector, it provides reliable and interpretable automated technical support for remote medical informed consent confirmation scenarios, promoting the development of medical information systems towards safer and more intelligent collaborative review.

[0062] Step S150: Based on the matching results, output the audit conclusion regarding whether the audio / video data to be audited has been answered by someone else or whether the customer's lip movements are abnormal.

[0063] It should be noted that the review conclusion output mechanism in step S150 is a hierarchical decision-making process based on dual threshold comparison and multi-condition logical judgment. Its execution logic is closely integrated with the risk tolerance and business requirements of different industries. In the fintech field, the compliance requirements for loan review are extremely high. Accidentally approving videos of proxy answers could lead to credit fund losses and regulatory penalties, while mistakenly rejecting legitimate customers would affect customer experience and business conversion rates. Therefore, the output of review conclusions must achieve a delicate balance between accurate identification and false positive rate control. In the healthcare field, verifying the authenticity of patient identities is directly related to the security of medical insurance funds and the compliance of medical services. Incorrect judgments could lead to successful insurance fraud or damage to patients' legitimate medical rights. Therefore, review conclusions need to have clear confidence levels and a manual fallback mechanism to ensure that high-risk cases are accurately intercepted and low-confidence cases receive timely manual intervention.

[0064] In some embodiments of the present invention, such as Figure 5 As shown, step S150 includes the following steps: Step S151: Compare the first average similarity value with a preset first threshold. Step S152: Compare the second similarity mean with a preset second threshold; Step S153: When the first average similarity value is greater than the first threshold and the first average similarity value is greater than the second average similarity value, output a normal review conclusion; Step S154: When the second similarity mean is greater than the second threshold and the second similarity mean is greater than the first similarity mean, output the audit conclusion that there is someone else answering on behalf of the customer or that the customer's lip-syncing is abnormal. Step S155: If neither of the above two conditions is met, output a conclusion that cannot be automatically determined and needs to be reviewed.

[0065] Specifically, in step S151, the process of comparing the first similarity mean with the preset first threshold is achieved by comparing the value of the first element in the matching result vector generated in step S144 with the preset first threshold. The first similarity mean represents the overall similarity between the audio / video to be reviewed and the normal reference representation set. The larger the value, the closer the lip-sync feature of the video to be reviewed is to the feature distribution of known compliant samples. The preset first threshold can be calibrated based on the performance indicators of the validation set of historical review data, for example, by maximizing the comprehensive score of the pass rate of normal samples and the interception rate of abnormal samples. In the field of fintech, the setting of the first threshold can be differentiated according to different loan product types. For high-risk loan products (such as large loans), a higher first threshold can be set to increase the strictness of the review, allowing only videos with lip-sync features that highly match the normal pattern to pass; for low-risk products (such as small consumer installment loans), the first threshold can be appropriately reduced to balance customer experience and risk control. In the field of healthcare, the first threshold can be set according to the characteristics of the patient group. For example, a relatively low first threshold can be set for elderly patients or patients with oral and maxillofacial diseases to reduce the probability of normal confirmation being mistakenly rejected due to unclear pronunciation or limited lip movement, and to ensure that patients have a fair opportunity to access telemedicine services.

[0066] Specifically, in step S152, the process of comparing the second similarity mean with the preset second threshold is similar to step S151. It is achieved by comparing the value of the second element in the matching result vector with the preset second threshold. The second similarity mean represents the overall similarity between the audio / video to be reviewed and the abnormal reference representation set. A larger value indicates that the lip-syncing features of the video to be reviewed are closer to the feature distribution of known proxy answers or lip-syncing violations. The preset second threshold can be calibrated according to the business requirements for anomaly detection sensitivity. In the fintech field, the calibration of the second threshold must fully consider the recall requirements of fraud detection. Typically, the optimal threshold point at the target false positive rate level is determined through the receiver operating characteristic curve, enabling the system to effectively capture various proxy answering and lip-syncing fraud patterns, while avoiding the situation where a large number of normal videos are marked as abnormal due to an excessively low threshold setting, thus increasing the burden of manual review. In the healthcare field, the setting of the second threshold can be combined with the distribution characteristics of historical fraud cases on the telemedicine platform. For high-incidence periods or departments of fraud, the second threshold can be dynamically lowered to enhance the sensitivity of risk perception, while for low-risk business during normal periods, the second threshold can be raised to reduce the false alarm rate in daily operations, thereby achieving intelligent adaptive risk management.

[0067] Specifically, in step S153, when the average first similarity value is greater than the first threshold and the average first similarity value is greater than the average second similarity value, a normal review conclusion is output. This condition indicates that the video to be reviewed not only has a similarity to the normal compliant mode exceeding the threshold for normal passage, but also that its similarity to the normal mode is significantly higher than its similarity to the abnormal mode. The consistency of the two dimensions both support a normal judgment, and the system can output a normal conclusion with high confidence. In the fintech field, when the system outputs a normal review conclusion, the business application corresponding to the loan review video can be automatically marked as compliant and enter the subsequent credit approval process. At the same time, the specific values ​​of the average first and second similarity values ​​and the fulfillment of the conditions for normal judgment are recorded in the review log, forming a traceable review evidence chain. In the healthcare field, the generation of a normal review conclusion means that the patient's remote confirmation video has passed the audio-visual lip-sync review, and the authenticity of their identity has been verified by the system. The subsequent medical insurance reimbursement process or electronic prescription circulation can continue to proceed according to the compliant path. The system can write the normal review conclusion and its confidence information into the patient's electronic medical record or review record in the form of structured data.

[0068] Specifically, in step S154, when the second similarity mean is greater than the second threshold and greater than the first similarity mean, an audit conclusion is output indicating that there is an issue with someone else answering on behalf of the user or the customer lip-syncing. This condition indicates that the similarity between the video to be audited and the abnormal violation pattern not only exceeds the threshold for anomaly alarms, but also that its similarity to the abnormal pattern is significantly higher than its similarity to the normal pattern. The system has sufficient grounds to determine that the video contains a violation of lip-syncing. In the fintech field, when the system outputs an abnormal audit conclusion, it can automatically trigger a risk handling process, including marking the loan application as high-risk, freezing the approval process, sending an alarm notification to risk control personnel, and pushing the abnormal video and its matching result vector to the manual review queue. Simultaneously, based on the matching situation in the anomaly reference representation centralized sub-database, the anomaly type can be further annotated and inferred, such as a tendency towards someone else answering on behalf of the user or a tendency towards the customer lip-syncing, providing clues for manual review. In the healthcare field, the generation of an anomaly review conclusion signifies a failure to verify the authenticity of the patient's identity. The system can automatically intercept the remote consultation or reimbursement application, issue an alert to the medical compliance department, and retain relevant videos and review records as evidence for subsequent investigations. If the frequency of anomaly conclusions is significantly correlated with a specific patient or medical institution, the system can further generate a risk aggregation report to support the optimization of the platform's risk control strategy.

[0069] Specifically, in step S155, when neither of the above two conditions is met, a review conclusion that cannot be automatically determined is output. This includes, but is not limited to: both the first and second similarity mean values ​​are below their respective thresholds; the first similarity mean is above the first threshold but the second similarity mean is also above the second threshold; both are above their respective thresholds but their relative magnitudes do not meet the aforementioned conditions. These situations fall within the system's low confidence range, indicating that the features of the video to be reviewed lack sufficient matching with both the normal and abnormal databases, or that the matching signals themselves are contradictory. Automatic judgment carries a high risk, requiring a transfer to the manual review process. In the fintech field, the introduction of review conclusions allows the system to relinquish control to human experts when uncertain, avoiding decision-making errors under low confidence conditions and effectively reducing the overall risk of mistakenly rejecting normal cases or releasing abnormal cases. Videos awaiting review can be uniformly aggregated in the manual review workbench and ranked by risk based on the difference or ratio between the first and second similarity mean values, allowing reviewers to prioritize cases with the highest degree of ambiguity and improving overall review efficiency. In the healthcare field, the pending review conclusion ensures the bottom line of security for telemedicine review. When the system cannot give a high confidence level judgment due to factors such as excessive environmental noise, blurry image, or rare lip-reading characteristics of the patient, it will not directly refuse the patient's service, but will transfer to the manual intervention process. Medical compliance personnel will make a comprehensive judgment based on multi-source information such as patient identity certificate, facial recognition results, and voiceprint comparison results. This not only prevents fraud risks, but also protects the legitimate rights and interests of patients to enjoy telemedicine services.

[0070] Understandably, the hierarchical review conclusion output mechanism based on dual threshold comparison and multi-condition logical judgment, compared to the traditional approach of using only a single threshold for binary judgment, constructs a three-state hierarchical decision-making system of compliance, abnormality, and pending review. This system can efficiently and automatically process high-confidence cases while providing manual fallback for low-confidence boundary cases, thus balancing review efficiency and decision accuracy from a technical perspective. In the fintech field, this mechanism provides financial institutions with an intelligent compliance review solution that allows for flexible configuration of risk preferences and quantifiable traceability of review basis. It effectively reduces the risk of underreporting of proxy answers and verbal fraud in large-scale credit reviews, while avoiding unnecessary customer complaints and manual review costs caused by false reports, thus promoting the large-scale application of automated financial risk control review systems. In the healthcare field, the tiered review mechanism fully meets the healthcare industry's multiple demands for safety, reliability, and humanistic care. Through adaptive threshold configuration and a fallback mechanism for pending review, it not only effectively deters medical insurance fraud but also ensures a smooth experience for compliant patients seeking remote medical treatment, providing robust technical support for the construction of the voice-lip synchronization identity authentication system of internet healthcare platforms.

[0071] In summary, the solution implemented in this embodiment of the invention can solve the problems of low accuracy and high false alarm rate in the prior art of short speech lip-phonetic synchronization review, which are caused by the difficulty in eliminating non-speech interference, insufficient robustness of lip-phonetic alignment, and the inability to dynamically adapt the matching method.

[0072] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0073] In one embodiment, a lip-phone synchronization verification device for short speech scenarios is provided, which corresponds one-to-one with the lip-phone synchronization verification method for short speech scenarios described in the above embodiments. For example... Figure 6 As shown, the lip-sync verification device for this short speech scenario includes an acquisition module 610, an extraction module 620, a generation module 630, a matching module 640, and an output module 650. Detailed descriptions of each functional module are as follows: The acquisition module 610 is used to acquire audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; Extraction module 620 is used to extract multiple audio local representations from the audio data segment and multiple video local representations from the video data segment; The generation module 630 is used to generate a target representation for characterizing the lip alignment state based on the plurality of audio local representations and the plurality of video local representations, wherein, in the process of generating the target representation, non-speech interference representations in the plurality of audio local representations and the plurality of video local representations are excluded; The matching module 640 is used to match the target representation with a preset reference representation set to obtain a matching result; The output module 650 is used to output a review conclusion based on the matching result, indicating whether the audio and video data to be reviewed has been answered by someone else or whether the customer's lip movements are abnormal.

[0074] In one embodiment, the generation module 630 is specifically used for: Initialize a preset number of cluster centers; Calculate the distance between each audio local representation and each video local representation and each cluster center; Each audio local representation and each video local representation is assigned to the nearest cluster center based on the calculated distance; The target representation is obtained by fusing all local representations assigned to the same cluster center.

[0075] In one embodiment, the generation module 630 is further configured to: Obtain a training dataset, which contains multiple labeled normal audio and video samples and abnormal audio and video samples; The cluster centers are used as trainable parameters to construct an end-to-end joint learning model with a feature extraction network used to extract local representations; The backpropagation algorithm is used to simultaneously update the weight parameters of the feature extraction network and the parameters of the cluster centers.

[0076] In one embodiment, the generation module 630 is further specifically used for: Based on optimal transmission theory, a transmission cost matrix is ​​constructed from each local audio representation and each local video representation to each cluster center; Solve for the optimal transmission plan of the transmission cost matrix to obtain the optimal allocation probability of each local representation to each cluster center; Local representations with an allocation probability below a preset threshold are marked as non-speech interference representations; Before generating the target representation, local representations marked as non-speech interference representations are removed from the dataset to be fused.

[0077] In one embodiment, the generation module 630 is further configured to: Initialize the transmission plan matrix and set the regularization coefficients; Iteratively perform row normalization and column normalization operations to update the transport plan matrix; The iteration stops when the difference between the transmission plan matrices of two consecutive iterations is less than a preset convergence threshold. During the iteration process, a preset trash can category is introduced, and non-voice interference is represented as the allocation entry for the trash can category.

[0078] In one embodiment, the matching module 640 is specifically used for: Obtain an offline constructed normal reference representation set, which contains target representations of multiple normal audio and video samples; Obtain an offline constructed set of abnormal reference representations, which contains target representations of multiple abnormal audio and video samples; Calculate the first average similarity between the target representation of the audio / video to be reviewed and each target representation in the normal reference representation set, and the second average similarity between the target representation and each target representation in the abnormal reference representation set. The first mean similarity and the second mean similarity are combined to form a matching result vector.

[0079] In one embodiment, the output module 650 is specifically used for: The first average similarity value is compared with a preset first threshold. The second similarity mean is compared with a preset second threshold. When the first average similarity value is greater than the first threshold and the first average similarity value is greater than the second average similarity value, a normal review conclusion is output. When the second similarity mean is greater than the second threshold and the second similarity mean is greater than the first similarity mean, the audit conclusion is output that there is someone else answering on behalf of the customer or the customer's lip-syncing is abnormal. If neither of the above two conditions is met, the output will be a conclusion that cannot be automatically determined and needs to be reviewed.

[0080] This invention provides a solution for a lip-phonetic synchronization verification device in short speech scenarios. By introducing the Sinkhorn algorithm based on optimal transmission theory and a built-in garbage can mechanism, it automatically identifies and removes non-speech interference representations such as silent segments, environmental noise, and irrelevant mouth movements before and after pronunciation during the generation of lip-phonetic alignment target representations. This enables the model to focus on real speech audio features and mouth movement features in short speech scenarios where effective information is scarce, significantly improving the robustness and purity of lip-phonetic alignment feature representations. At the same time, by using an end-to-end joint learning mechanism with cluster centers as network trainable parameters, adaptive clustering and aggregation of local audio and video representations are performed, transforming the local feature alignment problem into a clustering assignment problem. This reduces the model's dependence on large-scale training data and naturally achieves semantic alignment of similar lip shapes and pronunciation features, effectively overcoming the feature confusion and false alarm problems caused by multiple pronunciations corresponding to the same lip shape in short speech scenarios.

[0081] Furthermore, the core reason why lip-sync verification in related technologies is difficult to flexibly adapt to different business scenarios is that the model's judgment logic is rigid and lacks dynamically adjustable reference criteria. This solution, however, transforms lip-sync verification into a similarity-based feature retrieval and matching task by building normal and abnormal feature retrieval libraries offline. This allows for dynamic updates and expansion of the reference representation set based on actual business needs. This design enables the system to easily adapt to compliance verification requirements under different loan products and different dialogue protocols, significantly improving the method's scenario adaptability and scalability.

[0082] Furthermore, existing technologies that rely on fixed thresholds or single-feature comparisons for anomaly detection suffer from low accuracy and high false alarm rates in short-speech scenarios. This solution, however, calculates the bidirectional similarity between the target representation and both normal and abnormal reference representation sets, and combines this with a pre-defined dual-threshold comparison mechanism to output tiered judgment results. A review mechanism is triggered when automatic judgment fails. This tiered judgment design effectively controls the alarm rate while ensuring accurate identification of high-risk cases and fallback handling of low-confidence cases. Technically, it balances review efficiency and accuracy, making large-scale automated lip-sync fraud screening truly feasible in real-world financial risk control scenarios.

[0083] In summary, the above solution can solve the problems of low accuracy and high false alarm rate in the existing technology of short speech lip-phonetic synchronization verification, which are difficult to eliminate non-speech interference, lack of robustness of lip-phonetic alignment representation, and inability to dynamically adapt the matching method.

[0084] Specific limitations regarding the lip-phoneme synchronization verification device for short speech scenarios can be found in the above-mentioned limitations on the lip-phoneme synchronization verification method for short speech scenarios, and will not be repeated here. Each module in the aforementioned lip-phoneme synchronization verification device for short speech scenarios can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0085] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements a recommended method for an optimal strategy, representing the functions or steps on the server side.

[0086] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a recommended method for an optimal strategy, representing client-side functions or steps.

[0087] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; Multiple audio local representations are extracted from the audio data segment, and multiple video local representations are extracted from the video data segment; Based on the multiple audio local representations and the multiple video local representations, a target representation for characterizing the lip alignment state is generated, wherein, in the process of generating the target representation, non-speech interference representations in the multiple audio local representations and the multiple video local representations are excluded; The target representation is matched with a preset set of reference representations to obtain a matching result; Based on the matching results, output a conclusion on whether the audio and video data to be reviewed indicates that someone else answered on behalf of the user or that the customer's lip movements were abnormal.

[0088] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain the audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; Multiple audio local representations are extracted from the audio data segment, and multiple video local representations are extracted from the video data segment; Based on the multiple audio local representations and the multiple video local representations, a target representation for characterizing the lip alignment state is generated, wherein, in the process of generating the target representation, non-speech interference representations in the multiple audio local representations and the multiple video local representations are excluded; The target representation is matched with a preset set of reference representations to obtain a matching result; Based on the matching results, output a conclusion on whether the audio and video data to be reviewed indicates that someone else answered on behalf of the user or that the customer's lip movements were abnormal.

[0089] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0090] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0092] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0093] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for simultaneous lip-sync verification in short speech scenarios, characterized in that, include: Obtain the audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; Multiple audio local representations are extracted from the audio data segment, and multiple video local representations are extracted from the video data segment; Based on the multiple audio local representations and the multiple video local representations, a target representation for characterizing the lip alignment state is generated, wherein, in the process of generating the target representation, non-speech interference representations in the multiple audio local representations and the multiple video local representations are excluded; The target representation is matched with a preset set of reference representations to obtain a matching result; Based on the matching results, output a conclusion on whether the audio and video data to be reviewed indicates that someone else answered on behalf of the user or that the customer's lip movements were abnormal.

2. The method for simultaneous lip-phoneme verification in short speech scenarios according to claim 1, characterized in that, The step of generating a target representation for characterizing lip alignment state based on the multiple audio local representations and the multiple video local representations includes: Initialize a preset number of cluster centers; Calculate the distance between each audio local representation and each video local representation and each cluster center; Each audio local representation and each video local representation is assigned to the nearest cluster center based on the calculated distance; The target representation is obtained by fusing all local representations assigned to the same cluster center.

3. The method for simultaneous lip-phoneme verification in short speech scenarios according to claim 2, characterized in that, After initializing the preset number of cluster centers, the method further includes: Obtain a training dataset, which contains multiple labeled normal audio and video samples and abnormal audio and video samples; The cluster centers are used as trainable parameters to construct an end-to-end joint learning model with a feature extraction network used to extract local representations; The backpropagation algorithm is used to simultaneously update the weight parameters of the feature extraction network and the parameters of the cluster centers.

4. The method for simultaneous lip-phoneme verification in short speech scenarios according to claim 2, characterized in that, The exclusion of non-speech interference representations from the plurality of audio local representations and the plurality of video local representations includes: Based on optimal transmission theory, a transmission cost matrix is ​​constructed from each local audio representation and each local video representation to each cluster center; Solve for the optimal transmission plan of the transmission cost matrix to obtain the optimal allocation probability of each local representation to each cluster center; Local representations with an allocation probability below a preset threshold are marked as non-speech interference representations; Before generating the target representation, local representations marked as non-speech interference representations are removed from the dataset to be fused.

5. The method for simultaneous lip-phoneme verification in short speech scenarios according to claim 4, characterized in that, The optimal transmission plan for solving the transmission cost matrix includes: Initialize the transmission plan matrix and set the regularization coefficients; Iteratively perform row normalization and column normalization operations to update the transport plan matrix; The iteration stops when the difference between the transmission plan matrices of two consecutive iterations is less than a preset convergence threshold. During the iteration process, a preset trash can category is introduced, and non-voice interference is represented as the allocation entry for the trash can category.

6. The method for simultaneous lip-phoneme verification in short speech scenarios according to claim 1, characterized in that, The preset reference representation set includes a normal reference representation set and an abnormal reference representation set; the matching of the target representation with the preset reference representation set to obtain a matching result includes: Obtain an offline constructed normal reference representation set, which contains target representations of multiple normal audio and video samples; Obtain an offline constructed set of abnormal reference representations, which contains target representations of multiple abnormal audio and video samples; Calculate the first average similarity between the target representation of the audio / video to be reviewed and each target representation in the normal reference representation set, and the second average similarity between the target representation and each target representation in the abnormal reference representation set. The first mean similarity and the second mean similarity are combined to form a matching result vector.

7. The method for simultaneous lip-phoneme verification in short speech scenarios according to claim 6, characterized in that, The step of outputting a review conclusion based on the matching results regarding whether the audio / video data to be reviewed contains any instances of someone else answering on behalf of the user or abnormal lip-syncing by the customer includes: The first average similarity value is compared with a preset first threshold. The second similarity mean is compared with a preset second threshold. When the first average similarity value is greater than the first threshold and the first average similarity value is greater than the second average similarity value, a normal review conclusion is output. When the second similarity mean is greater than the second threshold and the second similarity mean is greater than the first similarity mean, the audit conclusion is output that there is someone else answering on behalf of the customer or the customer's lip-syncing is abnormal. If neither of the above two conditions is met, the output will be a conclusion that cannot be automatically determined and needs to be reviewed.

8. A device for simultaneous lip-sync verification in short speech scenarios, characterized in that, include: The acquisition module is used to acquire audio and video data to be reviewed, wherein the audio and video data includes audio data segments and video data segments; The extraction module is used to extract multiple audio local representations from the audio data segment and multiple video local representations from the video data segment; A generation module is used to generate a target representation for characterizing the lip alignment state based on the plurality of audio local representations and the plurality of video local representations, wherein, in the process of generating the target representation, non-speech interference representations in the plurality of audio local representations and the plurality of video local representations are excluded; The matching module is used to match the target representation with a preset reference representation set to obtain a matching result; The output module is used to output a review conclusion based on the matching results, indicating whether the audio and video data to be reviewed has been answered by someone else or whether the customer's lip movements are abnormal.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the lip-sync verification method in the short speech scenario as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the lip-sync verification method in the short speech scenario as described in any one of claims 1 to 7.