Key person identification passing and early warning system based on multi-modal biological characteristics

By using a multimodal biometric identification system, combined with time event sequence analysis of audio and video, the problem of mis-release or mis-blocking in single-modal identification has been solved, achieving efficient and reliable access management for key personnel and traceable decision-making.

CN121482912AInactive Publication Date: 2026-02-06TIANJIN XINGHE TIANJI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511643508.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, single-modal face or voice solutions are easily affected by occlusion, playback, asynchrony, and optical forgery in access control of rail transit, park access control, data centers, and confidential places, leading to mis-access or mis-blocking, increasing the cost of manual verification and prolonging queuing time, making it difficult to meet the requirements of high passage efficiency and high security.

Method used

A multimodal biometric identification system is adopted, which connects audio and video spatiotemporal through worldline objects. Copula joint modeling and topological data analysis are used to form joint evidence. Combined with audio-visual mutual prediction rebuttal, visual liveness rebuttal and audio playback rebuttal, an aligned time event sequence is generated, and the order probability ratio test and conformal prediction are performed to output the pass, attention or interception command.

Benefits of technology

It enables continuous tracking of multi-source data under a unified coordinate and time index, reduces the impact of misbinding and short-term noise, improves the accuracy and traceability of identification, and meets the comprehensive requirements of key personnel management for real-time performance, reliability and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482912A_ABST
    Figure CN121482912A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent security and protection, in particular to a key personnel identification passing and early warning system based on multi-modal biological characteristics, and the system is used for executing the flow: collecting and binding to generate a world line object; face features, speaker features and gait features are extracted through alignment modeling, an alignment time event sequence is obtained through optimal transmission, and combined evidence is formed through Copula combined modeling and topological data analysis; according to the consistent anti-verification, geometric and sound source beam consistency, audio-visual mutual prediction anti-verification, visual living body anti-verification and audio playback anti-verification are generated in alignment time, and evidences are accumulated into a logarithmic posteriori trajectory of world line identity; decision control implements confidence control through sequence probability ratio test and conformal prediction, outputs passing, attention or interception instructions, and gives consideration to real-time performance, accuracy and evidence obtaining traceability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent security technology, and in particular to a key personnel identification, access control, and early warning system based on multimodal biometrics. Background Technology

[0002] In rail transit, park access control, data centers, and confidential locations, access control faces the dual challenges of high efficiency and high security. Single-modal facial or voice recognition solutions are susceptible to occlusion, playback issues, asynchrony, and optical forgery, leading to false entry or exit, increasing manual verification costs, and extending queuing times. The industry urgently needs to integrate auditory and visual evidence along a unified timeline and enable online verification of forged inputs, forming an explainable and auditable decision-making process to meet the dual goals of security compliance and continuous operation. Summary of the Invention

[0003] To address the numerous problems existing in the prior art, this invention provides a key personnel identification, access control, and early warning system based on multimodal biometrics. This invention uses worldline objects to connect audio and video spatiotemporal data. First, it constructs an aligned time event sequence through optimal transmission. Then, it uses Copula joint modeling and topological data analysis to form joint evidence. Finally, it verifies the evidence through audiovisual mutual prediction, visual liveness verification, and audio playback verification. Combining sequential probability ratio testing and conformal prediction, it outputs access, attention, or interception, thereby improving accuracy and traceability.

[0004] This specification provides one or more embodiments of a key personnel identification, access control, and early warning system based on multimodal biometrics. The system includes:

[0005] The acquisition and binding module is used to acquire synchronous audio and video data, establish the spatiotemporal trajectory of the person and bind it to the speaker's trajectory, and generate a world line object containing face sequence, skeleton sequence, speech segment and sound source arrival direction sequence;

[0006] The alignment modeling module is used to extract facial features, speaker features, and gait features from worldline objects, calculate the log-likelihood ratio and perform quality gating, use optimal transmission to perform cross-modal alignment to generate alignment time event sequences, and form joint evidence based on Copula joint modeling and topological data analysis.

[0007] The consistent disconfirmation module is used to calculate the geometric and sound source beam consistency on the aligned time event sequence, perform audiovisual mutual prediction disconfirmation, visual liveness disconfirmation, and audio playback disconfirmation, and accumulate joint evidence and disconfirmation into the logarithmic posterior trajectory of worldline identity.

[0008] The decision control module is used to calculate the posterior probability of the log-posterior trajectory of the worldline identity output by the consistent proof of contradiction module, determine the judgment threshold based on the cost ratio, perform the order probability ratio test on the log-posterior trajectory of the worldline identity, and combine the conformal prediction method to implement confidence control on the inconsistency score composed of the negative logarithm of the posterior probability and the log-likelihood ratio of the proof of contradiction, and output the pass instruction, attention instruction or interception instruction.

[0009] According to one or more embodiments of the system described in this specification, the acquisition and binding module establishes the spatiotemporal trajectory of a person through target detection and multi-target tracking, obtains speech segments through speech activity detection and speaker segmentation, obtains the sound source arrival direction sequence through arrival direction estimation, and binds the spatiotemporal trajectory of the person with the speaker trajectory based on the head posture direction and geometric constraints, as well as the temporal correlation between lip movement and speech energy envelope, to generate a worldline object.

[0010] According to one or more embodiments of the system described in this specification, the alignment modeling module includes a log-likelihood ratio calibration unit and a quality gating unit. The log-likelihood ratio calibration unit is used to convert the similarity of face features, speaker features and gait features into a log-likelihood ratio. The quality gating unit is used to adjust the log-likelihood ratio by weighting based on sharpness, occlusion ratio, pose angle, signal-to-noise ratio, reverberation features and skeleton confidence.

[0011] According to one or more embodiments of the system described in this specification, the alignment modeling module uses optimal transmission to match lip movement segments, speech sub-band segments, and gait skeleton segments, generates alignment pairing weights, and forms an alignment time event sequence based on the convergence time index of the pairing weights.

[0012] According to one or more embodiments of the system described in this specification, the alignment modeling module performs Copula joint modeling after standardizing the edge log-likelihood ratio to a unified probability domain to generate relevant joint evidence, and performs topological data analysis on the aligned time event sequence to extract time shape features and form shape joint evidence.

[0013] According to one or more embodiments of the system described in this specification, the alignment modeling module updates the cost matrix of the optimal transmission based on relevant joint evidence and shape joint evidence, and performs the optimal transmission again to generate an updated alignment time event sequence.

[0014] According to one or more embodiments of the system described in this specification, the consistency rebuttal module generates geometric and sound source beam consistency evidence based on the angular consistency between the head posture direction and the sound source arrival direction and the temporal correlation between the lip opening sequence and the speech energy envelope, and converts the consistency evidence into a log-likelihood ratio.

[0015] According to one or more embodiments of the system described in this specification, the consistent rebuttal module includes an audiovisual mutual prediction unit, which predicts the speech energy envelope based on the lip movement sequence, predicts the mandibular and laryngeal displacements based on the speech fundamental frequency and speech subband energy, and generates the log-likelihood ratio of the mutual prediction rebuttal based on the prediction correlation and residual signal.

[0016] According to one or more embodiments of the system described in this specification, the consistent rebuttal module includes a visual liveness unit and an audio playback unit. The visual liveness unit generates the log-likelihood ratio of visual liveness evidence based on photoplethysmography and the consistency of three-dimensional facial deformation and illumination. The audio playback unit generates the log-likelihood ratio of audio playback rebuttal evidence based on spectral periodicity characteristics, resonance characteristics, and early reflection residual signals.

[0017] According to one or more embodiments of the system described in this specification, the decision control module sequentially determines the log-posterior trajectory of the world line identity based on the sequential probability ratio test, and performs group calibration on the inconsistency score composed of the negative logarithm of the posterior probability and the log-likelihood ratio of the proof by contradiction based on the conformity-preserving prediction method, and outputs a passage instruction, a attention instruction or an interception instruction according to the quantile threshold.

[0018] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0019] By using spatiotemporal binding and cross-modal association of worldline objects, continuous tracking of multi-source data under a unified coordinate and time index is achieved.

[0020] Precise registration of asynchronous audio / video with gait rhythms was achieved through optimal transmission cross-modal alignment and alignment time event sequence generation.

[0021] By constructing evidence through Copula joint modeling and topological data analysis, an integrated measurement of relevant dependencies and temporal shape was achieved, reducing the impact of misbinding and short-term noise.

[0022] By using audiovisual mutual prediction and rebuttal, the consistency verification of the physical coupling between lip movements and speech is achieved, and the matching of lip movements and dubbing scenes is identified.

[0023] By using visual liveness verification and audio playback verification, online identification of screen / photo forgery and speaker playback was achieved.

[0024] By combining sequential probability ratio testing and conformal prediction in a joint decision-making process, sequential optimal determination and confidence control are achieved under a given cost ratio.

[0025] By recording evidence and thresholds throughout the entire process, the traceability of decision-making results and the auditability of operations and maintenance are achieved. Attached Figure Description

[0026] Figure 1It is a structural block diagram of the system of the present invention;

[0027] Figure 2 It is a schematic diagram of the execution process of the system of the present invention. Specific embodiments

[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0031] In the scenarios of access control and on-site operation, "key personnel" refers to the set of personnel who need to be subject to differential management, usually including authorized registered personnel, restricted or blacklisted personnel, visiting personnel who need to be focused on, and personnel dynamically marked due to tasks, time periods, or regions. This set has timeliness and contextuality: the access strategies and handling actions of the same person are different under different regions, time periods, and task roles; at the same time, it is required that the records are traceable, the handling is explainable, and the linkage is executable. In actual operation, occlusion and backlight, voice overlap and reverberation, screen or speaker playback, and clock errors across cameras and microphones make it difficult for single-modal recognition to steadily meet the dual goals of low false release and low false interception, and it is also difficult to provide a continuous evidence chain for security audits.

[0032] Based on the aforementioned needs and challenges, this invention proposes a "Key Personnel Identification, Access Control, and Early Warning System Based on Multimodal Biometrics": It uses "worldline objects" to connect the spatiotemporal information of audio and video. In alignment modeling, optimal transmission is used to align faces, speakers, and gaits into a unified sequence of aligned temporal events. This is combined with Copula joint modeling and topological data analysis to form cross-modal joint evidence. In the consistency-based rebuttal stage, audiovisual mutual prediction, visual liveness detection, and audio playback rebuttal are introduced to eliminate forgery and inconsistent inputs. Finally, a confidence-based sequential decision is implemented using ordinal probability ratio testing and conformal prediction, outputting access, attention, or interception commands. This approach establishes "who," "when and where," and "how" are verified as a continuous and auditable chain of evidence, thereby meeting the comprehensive requirements of key personnel management for real-time performance, reliability, and compliance.

[0033] like Figure 1-2 As shown, a key personnel identification, access control, and early warning system based on multimodal biometrics is disclosed. The system includes:

[0034] The acquisition and binding module is used to acquire synchronous audio and video data, establish the spatiotemporal trajectory of the person and bind it to the speaker's trajectory, and generate a world line object containing face sequence, skeleton sequence, speech segment and sound source arrival direction sequence;

[0035] The acquisition and binding module establishes the spatiotemporal trajectory of the person through target detection and multi-target tracking, obtains speech segments through speech activity detection and speaker segmentation, obtains the sound source arrival direction sequence through arrival direction estimation, and binds the spatiotemporal trajectory of the person with the speaker trajectory based on the head posture direction and geometric constraints, as well as the temporal correlation between lip movement and speech energy envelope to generate a worldline object.

[0036] The acquisition and binding module of this invention establishes a one-to-one correspondence between the spatiotemporal trajectory of a person on the video side and the trajectory of a speaker on the audio side under a unified time and space reference, and outputs a world line object for direct use in subsequent alignment modeling and decision control. The system first completes device time synchronization and extrinsic parameter calibration to ensure that the timestamps of the camera and microphone array are consistent and the coordinates are convertible.

[0037] The video processing workflow involves obtaining human and face bounding boxes through object detection; extracting facial and skeletal keypoints to obtain the geometric positions of the upper and lower edges of the lips and the spatial positions of major joints; estimating the head orientation vector based on facial keypoints and camera extrinsic parameters; and using multi-object tracking to correlate consecutive frames in time series to form a spatiotemporal trajectory of the person, storing the face sequence and skeleton sequence for each trajectory. Simultaneously, the time series of lip opening and closing amplitude is calculated as an observation for correlation with the speech energy envelope.

[0038] The audio processing flow involves detecting speech activity on a continuous audio stream, dividing the audio into speech and non-speech segments; segmenting the speech segments by speaker to obtain speaker trajectories; estimating the direction of arrival (DOA) of the sound source based on the microphone array to form a DOA sequence; and extracting time-domain descriptions such as subband energy envelopes and fundamental frequencies from the speech segments, which are used as the basis for temporal correlation with lip movements. All the above results are recorded with a unified timestamp.

[0039] The binding of characters and speakers employs a hierarchical strategy. The first layer is temporal overlap screening, where only the spatiotemporal trajectories of characters and speakers that exist simultaneously within the same time window are included in the candidate set. The second layer is geometric consistency screening, comparing the head orientation vector and the sound source arrival direction in the same coordinate system; when the angle is large, the candidate pair is directly excluded, thus reducing subsequent computation. The third layer is audiovisual coherence verification, calculating the correlation between the lip opening / closing time series and the speech subband energy envelope within a small time offset, and selecting the candidate pair with the highest correlation as the binding result. When multiple candidate pairs have similar scores, the highest score is retained as the primary binding, and the second highest score is recorded as a backup candidate, which is then further verified and eliminated by the subsequent alignment modeling module and consistency verification module on a unified time event sequence. To prevent short-term misbinding, the system continuously monitors geometric consistency and audiovisual coherence for several frames after binding; if they continuously fall below the threshold, the binding is revoked and re-evaluated.

[0040] The worldline object encapsulation is based on a unified temporal index, writing face sequences, skeleton sequences, speech segments, sound source arrival direction sequences, and associated quality markers into the same data structure. Quality markers include face sharpness, occlusion ratio, pose angle range, speech signal-to-noise ratio, reverberation intensity, and skeleton keypoint confidence, used for subsequent quality gating and evidence weighting. Worldline objects support cross-camera and cross-array stitching. When the same person appears consecutively in adjacent viewpoints, worldlines are reconnected based on temporal continuity, spatial adjacency, and appearance similarity, maintaining a single identifier.

[0041] Example: Two fixed cameras and a microphone array are deployed at the entrance of the passage. The cameras and array are synchronized via a network clock. The server receives video and audio streams, performs real-time detection, key point extraction, and multi-target tracking, generating spatiotemporal trajectories of people, head orientation, and lip opening / closing sequences. In parallel, it performs speech activity detection, speaker segmentation, and direction of arrival estimation, generating speaker trajectories, sound source direction of arrival sequences, and speech sub-band energy envelopes. The system constructs candidate pairs according to time windows, first performing geometric consistency screening, then audiovisual coherence verification, determining the binding relationship, and writing it into the worldline object. This example can stably output continuous worldline objects under continuous pedestrian flow conditions. Each object carries a complete multimodal sequence and quality markers, providing direct input for cross-modal alignment and joint evidence construction in the alignment modeling module, while reducing the probability of subsequent misbinding and playback interference.

[0042] The alignment modeling module is used to extract facial features, speaker features, and gait features from worldline objects, calculate the log-likelihood ratio and perform quality gating, use optimal transmission to perform cross-modal alignment to generate alignment time event sequences, and form joint evidence based on Copula joint modeling and topological data analysis.

[0043] The alignment modeling module performs feature extraction, evidence calibration and quality gating, cross-modal temporal alignment, and joint evidence generation on worldline objects, outputting an evidence sequence that can be superimposed on a unified timeline. Inputs include face sequences, skeleton sequences, speech segments, and sound source arrival direction sequences; outputs are three quality-weighted log-likelihood ratios, aligned temporal event sequences, and joint evidence.

[0044] Facial features generate embedding vectors from aligned facial images; speaker features generate speaker embeddings from speech segments; gait features generate gait vectors from skeletal joint sequences as input. Similarity scores are calculated between each feature and the roster template, and log-likelihood ratio calibration is performed.

[0045] Let be the log-likelihood ratio for a certain mode. The similarity score for this modality. and These are calibration parameters obtained from offline training. To calibrate the slope, To calibrate the intercept, It is one of three modalities: face, speaker, and gait. To suppress the influence of low-quality samples, quality gating is introduced:

[0046]

[0047]

[0048] This is the quality-weighted log-likelihood ratio. For gating weights, For mass vectors, For gating parameters, This is the Sigmoid function. The quality vector includes facial sharpness and occlusion ratio, pose angle range, speech signal-to-noise ratio and reverberation intensity, skeleton keypoint confidence, and gait periodicity integrity.

[0049] The goal of cross-modal alignment is to align lip movement segments, speech energy segments, and gait segments on a unified time axis. Specifically, the three time series are divided into fixed-length segments, and a cross-modal cost matrix is ​​constructed, consisting of time offset, rhythm difference, and phase difference. The pairing weights are solved using entropy-regularized optimal transmission, and then the centroid of the event time is calculated using the pairing weights to obtain the aligned time event sequence.

[0050]

[0051] For event time indexing, For cross-modal pairing weights, The average of the three-modal time indices. As the time center of the lip segment, As the time center of the speech segment, The time center of the gait segment is defined. The solution employs iterative normalization (Sinkhorn) and numerical stabilization of the cost matrix.

[0052] Joint evidence comprises relevant joint evidence and shape joint evidence. Relevant joint evidence characterizes the dependencies among three log-likelihood ratios: within a sliding window, the three quality-weighted log-likelihood ratios are mapped to a unified probability domain; joint modeling is performed based on Copula, and the model family is automatically selected according to the information criterion; the output window-level joint log density serves as relevant joint evidence. Shape joint evidence characterizes the consistency between behavioral rhythms and event structures: time-delay embedding is performed on aligned time event sequences, and stable shape summaries are extracted. These summaries are compared with the shape summaries of the roster template to obtain difference scores, which are then mapped to shape joint evidence. These two types of joint evidence do not replace single-modal evidence but are used as superposition terms for subsequent unified posterior accumulation and anomaly suppression.

[0053] This module writes the three quality-weighted log-likelihood ratios, the alignment time event sequence, and the two types of joint evidence into a buffer for the subsequent consistent rebuttal module to read directly; at the same time, it writes back the low-confidence pairings as cost increments to improve the alignment stability of the next batch of data.

[0054] In the implementation of the passage entrance, the server performs feature generation, log-likelihood ratio calibration, and quality gating in parallel for each worldline object. Subsequently, three types of fragments are generated, and paired weights are calculated to determine the event temporal centroids, forming an aligned temporal event sequence. Finally, Copula joint modeling and shape summary extraction are completed within a two-second sliding window, outputting joint evidence. This implementation maintains a stable event sequence and evidence output during continuous personnel passage, significantly reducing the risk of misjudgment caused by cross-modal asynchrony, short-term occlusion, and reverberation, and providing interpretable input for subsequent decision-making.

[0055] The alignment modeling module includes a log-likelihood ratio calibration unit and a quality gating unit. The log-likelihood ratio calibration unit is used to convert the similarity of face features, speaker features and gait features into a log-likelihood ratio. The quality gating unit is used to adjust the log-likelihood ratio by weight based on sharpness, occlusion ratio, pose angle, signal-to-noise ratio, reverberation features and skeleton confidence.

[0056] The log-likelihood ratio calibration unit and quality gating unit in the alignment modeling module are used to unify the three types of similarity scores into comparable amounts of evidence and automatically reduce the weight when the sample quality is poor, preventing erroneous evidence from entering subsequent fusion. The inputs are facial features, speaker features, and gait features along with their corresponding similarity scores; the output is three quality-weighted log-likelihood ratio sequences. The calibration unit constructs positive and negative samples based on person identity on offline labeled data, and obtains the mapping from similarity to log-likelihood ratio through distribution fitting. In the online stage, this can be accomplished directly through table lookup or linear transformation.

[0057] The core calibration uses a linear mapping: The parameters are obtained by fitting the development set and their monotonicity and threshold stability are checked on an independent validation set. To suppress domain offset, the system saves parameter sets separately for each camera and microphone array, and loads them by device group at startup.

[0058] The quality gating unit generates a quality vector for each time slice and assigns gating weights, which are then multiplied by the log-likelihood ratio to obtain valid evidence. The quality vector consists of facial sharpness, occlusion ratio, pose angle range, speech signal-to-noise ratio, reverberation intensity, skeleton keypoint confidence, and gait periodicity integrity, and is first normalized to the zero-to-one interval. Gating uses a compression mapping: , When the gating weight is below the lower limit, evidence for that modality in that time slice is discarded; when it is in the middle range, it is scaled proportionally; high-quality samples maintain a weight close to one. To avoid fluctuations, the system smooths the gating weights with a short window and limits abnormal jumps.

[0059] In the engineering implementation, the calibration unit provides a unified interface, supporting multiple template libraries and incremental updates; the quality gating unit and feature extraction share a cache within the same process, avoiding repeated reading of image and speech segments. Both units output timestamped sequences, which are directly written to the buffer of the alignment modeling module for use in cross-modal alignment and joint evidence generation. For anomaly handling, when multiple consecutive time slices are judged to be of low quality, a quality event is recorded and the upstream acquisition end is notified to adjust the exposure or pickup gain.

[0060] Example: Two cameras and a microphone array are deployed at the access control channel. The system first uses historical on-site data to complete parameter fitting, generating three sets of calibration parameters and gating parameters for each device group: face, speaker, and gait. During operation, the face similarity, speaker similarity, and gait similarity are calibrated to obtain the log-likelihood ratio. Combined with real-time calculated sharpness, occlusion ratio, signal-to-noise ratio, reverberation intensity, and skeleton confidence, gating weights are generated, outputting three quality-weighted log-likelihood ratio sequences. This sequence serves as the main input in subsequent cross-modal alignment and joint evidence steps, ensuring consistent evidence scale and controlled noise impact, thereby improving the overall stability of recognition and early warning.

[0061] The alignment modeling module uses optimal transmission to match lip movement segments, speech sub-band segments, and gait skeleton segments, generates alignment pairing weights, and forms an alignment time event sequence based on the convergence time index of the pairing weights.

[0062] The alignment modeling module takes the face sequence, skeleton sequence, speech segment, and sound source arrival direction sequence of the worldline object as input, and completes fragmentation, cross-modal matching, and event time generation, outputting three quality-weighted log-likelihood ratio aligned time event sequences on a unified time axis. First, time slices are created for the lip opening / closing signal, speech subband energy, and skeleton joint motion. Lip motion segments, speech subband segments, and gait skeleton segments are generated using fixed lengths and fixed step sizes, and the time center and quality marker are recorded for each segment. Then, a cross-modal cost is constructed, consisting of time offset, rhythm difference, and phase difference: the time offset is obtained from the segment time center difference; the rhythm difference is measured by the difference in the dominant frequency or step frequency within the segment; and the phase difference is measured by the difference in the segment's first peak position. The cost is normalized to zero to one and a quality penalty and a geometric consistency penalty are added. Segments with low quality or inconsistent with the sound source arrival direction have their cost increased.

[0063] To avoid overly complex 3D solutions, pairwise optimal transmission is employed, followed by fusion to obtain trimodal pairing weights. First, the pairing matrix is ​​obtained by solving for the optimal transmission with entropy regularization between lip and speech, and between lip and gait. The objective is to minimize the sum of cost and entropy terms while satisfying segment edge distribution constraints.

[0064]

[0065] This is the pairing matrix for the two modes. For the corresponding cost matrix, The regularization coefficient is... For entropy, and The edge distribution of the two modal segments, Given a vector of all ones, an iterative normalization numerical method is used to solve the problem, and elements smaller than a threshold are set to zero to obtain sparse pairings.

[0066] After obtaining two pairs, using the same lip segment as a bridge, the two pairing matrices are multiplied and normalized to form a three-modal pairing weight; entries with extremely small weights or conflicting relationships are removed. Based on this weight, the three time indices are centroidally converged to generate an aligned time event sequence.

[0067]

[0068] For event time indexing, For cross-modal pairing weights, The average of the three-modal time indices. As the time center of the lip segment, As the time center of the speech segment, This serves as the time center for gait segments. Event-level evidence is obtained by weighted summation of three quality-weighted log-likelihood ratios near that time, and the segment indexes involved in the centroid calculation are retained for subsequent auditing and playback.

[0069] In system implementation, fragment generation and cost construction are performed using a parallel pipeline; the optimal transmission input matrix is ​​stored in a block-sparse structure to limit matching to a reasonable temporal neighborhood, reducing computational load; short-window median filtering is used for event temporal smoothing to suppress jumps; when a consecutive event lacks any modality, the event is marked as low confidence and downweighted in subsequent steps. To improve stability, the module writes back rejected high-cost matches as cost increments, and the matching threshold for similar fragments is automatically raised in the next batch of data.

[0070] Example: On the access control server, the system generates lip, speech, and gait segments with a fixed window and fixed step size, and calculates a cost matrix composed of time offset, dominant frequency difference, and phase difference. Optimal transmission is solved for lip and speech, and lip and gait separately, resulting in two sets of pairing matrices. The two sets of pairings are fused into a three-modal pairing weight using lip segments, and the time index is centroided to obtain an aligned time event sequence. Event-level evidence and segment indexes are simultaneously output to a buffer for direct use in subsequent consistent rebuttal and decision control. This example maintains event time stability even in scenarios with continuous personnel entry and exit and acoustic reverberation, significantly reducing misjudgments caused by cross-modal asynchrony.

[0071] The alignment modeling module performs Copula joint modeling after standardizing the edge log-likelihood ratio to a unified probability domain to generate relevant joint evidence, and performs topological data analysis on the aligned time event sequence to extract time shape features and form shape joint evidence.

[0072] The alignment modeling module standardizes, performs co-modeling, and temporal shape modeling on three types of marginal evidence along a unified timeline, producing "co-related co-evidence" and "shape co-evidence" to enhance multimodal consistency and robustness. The input consists of an aligned temporal event sequence and the corresponding log-likelihood ratios for face, speaker, and gait quality weights; the output is a time series of the two types of co-evidence.

[0073] First, the log-likelihood ratios weighted by the three quality classes are mapped to a unified probability domain to facilitate cross-modal dependency characterization. In the online phase, monotonic standardization is performed using a logistic function. ,in, For the first The standardized probability of a mode. For the first The quality-weighted log-likelihood ratio of the modalities For logic functions, The parameters represent the face, speaker, and gait in that order. The above standardization shares a set of offline calibration parameters within each device group and performs median smoothing within a sliding window to suppress short-term jitter.

[0074] In the relevant joint modeling, the module reads the ternary probability quantities with a fixed-length window, fits the Copula joint model to obtain the joint density, and outputs the joint score:

[0075] For relevant joint evidence, For Copula density function, Face probability measure For speaker probability, This represents the gait probability. The model family is selected from offline data using information criteria and then embedded in the field; online processing only involves parameter lookup and numerical evaluation. When a certain mode is temporarily missing, the module degenerates into a bivariate copula, and the confidence level is marked at the output to avoid ineffective amplification.

[0076] In time-shape modeling, the module constructs time-amplitude trajectories for aligned time event sequences, extracting stable rhythmic and structural features. Specifically, it embeds three types of probability quantities with a time delay within a local window to obtain a shape summary of the event trajectory, which is then compared with the shape summaries of the corresponding individuals in the roster template to generate shape scores.

[0077] For shape-related joint evidence, A distance metric between shape summaries. A shape summary of the current event trajectory. This is a shape summary for the roster template. Smaller distances indicate closer dynamic patterns; to enhance robustness, the module performs a short-window averaging of the shape scores and sets a weighting flag for significant outliers.

[0078] The two types of joint evidence maintain the same time index on the timeline as the log-likelihood ratio after weighting the three quality levels, facilitating subsequent weighted accumulation of the unified posterior. The module also maintains an operation log and a fragment reference index, ensuring that each joint score is traceable to a specific time window and original fragment, meeting audit and review requirements. In abnormal situations (e.g., insufficient valid samples within the window or probability saturation), the module outputs a degraded score with a "for reference only" flag for subsequent decision-making processes according to rules.

[0079] Example: On the access control server, the alignment modeling module reads the alignment time event sequence with a two-second window and a half-second step. First, the log-likelihood ratios of the three quality-weighted classes are standardized and smoothed using a logistic function to obtain a ternary probability flow. Then, a pre-set Copula density is evaluated in each window, and relevant joint evidence is output. In parallel, a time-delay embedding is constructed for the probability flow of the window to generate a shape summary of the event trajectory. The distance between this summary and the shape summary of the corresponding personnel template is calculated to obtain joint shape evidence. Both types of evidence are written to a shared buffer over time for direct reading by the consistency rebuttal module. Field operation shows that this process can stably provide cross-modal consistency measurement in scenarios with short-term occlusion, speech overlap, and incomplete gait, significantly reducing the risk of misjudgment caused by asynchrony and misalignment.

[0080] The alignment modeling module updates the cost matrix of the optimal transmission with weights based on relevant joint evidence and shape joint evidence, and then executes the optimal transmission again to generate an updated alignment time event sequence.

[0081] After obtaining relevant joint evidence and shape joint evidence, the alignment modeling module performs an evidence-driven weighted update on the initial cross-modal cost, and then re-solves for the optimal transport on the updated cost to generate an updated alignment time event sequence. The inputs to this step are the initial cost matrix, the time indices of the three-modal segments, the time series of relevant joint evidence, and the time series of shape joint evidence; the outputs are the updated three-modal pairing weights and the new event time series. The overall process follows the sequence of "evidence sampling—cost correction—secondary matching—event reconstruction".

[0082] First, the joint evidence is mapped to the time center of the segment triples according to the segment time, and then smoothed with a short window to obtain the evidence values ​​that correspond one-to-one with the segment triples. Then, the cost matrix elements are updated according to the following formula:

[0083] For the updated value; For initial generation value; and Weight of evidence; This is related joint evidence at the time center of the fragment triplet; Joint evidence of shape at the time center of the fragment triplet; Regularization weights; The regularization penalty term is derived from the sum of three types of constraints: time span exceeding the limit, rhythm discontinuity, and inconsistency with the direction of arrival of the sound source. This is the average of the time centers of the lip segment, speech segment, and gait segment. For triples that are significantly low quality or marked as conflicting in the previous step, their value is directly increased to the upper limit to achieve hard masking.

[0084] After obtaining the updated cost, the module re-executes the optimal transmission solution within the constrained time neighborhood. To ensure engineering feasibility, a pairwise matching and fusion approach is adopted: first, the optimal transmission with entropy regularization is solved separately for lip and speech, and lip and gait, resulting in two sets of pairing matrices; then, using lip segments as a bridge, the two sets of pairings are fused into a three-modal pairing weight. To improve numerical stability, the cost is standardized before solving, and threshold sparsity is implemented after each iteration to remove unreliable small weight terms. After completing the second matching, the three-modal time indices are centroidally converged according to the new pairing weights to obtain an updated aligned time event sequence. Event-level evidence is extracted from the three quality-weighted log-likelihood ratios according to the new time index and weighted and synthesized, while retaining the segment identifiers involved in the event for subsequent auditing.

[0085] The effects of this step are twofold: first, it transforms "moments of strong evidence" into "lower matching costs," making it easier for optimal transmission to achieve trimodal consistency at critical moments; second, it increases the cost of unreasonable pairings through regularization penalties, reducing cross-speaker mispairing and false synchronization. In practice, the module performs only one cost update and one secondary match, which can significantly stabilize event timing and reduce misjudgment rates in most channel scenarios. If there are significant domain changes on-site, the above process can be repeated in fixed rounds, but the source of the update and the magnitude of the change are recorded each time for backtracking.

[0086] Example: On the access control server, relevant joint evidence and shape joint evidence within the window are projected onto the time center of the segment triple. The system calculates the update cost according to the aforementioned formula and adds a regularization penalty consisting of time span, dominant frequency abrupt change, and inconsistency with the direction of arrival of the sound source. Subsequently, the optimal transmission is solved for lip and speech, and lip and gait respectively within the constrained time neighborhood, and the three-modal pairing weights are obtained through lip segment fusion. Finally, the event time centroid is calculated based on the new pairing weights, forming an updated aligned time event sequence and writing it into the shared buffer. Field tests show that this step can maintain event time stability under conditions of speech reverberation and gait frame loss, further improving the reliability of subsequent consistent counter-evidence and decision control.

[0087] The consistent disconfirmation module is used to calculate the geometric and sound source beam consistency on the aligned time event sequence, perform audiovisual mutual prediction disconfirmation, visual liveness disconfirmation, and audio playback disconfirmation, and accumulate joint evidence and disconfirmation into the logarithmic posterior trajectory of worldline identity.

[0088] The consistent rebuttal module generates identity-independent coherence and rebuttal metrics on aligned temporal event sequences to verify the reliability of preceding joint evidence, and accumulates them chronologically to form the log-posterior trajectory of the worldline identity. Inputs include the aligned temporal event sequence, head pose vectors and sound source arrival direction sequences from the worldline object, lip opening / closing time sequences, speech subband energy envelopes, face region video clips, and three quality-weighted log-likelihood ratios. Outputs are the calibrated and fused log-posterior trajectory of the worldline identity and the evidence source index.

[0089] Geometric and source beam consistency is indexed by event time. Angular consistency is calculated using head pose vectors and the direction of sound source arrival. After short-window smoothing, it is mapped to an evidence score, which is then converted into a log-likelihood ratio through offline calibration for accumulation. Audiovisual mutual prediction for rebuttal is trained within each event window: one predicts the speech energy envelope based on lip opening and closing, and the other predicts jaw and laryngeal micro-movements based on speech features. A rebuttal score is generated when prediction correlation decreases or residuals increase, and after monotonic mapping, it is recorded as a negative log-likelihood ratio. Visual liveness rebuttal extracts remote photoplethysmography signals from multiple facial regions and combines them with facial 3D deformation and illumination consistency checks; a log-likelihood ratio of liveness evidence is generated when cardiac modulation and geometric consistency are satisfied. Audio playback rebuttal analyzes the spectral periodicity, resonance characteristics, and early reflection remnants of speech. If these characteristics are inconsistent with the acoustic properties of natural near-speech, the log-likelihood ratio of the audio playback rebuttal is output. All of the above branches are subject to quality gating and visibility marking; low-quality events are only recorded and not included in the accumulation.

[0090] To achieve unified accumulation with prior evidence, the module adopts the following update formula:

[0091] For a moment The logarithmic posterior trajectory value of the worldline identity; The trajectory value at the previous moment; The log-likelihood ratio is the ratio of geometrical consistency with the sound source beam. To prove the log-likelihood ratio after monotonic mapping for audiovisual mutual prediction (the lower the correlation, the larger the value); For visual evidence of liveness, the log-likelihood ratio is used. The log-likelihood ratio is used to prove the proof by contradiction for audio playback. These are the fusion weights obtained through offline learning. If necessary, adjustments can be made to... Upper and lower limit amplitudes are applied to suppress anomalous jumps. For ease of implementation, the geometric and source beam consistency can be obtained using the following formula:

[0092]

[0093] For a moment The basic score for geometric and acoustic source beam consistency; This is the angle between the head posture vector and the direction of sound source arrival. This score is mapped to the equipment group calibration table as follows: .

[0094] The engineering process is as follows: Read the required multimodal segments according to event time; calculate the scores of the four types of branches and complete the monotonic mapping; perform gating based on the consistency of quality markers and arrival directions; accumulate the logarithmic posterior trajectory of the worldline identity in an update-order manner; simultaneously, write the source segment index, branch score, and gating results of each event into the evidence index for auditing and review. In abnormal situations (such as the absence of one modality), the available branches are retained and their corresponding weights are reduced, while the trajectory continues to be updated continuously.

[0095] Example: In an access control channel, the server processes aligned temporal event sequences using a fixed window. Each event first generates a consistent log-likelihood ratio based on head posture and sound source arrival direction; then, two lightweight temporal models—lip movement predicting speech and speech predicting jaw movement—are run to generate mutual predictive evidence; in parallel, remote optical volumetric mapping and 3D deformation consistency are extracted to generate liveness evidence; periodic stripes and early reflections are detected on the audio side for playback evidence reversal. Each branch, after gating, writes the cumulatively updated log-posterior trajectory of the worldline identity and records the evidence index and gating state in a cache. Field results show that this module can effectively suppress misjudgments and maintain a stable and traceable trajectory even in situations with overlapping speech, screen playback, or mask occlusion.

[0096] The consistency-contrast module generates geometric and source beam consistency evidence based on the angular consistency between the head posture direction and the sound source arrival direction, as well as the temporal correlation between the lip opening sequence and the speech energy envelope, and converts this consistency evidence into a log-likelihood ratio.

[0097] The consistency-contrast proof module measures the geometric pointing and audiovisual temporal coherence based on the aligned time event sequence, generating geometric and source beam consistency evidence. This consistency evidence is then converted into a log-likelihood ratio to verify the reliability of the preceding aligned modeling output. Inputs include event time indices, head pose direction, source arrival direction, lip opening / closing time sequence, speech energy envelope, and quality markers such as visibility and signal-to-noise ratio. Outputs are a log-likelihood ratio sequence arranged by event time and corresponding segment indices.

[0098] Head pose orientation was determined by facial key points and camera calibration, while the sound source arrival direction was estimated by the microphone array; both were compared in a unified coordinate system to obtain angular consistency. The lip opening and closing time series was obtained from the vertical distance variation of lip key points, and the speech energy envelope was obtained by smoothing the speech sub-band energy or full-band energy. All signals were aligned according to event time, and low-visibility and low signal-to-noise ratio segments were annotated and weighted.

[0099] The module calculates geometric consistency and audiovisual coherence in each event window, forming a single score and mapping it to evidence. The core calculations are as follows:

[0100] in, For consistency scoring, The angle between the head posture direction and the direction of sound source arrival. and As weight, The Pearson correlation coefficient is used. This is a time series of lip opening and closing. For the speech energy envelope, To allow for minute differences in audio and video frequencies, This is for time offset tolerance. The score is mapped to the log-likelihood ratio via linear calibration:

[0101]

[0102] in, For log-likelihood ratio, and These are the parameters obtained from offline calibration.

[0103] In terms of engineering implementation, the event window adopts a fixed length and a fixed step size; The result is obtained by the dot product of the unit vectors of the head pose direction and the direction of sound source arrival, followed by short-window smoothing. The audiovisual correlation is maximized within a small time offset to absorb acquisition clock jitter. Quality gating is applied before generating scores: when visibility or signal-to-noise ratio is below a threshold, the event is only logged and not included in the cumulative score; in boundary cases such as mask obstruction or the simultaneous presence of multiple sound sources, the weight is reduced while retaining the segment index for subsequent auditing. The log-likelihood ratio sequence is written to a shared buffer in chronological order for use in the subsequent accumulation of the log-posterior trajectory of worldline identities.

[0104] Example: On the access control server, the system reads the head posture direction, sound source arrival direction, lip opening and closing time sequence, and speech energy envelope for each event; calculates the correlation between angular consistency and maximum time shift, obtains a consistency score according to the above formula, and linearly calibrates it as a log-likelihood ratio; performs weight reduction and labeling on low-quality events; finally, it outputs a log-likelihood ratio sequence arranged by time and the corresponding segment index. Field tests show that this method can still stably provide consistency evidence even when the angle between the speaker and the camera changes, there is mild reverberation, and short-term occlusion, significantly reducing misjudgments caused by misbinding and playback interference.

[0105] The consistent rebuttal module includes an audiovisual mutual prediction unit, which predicts the speech energy envelope based on the lip movement sequence and predicts the mandibular and laryngeal displacements based on the speech fundamental frequency and speech subband energy. It generates the log-likelihood ratio of the mutual prediction rebuttal based on the prediction correlation and residual signal.

[0106] The audiovisual cross-prediction unit of the consistency proof module performs consistency verification on the aligned time event sequence using the bidirectional predictability relationship between vision and audio, and outputs the log-likelihood ratio of the cross-prediction proof. The inputs include the event time index, lip movement sequence, speech energy envelope, speech fundamental frequency and speech subband energy, mandibular and laryngeal displacement sequences, and quality markers such as visibility and signal-to-noise ratio; the outputs are the log-likelihood ratio of the cross-prediction proof arranged by event time and the segment index.

[0107] The engineering process includes segment construction, bidirectional prediction, metric generation, and evidence labeling. In the segment construction phase, lip movements and speech signals are extracted near the event time using fixed lengths and step sizes, while simultaneously extracting jaw and laryngeal displacements (estimated from facial landmarks and optical flow in the neck region). Bidirectional prediction employs a lightweight temporal model: one "visual → audio" path uses lip movements as input to predict the speech energy envelope; the other "audio → visual" path uses the speech fundamental frequency and speech subband energy as input to predict jaw and laryngeal displacements. The model is trained offline using paired data recorded live, and only performs inference online, without introducing additional annotation costs.

[0108] The core calculation involves obtaining the correlation and residuals for each event window and summing them into a cross-prediction score:

[0109] in, For mutual prediction scoring; The Pearson correlation coefficient; This is the root mean square error; For speech energy envelope; The speech energy envelope predicted by lip movements; This refers to the displacement of the mandible and larynx; The mandibular and laryngeal displacements are predicted by the fundamental frequency and subband energy of speech. The weights are calibrated offline. To unify them to the evidence domain, linear calibration is used to obtain the log-likelihood ratio of mutual prediction of proof by contradiction:

[0110]

[0111] in, The log-likelihood ratio for cross-predictive proof by contradiction; These are calibration parameters estimated offline. Lower scores or larger residuals indicate better performance. The more negative the value, the stronger the "proof of contradiction".

[0112] To ensure feasibility, the unit performs quality gating before generating metrics: skipping events when visibility or signal-to-noise ratio falls below a threshold; reducing weights and recording markers for boundary scenarios such as mask obstruction or strong reverberation. For numerical stability, amplitude normalization and short-window smoothing are applied to the input sequence; amplitude limiting is applied to the predicted output to prevent abnormal peaks from dominating the residuals; and... Slight time smoothing is applied to reduce jitter. All results, along with the fragment indexes used and quality markers, are written into the evidence index for easy auditing and playback later.

[0113] Example: On the access control server, the system processes events with a one-second window and a half-second step. The "visual → audio" path uses a combination of one-dimensional convolution and gated recurrent units to predict the speech energy envelope; the "audio → visual" path uses a one-dimensional convolutional network to predict jaw and laryngeal displacement. During the online phase, the correlation coefficient and root mean square error of the two paths are calculated for each event to form a mutual prediction score, which is linearly calibrated as the log-likelihood ratio of the mutual prediction proof and written to a shared buffer for subsequent trajectory accumulation. Field results show that when screen playback, dubbing, or lip-syncing occurs, the mutual prediction correlation significantly decreases and the residuals increase. It shows a clear negative trend, effectively improving the ability to identify abnormal audiovisual combinations.

[0114] The consistent rebuttal module includes a visual liveness unit and an audio playback unit. The visual liveness unit generates the log-likelihood ratio of visual liveness evidence based on photoplethysmography and the consistency of facial 3D deformation and illumination. The audio playback unit generates the log-likelihood ratio of audio playback rebuttal evidence based on spectral periodicity characteristics, resonance characteristics, and early reflection residual signals.

[0115] The consistency-based rebuttal module consists of a visual liveness unit and an audio playback unit. It independently generates two evidence sequences based on the aligned temporal event sequence and writes them to a shared buffer for subsequent accumulation. The visual liveness unit establishes tracking masks in stable skin areas such as the forehead and cheeks, extracts color components according to event windows, and obtains the pulsation components of remote photovolume mapping after de-trending and bandpassing. Simultaneously, it fits a 3D face mesh and head pose based on facial key points, estimates the principal light direction by combining image brightness and facial normals, and calculates the photometric error of illumination consistency. To avoid motion artifacts, motion suppression is performed using the relative optical flow between the face and background, and low-confidence markers are added to strongly saturated, strongly compressed, or occluded segments. The visual liveness unit's core scoring... The evidence is mapped as follows:

[0116]

[0117] in, For long-range optical volumetric mapping, the signal-to-noise ratio is recorded. To account for photometric errors due to inconsistent illumination, For offline weight calibration, The log-likelihood ratio for visual liveness evidence. These are calibration parameters. Higher scores indicate stronger reliability of the liveness detection; low-quality windows are only recorded, not accumulated.

[0118] The audio playback unit calculates the short-time spectrum and direction-of-arrival trajectory on the speech segment, extracting three types of indicators: first, spectral periodicity intensity, used to identify comb-like stripes caused by the speaker; second, resonance anomaly indicators, amplifying unnatural narrowband resonances based on spectral peak bandwidth and stability; and third, early reflection residue, measured by beam pointing to the direct direction of arrival and deconvolutional residual energy to measure room feedback. If the direction of arrival deviates from the head posture and changes over time, the weight of early reflection residue is increased. The core score and evidence mapping of the audio playback unit are as follows:

[0119]

[0120] in, For the periodic intensity of the spectrum, This is an indicator of resonance anomaly. For early reflection of residual energy For offline weight calibration, The log-likelihood ratio is used to prove the proof by contradiction for audio playback. These are calibration parameters. A higher score indicates a stronger suspicion of replay, which is then converted into negative evidence and included in the cumulative calculation.

[0121] Both units first perform quality gating: events with visibility, signal-to-noise ratio, or direction-of-arrival stability below a threshold are downweighted or skipped; the scores and log-likelihood ratios of retained events are smoothed and limited using a short window to avoid misjudgments caused by sudden peaks. The module saves the segment index, score, weight, and gating status of each event, supporting auditing and playback review.

[0122] Example: Two cameras and a microphone array are deployed at the entrance of the passage. The server runs the two units in parallel according to the event window: the visual side outputs the log-likelihood ratio of visual liveness evidence, along with skin masking and photometric errors; the audio side outputs the log-likelihood ratio of audio playback evidence, along with periodic intensity, direction of arrival, and early reflection residue. Experimental results show that it is difficult for the screen and photograph to simultaneously satisfy both pulsation and illumination consistency, while the speaker playback exhibits significant characteristics in periodicity and early reflection. The combined use of the two units effectively reduces misjudgments caused by forged input.

[0123] The decision control module is used to calculate the posterior probability of the log-posterior trajectory of the worldline identity output by the consistent proof of contradiction module, determine the judgment threshold based on the cost ratio, perform the order probability ratio test on the log-posterior trajectory of the worldline identity, and combine the conformal prediction method to implement confidence control on the inconsistency score composed of the negative logarithm of the posterior probability and the log-likelihood ratio of the proof of contradiction, and output the pass instruction, attention instruction or interception instruction.

[0124] The decision control module sequentially determines the log-posterior trajectory of the worldline identity based on the sequential probability ratio test, and performs group calibration on the inconsistency score composed of the negative logarithm of the posterior probability and the log-likelihood ratio of the proof by contradiction based on the conformity-preserving prediction method. Based on the quantile threshold, it outputs the passage instruction, attention instruction or interception instruction.

[0125] The decision control module takes the log-posterior trajectory of the worldline identity output by the consistent disproportionation module as input, performs sequential determination and confidence control on the event time series, and maps the results into passage commands, attention commands, or interception commands. The inputs include the log-posterior trajectory of the worldline identity, the log-likelihood ratio of the disproportionation from the consistent disproportionation module, and the device grouping label; the outputs are three types of discretized control commands and their timestamps.

[0126] First, the logarithmic posterior trajectory of worldline identities is converted into posterior probabilities, which serve as a unified measure for subsequent statistics:

[0127] For a moment The posterior probability, For a moment The logarithmic posterior trajectory of the worldline identity It is a logical function.

[0128] Based on the cost ratio, set upper and lower thresholds for the order probability ratio test, and directly perform sequential determination in the log-posterior domain: when When the data continuously enters the upper threshold range, a "pass" decision is made; when... When the device enters the lower threshold range, an "interception" decision is made; when it is in the middle range, it continues to accumulate until it crosses the boundary or reaches the timeout limit. The threshold is calibrated offline based on the cost of false alarms and the cost of missed alarms and is managed by device group. After being loaded on-site, it does not rely on manual intervention.

[0129] To control uncertainty and suppress anomalous combinations, an inconsistency score composed of posterior probability and proof-of-contradiction index is constructed, and confidence control is performed using a conformal prediction method:

[0130] For a moment Inconsistent scores, It is a negative logarithmic posterior term. For a moment The aggregated log-likelihood ratio of the proof by contradiction (including the weighted sum of the audio-visual cross-prediction proof by contradiction and the audio playback proof by contradiction). Weights are calibrated offline. The module maintains a sliding calibration set within the device group and calculates the quantile threshold. ,when At that time, it was considered sufficiently confident; when The confidence level is lowered and "attention" is triggered, or "interception" is given in conjunction with sequential decision.

[0131] The decision rule is: if the order probability ratio test enters the upper threshold and... Output the passage command; if the order probability ratio is higher than the threshold, proceed to the next threshold. If the threshold for high percentiles is exceeded, an interception command is output; otherwise, a concern command is output and continues to accumulate. Module for... and Short-window smoothing and amplitude limiting are applied to avoid frequent switching caused by single-frame jitter; for missing modalities, the most recent valid value is used and the weight is reduced to ensure the temporal continuity of the trajectory and score. All decisions simultaneously record the evidence source index and threshold version for easy post-audit.

[0132] Example: On the server at the channel entrance, the decision control module reads the log-posterior and log-likelihood ratios with a one-second window and a half-second step size. The system calculates the posterior probability and inconsistency score online, performs ordinal probability ratio testing and conformal confidence control according to preset thresholds; when the passage conditions are met, it sends an opening command to the access controller and records it in the log. , , , This system, along with counter-evidence, triggers an alert and alarm system when a trigger is established, and pushes the worldline object index. Field operation shows that this process maintains stable judgment in complex audio-visual environments, reduces false releases and false blocks, and provides traceable thresholds and chains of evidence.

[0133] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0134] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A key personnel identification, access control, and early warning system based on multimodal biometrics, characterized in that, The system includes: The acquisition and binding module is used to acquire synchronous audio and video data, establish the spatiotemporal trajectory of the person and bind it to the speaker's trajectory, and generate a world line object containing face sequence, skeleton sequence, speech segment and sound source arrival direction sequence; The alignment modeling module is used to extract facial features, speaker features, and gait features from worldline objects, calculate the log-likelihood ratio and perform quality gating, use optimal transmission to perform cross-modal alignment to generate alignment time event sequences, and form joint evidence based on Copula joint modeling and topological data analysis. The consistent disconfirmation module is used to calculate the geometric and sound source beam consistency on the aligned time event sequence, perform audiovisual mutual prediction disconfirmation, visual liveness disconfirmation, and audio playback disconfirmation, and accumulate joint evidence and disconfirmation into the logarithmic posterior trajectory of worldline identity. The decision control module is used to calculate the posterior probability of the log-posterior trajectory of the worldline identity output by the consistent proof of contradiction module, determine the judgment threshold based on the cost ratio, perform the order probability ratio test on the log-posterior trajectory of the worldline identity, and combine the conformal prediction method to implement confidence control on the inconsistency score composed of the negative logarithm of the posterior probability and the log-likelihood ratio of the proof of contradiction, and output the pass instruction, attention instruction or interception instruction.

2. The system according to claim 1, characterized in that, The acquisition and binding module establishes the spatiotemporal trajectory of the person through target detection and multi-target tracking, obtains speech segments through speech activity detection and speaker segmentation, obtains the sound source arrival direction sequence through arrival direction estimation, and binds the spatiotemporal trajectory of the person with the speaker trajectory based on the head posture direction and geometric constraints, as well as the temporal correlation between lip movement and speech energy envelope to generate a worldline object.

3. The system according to claim 1, characterized in that, The alignment modeling module includes a log-likelihood ratio calibration unit and a quality gating unit. The log-likelihood ratio calibration unit is used to convert the similarity of face features, speaker features and gait features into a log-likelihood ratio. The quality gating unit is used to adjust the log-likelihood ratio by weight based on sharpness, occlusion ratio, pose angle, signal-to-noise ratio, reverberation features and skeleton confidence.

4. The system according to claim 1, characterized in that, The alignment modeling module uses optimal transmission to match lip movement segments, speech sub-band segments, and gait skeleton segments, generates alignment pairing weights, and forms an alignment time event sequence based on the convergence time index of the pairing weights.

5. The system according to claim 1, characterized in that, The alignment modeling module performs Copula joint modeling after standardizing the edge log-likelihood ratio to a unified probability domain to generate relevant joint evidence, and performs topological data analysis on the aligned time event sequence to extract time shape features and form shape joint evidence.

6. The system according to claim 5, characterized in that, The alignment modeling module updates the cost matrix of the optimal transmission with weights based on relevant joint evidence and shape joint evidence, and then executes the optimal transmission again to generate an updated alignment time event sequence.

7. The system according to claim 1, characterized in that, The consistency-contrast module generates geometric and source beam consistency evidence based on the angular consistency between the head posture direction and the sound source arrival direction, as well as the temporal correlation between the lip opening sequence and the speech energy envelope, and converts this consistency evidence into a log-likelihood ratio.

8. The system according to claim 1, characterized in that, The consistent rebuttal module includes an audiovisual mutual prediction unit, which predicts the speech energy envelope based on the lip movement sequence and predicts the mandibular and laryngeal displacements based on the speech fundamental frequency and speech subband energy. It generates the log-likelihood ratio of the mutual prediction rebuttal based on the prediction correlation and residual signal.

9. The system according to claim 1, characterized in that, The consistent rebuttal module includes a visual liveness unit and an audio playback unit. The visual liveness unit generates the log-likelihood ratio of visual liveness evidence based on photoplethysmography and the consistency of facial 3D deformation and illumination. The audio playback unit generates the log-likelihood ratio of audio playback rebuttal evidence based on spectral periodicity characteristics, resonance characteristics, and early reflection residual signals.

10. The system according to claim 1, characterized in that, The decision control module sequentially determines the log-posterior trajectory of the worldline identity based on the sequential probability ratio test, and performs group calibration on the inconsistency score composed of the negative logarithm of the posterior probability and the log-likelihood ratio of the proof by contradiction based on the conformity-preserving prediction method. Based on the quantile threshold, it outputs the passage instruction, attention instruction or interception instruction.