Dispensing authority control method and system based on multi-modal biological characteristic comparison

By using a multimodal biometric comparison method, infrared images of the face and visible light images of the iris are simultaneously acquired and analyzed to generate live physiological resonance feature values. This solves the problems of easy leakage of identity verification and insufficient robustness of single-modal recognition in existing technologies, and achieves high-precision identity verification and accurate drug dispensing.

CN121921853AActive Publication Date: 2026-04-24YUEYANG FEIPENG NETWORK TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUEYANG FEIPENG NETWORK TECH CO LTD
Filing Date
2026-03-27
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing medication dispensing equipment's authentication methods are easily leaked and misused. Single-modal biometric recognition is not robust enough in complex environments and lacks effective liveness detection, leading to security vulnerabilities and medication dispensing errors.

Method used

A multimodal biometric comparison method is adopted to simultaneously acquire infrared image sequences of the face and visible light image sequences of the iris. Pupil change curves and micro-expression muscle tremor frequencies are extracted through inter-frame differential motion analysis to generate live physiological resonance feature values. Feature-level heterogeneous fusion is performed through cross-modal joint encoder, and authorization biometric tensor is combined for permission verification.

Benefits of technology

It achieves high-precision and high-stability identity verification, effectively resists attacks using counterfeit media, ensures the accuracy of drug dispensing, and eliminates the security risk of correct identity but incorrect drug dispensing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921853A_ABST
    Figure CN121921853A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological recognition, and particularly discloses a medicine dispensing authority control method and system based on multi-mode biological characteristic comparison, and the method comprises the steps: collecting a face infrared image sequence and an iris visible light image sequence, extracting a pupil change curve and micro-expression muscle vibration frequency, and carrying out the time domain alignment, generating a living body physiological resonance characteristic value; performing wavelet transformation on the key frame iris image to obtain an iris phase feature tensor, and performing deep convolution feature extraction on the key frame face image to obtain a face semantic feature tensor; performing feature-level heterogeneous fusion on the two feature tensors to generate a high-dimensional joint biological feature embedding vector; mapping the permission identifier into an authorized biological feature tensor base, performing distance measurement on the high-dimensional joint biological feature embedded vector and the authorized biological feature tensor base, and if the measured distance is smaller than an intra-class aggregation threshold, generating a verification pass signal; according to the invention, the anti-counterfeiting capability and the identification precision of medicine dispensing authority control can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biometrics, and in particular to a method and system for drug dispensing access control based on multimodal biometric comparison. Background Technology

[0002] In existing technologies, identity verification methods for dispensing equipment mainly rely on operator-entered employee ID passwords, contact IC card recognition, or single biometric identification such as fingerprints or facial recognition. However, employee ID passwords and IC cards are easily leaked, misused, or stolen, posing a security risk of unauthorized personnel operating the dispensing equipment. Single-modal biometric identification methods lack robustness in complex environments; when the operator's fingers are wet, peeling, or their face is obscured, or when lighting changes, the recognition accuracy drops significantly, failing to meet the requirement of absolute uniqueness of operator identity for high-risk drug management. Furthermore, traditional facial recognition methods mostly rely on static image comparison and lack effective liveness detection mechanisms, making them vulnerable to sophisticated forgery attacks using photos, video playback, or silicone masks, resulting in serious security vulnerabilities in the dispensing authorization verification process.

[0003] In existing technologies, multimodal biometric recognition methods typically employ a decision-level fusion strategy, which involves scoring the recognition results of different modalities separately and then weighting and fusing them. However, this method fails to fully exploit the inherent correlations between features at the data level and is ill-suited for scenarios with significant differences in acquisition quality between modalities or partial loss of information from a single modality. Furthermore, existing technologies lack authenticity verification of the operator's physiological activity state, making it difficult to distinguish the dynamic physiological characteristics of a real living person from those of a counterfeit medium. At the authorization verification level, existing methods only verify the legitimacy of the operator's identity, failing to correlate the identity verification result with the prescription authorization and specific medication compartment location for the current dispensing task. This can lead to situations where identity verification is successful but medication is dispensed incorrectly. Therefore, there is an urgent need to develop a dispensing authorization control method that can deeply integrate multimodal biometrics, possess dynamic liveness detection capabilities, and form a closed-loop verification system with prescription authorization and compartment location operations to improve the anti-counterfeiting capabilities and medication security of dispensing operations. Summary of the Invention

[0004] This invention provides a method and system for drug dispensing access control based on multimodal biometric comparison, in order to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a drug dispensing access control method based on multimodal biometric comparison, comprising:

[0006] S1: Simultaneously acquire infrared image sequences of the faces and visible light image sequences of the irises of the dispensing personnel to generate a temporal multimodal image set;

[0007] S2: Perform inter-frame difference motion analysis on the time-domain multimodal image set to extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence;

[0008] S3: Align the pupil change curve with the muscle tremor frequency in the time domain to generate in vivo physiological resonance characteristic values;

[0009] S4: When the in vivo physiological resonance feature value exceeds the preset threshold, wavelet transform encoding is performed on the key frame iris image in the infrared image sequence to obtain the iris phase feature tensor, and deep convolution feature extraction is performed on the key frame face image in the visible light image sequence to obtain the face semantic feature tensor.

[0010] S5: Input the iris phase feature tensor and the face semantic feature tensor into the cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector;

[0011] S6: Map the authorization identifier of the prescription issuer to the authorized biometric tensor base pre-stored in the security chip, measure the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base, and generate a verification pass signal if the measured distance is less than the intra-class aggregation threshold.

[0012] In a preferred embodiment, the simultaneous acquisition of infrared image sequences of the face and visible light image sequences of the iris of the dispensing operator to generate a temporal multimodal image set includes:

[0013] Based on the permission verification command triggered by the dispensing terminal, the infrared camera module and the visible light camera module in the binocular camera are activated;

[0014] Based on a preset synchronous sampling clock, infrared light reflection images containing the iris region and visible light reflection images containing the face region are acquired in parallel using an infrared camera module and a visible light camera module.

[0015] Infrared and visible light images are paired into synchronized image frame pairs and arranged in chronological order to generate a time-domain multimodal image set.

[0016] The step of performing inter-frame difference motion analysis on a time-domain multimodal image set to extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence includes:

[0017] The infrared image sequence in the time-domain multimodal image set is segmented frame by frame to extract the pupil boundary contour of the iris region in the infrared image, and the pupil area time series data changing over time is determined based on the number of pixels in the area surrounded by the pupil boundary contour.

[0018] Perform sliding window difference operation on the pupil area time series data to extract the expansion and contraction of pupil area between adjacent frames and generate pupil change curve;

[0019] Facial key points are detected frame by frame in the visible light image sequence of the temporal multimodal image set. The distribution area of ​​facial muscle groups in the face region of the visible light image is located. Optical flow tracing is performed on the displacement vector of the pixel points in the distribution area of ​​facial muscle groups between adjacent frames. The number of times the displacement vector direction is periodically reversed per unit time is counted to generate the frequency of micro-expression muscle tremors.

[0020] In a preferred embodiment, the step of aligning the pupil change curve with the muscle tremor frequency in the time domain to generate in vivo physiological resonance characteristic values ​​includes:

[0021] Peak detection is performed on the pupil change curve to extract the alternating rhythm of pupil dilation and constriction, and a first physiological rhythm signal is generated.

[0022] The frequency of the micro-expression muscle tremors is interpolated in time to generate a second physiological rhythm signal;

[0023] The first and second physiological rhythm signals are input into a cross-correlation analyzer. The maximum overlap position is determined by sliding matching, and the offset of the maximum overlap position is used as an alignment reference. The first and second physiological rhythm signals are aligned on the time axis according to the alignment reference to generate an aligned dual-modal physiological signal pair.

[0024] The phase-locking degree of the aligned dual-modal physiological signal pair within a preset time window is determined, and the phase-locking degree is used as the in vivo physiological resonance characteristic value.

[0025] Determining the phase-locking degree of the aligned dual-modal physiological signal pair within a preset time window, and using the phase-locking degree as the in vivo physiological resonance characteristic value, includes:

[0026] Hilbert transforms are performed on the first and second physiological rhythm signals in the aligned dual-modal physiological signal pair respectively to extract the instantaneous phase values ​​of the first and second physiological rhythm signals.

[0027] The phase difference between the instantaneous phase value of the first physiological rhythm signal and the instantaneous phase value of the second physiological rhythm signal is determined to obtain an instantaneous phase difference sequence;

[0028] The orderly distribution of the instantaneous phase difference sequence within a preset time window is quantified to obtain a phase synchronization index, which is calculated using the following formula:

[0029]

[0030] in, For phase synchronization index, The maximum Shannon entropy value is given by the instantaneous phase difference under uniform distribution conditions. The Shannon entropy value is the instantaneous phase difference sequence within a preset time window. The preset dispersion suppression coefficient, It represents the standard deviation of the instantaneous phase difference sequence within the same time window.

[0031] In a preferred embodiment, when the in vivo physiological resonance feature value exceeds a preset threshold, wavelet transform encoding is performed on the keyframe iris image in the infrared image sequence to obtain the iris phase feature tensor, and deep convolution feature extraction is performed on the keyframe face image in the visible light image sequence to obtain the face semantic feature tensor, including:

[0032] The image frame with the highest iris texture clarity score is selected from the infrared image sequence as the keyframe iris image;

[0033] The iris region is normalized and expanded on the keyframe iris image to obtain the iris ring band expansion diagram in polar coordinates.

[0034] Extract the complex coefficient response of the iris texture in the unfolded iris ring band image, and binarize the phase component of the complex coefficient response to generate the iris phase feature tensor;

[0035] The image frames with the smallest deviation between the face pose angle and the standard frontal template are selected from the visible light image sequence as keyframe face images;

[0036] The keyframe face image is aligned and normalized for cropping to obtain a face-aligned image;

[0037] The face-aligned image is input into a pre-trained residual convolutional neural network for forward propagation, and feature maps of the pooling layers in the residual convolutional neural network are extracted.

[0038] Global average pooling is performed on the feature map to generate a face semantic feature tensor.

[0039] In a preferred embodiment, the step of inputting the iris phase feature tensor and the face semantic feature tensor into a cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector includes:

[0040] A cross-modal joint encoder is constructed, which includes an iris feature coding branch, a face feature coding branch, and a cross-modal attention fusion component;

[0041] The iris phase feature tensor is input into the iris feature encoding branch to generate the iris intermediate feature vector;

[0042] The facial semantic feature tensor is input into the facial feature encoding branch, and after dimensionality reduction and nonlinear activation by a fully connected layer, a facial intermediate feature vector is generated.

[0043] The iris intermediate feature vector and the face intermediate feature vector are input into the cross-modal attention fusion component;

[0044] Determine a first attention weight of the iris intermediate feature vector relative to the face intermediate feature vector, and determine a second attention weight of the face intermediate feature vector relative to the iris intermediate feature vector;

[0045] The iris intermediate feature vector is weighted according to the first attention weight, and the face intermediate feature vector is weighted according to the second attention weight. The weighted iris intermediate feature vector and the weighted face intermediate feature vector are concatenated to generate a high-dimensional joint biometric embedding vector.

[0046] In a preferred embodiment, mapping the prescription issuer's authorization identifier to an authorized biometric tensor pre-stored in the security chip, and measuring the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor, generating a verification pass signal if the measured distance is less than the intra-class aggregation threshold, includes:

[0047] Obtain the prescription for the current dispensing task, parse the employee ID of the prescription issuer in the prescription, and use the employee ID of the prescription issuer as an access identifier;

[0048] The permission identifier is input into the index mapper of the security chip, and the authorized biometric tensor base bound to the permission identifier is read from the tamper-proof storage area of ​​the security chip.

[0049] The authorized biometric tensor base is a template feature vector generated after being collected during the pre-registration stage and processed by the cross-modal joint encoder;

[0050] The high-dimensional joint biometric embedding vector and the authorized biometric tensor basis are mapped to a pre-constructed Riemannian manifold space, and the geodesic distance in the Riemannian manifold space is determined as follows:

[0051]

[0052] in, The geodesic distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis is... Let be the tensor representation of the high-dimensional joint biometric embedding vector in the Riemannian manifold space. As a reference point, Let the authorized biometric tensor basis be represented by a tensor in the Riemannian manifold space. To make tensor Projected onto the reference point via a logarithmic map The tangent vector obtained after tangent space. To make tensor Projected onto the reference point via a logarithmic map The tangent vector obtained after tangent space. For reference point The metric norm defined at that location;

[0053] The geodesic distance is compared with a preset intra-class aggregation threshold. If the geodesic distance is less than the intra-class aggregation threshold, it is determined that the dispensing operator and the prescription issuer are the same natural person, and a verification pass signal is generated.

[0054] In a preferred embodiment, after mapping the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis to a pre-constructed Riemannian manifold space, the method further includes:

[0055] Obtain the historical verification records of the dispensing personnel, and extract the historical high-dimensional joint biometric feature embedding vector from the historical verification records;

[0056] The historical high-dimensional joint biofeature embedding vector is mapped to the Riemannian manifold space, the Fraser mean of the historical high-dimensional joint biofeature embedding vector in the Riemannian manifold space is determined, and the Fraser mean is used as a reference base point for dynamic updating.

[0057] Based on the dynamically updated reference base point, the geodesic distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis at the current moment is re-determined to obtain the corrected metric distance;

[0058] The corrected metric distance is compared with the intra-class aggregation threshold. If the corrected metric distance is less than the intra-class aggregation threshold, an enhanced verification pass signal is generated.

[0059] To address the aforementioned problems, this invention also provides a multimodal biometrics-based drug dispensing authorization control system, the system comprising:

[0060] The multimodal image synchronous acquisition module is used to synchronously acquire infrared image sequences of the face and visible light image sequences of the iris of the dispensing operator, and generate a time-domain multimodal image set;

[0061] The physiological dynamic feature extraction module is used to perform inter-frame difference motion analysis on the time-domain multimodal image set, and extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the human face in the visible light image sequence.

[0062] The physiological resonance feature generation module is used to time-domain align the pupil change curve with the muscle tremor frequency to generate in vivo physiological resonance feature values.

[0063] The single-modal feature encoding module is used to perform wavelet transform encoding on the key frame iris image in the infrared image sequence to obtain the iris phase feature tensor when the live physiological resonance feature value exceeds the preset threshold, and to perform deep convolution feature extraction on the key frame face image in the visible light image sequence to obtain the face semantic feature tensor.

[0064] The cross-modal feature fusion module is used to input the iris phase feature tensor and the face semantic feature tensor into the cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric feature embedding vector.

[0065] The permission verification and unlocking module is used to map the permission identifier of the prescription issuer to the authorized biometric tensor base pre-stored in the security chip, and to measure the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base. If the measured distance is less than the intra-class aggregation threshold, a verification pass signal is generated.

[0066] Compared with the prior art, the present invention has the following beneficial effects:

[0067] 1. This invention simultaneously acquires infrared image sequences of the face and visible light image sequences of the iris, performs inter-frame differential motion analysis on the temporal multimodal image set, and extracts the pupil change curve regulated by the autonomic nervous system and the frequency of micro-expression muscle tremors regulated by the non-autonomic nervous system, respectively. These two are then temporally aligned to generate a live physiological resonance feature value. This feature value quantifies the intrinsic coupling degree of the two physiological rhythm signals. Since forged media cannot simultaneously simulate the involuntary contraction and expansion of the pupil and the high-frequency tremors of facial micro-expressions and their temporal correlation, the possibility of forged media impersonation is eliminated from a physiological perspective. After successful liveness verification, this invention further performs wavelet transform encoding on keyframe iris images to extract a refined iris phase feature tensor, and simultaneously performs deep convolution feature extraction on keyframe face images to obtain a high-level semantic feature tensor. Furthermore, feature-level heterogeneous fusion is achieved through the attention mechanism in the cross-modal joint encoder to generate a high-dimensional joint biometric embedding vector. This fusion strategy not only preserves the complementary information of the two modalities, but also strengthens the correspondence between the iris and the face in geometric space through attention weights. It effectively solves the problem of a sharp drop in recognition rate of a single modality in complex environments such as occlusion, changes in lighting, and defocus blur, and achieves high-precision and high-stability dual identity verification.

[0068] 2. This invention maps the generated high-dimensional joint biometric embedding vector and the authorized biometric tensor basis pre-stored in the secure chip onto a Riemannian manifold space, and calculates the shortest path length of the two on the manifold using the geodesic distance formula. This metric fully considers the nonlinear geometric structure of the feature space, resulting in higher intra-class aggregation and more obvious inter-class separation, significantly reducing the false recognition and false rejection rates. Simultaneously, this invention associates identity verification with the storage location code of the medication dispensing object through a signal. By obtaining the target storage location code from the prescription information and performing a bitwise XOR verification with the expected storage location code of the medication to be dispensed, it ensures that the medication box storage location to be opened is completely consistent with the prescription. This mechanism extends the verification of "personal" permissions to the precise verification of "medication," forming a complete control closed loop from operator identity confirmation to specific medication dispensing, effectively eliminating the security risk of correct identity but incorrect medication dispensing, and ensuring patient medication safety from the source. Attached Figure Description

[0069] Figure 1 This is a flowchart illustrating a drug dispensing access control method based on multimodal biometric comparison, provided in an embodiment of the present invention.

[0070] Figure 2 This is a functional block diagram of a multimodal biometric drug dispensing authority control system provided in an embodiment of the present invention;

[0071] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0072] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0073] This application provides a method for medication dispensing access control based on multimodal biometric comparison. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0074] Reference Figure 1 The diagram shown is a flowchart illustrating a multimodal biometrics-based drug dispensing access control method according to an embodiment of the present invention. In this embodiment, the multimodal biometrics-based drug dispensing access control method includes:

[0075] S1: Simultaneously acquire infrared image sequences of the faces and visible light image sequences of the irises of the dispensing personnel to generate a temporal multimodal image set;

[0076] In this embodiment of the invention, the simultaneous acquisition of facial infrared image sequences and iris visible light image sequences of the dispensing operator to generate a temporal multimodal image set includes:

[0077] Based on the permission verification command triggered by the dispensing terminal, the infrared camera module and the visible light camera module in the binocular camera are activated;

[0078] Based on a preset synchronous sampling clock, infrared light reflection images containing the iris region and visible light reflection images containing the face region are acquired in parallel using an infrared camera module and a visible light camera module.

[0079] Infrared and visible light images are paired into synchronized image frame pairs and arranged in chronological order to generate a time-domain multimodal image set.

[0080] The dispensing terminal is the control interface installed on the dispensing equipment. When the dispensing operator stands in front of the dispensing terminal and presses the physical button to start verification or touches the verification area on the screen, the control circuit inside the dispensing terminal generates an electrical signal as an authorization verification instruction. This instruction is transmitted to the control unit of the binocular camera through the data bus.

[0081] Using a preset synchronous sampling clock as a reference, infrared light reflection images containing the iris region and visible light reflection images containing the face region are acquired in parallel using infrared and visible light camera modules. The synchronous sampling clock is generated by a high-precision crystal oscillator inside the dispensing terminal, outputting a square wave signal of a fixed frequency. This signal is simultaneously input to the trigger inputs of both the infrared and visible light camera modules. At each rising edge of the clock, the two modules synchronously perform one exposure and readout operation. The infrared camera module receives infrared light reflected from the operator's eyes through an infrared photosensitive element, converts the light signal into an electrical signal, and generates a digital image after analog-to-digital conversion. In this image, the ring-shaped texture of the iris and the pupil boundary are clearly presented. The visible light camera module receives visible light reflected from the operator's face through a color photosensitive element, similarly converting it into an electrical signal and digitizing it to generate a digital image containing the complete facial contours, the position of facial features, and skin texture.

[0082] Infrared and visible light images are paired into synchronized image frame pairs and arranged chronologically to generate a temporal multimodal image set. The image acquisition processor of the dispensing terminal reads image data from the output buffers of the infrared and visible light camera modules. Based on the acquisition time tag attached to each frame, it combines one infrared image and one visible light image with the same timestamp to form a synchronized image frame pair. The processor allocates a circular buffer in memory and stores the synchronized image frame pairs sequentially according to the acquisition time. As the acquisition process continues, multiple synchronized image frame pairs from consecutive moments accumulate in the buffer, and these frame pairs together constitute the temporal multimodal image set.

[0083] The beneficial effects are as follows: the above steps achieve synchronous acquisition and temporal alignment of face and iris images, providing accurate input data for subsequent liveness detection and identity recognition based on multimodal biometrics. The binocular camera modules operate in parallel under a unified clock control, ensuring strict temporal correspondence between infrared and visible light images, eliminating potential time misalignment issues caused by asynchronous acquisition. The active illumination design of the infrared camera module enhances the ability to capture iris texture, making the iris region clearly discernible in the image, while the visible light camera module fully preserves the color and detail information of the face. Pairing the two types of images acquired simultaneously into synchronous frame pairs and organizing them into an image set in chronological order lays the data foundation for analyzing the temporal correlation between pupil changes and facial micro-expressions, thus more effectively resisting attacks from forged media.

[0084] S2: Perform inter-frame difference motion analysis on the time-domain multimodal image set to extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence;

[0085] In this embodiment of the invention, the step of performing inter-frame difference motion analysis on a time-domain multimodal image set to extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence includes:

[0086] The infrared image sequence in the time-domain multimodal image set is segmented frame by frame to extract the pupil boundary contour of the iris region in the infrared image, and the pupil area time series data changing over time is determined based on the number of pixels in the area surrounded by the pupil boundary contour.

[0087] Perform sliding window difference operation on the pupil area time series data to extract the expansion and contraction of pupil area between adjacent frames and generate pupil change curve;

[0088] Facial key points are detected frame by frame in the visible light image sequence of the temporal multimodal image set. The distribution area of ​​facial muscle groups in the face region of the visible light image is located. Optical flow tracing is performed on the displacement vector of the pixel points in the distribution area of ​​facial muscle groups between adjacent frames. The number of times the displacement vector direction is periodically reversed per unit time is counted to generate the frequency of micro-expression muscle tremors.

[0089] The iris region is segmented frame by frame in the infrared image sequence from the temporal multimodal image set. The pupil boundary contour of the iris region in the infrared image is extracted, and the temporal data of pupil area changing over time is determined based on the number of pixels in the area surrounded by the pupil boundary contour. The image processor extracts the infrared image sequence from the temporal multimodal image set and calls the iris region segmentation algorithm for each frame of the infrared image.

[0090] The algorithm first utilizes the grayscale difference between the iris and the surrounding sclera under infrared illumination to separate the iris region from the background through thresholding. Within the iris region, the pupil region is further identified. Since the pupil appears as a low-grayscale circular dark area in infrared images, an edge detection operator is used to extract pixels at the boundary between the pupil and iris. These pixels are then fitted with an ellipse to form a closed pupil boundary contour. After contour extraction, the total number of pixels covered within this contour is counted; this value represents the pupil area at the current moment.

[0091] The processor records the pupil area value corresponding to each frame of the infrared image sequence in the original time order, forming a sequence of pupil area changes over time.

[0092] The processor moves a fixed-size sliding window across the pupil area time-series data, containing pupil area values ​​for several consecutive frames. For two adjacent frames within the window, the pupil area of ​​the later frame is subtracted from the pupil area of ​​the earlier frame, yielding the difference. A positive difference indicates pupil dilation during that time interval, and the amount of dilation is recorded as the difference. A negative difference indicates pupil constriction, and its absolute value is taken as the amount of constriction. The processor iterates through the entire time-series data, obtaining a series of continuous dilation and constriction data points. Connecting these data points in chronological order, with time on the horizontal axis and the magnitude of dilation or constriction on the vertical axis, generates a pupil change curve reflecting the dynamic changes in the pupil.

[0093] Facial key points are detected frame by frame in the visible light image sequence of the temporal multimodal image set. The distribution area of ​​facial muscle groups in the face region of the visible light image is located. Optical flow tracing is performed on the displacement vector of the pixel points in the distribution area of ​​facial muscle groups between adjacent frames. The number of times the displacement vector direction is periodically reversed per unit time is counted to generate the frequency of micro-expression muscle tremors.

[0094] The processor extracts visible light image sequences from a temporal multimodal image set. For each frame, it calls a facial keypoint detection model, which identifies key facial features, including the corners of the eyes, eyebrows, nostrils, and corners of the mouth, using a pre-trained convolutional neural network. Based on the coordinates of these keypoints, the distribution areas of major muscle groups such as the cheeks, orbicularis oculi, and orbicularis oris are delineated according to facial anatomy, and these areas are designated as regions of interest for subsequent analysis.

[0095] For two adjacent image frames, optical flow is used to track the motion of pixels within each region of interest (ROI). Optical flow calculates the grayscale and spatial position changes of pixels between frames to obtain a two-dimensional displacement vector for each pixel. The magnitude of this vector represents the intensity of the motion, and its direction represents the direction of motion. The processor continuously processes multiple image frames, recording the directional changes of the displacement vectors within each ROI. A change in direction from positive to negative or vice versa is considered a direction reversal. The processor counts the total number of displacement vector direction reversals occurring across all ROI regions within a one-second time unit; this number represents the frequency of micro-expression muscle twitches, reflecting the activity level of subtle facial muscle movements.

[0096] The beneficial effects are that the above steps enable the extraction of pupil dynamics and facial micro-expression dynamics from infrared and visible light images, respectively. The extraction of temporal pupil area data is based on precise iris region segmentation and pupil boundary fitting, ensuring the accuracy and reliability of the data source. Sliding window difference operations convert the static pupil area into dynamic changes in expansion and contraction, highlighting the physiological rhythmic characteristics of the pupil. The localization of facial muscle group distribution areas combines keypoint detection and facial anatomy knowledge, allowing optical flow tracing to focus on the areas where micro-expression movements actually occur, avoiding interference from background areas. Optical flow tracing technology accurately measures pixel-level displacement vectors and quantifies the tremor frequency of micro-expressions by statistically counting the number of direction reversals. These two dimensions of dynamic features together form the basis for distinguishing between real living organisms and forged media, because forged media cannot simultaneously present the involuntary rhythmic changes of the pupil and the high-frequency tremors of facial micro-expressions.

[0097] S3: Align the pupil change curve with the muscle tremor frequency in the time domain to generate in vivo physiological resonance characteristic values;

[0098] In this embodiment of the invention, the step of aligning the pupil change curve with the muscle tremor frequency in the time domain to generate in vivo physiological resonance feature values ​​includes:

[0099] Peak detection is performed on the pupil change curve to extract the alternating rhythm of pupil dilation and constriction, and a first physiological rhythm signal is generated.

[0100] The frequency of the micro-expression muscle tremors is interpolated in time to generate a second physiological rhythm signal;

[0101] The first and second physiological rhythm signals are input into a cross-correlation analyzer. The maximum overlap position is determined by sliding matching, and the offset of the maximum overlap position is used as an alignment reference. The first and second physiological rhythm signals are aligned on the time axis according to the alignment reference to generate an aligned dual-modal physiological signal pair.

[0102] The phase-locking degree of the aligned dual-modal physiological signal pair within a preset time window is determined, and the phase-locking degree is used as the in vivo physiological resonance characteristic value.

[0103] Determining the phase-locking degree of the aligned dual-modal physiological signal pair within a preset time window, and using the phase-locking degree as the in vivo physiological resonance characteristic value, includes:

[0104] Hilbert transforms are performed on the first and second physiological rhythm signals in the aligned dual-modal physiological signal pair respectively to extract the instantaneous phase values ​​of the first and second physiological rhythm signals.

[0105] The phase difference between the instantaneous phase value of the first physiological rhythm signal and the instantaneous phase value of the second physiological rhythm signal is determined to obtain an instantaneous phase difference sequence;

[0106] The orderly distribution of the instantaneous phase difference sequence within a preset time window is quantified to obtain a phase synchronization index, which is calculated using the following formula:

[0107]

[0108] in, For phase synchronization index, The maximum Shannon entropy value is given by the instantaneous phase difference under uniform distribution conditions. The Shannon entropy value is the instantaneous phase difference sequence within a preset time window. The preset dispersion suppression coefficient, It represents the standard deviation of the instantaneous phase difference sequence within the same time window.

[0109] Peak detection is performed on the pupillary change curve to extract the alternating rhythm of pupillary dilation and constriction, generating the first physiological rhythm signal. The processor treats the pupillary change curve as a signal sequence fluctuating over time, sliding a detection window across this sequence to identify local maxima and minima. Local maxima correspond to the moment when pupillary dilation reaches its maximum, and local minima correspond to the moment when pupillary constriction reaches its maximum. Adjacent maxima and minima are connected sequentially; the falling segment between a maximum and a minimum represents the constriction process, and the rising segment between a minimum and the next maximum represents the dilation process. The processor records the start and end times of each dilation and each constriction process in chronological order, connecting these time points to form a waveform composed of alternating dilation and constriction, which is the first physiological rhythm signal.

[0110] The frequency of micro-expression muscle tremors is interpolated temporally to generate a second physiological rhythm signal. Initially, the frequency of micro-expression muscle tremors is a discrete point sequence, a statistical value per second, with each integer second corresponding to a frequency value. To match this signal with the pupillary change curve in temporal density, the processor uses an interpolation method to supplement the values ​​at intermediate moments between two adjacent integer second frequency values. The interpolation process assumes that the muscle tremor frequency changes continuously within a short period, calculating an approximate value for each frame based on the actual measurements at two consecutive integer second moments. After interpolation, the original discrete point sequence with one value per second is expanded into a continuous sequence corresponding one-to-one with the sampling points of the pupillary change curve; this sequence is the second physiological rhythm signal.

[0111] The first and second physiological rhythm signals are input into a cross-correlation analyzer. The maximum overlap position is determined through sliding matching, and the offset of the maximum overlap position is used as an alignment reference. Based on this alignment reference, the first and second physiological rhythm signals are aligned on the time axis, generating an aligned bimodal physiological signal pair. The cross-correlation analyzer uses the second physiological rhythm signal as a sliding window, gradually moving the second physiological rhythm signal backward from the starting position of the first physiological rhythm signal. For each time unit of movement, the sum of the products of the two signal values ​​within the overlapping region is calculated.

[0112] When the second physiological rhythm signal slides to a certain position, the sum of the products reaches its maximum value; this position represents the moment when the waveforms of the two signals are most similar. At this point, the time offset of the second physiological rhythm signal relative to the first physiological rhythm signal is recorded as an alignment reference. Based on this offset, the processor moves the second physiological rhythm signal forward or backward as a whole, aligning the peaks and troughs of the two signals in time. After the movement is complete, the two signals have a corresponding relationship at every moment, together forming an aligned bimodal physiological signal pair.

[0113] The phase-locking degree of the aligned bimodal physiological signal pair within a preset time window is determined, and this phase-locking degree is used as a live physiological resonance characteristic value. The processor extracts a fixed-length time window from the aligned bimodal physiological signal pair and analyzes the two signals within this window. Hilbert transforms are performed on both signals. The Hilbert transform is a method for converting real-valued signals into complex-valued signals. Through this transform, the instantaneous phase value corresponding to each moment can be extracted from the original waveform, i.e., the oscillation position of the signal at that moment.

[0114] The first physiological rhythm signal undergoes a Hilbert transform to obtain a first instantaneous phase value sequence, and the second physiological rhythm signal undergoes a Hilbert transform to obtain a second instantaneous phase value sequence. The processor subtracts the first and second instantaneous phase values ​​at the same moment to obtain the instantaneous phase difference at that moment. This operation is repeated for all moments to form an instantaneous phase difference sequence. Analyzing the distribution of values ​​in this sequence, if the two signals are perfectly synchronized, the instantaneous phase difference will always remain around a fixed value; if the two signals are not synchronized, the instantaneous phase difference will be scattered over a large range. The processor quantifies the orderliness of the distribution of the instantaneous phase difference sequence and calculates the Shannon entropy value of the sequence within a preset time window. The Shannon entropy value reflects the degree of disorder in the phase difference values. At the same time, the standard deviation of the instantaneous phase difference sequence within the same time window is calculated. The standard deviation reflects the fluctuation range of the phase difference around the central value. Substituting the Shannon entropy and standard deviation into the phase synchronization index calculation formula yields a value between zero and one. This value is the in vivo physiological resonance characteristic value. The closer the value is to one, the higher the synchronicity between pupil changes and facial micro-expressions. The closer the value is to zero, the lower the synchronicity.

[0115] The instantaneous phase difference sequence within the preset time window is used to calculate the actual Shannon entropy value. The theoretical maximum Shannon entropy value of all possible values ​​in the sequence under uniform distribution conditions is predetermined by the system based on the number of instantaneous phase difference values. The instantaneous phase difference sequence within the same time window is used to calculate the standard deviation value reflecting the degree of dispersion. The dispersion suppression coefficient is a constant set based on experience to adjust the influence of the standard deviation on the result. Substituting these values ​​into the formula yields the phase synchronization index.

[0116] The significance of this formula lies in comprehensively assessing the phase synchronization between the pupil change curve and the muscle tremor frequency through the product of two factors. The first factor is composed of the maximum Shannon entropy value minus the actual Shannon entropy value and then divided by the maximum Shannon entropy value, which is used to measure the degree to which the instantaneous phase difference distribution deviates from the uniform distribution, i.e., the concentration of the distribution. The second factor is composed of one divided by one plus the product of the dispersion suppression coefficient and the standard deviation, which is used to measure the stability of the instantaneous phase difference fluctuation, i.e., the dispersion of the distribution. The two factors work together to enable the phase synchronization index to simultaneously reflect the phase lock strength of the two physiological rhythm signals.

[0117] The trend of this formula is that when the distribution of the instantaneous phase difference sequence is highly concentrated and the standard deviation is extremely small, the first factor approaches one and the second factor also approaches one. The phase synchronization index approaches one, indicating that the two physiological signals are completely synchronized. When the instantaneous phase difference distribution is close to uniform and the standard deviation is extremely large, the first factor approaches zero or the second factor approaches zero. The phase synchronization index approaches zero, indicating that the two physiological signals are completely out of sync. In practice, the phase synchronization index increases with the increase of concentration and stability, and decreases with the increase of dispersion and volatility.

[0118] The beneficial effects are that, through the above steps, pupil dynamics and facial micro-expression dynamics are converted into quantifiable physiological rhythm signals. Cross-correlation analysis is used to eliminate the time delay between the two, and the synchronization strength of the two physiological rhythm signals is quantified by phase locking analysis. The final generated in vivo physiological resonance characteristic value provides a reliable basis for distinguishing between real living organisms and fake media. The pupils and micro-expressions of real living organisms are regulated by the same nervous system and will inevitably show stable phase locking, while no fake media can simultaneously simulate the fine temporal correlation of these two physiological signals.

[0119] S4: When the in vivo physiological resonance feature value exceeds the preset threshold, wavelet transform encoding is performed on the key frame iris image in the infrared image sequence to obtain the iris phase feature tensor, and deep convolution feature extraction is performed on the key frame face image in the visible light image sequence to obtain the face semantic feature tensor.

[0120] In this embodiment of the invention, when the in vivo physiological resonance feature value exceeds a preset threshold, wavelet transform encoding is performed on the keyframe iris image in the infrared image sequence to obtain the iris phase feature tensor, and deep convolution feature extraction is performed on the keyframe face image in the visible light image sequence to obtain the face semantic feature tensor, including:

[0121] The image frame with the highest iris texture clarity score is selected from the infrared image sequence as the keyframe iris image;

[0122] The iris region is normalized and expanded on the keyframe iris image to obtain the iris annular band expansion diagram in polar coordinates.

[0123] Extract the complex coefficient response of the iris texture in the unfolded iris ring band image, and binarize the phase component of the complex coefficient response to generate the iris phase feature tensor;

[0124] The image frames with the smallest deviation between the face pose angle and the standard frontal template are selected from the visible light image sequence as keyframe face images;

[0125] The keyframe face image is aligned and normalized for cropping to obtain a face-aligned image;

[0126] The face-aligned image is input into a pre-trained residual convolutional neural network for forward propagation, and feature maps of the pooling layers in the residual convolutional neural network are extracted.

[0127] Global average pooling is performed on the feature map to generate a facial semantic feature tensor.

[0128] The image processor selects the frame with the highest iris texture sharpness score from the infrared image sequence as the keyframe iris image. Each frame of the infrared image sequence is sequentially read into memory, and the sharpness of the segmented iris region in each frame is evaluated. Sharpness evaluation is achieved by analyzing the high-frequency components in the image's frequency domain; the richer the high-frequency components, the clearer the details of the iris texture. The processor records the sharpness score of each frame. After traversing the entire infrared image sequence, the frame with the highest score is marked as the keyframe iris image, containing the clearest iris texture details for subsequent feature extraction.

[0129] The iris region is normalized and unfolded in the keyframe iris image to obtain the iris annular band unfolded image in polar coordinates. The processor maps the iris region in the keyframe iris image from Cartesian coordinates to polar coordinates. First, the inner and outer circular boundaries of the iris are determined; the inner boundary is the pupil edge, and the outer boundary is the junction of the iris and sclera. Using the iris center as the pole, the image extends radially from the inner boundary to the outer boundary, while sampling at fixed angular intervals along the circumference. At each sampling point, the corresponding pixel grayscale value is extracted. Using the circumferential direction as the x-axis and the radial direction as the y-axis, these pixels are rearranged to form a rectangular unfolded image. This unfolded image is the iris annular band unfolded image, which stretches the annular iris region into a rectangle, eliminating the size changes caused by pupil dilation and contraction.

[0130] The processor extracts the complex coefficient response of the iris texture from the unfolded iris ring band image and binarizes the phase component of the complex coefficient response to generate an iris phase feature tensor. The processor convolves the unfolded iris ring band image with a set of two-dimensional Gabor filters. Each Gabor filter has specific scale and orientation parameters, capable of capturing local features of the iris texture at different frequencies and directions. The result of the convolution operation is a complex number containing real and imaginary parts. The ratio of the real to imaginary parts is used to obtain the phase value through arctangent operation. The processor extracts the phase value for each pixel position, each scale, and each orientation, forming a multi-layered three-dimensional data structure. Each phase value is binarized; phase values ​​greater than zero are encoded as one, and phase values ​​less than or equal to zero are encoded as zero. The resulting three-dimensional data structure composed of zeros and ones after binarization is the iris phase feature tensor, which compactly encodes the fine structural information of the iris texture.

[0131] The processor selects the image frame with the smallest deviation from the standard frontal template from the visible light image sequence as the keyframe face image. For each frame in the visible light image sequence, the processor performs face detection, locating the face region and further detecting facial key points, including the inner and outer corners of the left and right eyes, the tip of the nose, and the corners of the mouth. Based on the two-dimensional coordinates of these key points, the yaw, pitch, and roll angles of the current face in three-dimensional space are estimated by solving the perspective transformation relationship. These three angles are compared with the zero-degree angle of the standard frontal template, and the sum of the absolute values ​​of the deviations for each angle is calculated. The processor traverses the entire visible light image sequence, finding the frame with the smallest sum of absolute deviations and marking it as the keyframe face image. The face pose in this image is closest to a frontal view, which is beneficial for extracting stable facial features.

[0132] Face alignment and normalized cropping are performed on keyframe face images to obtain a face-aligned image. The processor calculates the positions of the left and right eye centers based on the coordinates of facial key points detected in the keyframe face images. Using the line connecting the two eye centers as a reference, the image is rotated through an affine transformation to keep the line horizontal. Simultaneously, the face region is scaled to a preset standard size based on the distance between the eyes, ensuring that faces in all input images have the same scale and eye positions. After rotation and scaling, a rectangular region containing the forehead, eyes, nose, mouth, and chin is cropped from the transformed image; this region is the normalized face-aligned image.

[0133] The face-aligned image is input into a pre-trained residual convolutional neural network (RNN) for forward propagation, and feature maps from the pooling layers are extracted. The residual RNN consists of multiple convolutional layers, batch normalization layers, activation function layers, and residual connections. The network parameters have been pre-trained on a massive dataset of face images. The face-aligned image serves as input data, passing through each layer of the network sequentially. Convolutional layers extract local features using sliding filters, batch normalization layers normalize the features, activation function layers introduce non-linear transformations, and residual connections directly pass shallow features to deeper layers. When the computation reaches the last pooling layer, the output is a three-dimensional feature map. This feature map contains high-level semantic information obtained after layer-by-layer abstraction of the input image; each channel corresponds to a feature pattern, and each spatial location corresponds to a region of the original image.

[0134] Global average pooling is performed on the feature map to generate a facial semantic feature tensor. The processor performs global average pooling on the 3D feature map output by the pooling layer, calculating the average value of all spatial locations within each channel of the feature map, resulting in a single value. The number of average values ​​corresponds to the number of channels, and these average values ​​are arranged in channel order to form a one-dimensional vector. This vector is the facial semantic feature tensor extracted from the face image. This tensor compactly represents the overall attributes of the face, such as facial contours, feature distribution, and skin texture—high-level semantic information.

[0135] The beneficial effects are as follows: high-quality iris and facial features were extracted from infrared and visible light image sequences, respectively, through the above steps. Iris texture sharpness scoring ensured rich detail in the iris images used for encoding, and normalized unrolling eliminated distortion interference caused by pupil dilation on the iris texture. Multi-scale, multi-directional convolution of the Gabor filter bank comprehensively captured the fine structure of the iris texture, and binarized encoding of the phase components made the features illumination-invariant. Facial pose angle selection ensured that the facial images used for feature extraction were close to frontal, and face alignment and normalized cropping eliminated differences in pose and scale. The deep structure of the residual convolutional neural network extracted highly discriminative facial semantic features, and global average pooling compressed the 3D feature map into a compact feature vector. These two types of features describe the operator's biological characteristics from two dimensions: the fine texture of the iris and the overall appearance of the face, respectively, providing high-quality input data for subsequent multimodal fusion recognition.

[0136] S5: Input the iris phase feature tensor and the face semantic feature tensor into the cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector;

[0137] In this embodiment of the invention, the step of inputting the iris phase feature tensor and the face semantic feature tensor into a cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector includes:

[0138] A cross-modal joint encoder is constructed, which includes an iris feature coding branch, a face feature coding branch, and a cross-modal attention fusion component;

[0139] The iris phase feature tensor is input into the iris feature encoding branch to generate the iris intermediate feature vector;

[0140] The facial semantic feature tensor is input into the facial feature encoding branch, and after dimensionality reduction and nonlinear activation by a fully connected layer, a facial intermediate feature vector is generated.

[0141] The iris intermediate feature vector and the face intermediate feature vector are input into the cross-modal attention fusion component;

[0142] Determine a first attention weight of the iris intermediate feature vector relative to the face intermediate feature vector, and determine a second attention weight of the face intermediate feature vector relative to the iris intermediate feature vector;

[0143] The iris intermediate feature vector is weighted according to the first attention weight, and the face intermediate feature vector is weighted according to the second attention weight. The weighted iris intermediate feature vector and the weighted face intermediate feature vector are concatenated to generate a high-dimensional joint biometric embedding vector.

[0144] A cross-modal joint encoder is constructed, comprising an iris feature encoding branch, a face feature encoding branch, and a cross-modal attention fusion component. The processor creates three independent computation paths in memory. The iris feature encoding branch consists of sequentially connected fully connected layers, as does the face feature encoding branch. The cross-modal attention fusion component contains two parallel attention computation units and one concatenation unit. These three parts together constitute the complete cross-modal joint encoder structure.

[0145] The iris phase feature tensor is input into the iris feature encoding branch to generate the iris intermediate feature vector. The iris phase feature tensor is a multi-dimensional data structure, first flattened into a one-dimensional sequence, and then fed into the fully connected layer of the iris feature encoding branch. Each neuron in the fully connected layer is connected to all input values, and the output value of that neuron is calculated through a weighted summation. The iris phase feature tensor undergoes successive transformations through multiple fully connected layers, each layer reorganizing and abstracting the data, ultimately outputting a compact one-dimensional vector, which is the iris intermediate feature vector, condensing the key information from the original iris phase feature tensor.

[0146] The facial semantic feature tensor is input into the facial feature encoding branch. After dimensionality reduction and nonlinear activation by fully connected layers, a facial intermediate feature vector is generated. The facial semantic feature tensor itself is already a one-dimensional vector, and it is directly fed into the first fully connected layer of the facial feature encoding branch. The fully connected layer performs a linear transformation on this vector, reducing the dimensionality of the input data to a preset small size. The transformed result is then fed into a nonlinear activation function, which adjusts all negative values ​​to near zero while retaining positive values, enabling the network to learn more complex nonlinear relationships. Through alternating processing by multiple fully connected layers and activation functions, the final output facial intermediate feature vector has lower dimensionality and stronger discriminative ability while preserving semantic information.

[0147] The intermediate iris feature vector and the intermediate face feature vector are input into the cross-modal attention fusion component. The processor simultaneously feeds the two intermediate feature vectors into two parallel attention computation units in the cross-modal attention fusion component. One unit focuses on the intermediate face feature vector based on the intermediate iris feature vector, and the other unit focuses on the intermediate iris feature vector based on the intermediate face feature vector. The computation processes of the two units are performed simultaneously without interference.

[0148] A first attention weight is determined relative to the middle feature vector of the iris, and a second attention weight is determined relative to the middle feature vector of the face. In the first attention calculation unit, the processor calculates the matching degree between each element of the iris middle feature vector and each element of the middle feature vector of the face; elements of the iris feature vector with a high correlation to the face feature receive a larger weight value. In the second attention calculation unit, the processor calculates the matching degree between each element of the middle feature vector of the face and each element of the middle feature vector of the iris; elements of the face feature vector with a high correlation to the iris feature receive a larger weight value. All weight values ​​are normalized to ensure that the sum of all weights within each feature vector remains a fixed value.

[0149] The iris intermediate feature vector is weighted according to the first attention weight, and the face intermediate feature vector is weighted according to the second attention weight. The weighted iris intermediate feature vector and the weighted face intermediate feature vector are concatenated to generate a high-dimensional joint biometric embedding vector. The processor multiplies each weight value in the first attention weight with the corresponding element in the iris intermediate feature vector to obtain the weighted iris intermediate feature vector. Similarly, each weight value in the second attention weight is multiplied with the corresponding element in the face intermediate feature vector to obtain the weighted face intermediate feature vector. These two weighted vectors are then concatenated end-to-end to form a longer vector. This merged vector is the high-dimensional joint biometric embedding vector, which simultaneously contains feature information from both the iris and face, and strengthens the correspondence between the two modalities through attention weights.

[0150] The beneficial effects are that the above steps achieve deep fusion of iris and facial features. The iris and facial feature encoding branches respectively reduce the dimensionality and abstract the features of the two modalities, removing redundant information and enhancing the expression of key features. The cross-modal attention fusion component strengthens the face-related parts of the iris features and vice versa by calculating the mutual attention weights between the two feature vectors. The weighting operation applies the effect of the attention mechanism to the original features, and the concatenation operation merges the features of the two modalities into a unified joint representation. The generated high-dimensional joint biometric embedding vector not only contains complementary information from the two modalities but also reflects their intrinsic relationship, providing a more comprehensive and accurate feature foundation for subsequent identity comparison.

[0151] S6: Map the authorization identifier of the prescription issuer to the authorized biometric tensor base pre-stored in the security chip, measure the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base, and generate a verification pass signal if the measured distance is less than the intra-class aggregation threshold.

[0152] In this embodiment of the invention, mapping the authorization identifier of the prescription issuer to an authorized biometric tensor base pre-stored in the security chip, and measuring the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base, generating a verification pass signal if the measured distance is less than the intra-class aggregation threshold, includes:

[0153] Obtain the prescription for the current dispensing task, parse the employee ID of the prescription issuer in the prescription, and use the employee ID of the prescription issuer as an access identifier;

[0154] The permission identifier is input into the index mapper of the security chip, and the authorized biometric tensor base bound to the permission identifier is read from the tamper-proof storage area of ​​the security chip.

[0155] The authorized biometric tensor base is a template feature vector generated after being collected during the pre-registration stage and processed by the cross-modal joint encoder;

[0156] The high-dimensional joint biometric embedding vector and the authorized biometric tensor basis are mapped to a pre-constructed Riemannian manifold space, and the geodesic distance in the Riemannian manifold space is determined as follows:

[0157]

[0158] in, The geodesic distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis is... Let be the tensor representation of the high-dimensional joint biometric embedding vector in the Riemannian manifold space. As a reference point, Let the authorized biometric tensor basis be represented by a tensor in the Riemannian manifold space. To make tensor Projected onto the reference point via a logarithmic map The tangent vector obtained after tangent space. To make tensor Projected onto the reference point via a logarithmic map The tangent vector obtained after tangent space. For reference point The metric norm defined at that location;

[0159] The geodesic distance is compared with a preset intra-class aggregation threshold. If the geodesic distance is less than the intra-class aggregation threshold, it is determined that the dispensing operator and the prescription issuer are the same natural person, and a verification pass signal is generated.

[0160] After mapping the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis to the pre-constructed Riemannian manifold space, the method further includes:

[0161] Obtain the historical verification records of the dispensing personnel, and extract the historical high-dimensional joint biometric feature embedding vector from the historical verification records;

[0162] The historical high-dimensional joint biofeature embedding vector is mapped to the Riemannian manifold space, the Fraser mean of the historical high-dimensional joint biofeature embedding vector in the Riemannian manifold space is determined, and the Fraser mean is used as a reference base point for dynamic updating.

[0163] Based on the dynamically updated reference base point, the geodesic distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis at the current moment is re-determined to obtain the corrected metric distance;

[0164] The corrected metric distance is compared with the intra-class aggregation threshold. If the corrected metric distance is less than the intra-class aggregation threshold, an enhanced verification pass signal is generated.

[0165] The system retrieves the prescription for the current dispensing task and parses the employee ID of the prescribing personnel from it, using this employee ID as an access identifier. The dispensing terminal establishes a communication connection with the hospital information system and requests the corresponding electronic prescription from the system based on the dispensing task number handled by the currently logged-in dispensing operator. The electronic prescription data returned by the hospital information system contains multiple data fields. The processor parses the prescription data, locates the prescribing doctor's information field, and extracts the specific employee ID string from this field. This string serves as the access identifier for the prescribing personnel.

[0166] The authorization identifier is input into the index mapper of the security chip, and the authorized biometric tensor base bound to the authorization identifier is read from the tamper-proof storage area of ​​the security chip. The security chip is an independent hardware module installed on the mainboard of the dispensing terminal, with a physical anti-attack design. The processor sends the authorization identifier to the input interface of the security chip via the internal bus. After receiving the identifier, the index mapper inside the security chip uses it as a key to look up the identifier in the index table of the tamper-proof storage area. The tamper-proof storage area pre-stores biometric templates of multiple authorized personnel, each template being bound to a corresponding employee number. After the index mapper finds a record that matches the input employee number, it reads the data content contained in that record from the storage area; this data is the authorized biometric tensor base.

[0167] The authorized biometric tensor base is a template feature vector generated after being collected during the pre-registration phase and processed by a cross-modal joint encoder. During system initialization, medical personnel with prescription-issuing authority complete the registration process at the dispensing terminal. The terminal collects their facial infrared image sequence and iris visible light image sequence, which undergo the same in vivo physiological resonance feature value calculation, keyframe screening, wavelet transform encoding, deep convolution feature extraction, and cross-modal joint encoder fusion processing to finally generate a high-dimensional joint biometric embedding vector. This vector is written into the tamper-proof storage area of ​​the secure chip as the personnel's biometric template and permanently bound to the personnel's employee ID.

[0168] The high-dimensional joint biometric embedding vector and the authorized biometric tensor basis are mapped to a pre-constructed Riemannian manifold space, and the geodesic distance in the Riemannian manifold space is determined. The processor first constructs a Riemannian manifold space, a nonlinear space with a specific geometric structure, on which all valid biometric embedding vectors are distributed. The high-dimensional joint biometric embedding vector and the authorized biometric tensor basis are placed as two points in this manifold space. To measure the shortest path length between these two points, the processor selects a reference point on the manifold and projects the two points onto the tangent space of the reference point using a logarithmic mapping. The logarithmic mapping process involves finding the shortest path to the target point along the manifold surface from the reference point and converting the direction and length of this path into a vector in the tangent space. The two tangent vectors obtained by projecting the two target points into their tangent spaces are subtracted, and the length of the resulting difference vector is calculated under the metric norm defined at the reference point. This length value is the geodesic distance between the two points in the original manifold space.

[0169] The geodesic distance is compared with a preset intra-class aggregation threshold. If the geodesic distance is less than the intra-class aggregation threshold, the dispensing operator and the prescription issuer are determined to be the same person, and a verification pass signal is generated. The processor compares the calculated geodesic distance value with a preset threshold in the system, which is determined based on the distribution range of multiple data collections of a large number of individuals. If the geodesic distance is less than the threshold, it indicates that the currently collected biometric vector and the template vector are very close in manifold space, falling within the normal fluctuation range of the same person. Based on this, the processor determines that the current dispensing operator and the prescription issuer on the prescription form are the same person, generating a high-level verification pass signal, which is sent to the control circuit of the dispensing execution mechanism.

[0170] The system acquires the historical verification records of dispensing operators and extracts the historical high-dimensional joint biometric embedding vectors from these records. After each successful identity verification, the dispensing terminal stores the newly acquired and generated high-dimensional joint biometric embedding vector, along with the verification time, into local storage, forming the operator's historical verification record. When dynamic updates to the reference point are required, the processor reads the historical high-dimensional joint biometric embedding vectors corresponding to the operator's most recent successful verifications from the storage medium and uses these vectors as the basis for the update.

[0171] The historical high-dimensional joint biometric embedding vectors are mapped onto a Riemannian manifold space. The Fraser mean of these vectors in the Riemannian manifold space is determined, and this Fraser mean is used as a dynamically updated reference point. The processor maps all the read historical high-dimensional joint biometric embedding vectors onto the Riemannian manifold space, making each vector a point on the manifold. A point is found on the manifold that minimizes the sum of the squares of the geodesic distances to all historical feature points; this point is the Fraser mean. The search process is iterative. First, an initial point is selected, and the geodesic distances and directions from this point to all historical points are calculated. Then, the point is moved along the direction that minimizes the sum of the distances, repeating this process multiple times until convergence. The final Fraser mean point is the center point that best represents the recent biometric distribution of the operator, and the processor sets this point as the new reference point.

[0172] Based on the dynamically updated reference base point, the geodesic distance between the current high-dimensional joint biometric embedding vector and the authorized biometric tensor basis is redefined, resulting in the corrected metric distance. The processor uses the newly determined Friesian mean point as the reference base point and repeats the logarithmic mapping and norm calculation process. The current high-dimensional joint biometric embedding vector and the authorized biometric tensor basis are projected onto the tangent space of the new reference base point through logarithmic mapping, resulting in two new tangent vectors. The difference between these two new tangent vectors is calculated, and the length of the difference vector is calculated under the metric norm defined at the new reference base point. This new length value is the corrected metric distance.

[0173] The corrected metric distance is compared with the intra-class aggregation threshold. If the corrected metric distance is less than the intra-class aggregation threshold, an enhanced verification pass signal is generated. The processor compares the corrected metric distance with the same intra-class aggregation threshold. If the corrected distance is still less than the threshold, it means that even with a different reference point that better reflects the current biometric distribution, the current feature remains highly similar to the template feature. Based on this, the processor generates an enhanced verification pass signal, which not only indicates that authentication has passed but also that the verification process used a dynamically updated reference point, resulting in higher adaptability and accuracy.

[0174] The beneficial effects are as follows: High-precision biometric matching based on Riemannian manifolds is achieved through the above steps. Mapping feature vectors to the Riemannian manifold space and using geodesic distance for measurement fully considers the nonlinear distribution characteristics of high-dimensional biometric data, reflecting the true geometric relationship between features more accurately than traditional Euclidean distance or cosine similarity. The tamper-proof storage area of ​​the secure chip ensures the physical security of the template feature vectors, preventing malicious reading or tampering. The mechanism for dynamically updating the reference base point uses historical verification data to calculate the Friesian mean, enabling the reference base point to adaptively adjust to gradual changes in biometric features, such as aging or minor scarring, avoiding an increase in false rejection rate due to template feature aging. The corrected distance measurement is calculated based on a new base point that better represents the current state, further improving recognition accuracy over long-term use.

[0175] like Figure 2 The diagram shown is a functional block diagram of a multimodal biometric comparison drug dispensing authority control system provided in an embodiment of the present invention.

[0176] The multimodal biometrics-based drug dispensing access control system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the multimodal biometrics-based drug dispensing access control system 100 may include a multimodal image synchronous acquisition module 101, a physiological dynamic feature extraction module 102, a physiological resonance feature generation module 103, a single-modal feature encoding module 104, a cross-modal feature fusion module 105, and an access verification and unlocking module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.

[0177] In this embodiment, the functions of each module / unit are as follows:

[0178] The multimodal image synchronous acquisition module 101 is used to synchronously acquire the infrared image sequence of the face and the visible light image sequence of the iris of the dispensing operator, and generate a time-domain multimodal image set;

[0179] The physiological dynamic feature extraction module 102 is used to perform inter-frame differential motion analysis on the time-domain multimodal image set, and extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence.

[0180] The physiological resonance feature generation module 103 is used to time-domain align the pupil change curve with the muscle tremor frequency to generate in vivo physiological resonance feature values.

[0181] The single-modal feature encoding module 104 is used to perform wavelet transform encoding on the key frame iris image in the infrared image sequence to obtain the iris phase feature tensor when the live physiological resonance feature value exceeds a preset threshold, and to perform deep convolution feature extraction on the key frame face image in the visible light image sequence to obtain the face semantic feature tensor.

[0182] The cross-modal feature fusion module 105 is used to input the iris phase feature tensor and the face semantic feature tensor into the cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector.

[0183] The permission verification and unlocking module 106 is used to map the permission identifier of the prescription issuer to the authorized biometric tensor base pre-stored in the security chip, and to measure the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base. If the measured distance is less than the intra-class aggregation threshold, a verification pass signal is generated.

[0184] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0185] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0186] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0187] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0188] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for drug dispensing access control based on multimodal biometric comparison, characterized in that, The method includes: S1: Simultaneously acquire infrared image sequences of the faces and visible light image sequences of the irises of the dispensing personnel to generate a time-domain multimodal image set; S2: Perform inter-frame difference motion analysis on the time-domain multimodal image set to extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence; S3: Align the pupil change curve with the muscle tremor frequency in the time domain to generate in vivo physiological resonance characteristic values; S4: When the in vivo physiological resonance feature value exceeds the preset threshold, wavelet transform encoding is performed on the key frame iris image in the infrared image sequence to obtain the iris phase feature tensor, and deep convolution feature extraction is performed on the key frame face image in the visible light image sequence to obtain the face semantic feature tensor. S5: Input the iris phase feature tensor and the face semantic feature tensor into the cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector; S6: Map the authorization identifier of the prescription issuer to the authorized biometric tensor base pre-stored in the security chip, measure the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base, and generate a verification pass signal if the measured distance is less than the intra-class aggregation threshold.

2. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 1, characterized in that, The simultaneous acquisition of infrared facial image sequences and visible light iris image sequences of the dispensing operator generates a temporal multimodal image set, including: Based on the permission verification command triggered by the dispensing terminal, the infrared camera module and the visible light camera module in the binocular camera are activated; Based on a preset synchronous sampling clock, infrared light reflection images containing the iris region and visible light reflection images containing the face region are acquired in parallel using an infrared camera module and a visible light camera module. Infrared and visible light images are paired into synchronized image frame pairs and arranged in chronological order to generate a time-domain multimodal image set.

3. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 1, characterized in that, The step of performing inter-frame difference motion analysis on a time-domain multimodal image set to extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the face in the visible light image sequence includes: The infrared image sequence in the time-domain multimodal image set is segmented frame by frame to extract the pupil boundary contour of the iris region in the infrared image, and the pupil area time series data changing over time is determined based on the number of pixels in the area surrounded by the pupil boundary contour. Perform sliding window difference operation on the pupil area time series data to extract the expansion and contraction of pupil area between adjacent frames and generate pupil change curve; Facial key points are detected frame by frame in the visible light image sequence of the temporal multimodal image set. The distribution area of ​​facial muscle groups in the face region of the visible light image is located. Optical flow tracing is performed on the displacement vector of the pixel points in the distribution area of ​​facial muscle groups between adjacent frames. The number of times the displacement vector direction is periodically reversed per unit time is counted to generate the frequency of micro-expression muscle tremors.

4. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 1, characterized in that, The step of aligning the pupil change curve with the muscle tremor frequency in the time domain to generate in vivo physiological resonance feature values ​​includes: Peak detection is performed on the pupil change curve to extract the alternating rhythm of pupil dilation and constriction, and a first physiological rhythm signal is generated. The frequency of the micro-expression muscle tremors is interpolated in time to generate a second physiological rhythm signal; The first and second physiological rhythm signals are input into a cross-correlation analyzer. The maximum overlap position is determined by sliding matching, and the offset of the maximum overlap position is used as an alignment reference. The first and second physiological rhythm signals are aligned on the time axis according to the alignment reference to generate an aligned dual-modal physiological signal pair. The phase-locking degree of the aligned dual-modal physiological signal pair within a preset time window is determined, and the phase-locking degree is used as the in vivo physiological resonance characteristic value.

5. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 4, characterized in that, Determining the phase-locking degree of the aligned dual-modal physiological signal pair within a preset time window, and using the phase-locking degree as the in vivo physiological resonance characteristic value, includes: Hilbert transforms are performed on the first and second physiological rhythm signals in the aligned dual-modal physiological signal pair respectively to extract the instantaneous phase values ​​of the first and second physiological rhythm signals. The phase difference between the instantaneous phase value of the first physiological rhythm signal and the instantaneous phase value of the second physiological rhythm signal is determined to obtain an instantaneous phase difference sequence; The orderly distribution of the instantaneous phase difference sequence within a preset time window is quantified to obtain a phase synchronization index, which is calculated using the following formula: ; in, For phase synchronization index, The maximum Shannon entropy value is given by the instantaneous phase difference under uniform distribution conditions. The Shannon entropy value is the instantaneous phase difference sequence within a preset time window. The preset dispersion suppression coefficient, It represents the standard deviation of the instantaneous phase difference sequence within the same time window.

6. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 1, characterized in that, When the in vivo physiological resonance feature value exceeds a preset threshold, wavelet transform encoding is performed on the keyframe iris image in the infrared image sequence to obtain the iris phase feature tensor, and deep convolution feature extraction is performed on the keyframe face image in the visible light image sequence to obtain the face semantic feature tensor, including: The image frame with the highest iris texture clarity score is selected from the infrared image sequence as the keyframe iris image; The iris region is normalized and expanded on the keyframe iris image to obtain the iris annular band expansion diagram in polar coordinates. Extract the complex coefficient response of the iris texture in the unfolded iris ring band image, and binarize the phase component of the complex coefficient response to generate the iris phase feature tensor; The image frames with the smallest deviation between the face pose angle and the standard frontal template are selected from the visible light image sequence as keyframe face images; The keyframe face image is aligned and normalized for cropping to obtain a face-aligned image; The face-aligned image is input into a pre-trained residual convolutional neural network for forward propagation, and feature maps of the pooling layers in the residual convolutional neural network are extracted. Global average pooling is performed on the feature map to generate a facial semantic feature tensor.

7. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 1, characterized in that, The step of inputting the iris phase feature tensor and the face semantic feature tensor into a cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric embedding vector includes: A cross-modal joint encoder is constructed, which includes an iris feature coding branch, a face feature coding branch, and a cross-modal attention fusion component; The iris phase feature tensor is input into the iris feature encoding branch to generate the iris intermediate feature vector; The facial semantic feature tensor is input into the facial feature encoding branch, and after dimensionality reduction and nonlinear activation by a fully connected layer, a facial intermediate feature vector is generated. The iris intermediate feature vector and the face intermediate feature vector are input into the cross-modal attention fusion component; Determine a first attention weight of the iris intermediate feature vector relative to the face intermediate feature vector, and determine a second attention weight of the face intermediate feature vector relative to the iris intermediate feature vector; The iris intermediate feature vector is weighted according to the first attention weight, and the face intermediate feature vector is weighted according to the second attention weight. The weighted iris intermediate feature vector and the weighted face intermediate feature vector are concatenated to generate a high-dimensional joint biometric embedding vector.

8. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 1, characterized in that, The process of mapping the prescription issuer's authorization identifier to an authorized biometric tensor pre-stored in the security chip, and measuring the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor, generates a verification pass signal if the measured distance is less than the intra-class aggregation threshold, including: Obtain the prescription for the current dispensing task, parse the employee ID of the prescription issuer in the prescription, and use the employee ID of the prescription issuer as an access identifier; The permission identifier is input into the index mapper of the security chip, and the authorized biometric tensor base bound to the permission identifier is read from the tamper-proof storage area of ​​the security chip. The authorized biometric tensor base is a template feature vector generated after being collected during the pre-registration stage and processed by the cross-modal joint encoder; The high-dimensional joint biometric embedding vector and the authorized biometric tensor basis are mapped to a pre-constructed Riemannian manifold space, and the geodesic distance in the Riemannian manifold space is determined as follows: ; in, The geodesic distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis is... Let be the tensor representation of the high-dimensional joint biometric embedding vector in the Riemannian manifold space. As a reference point, Let the authorized biometric tensor basis be represented by a tensor in the Riemannian manifold space. To make tensor Projected onto the reference point via a logarithmic map The tangent vector obtained after tangent space. To make tensor Projected onto the reference point via a logarithmic map The tangent vector obtained after tangent space. For reference point The metric norm defined at that location; The geodesic distance is compared with a preset intra-class aggregation threshold. If the geodesic distance is less than the intra-class aggregation threshold, it is determined that the dispensing operator and the prescription issuer are the same natural person, and a verification pass signal is generated.

9. The method for drug dispensing access control based on multimodal biometric comparison as described in claim 8, characterized in that, After mapping the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis to the pre-constructed Riemannian manifold space, the method further includes: Obtain the historical verification records of the dispensing personnel, and extract the historical high-dimensional joint biometric feature embedding vector from the historical verification records; The historical high-dimensional joint biofeature embedding vector is mapped to the Riemannian manifold space, the Fraser mean of the historical high-dimensional joint biofeature embedding vector in the Riemannian manifold space is determined, and the Fraser mean is used as a reference base point for dynamic updating. Based on the dynamically updated reference base point, the geodesic distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor basis at the current moment is re-determined to obtain the corrected metric distance; The corrected metric distance is compared with the intra-class aggregation threshold. If the corrected metric distance is less than the intra-class aggregation threshold, an enhanced verification pass signal is generated.

10. A multimodal biometric comparison-based drug dispensing access control system, characterized in that, The system is used to implement the drug dispensing access control method for multimodal biometric comparison as described in claim 1, the system comprising: The multimodal image synchronous acquisition module is used to synchronously acquire infrared image sequences of the face and visible light image sequences of the iris of the dispensing operator, and generate a time-domain multimodal image set; The physiological dynamic feature extraction module is used to perform inter-frame difference motion analysis on the time-domain multimodal image set, and extract the pupil change curve of the iris in the infrared image sequence and the frequency of micro-expression muscle tremors of the human face in the visible light image sequence. The physiological resonance feature generation module is used to time-domain align the pupil change curve with the muscle tremor frequency to generate in vivo physiological resonance feature values. The single-modal feature encoding module is used to perform wavelet transform encoding on the key frame iris image in the infrared image sequence to obtain the iris phase feature tensor when the live physiological resonance feature value exceeds the preset threshold, and to perform deep convolution feature extraction on the key frame face image in the visible light image sequence to obtain the face semantic feature tensor. The cross-modal feature fusion module is used to input the iris phase feature tensor and the face semantic feature tensor into the cross-modal joint encoder for feature-level heterogeneous fusion to generate a high-dimensional joint biometric feature embedding vector. The permission verification and unlocking module is used to map the permission identifier of the prescription issuer to the authorized biometric tensor base pre-stored in the security chip, and to measure the distance between the high-dimensional joint biometric embedding vector and the authorized biometric tensor base. If the measured distance is less than the intra-class aggregation threshold, a verification pass signal is generated.

Citation Information

Patent Citations

  • Prescription behavior supervision method based on biometric identification and related equipment

    CN109559793A

  • Drug taking detection method and device, storage medium and equipment

    CN119399804A

  • Informationized wound ticket endowing system based on facial recognition

    CN120387156A

  • Client identity real-time checking method based on image recognition

    CN120932282A

  • Face depth detection method and system based on multi-mode double shooting

    CN121096003A