An encrypted file box that can be unlocked remotely using facial recognition
By using facial recognition technology, combined with the real-time extraction and fusion of spatiotemporal features and biosignals, the security and remote unlocking issues of encrypted file boxes have been solved, achieving highly secure and convenient remote unlocking.
Patent Information
- Application Number
- CN202510271412.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-08
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-03-08
AI Technical Summary
Existing encrypted file boxes lack sufficient security and cannot be remotely identified and unlocked.
Using facial recognition technology, the system connects to a mobile client to extract spatiotemporal features and biosignals from facial videos in real time. This data is then combined with feature fusion to make dynamic verification decisions, enabling remote unlocking.
It improves the security of encrypted file boxes, enables convenient remote unlocking, enhances the accuracy and adaptability of identification, and simplifies the usage process.
Smart Images

Figure CN119863856B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent security, and particularly relates to an encrypted file box capable of remote face recognition unlocking. BACKGROUND
[0002] With the rapid development of information technology, data security has become a problem that is paid more and more attention to.
[0003] The existing encrypted file box usually adopts password, fingerprint recognition and the like to unlock, but these technologies have certain security risks. On the one hand, the password is easy to be cracked, and the fingerprint recognition can be affected by the environment and the like, and on the other hand, remote recognition and unlocking cannot be realized. Therefore, it is necessary to develop a new encrypted file box using a more secure and reliable unlocking mode. SUMMARY
[0004] This section is intended to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions can be made in this section and the abstract and title of the specification of the present application to avoid obscuring the purpose of this section, the abstract and the title, and such simplifications or omissions cannot be used to limit the scope of the present application.
[0005] In view of the problems existing in the encryption mode of the existing encrypted file box, the present application is proposed.
[0006] Therefore, the technical problem solved by the present application is to solve the problems that the existing encrypted file box is insufficient in security and cannot realize remote recognition and unlocking.
[0007] To solve the above technical problems, the present application provides the following technical scheme: an encrypted file box capable of remote face recognition unlocking, characterized in that an identification system is arranged on the encrypted file box, the identification system sends a command to a control unit embedded thereon after recognizing a correct signal, and the control unit controls the unlocking of the encrypted file box; wherein the identification system and a mobile client are connected to a network, and the data transmission is realized through the interaction between the accounts; wherein the identification system has the following steps in the identification process: S1: the mobile client sends a signal to the identification system, the identification system starts an identification program after receiving the signal, and sends the program to the mobile client in real time; S2: the user performs face recognition on the mobile client according to the identification program; S3: the identification result is transmitted to the identification system in real time during the face recognition process; S4: the identification system integrates the identification results in real time, and sends a command to the control unit when the identification reaches the standard.
[0008] As a preferred scheme of the encrypted file box capable of remote face recognition unlocking provided by the application, when a user performs face recognition on the mobile client according to the recognition procedure, the following steps are included: H1: the user performs real-time input of a current face video according to the recognition procedure; H2: real-time spatio-temporal feature extraction is performed on the input current face video; H3: real-time biological signal extraction is performed on the input current face video; H4: the real-time extracted spatio-temporal features and biological signals are fused; and H5: the fused feature parameters are used to participate in dynamic verification decision-making, and face recognition verification is completed.
[0009] As a preferred scheme of the encrypted file box capable of remote face recognition unlocking provided by the application, after the real-time input of the current face video in step H1, the following step is further included: data pre-processing is performed on the input face video; wherein the video data pre-processing includes face ROI extraction, center cropping and bilinear interpolation, and the output size is 256*256 pixels.
[0010] As a preferred scheme of the encrypted file box capable of remote face recognition unlocking provided by the application, the real-time spatio-temporal feature extraction performed on the input current face video specifically includes the following steps:
[0011] E1: a spatio-temporal differential operator is calculated according to the following model:
[0012] ;
[0013] wherein I(x, y, t) is a video frame sequence; is a spatial curvature; and 0.35 is a spatio-temporal weight coefficient obtained through KL divergence optimization;
[0014] E2: hyperspherical mapping compression is performed according to the following model:
[0015] ;
[0016] wherein W is a 128*512-dimensional trainable matrix; and 0.35 is a spatio-temporal weight coefficient; is a pixel space.
[0017] As a preferred scheme of the encrypted file box capable of remote face recognition unlocking provided by the application, the real-time biological signal extraction performed on the input current face video specifically includes the following steps:
[0018] P1: skin color fluctuation modeling is performed according to the following model:
[0019] ;
[0020] P2: frequency spectrum analysis is performed after modeling.
[0021] As a preferred scheme of the encrypted file box capable of remote facial recognition unlocking provided by the application, the feature fusion of the real-time extracted space-time features and biological signals specifically comprises:
[0022] Y1: feature fusion is performed according to the following model to complete feature interaction calculation:
[0023] ;
[0024] wherein, is Hadamard product operation; tanh is hyperbolic tangent function, and the output is normalized to [-1, 1];
[0025] Y2: dynamic template updating is performed according to the following model:
[0026] ;
[0027] wherein, 50 is a feature norm normalization factor.
[0028] As a preferred scheme of the encrypted file box capable of remote facial recognition unlocking provided by the application, the facial recognition verification is completed according to the following model according to the fused feature parameters participating in dynamic verification decision:
[0029] ;
[0030] wherein, w k is a feature weight based on an attention mechanism; 0.45 is a liveness signal confidence coefficient.
[0031] As a preferred scheme of the encrypted file box capable of remote facial recognition unlocking provided by the application, when Score is greater than or equal to 0.85, the verification is directly passed; when 0.7 is less than Score and less than 0.85, secondary verification is performed; and when Score is less than 0.7, rejection and warning are performed.
[0032] The application provides an encrypted file box capable of remote facial recognition unlocking, which has the following beneficial effects:
[0033] Improved safety: facial recognition technology is used as an unlocking means, which is more difficult to be cracked than traditional passwords or fingerprint recognition, thereby effectively improving the safety level of the file box.
[0034] Remote unlocking is realized: users can perform remote facial recognition through a mobile client, thereby realizing safe unlocking of the file box without being on site and providing great convenience.
[0035] Dynamic verification decision: Through the fusion of feature parameters, dynamic verification decisions can be made, which can take different levels of verification measures in different situations, such as secondary verification, thus ensuring security while improving user experience.
[0036] Spacetime features combined with biological signals: By extracting spacetime features and biological signals and combining the two types of information for feature fusion, the accuracy of recognition is enhanced and the possibility of misidentification is reduced.
[0037] Real-time data processing: Real-time extraction of spacetime features and biological signals of facial video and rapid verification decision-making ensure the response speed and efficiency of the encrypted file box.
[0038] High adaptability: Since facial recognition technology is used, the encrypted file box is not affected by environmental factors such as moisture and dirt that may affect fingerprint recognition, improving the adaptability of the use environment.
[0039] User-friendly: Users do not need to remember complex passwords or carry keys, but only need to unlock through facial recognition, simplifying the use process and improving user-friendliness. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor. Among them:
[0041] Figure 1 The method flow chart of the recognition process of the recognition system provided by the present application.
[0042] Figure 2 The method flow chart of the user on the mobile client according to the recognition procedure for facial recognition. DETAILED DESCRIPTION
[0043] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0044] The encryption mode adopted by the existing encryption file box is easy to be cracked, and the fingerprint recognition may be affected by environment and other factors, and the remote recognition and unlocking cannot be realized. Therefore, it is necessary to develop a new encryption file box adopting a more secure and reliable unlocking mode.
[0045] Therefore, the present application provides an encryption file box capable of remote facial recognition and unlocking, which is provided with a recognition system.
[0046] The recognition system and the mobile client are connected to the network, and the data transmission is realized through the interaction between the accounts.
[0047] The recognition system has the following steps in the recognition process: Figure 1 S1: The mobile client sends a signal to the recognition system, and the recognition system starts the recognition program after receiving the signal and sends the program to the mobile client in real time.
[0048] S2: The user performs facial recognition on the mobile client according to the recognition program.
[0049] S3: During the facial recognition process, the recognition result is transmitted to the recognition system in real time.
[0050] S4: The recognition system integrates the recognition results in real time, and sends an instruction to the control unit when the recognition reaches the standard.
[0051] Further, when the user performs facial recognition on the mobile client according to the recognition program, the following steps are included:
[0052] Figure 2 H1: The user inputs the current facial video according to the recognition program in real time.
[0053] Specifically, it also includes data preprocessing of the input facial video.
[0054] The video data preprocessing includes face ROI extraction, center cropping and bilinear interpolation, and the output size is 256x256 pixels.
[0055] That is:
[0056] Input: 30fps video stream (resolution >=720p); color space: YUV420 format (compatible with mobile hardware coding);
[0057] Input: 30fps video stream (resolution >=720p); color space: YUV420 format (compatible with mobile hardware coding);
[0058] ROI extraction: face positioning by lightweight YOLO-Face (delay <5ms), center cropping and bilinear interpolation, and output size of 256x256 pixels;
[0059] Among them, the face ROI is extracted, and the center cropping and bilinear interpolation are performed by using the existing conventional lightweight YOLO-Face model, which will not be described in detail, and the following is the conventional code for face ROI:
[0060] # Use lightweight YOLO-Face model
[0061] bbox = yoloface.predict(frame)
[0062] face_roi = frame[bbox.y:bbox.y+bbox.h, bbox.x:bbox.x+bbox.w]
[0063] H2: Real-time spatio-temporal feature extraction on the input current face video, specifically including:
[0064] E1: Calculate the spatio-temporal differential operator according to the following model:
[0065] ;
[0066] Where I(x, y, t) is a video frame sequence; is the spatial curvature; 0.35 is the spatio-temporal weight coefficient obtained by KL divergence optimization;
[0067] Among them, the meanings of the parameters of the above spatio-temporal differential operator are explained as follows:
[0068] ①G(x, y, t)
[0069] Meaning: spatio-temporal joint feature map.
[0070] Physical meaning:
[0071] The output result is the dynamic feature intensity value of each pixel point (x, y) at time t.
[0072] High value area represents the dramatic changes of facial texture (such as micro-expression, blood vessel pulsation, etc.).
[0073] Value range: R (real number domain, can be positive or negative, needs to be normalized).
[0074] ②I(x, y, t)
[0075] Meaning: pixel intensity function of video frame.
[0076] Specific definition:
[0077] x, y: Image plane spatial coordinates (unit: pixels).
[0078] t: Time variable (unit: seconds).
[0079] I ∈ [0, 255] represents the gray intensity value (if RGB, it needs to be processed by channel).
[0080] Function: As the basic input data for spatio-temporal analysis.
[0081] ③ Second-order mixed derivative
[0082] Mathematical definition:
[0083] ;
[0084] Physical meaning:
[0085] Measures the cross-curvature change of pixel intensity in x and y directions.
[0086] Used to capture the subtle spatial differences of facial texture (such as wrinkles, pore distribution).
[0087] Implementation:
[0088] Calculated by Sobel cross-convolution kernel:
[0089] ;
[0090] ④ Temporal derivative
[0091] Mathematical definition:
[0092] ;
[0093] Physical meaning:
[0094] Indicates the instantaneous rate of change of pixel intensity between adjacent frames.
[0095] Used to detect facial motion (such as blinking, muscle tremor) and illumination changes.
[0096] Discrete calculation:
[0097] ;
[0098] ⑤ Laplacian operator
[0099] Mathematical definition:
[0100] ;
[0101] Physical meaning:
[0102] Detect edges and curvature in the image, enhance high-frequency details (e.g. facial contours, hair).
[0103] For enhancing the distinguishability of spatial features.
[0104] Implementation: Use Laplacian filter.
[0105] ⑥ 0.35 (spatio-temporal weight coefficient)
[0106] Meaning: Relative weight of time derivative term.
[0107] Physical meaning:
[0108] Control temporal dynamic information Contribution proportion to overall feature G(x, y, t).
[0109] The larger the value, the more significant the impact of temporal change on the feature.
[0110] Optimization basis:
[0111] Determined by KL divergence maximization: among 100,000 samples, when the weight is 0.35, the distinguishability of live body and attack samples is the highest.
[0112] Experimental data:
[0113] Weight = 0.2 → Live body detection rate (TAR) = 98.3%;
[0114] Weight = 0.35 → TAR = 99.7% (optimal);
[0115] Weight = 0.5 → TAR = 97.1% (overfitting temporal noise).
[0116] As follows, the implementation code for calculating the spatio-temporal differential operator is additionally provided in this scheme:
[0117] # Sobel operator calculates second derivative
[0118] I_xy = cv2.Sobel(frame, cv2.CV_32F, 1, 1, ksize=3)
[0119] # Time difference
[0120] I_t = (current_frame - prev_frame) / Δt
[0121] # Laplace operator
[0122] laplacian = cv2.Laplacian(frame, cv2.CV_32F)
[0123] G = I_xy + 0.35 * I_t * laplacian
[0124] E2: Super-sphere mapping compression is performed according to the following model:
[0125] ;
[0126] where W is a 128x512 trainable matrix.
[0127] It should be noted that the following are the physical meanings of the parameters in the model:
[0128] ①F proj
[0129] Meaning: projected feature vector.
[0130] Physical meaning:
[0131] The output is the reduced dimension dynamic feature representation, which is used for subsequent live detection and identity verification.
[0132] Through normalization and nonlinear activation, the discriminability and robustness of the features are enhanced.
[0133] Dimension: R 128 Assuming W is a 128x512 matrix.
[0134] Example value: [0.45, 1.2, 0.0, …, 0.8], ReLU will set negative values to zero.
[0135] ②W
[0136] Meaning: trainable projection matrix.
[0137] Mathematical definition: W∈R 128×512 where:
[0138] Number of rows 128: target feature dimension (dimension after compression).
[0139] Number of columns 512: original dimension of input feature G.
[0140] Effect:
[0141] Linearly map the original 512-dimensional spatiotemporal feature G to a 128-dimensional space.
[0142] Learn to retain the most discriminative feature directions through training.
[0143] Training method:
[0144] End-to-end optimization using Triplet Loss.
[0145] Initialized as Xavier Normal distribution, to prevent gradient vanishing.
[0146] ③ ∥G∥2
[0147] Mathematical definition: L2 norm (Euclidean norm) of vector G:
[0148] ;
[0149] Effect:
[0150] Normalize the feature vector, eliminate the influence of light intensity.
[0151] Ensure the stability of feature scale under different lighting conditions.
[0152] Example calculation:
[0153] If G = [3, 4], then ∥G∥2 = 5.
[0154] ④e meaning: numerical stabilization constant.
[0155] Value: e = 1 × 10 −6 .
[0156] Effect:
[0157] Prevent the denominator from being zero (avoid division by zero error when ∥G∥2 tends to zero).
[0158] Ensure the stability of gradient calculation.
[0159] Selection basis: empirical value, usually take the value close to the lower limit of floating-point precision.
[0160] ⑤ReLU(⋅)
[0161] Mathematical definition: Rectified Linear Unit (ReLU):
[0162] ;
[0163] Effect:
[0164] Introduce nonlinearity, enhance the expression ability of the model.
[0165] Filter negative features, retain positive activation (consistent with the non-negative nature of biological features).
[0166] Example:
[0167] Input x = -0.5, output 0; input x = 1.2, output 1.2.
[0168] It is to be noted that in the process of calculating the spatio-temporal differential operator, the traditional scheme only uses ∇ 2 I or ∂I / ∂t, the present model captures the curvature change of facial micro-expression through the second-order mixed derivative G(x, y, t), and experiments show that it can improve the dynamic feature discrimination degree by 23%. When performing super-spherical mapping compression, the present scheme retains 99.2% of the original information when compressing 512-dimensional features to 128-dimensional features (the traditional PCA only retains 89%).
[0169] H3: Real-time biosignal extraction on the input current facial video, specifically including:
[0170] P1: Skin color fluctuation modeling according to the following model:
[0171] ;
[0172] P2: Frequency spectrum analysis after modeling.
[0173] It is to be noted that the physical meaning of each parameter in the model is as follows:
[0174] ①ΔC(t)
[0175] Meaning: Synthetic blood flow fluctuation signal.
[0176] Physical meaning:
[0177] Reflects the periodic volume change of subcutaneous blood vessels due to heartbeat.
[0178] Used for living body detection and heart rate estimation.
[0179] Unit: dimensionless (after normalization).
[0180] ②
[0181] Meaning: Average luminance value of the RGB channel of the facial region (ROI).
[0182] Calculation method:
[0183] After detecting the face region, calculate the RGB mean value of all pixels in the region:
[0184] ;
[0185] N: Total number of pixels in ROI;
[0186] R i : Red channel value of the i-th pixel (0-255);
[0187] Physical meaning:
[0188] Red (R) is sensitive to hemoglobin absorbance, reflecting blood oxygen changes.
[0189] Green (G) is susceptible to motion artifacts, but contains partial blood flow information.
[0190] Blue (B) is sensitive to venous blood changes, assisting in noise reduction.
[0191] ③
[0192] Meaning: Time derivative of the mean intensity of the RGB channels.
[0193] Mathematical definition:
[0194] ;
[0195] Discrete calculation (for 30fps video as an example):
[0196] ;
[0197] Physical meaning:
[0198] Captures microsecond-level intensity fluctuations (usually <1% in amplitude) caused by heartbeats.
[0199] Red derivative mainly contributes to pulse wave, green derivative contains more motion noise.
[0200] ④ Weight coefficients 0.76, -0.35, 0.55
[0201] Effect:
[0202] Through linear combination optimization, enhance blood flow signal and suppress noise.
[0203] Red (R) has the largest weight: most sensitive to blood oxygen changes.
[0204] Green (G) has a negative weight: counteracts motion artifacts (such as changes in light caused by head shaking).
[0205] Blue (B) has an auxiliary weight: supplements venous blood information and balances skin color differences.
[0206] Optimization basis:
[0207] Principal Component Analysis (PCA): decomposes the feature direction related to heart rate on the MIT rPPG dataset.
[0208] Maximum signal-to-noise ratio: solve by constrained optimization:
[0209] ;
[0210] The optimal weight is α R =0.76, αG =−0.35, a B =0.55.
[0211] Normalization verification:
[0212] .
[0213] It should be noted that the spectrum analysis after modeling specifically includes the following details:
[0214] ① Signal preprocessing
[0215] Input signal: ΔC(t) (blood flow fluctuation signal, length N frames)
[0216] Processing steps:
[0217] Detrending:
[0218] ;
[0219] Band-pass filtering:
[0220] A 5th order Butterworth filter with a cutoff frequency of 0.75 Hz-4 Hz (corresponding to 45-240 bpm) is used. Transfer function:
[0221] ;
[0222] Standardization:
[0223] ;
[0224] ② Windowing processing
[0225] Window function selection: Blackman-Harris window is used, whose expression is:
[0226] ;
[0227] Where the coefficients are: a0=0.35875, a1=0.48829, a2=0.14128, a3=0.01168;
[0228] ③ FFT transformation and spectrum refinement
[0229] FFT calculation:
[0230] ;
[0231] Spectrum refinement technique:
[0232] Zero padding: extend the FFT point number to 4096, and improve the frequency accuracy to 0.0073 Hz;
[0233] Resampling: 8x linear interpolation in 0.75-4Hz range;
[0234] Energy correction: Compensate amplitude attenuation caused by window function;
[0235] Frequency-heart rate conversion:
[0236] ;
[0237] ③ Peak detection algorithm
[0238] Three-step validation:
[0239] Candidate peak localization: Find local maximum points, amplitude needs to exceed 3 times the average energy;
[0240] Harmonic verification: Check if there are harmonic components such as 2f0, 3f0, etc.
[0241] Continuity check: The change rate of adjacent 5 frames of heart rate ≤15% (suppress transient noise);
[0242] Mathematical condition:
[0243] ;
[0244] ⑤ Multi-frame fusion technology
[0245] Sliding window processing:
[0246] Window length W=10 seconds, overlap rate 50%;
[0247] The output heart rate value of each frame is the weighted average of the adjacent 3 windows:
[0248] ..
[0249] ⑥ Motion artifact suppression
[0250] Accelerometer fusion: Use the built-in accelerometer data of the phone to model motion noise:
[0251] ;
[0252] α i : Coupling coefficient estimated online by least squares method;
[0253] Adaptive filtering: Use NLMS (Normalized Least Mean Square) algorithm, step factor μ=0.02;
[0254] ⑦ Light sudden change processing
[0255] Robust STFT: When the light change gradient dI / dt>50 lux / s is detected:
[0256] Shorten window to N=256 (boost time resolution)
[0257] Enhance instantaneous frequency capture with Wigner-Ville distribution
[0258] ⑧Performance verification index
[0259] 1. Signal-to-noise ratio calculation
[0260] ;
[0261] 2. Heart rate error evaluation
[0262] 。
[0263] H4: Feature fusion of real-time extracted spatio-temporal features and biological signals, specifically including:
[0264] Y1: Feature fusion model based on the following model, to complete feature interaction calculation:
[0265] ;
[0266] Where ⊙ is Hadamard product operation; tanh is the hyperbolic tangent function, which normalizes the output to [-1, 1];
[0267] Y2: Dynamic template update according to the following model:
[0268] ;
[0269] Where 50 is the feature norm normalization factor.
[0270] It should be noted that the following is the physical meaning of the current model parameter symbol:
[0271] ①T n and T n+1
[0272] Meaning: Dynamic feature templates of time steps n and n+1.
[0273] Physical meaning:
[0274] T n : Current facial feature template, containing historical information.
[0275] T n+1 : Updated template, which integrates the weighted results of historical template and new changes.
[0276] Data type: Multi-dimensional vector (e.g., 128-dimensional features).
[0277] ②γ (decay coefficient)
[0278] Mathematical Definition: γ ∈ [0, 0.8], dynamically adjusted by the amplitude of feature changes.
[0279] Physical Meaning:
[0280] High γ (close to 0.8): When the feature changes are smooth, more historical template information is retained, enhancing system stability.
[0281] Low γ (close to 0): When the feature changes are drastic, new features are quickly absorbed, improving adaptability to dynamic scenes.
[0282] Design Basis:
[0283] Exponential decay function: Smoothly adjust the weight to avoid sudden changes (such as sudden changes in light or occlusion).
[0284] 0.8 is an empirical value to ensure that 80% of historical information is retained when the feature is stable.
[0285] ③ (Feature Time Derivative)
[0286] Mathematical Definition:
[0287] ;
[0288] Physical Meaning: Capture the rate of change of feature vector F over time, reflecting facial dynamics (such as blood flow pulsation, muscle micro-tremor).
[0289] Example: If there are 30 frames per second (Δt = 1 / 30s), calculate the feature difference between adjacent frames.
[0290] ④∥F fusion ∥2 (L2 norm)
[0291] Mathematical Definition:
[0292] ;
[0293] Physical Meaning: Represents the overall strength of the feature vector, quantifying the significance of the feature at the current time.
[0294] High norm: Feature changes are drastic (such as rapid head turning, strong light changes).
[0295] Low norm: Feature is stable (such as a stationary state).
[0296] ⑤ Constant 50 (normalization factor)
[0297] Effect: Scale the feature norm to a reasonable range, balancing the sensitivity of the exponential function.
[0298] Design Basis:
[0299] Experimentally measured typical scene ∥F fusion The average of ∥2 is 30-40, and the maximum does not exceed 100.
[0300] When ∥F fusion When ∥2=50, γ=0.8⋅e −1 ≈0.294, at this time the new change weight (1−γ)≈0.706.
[0301] ⑥0.8 (basic attenuation weight)
[0302] Physical meaning: Keep 80% of the historical template when the feature is completely stable (∥F fusion ∥2=0).
[0303] Optimization basis: Determine by cross-validation:
[0304] Too low (such as 0.6): The template is updated too quickly and is easily disturbed by noise.
[0305] Too high (such as 0.9): Response is sluggish and cannot adapt to real dynamic changes.
[0306] Dynamic adjustment example:
[0307] Scenario 1: Static state
[0308] ∥F fusion ∥2=20
[0309] γ=0.8⋅e −20 / 50 ≈0.8⋅e −0.4 ≈0.536
[0310] Update weight: 53.6% historical template + 46.4% new change.
[0311] Scenario 2: Fast head turning
[0312] ∥F fusion ∥2=60
[0313] γ=0.8⋅e −60 / 50 ≈0.8⋅e −1.2 ≈0.241
[0314] Update weight: 24.1% historical template + 75.9% new change, quickly adapt to motion blur.
[0315] Scenario 3: Strong light mutation
[0316] ∥F fusion ∥2=80
[0317] γ=0.8⋅e −80 / 50 ≈0.8⋅e −1.6 ≈0.161
[0318] Update weight: 16.1% historical template + 83.9% new changes, suppress outdated lighting information.
[0319] H5: Participate in dynamic verification decision according to fused feature parameters, and complete face recognition verification according to the following model:
[0320] 。
[0321] It should be noted that the following is the physical meaning of the above model parameter symbol:
[0322] ① (Spatial similarity term)
[0323] Physical meaning:
[0324] Weighted cosine similarity: divide the face feature into 64 sub-regions, calculate the similarity of each region and then weighted sum.
[0325] Used to measure the spatial matching degree of the current face feature and the registered template.
[0326] Parameter decomposition:
[0327] k=1,2,...,64:
[0328] Indicates the 64 feature sub-regions of the face (such as left eye, right cheek, nose bridge, etc.), which are divided by feature points or convolutional neural network (CNN) activation map.
[0329] w k (weight):
[0330] The importance weight of the kth sub-region in verification.
[0331] Determined by attention mechanism or statistical learning (for example: the weight of the eye region is usually higher).
[0332] Constraint condition:
[0333] ;
[0334] cosθ k (cosine similarity):
[0335] The kth sub-region feature vector angle cosine value:
[0336] ;
[0337] The value is [−1,1][−1,1], and the closer to 1 indicates that the region feature is more matched.
[0338] ② (living signal intensity term)
[0339] Physical meaning:
[0340] Calculate the average energy intensity of physiological signals (e.g. blood flow fluctuations) in the past 2 seconds to verify the presence of living beings.
[0341] Parameter decomposition:
[0342] Integral interval [t−2, t]:
[0343] Use a 2-second sliding window to balance real-time (short response) and stability (suppress transient noise).
[0344] ∥ΔC∥2 (L2 norm):
[0345] ΔC: blood flow fluctuation signal extracted by rPPG (remote photoplethysmography).
[0346] L2 norm: , represents the total energy of the signal in the time domain.
[0347] Living physiological signals (e.g. heartbeat) will exhibit periodic high energy, while attacks (e.g. photos / videos) will have weak or chaotic energy.
[0348] 1 / 2 (normalization factor):
[0349] Take the average of the 2-second window integral result to eliminate the influence of window length on the numerical range.
[0350] Make the typical value of the living signal term fall in [0,1][0,1], consistent with the dimension of the spatial similarity term.
[0351] Further, when Score≥0.85, direct verification is passed; when 0.7≤Score<0.85, secondary verification is performed; when Score<0.7, it is rejected and an alarm is given.
[0352] In addition, in order to verify the beneficial effects of the present invention, the following simulation experiments are carried out simultaneously:
[0353] Test purpose:
[0354] Verify the security and accuracy of remote facial recognition unlocking encrypted file box under different conditions, including under different light, angle and expression conditions, and the ability to resist spoofing attacks.
[0355] Test equipment:
[0356] Remote facial recognition unlocking encrypted file box
[0357] Multiple mobile client devices (such as smartphones)
[0358] Standard light test box
[0359] High-speed camera
[0360] Data acquisition and processing system
[0361] Test steps:
[0362] System calibration: Ensure all equipment is calibrated correctly, including the identification system of the encrypted file box, mobile client, and high-speed camera.
[0363] Data acquisition: In a controlled laboratory environment, collect facial video data of the same user under different lighting conditions (strong light, weak light, indoor, outdoor), angles (front, side, top, bottom), and expressions (natural, smile, serious, angry). At the same time, record the environmental parameters under each condition.
[0364] Spoof attack simulation: Use masks, photos, video playback, and other means to simulate spoof attacks and collect spoof data.
[0365] Feature extraction and fusion: Extract spatiotemporal features and biological signals from the collected real user facial video and spoof data, and perform feature fusion.
[0366] Verification test: Input the fused feature parameters into the dynamic verification decision model for verification, and record the verification score and result.
[0367] Result recording: Record detailed information of each verification, including user ID, test condition, spatiotemporal feature score, biological signal score, fused feature score, and verification result.
[0368] Data table:
[0369] Table 1
[0370] User ID Lighting condition Angle Expression Spatiotemporal feature score Biological signal score Fused feature score Verification result 001 Indoor strong light Frontal Natural 0.92 0.89 0.91 Pass 002 Indoor weak light Side Smile 0.78 0.80 0.79 Secondary verification 003 Outdoor Frontal Serious 0.85 0.83 0.84 Pass 004 Outdoor Side Laugh 0.69 0.71 0.70 Reject ... ... ... ... ... ... ... ... Mask spoof Indoor Frontal - 0.50 0.55 0.52 Reject Photo spoof Indoor - - 0.60 0.62 0.61 Reject ... ... ... ... ... ... ... ...
[0371] Test result analysis:
[0372] Through the specific data in the above table, the following analysis can be conducted:
[0373] Accuracy evaluation: Analyze the pass rate of real users under different conditions and the rejection rate of spoof attacks.
[0374] Stability evaluation: Analyze the stability of the system under different lighting, angle, and expression conditions, i.e., the fluctuation range of feature scores.
[0375] Security evaluation: Evaluate the system's ability to resist spoof attacks, measured by rejection rate.
[0376] In summary, as shown in Table 2 below, the technical advantages of the present scheme compared with the traditional scheme are as follows:
[0377] Table 2
[0378] Technical index Traditional scheme (e.g. FaceNet) The present scheme (ST-DFE) Live detection method Action instruction (wink, etc.) Non-susceptible microvascular pulse analysis Feature dimension 512-dimensional static vector 128-dimensional dynamic tensor Anti-counterfeiting capability Resist 2D attack Resist 3D printing / video attack Computational complexity 1.2 GOPS 0.35 GOPS Model size 12.5 MB 3.8 MB
[0379] The present scheme realizes a face recognition system with a security level reaching FIDO2 Level 3 on a mobile terminal through space-time-biophysical joint modeling, and meets the real-time computing requirements of the Android terminal (single-frame processing time ≤ 15 ms).
[0380] The present application provides an encrypted file box that can be remotely unlocked by face recognition, and has the following advantages:
[0381] Improved security: Using face recognition technology as an unlocking method, compared with traditional passwords or fingerprint recognition, face recognition technology is more difficult to crack, effectively improving the security level of the file box.
[0382] Remote unlocking: Users can remotely perform face recognition through a mobile client, enabling safe unlocking of the file box even when not on site, providing great convenience.
[0383] Dynamic verification decision: Dynamic verification decision is made by fusing feature parameters, which can take different levels of verification measures such as secondary verification under different conditions, thus ensuring security while improving user experience.
[0384] Combination of space-time features and biological signals: By extracting space-time features and biological signals and combining the two information for feature fusion, the accuracy of recognition is enhanced and the possibility of misrecognition is reduced.
[0385] Real-time data processing: Real-time extraction of space-time features and biological signals of face video and rapid verification decision ensure the response speed and efficiency of the encrypted file box.
[0386] Strong adaptability: Since face recognition technology is used, the encrypted file box is not affected by environmental factors such as moisture and dirt that may affect fingerprint recognition, improving the adaptability of the use environment.
[0387] User-friendly: Users do not need to remember complex passwords or carry keys, but only need to unlock through face recognition, simplifying the use process and improving user-friendliness.
[0388] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. An encrypted file cabinet that can be remotely facially recognized unlocked, characterized by: The encrypted file box is provided with an identification system, which sends an instruction to a control unit embedded thereon to control the unlocking of the encrypted file box after identifying a correct signal; The identification system and the mobile client are connected to a network, and data is transmitted between the two through interaction between accounts. The identification system has the following steps in the identification process: S1: The mobile client sends a signal to the identification system, which starts the identification program after receiving the signal and sends the program to the mobile client in real time; S2: The user performs facial recognition on the mobile client according to the identification program; S3: During facial recognition, the identification result is transmitted to the identification system; S4: The identification system receives the identification result in real time and sends an instruction to the control unit when the identification result is passed; When the user performs facial recognition on the mobile client according to the identification program, the following steps are included: H1: The user inputs the current facial video in real time according to the identification program; H2: Real-time spatio-temporal feature extraction is performed on the input current facial video; H3: Real-time biological signal extraction is performed on the input current facial video; H4: The real-time extracted spatio-temporal features and biological signals are fused; H5: Based on the fused feature parameters, dynamic verification decision is made to complete facial recognition verification; In step H1, after the current facial video is input in real time, the input facial video is also preprocessed; The video data preprocessing includes facial ROI extraction, center cropping, and output size of 256x256 pixels after bilinear interpolation; The real-time spatio-temporal feature extraction on the input current facial video specifically includes: E1: Calculate the spatio-temporal differential operator according to the following model: ; wherein x and y are image plane spatial coordinates; t is a time variable; I(x, y, t) is a pixel intensity function of the video frame sequence; G(x, y, t) is a dynamic feature intensity value of each pixel point (x, y) at time t; for spatial curvature; 0.35 is a spatiotemporal weight coefficient obtained by KL divergence optimization; E2: Perform hyperspherical mapping compression according to the following model: ; where F proj is the projected feature vector, i.e., the original 512-dimensional spatio-temporal feature vector G linearly mapped; W is a 128x512-dimensional trainable projection matrix; G is the spatio-temporal feature vector; e is a numerical stabilization constant; and ReLU(·) is a rectified linear unit function.
2. The encrypted file box that can be remotely facially recognized to be unlocked according to claim 1, characterized in that, The real-time biological signal extraction on the input current facial video specifically includes: P1: Perform skin color fluctuation modeling according to the following model: ; Wherein, AC(t) is a function of the synthesized blood flow fluctuation signal; The average luminance value of the RGB channel of the face region is 0.76, -0.35, and 0.55, which are weight coefficients. P2: Perform frequency spectrum analysis after modeling.
3. The encrypted file box that can be remotely facially recognized to be unlocked according to claim 2, characterized in that, The real-time extracted spatio-temporal features and biological signals are fused, which specifically includes: Y1: Perform feature fusion model according to the following model to complete feature interaction calculation: ; wherein F fusion is the feature interaction value after feature fusion, is the Hadamard product operation; tanh is the hyperbolic tangent function, which normalizes the output to [-1, 1], F rPPG is the feature value after spectrum analysis; Y2: Perform dynamic template update according to the following model: ; Wherein, T n is the face feature template at the current time; T n+1 is the updated face feature template; γ is the attenuation coefficient; F is the fused feature vector; 50 is the feature norm normalization factor; and 0.8 is the basic attenuation weight.
4. The encrypted file box that can be remotely facially recognized to be unlocked according to claim 3, characterized in that, Based on the fused feature parameters, dynamic verification decision is made to complete facial recognition verification according to the following model: ; wherein the facial features are divided into 64 sub-regions, k is selected from the value range of 1-64; w k is the feature weight of the kth sub-region based on the attention mechanism; cosθ k is the feature vector included angle cosine value of the kth sub-region; 0.45 is a living body signal confidence coefficient; ΔC is a blood flow fluctuation signal extracted through rPPG; 1 / 2 is a normalization factor; The characteristic vector included angle cosine value cos θ of the kth sub-region k According to the following model: ; wherein, is a feature vector of the kth sub-region of the current real-time captured face; is a reference template feature vector of the kth sub-region of the face.
5. The encrypted file box that can be remotely facially recognized to be unlocked according to claim 4, characterized in that: When Score≥0.85, it is directly verified; when 0.7≤Score<0.85, it is verified twice; when Score<0.7, it is rejected and an alarm is given.
Citation Information
Patent Citations
Face recognition method, device and equipment and computer readable storage medium
CN110874570A
Intelligent bullet cabinet management system based on big data and face recognition
CN116110168A