Face recognition type endemic area access control system

Through the face recognition ward access control system, the image motion blur is evaluated in real time, the full face and periophthalmic feature vectors are extracted, and the problem of misrecognition of face recognition in dynamic environments is solved, and the accuracy and security of face recognition in medical scenarios are improved.

CN120564302AInactive Publication Date: 2025-08-29JIANGSU PROVINCE HOSPITAL (THE FIRST AFFILIATED HOSPITAL OF NANJING MEDICAL UNIVERSITY)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510706417.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In medical scenarios, the facial recognition system is susceptible to dynamic environments, resulting in motion blur and misrecognition, and lacks an effective image quality screening mechanism, making it difficult to obtain effective feature data in low-light or backlight environments, and there are security loopholes.

Method used

The face recognition ward access control system is adopted to screen qualified images through image motion blur assessment, extract the full face and periophthalmic feature vectors, combine the orthogonal projection and standardization processing of multi-dimensional feature space, and perform double judgments to generate fused face recognition features, and use a two-way interval comparison mechanism in the authorization threshold determination.

Benefits of technology

It improves the accuracy of feature extraction, enhances the recognition accuracy of camouflage and forged faces, reduces the misjudgment rate, improves the adaptability and security of the system, and ensures the reliability and timeliness of authority judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564302A_ABST
    Figure CN120564302A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of face recognition, in particular to a face recognition type endemic area access control system, which comprises a face image acquisition module for starting an image sensor for endemic area access control to capture a face image sequence based on personnel in a visual field, carrying out image motion ambiguity value evaluation and screening a single-frame image to obtain a qualified face image. According to the method, the face image sequence is collected in real time, the motion ambiguity value is dynamically evaluated, low-quality images caused by personnel movement or environment interference are effectively eliminated through a single-frame screening mechanism, and the accuracy of subsequent feature extraction is improved. According to the method, double judgment logic of a full face feature vector and an unauthorized feature space boundary distance is adopted, and orthogonal projection and standardization processing of a multi-dimensional feature space are combined, so that the recognition precision of abnormal access behaviors such as disguise and forgery is enhanced, and the misjudgment rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of face recognition technology, and in particular to a face recognition-based ward access control system. Background Art

[0002] The facial recognition ward access control system is an intelligent access control system for medical areas developed based on facial recognition technology. It is used for access management of sensitive areas such as hospital wards and isolation wards.

[0003] Existing technologies are highly dependent on image quality in medical applications. In dynamic environments, motion blur can easily occur due to walking or equipment vibrations, leading to biased feature extraction and misidentification. Furthermore, single-shot image processing is often used, and effective front-end image quality screening mechanisms are not established. This makes it difficult to obtain valid feature data in low-light or backlit environments. The feature comparison process often relies on fixed thresholds and lacks the ability to dynamically model unauthorized features. This creates security vulnerabilities when facing intentionally forged or similar faces. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings of the existing technology and propose a face recognition type ward access control system.

[0005] In order to achieve the above objectives, the present invention adopts the following technical solutions: The face recognition ward access control system includes:

[0006] The facial image acquisition module activates the image sensor of the ward access control to capture facial image sequences based on the people in the field of view, evaluates the image motion blur value, and screens single-frame images to obtain qualified facial images;

[0007] The forbidden feature comparison module extracts the corresponding full-face feature vector value based on the qualified face image, establishes the current face feature vector, calculates the distance value to the preset unauthorized feature space boundary based on the current face feature vector, and compares it with the preset exclusion judgment threshold to obtain the person exclusion status;

[0008] A multi-region feature fusion module extracts a full-face region feature vector set based on the qualified face image and simultaneously locates and crops the periocular region image block to extract the periocular region feature vector set, establishes a full-face and periocular region dual feature vector set, and based on the full-face and periocular region dual feature vector set, splices the two vector sets within the group to obtain a fused face recognition feature;

[0009] The access control authority judgment module, based on the personnel exclusion status, determines that if the Boolean value is true, it outputs the access denial instruction text and establishes a preliminary access instruction. Based on the preliminary access instruction, if the text content is not rejection, the fused facial recognition feature is compared and retrieved with the authorized personnel feature library and an access control pass instruction is generated based on the relationship between the comparison score and the authorization threshold.

[0010] Preferably, the steps of obtaining the qualified face image are:

[0011] Based on the facial image sequence captured by the image sensor, the grayscale matrix data of each frame is extracted, and a two-dimensional discrete Fourier transform is performed on the grayscale matrix of each frame to generate a set of frequency domain components of each frame;

[0012] Calculate the motion blur value based on the frequency domain component set of each frame;

[0013] Based on the motion blur values ​​of all frames, the mean and standard deviation are calculated, and the mean minus the standard deviation is used as the dynamic screening threshold. The motion blur values ​​of all frames are traversed, and frames with motion blur values ​​less than or equal to the dynamic screening threshold are selected to generate qualified face images.

[0014] Preferably, the steps of obtaining the current facial feature vector are:

[0015] Based on the qualified face image, geometric correction is performed on the face area, facial key points are aligned to preset positions in a standard coordinate system, rotation and scaling differences are eliminated, and histogram equalization is applied to eliminate lighting interference, thereby generating full-face image data after geometric and lighting normalization;

[0016] Based on the geometrically and illumination-normalized full-face image data, hierarchical feature extraction is performed on the full-face image data using a multi-scale pyramid structure. Three convolution kernels, 3×3, 5×5, and 7×7, are used to scan the full-face image data, locate facial key points, and extract a high-dimensional texture descriptor fused with a histogram of oriented gradients and a local binary pattern based on a 16×16 pixel region around each key point to generate a set of high-dimensional texture descriptors for the full-face key points.

[0017] Based on the high-dimensional texture descriptor set of all facial key points, principal component analysis is used to independently reduce the dimensionality of the high-dimensional texture descriptor of each key point, and then the descriptors of all key points are spliced ​​through a fully connected network layer to establish the current facial feature vector.

[0018] Preferably, the steps for obtaining the personnel exclusion status are:

[0019] Based on the current facial feature vector, performing spectral decomposition on the covariance matrix of the preset unauthorized feature space, extracting eigenvalues ​​and orthogonal basis vectors, and performing normalized projection on the current facial feature vector along the orthogonal basis direction to generate a normalized projection feature component set;

[0020] Calculating a dynamic distance value to the boundary of the unauthorized feature space based on the set of standardized projected feature components;

[0021] Based on the dynamic distance value, the dynamic distance value is compared with the preset exclusion judgment threshold. If the logarithmic transformation value of the dynamic distance value is greater than or equal to 1.5 times the preset exclusion judgment threshold, the personnel exclusion status is determined to be true, otherwise it is false, and the personnel exclusion status is generated.

[0022] Preferably, the steps for obtaining the full face and eye area dual feature vector groups are:

[0023] Based on the qualified face image, identifying the coordinates of key points of the full face area through a convolutional neural network, cropping a minimum circumscribed rectangular area image covering the full face according to the key point coordinates, extracting feature vectors of the full face area image through a pre-trained ResNet-50 model, and generating a full face area feature vector set;

[0024] Based on the qualified face image, the coordinates of six key points, including the inner corners of the eyes, the outer corners of the eyes, and the center of the pupil, are located. An image block of the periocular area is generated by expanding 10 pixels outward according to the coordinates of the key points. Feature extraction is performed on the image block of the periocular area to generate a feature vector set of the periocular area. Combined with the feature vector set of the full face area, a dual feature vector group of the full face and periocular area is obtained.

[0025] Preferably, the steps of obtaining the fused face recognition features are:

[0026] Based on the full-face and peri-eye dual feature vector group, the vector of the full-face region feature vector and the vector of the peri-eye region feature vector are spliced ​​into a composite vector in dimensional order to generate a fused face recognition feature.

[0027] Preferably, the steps of obtaining the preliminary access instruction are:

[0028] Based on the personnel exclusion status, analyzing the truth state of the Boolean determination result, and if the truth state is true, generating original access denial instruction data including a timestamp, a device number and a personnel exclusion status code;

[0029] Based on the original access denial instruction data, a predefined instruction template library is called to match the text description rules of the corresponding status code, the timestamp is converted to a standard format, the device number is converted to a hexadecimal code, and the result is combined with the status code into a structured string to generate a standardized access denial instruction text;

[0030] Based on the standardized access denial instruction text, a protocol header, a checksum and a terminator are added according to the access control system communication protocol requirements, and the text is encapsulated into a binary data frame that meets the RS-485 bus transmission requirements to generate a preliminary access instruction.

[0031] Preferably, the steps for obtaining the access control instruction are:

[0032] Based on the preliminary access instruction, calculating a feature similarity score according to the fused facial recognition feature and the authorized personnel feature library;

[0033] Based on the feature similarity score, a two-way interval comparison is performed between the feature similarity score and the preset authorization threshold. If the feature similarity score falls within the interval, a pass permission instruction is generated; otherwise, a secondary verification request instruction is generated. When encapsulating the instruction, a CRC-32 checksum and a timestamp suffix are added to generate an access control pass instruction.

[0034] Compared with the prior art, the advantages and positive effects of the present invention are:

[0035] In the present invention, by real-time acquisition of facial image sequences and dynamic evaluation of motion blur values, a single-frame screening mechanism effectively eliminates low-quality images caused by human movement or environmental interference, thereby improving the accuracy of subsequent feature extraction. The dual judgment logic of the full-face feature vector and the boundary distance of the unauthorized feature space is adopted, combined with the orthogonal projection and standardization processing of the multi-dimensional feature space, to enhance the recognition accuracy of abnormal access behaviors such as disguise and forgery, and reduce the false positive rate. The dual feature vector groups of the full-face area and the eye area are extracted simultaneously, and the complementary information of the global structural features and local detail features of the face is retained through composite vector splicing and normalization processing, thereby enhancing the system's adaptability to common obstructions in medical scenarios such as wearing masks and goggles. After dynamically fusing multi-region features, the two-way interval comparison mechanism is combined in the authorization threshold judgment to effectively distinguish between legitimate personnel and unauthorized personnel with a high degree of similarity, thereby avoiding the boundary blur problem caused by the traditional single threshold judgment. The protocol encapsulation and real-time verification mechanism are integrated into the access control command generation process to ensure the reliability and timeliness of the permission judgment results to hardware execution. At the same time, a secondary verification process in a non-rejection state is realized through multi-level logic branches, taking into account both security protection and passage efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0038] See also Figure 1 The present invention provides a technical solution: a face recognition ward access control system includes:

[0039] The facial image acquisition module activates the image sensor of the ward access control to capture facial image sequences based on the people in the field of view, evaluates the image motion blur value, and screens single-frame images to obtain qualified facial images;

[0040] The forbidden feature comparison module extracts the corresponding full-face feature vector value based on the qualified face image, establishes the current face feature vector, calculates the distance value to the preset unauthorized feature space boundary based on the current face feature vector, and compares it with the preset exclusion judgment threshold to obtain the person's exclusion status;

[0041] The multi-region feature fusion module extracts the full-face region feature vector set based on qualified facial images and simultaneously locates and crops the periocular region image blocks to extract the periocular region feature vector set, establishing a full-face and periocular region dual feature vector set. Based on the full-face and periocular region dual feature vector set, the two vector sets within the group are spliced ​​to obtain fused facial recognition features;

[0042] The access control authority judgment module, based on the personnel exclusion status, determines that if the Boolean value is true, it outputs the access denial instruction text and establishes a preliminary access instruction. Based on the preliminary access instruction, if the text content is not a rejection, the fused facial recognition features are compared and retrieved with the authorized personnel feature library, and the access control pass instruction is generated based on the relationship between the comparison score and the authorization threshold.

[0043] The steps to obtain qualified face images are:

[0044] Based on the facial image sequence captured by the image sensor, the grayscale matrix data of each frame is extracted, and a two-dimensional discrete Fourier transform is performed on the grayscale matrix of each frame to generate a set of frequency domain components of each frame;

[0045] According to the frequency domain component set of each frame, the motion blur value is calculated using the following formula:

[0046]

[0047] Among them, B j is the motion blur value of the jth frame, F k (u, v) is the complex amplitude of the frequency domain component of the kth frame at the frequency coordinate (u, v), G(u, v) is the reference frequency domain component template without motion blur, N u and N v is the horizontal and vertical resolution of the frequency domain component, ∈ = 10 -6 For a very small amount;

[0048] Based on the motion blur values ​​of all frames, the mean and standard deviation are calculated, and the mean minus the standard deviation is used as the dynamic screening threshold. The motion blur values ​​of all frames are traversed, and frames with motion blur values ​​less than or equal to the dynamic screening threshold are selected to generate qualified face images.

[0049] Specifically, based on the facial image sequence captured by the image sensor, color space conversion is first performed on each frame of the sequence, and the RGB (red, green, and blue) color image is converted into a grayscale image using the standard weighted average method. The calculation formula is grayscale value = 0.299×R+0.587×G+0.114×B, where R, G, and B are the red, green, and blue channel values ​​of the pixel, respectively. The grayscale matrix data of each frame is obtained, and the dimension of the matrix data is the resolution of the image sensor, such as 640×480 pixels. Then, for the grayscale matrix of each frame, the fast Fourier transform (FFT) algorithm is applied to perform a two-dimensional discrete Fourier transform. This process converts the image from the spatial domain to the frequency domain, that is, the image is represented as a superposition of sine and cosine waves of different frequencies. The transformation process processes the entire grayscale matrix and calculates the complex Fourier coefficient F(u,v) corresponding to each frequency coordinate (u,v), where u ranges from 0 to N u -1, v ranges from 0 to N v -1, N u and N v are the width and height of the image respectively (e.g. N u =640,N v =480), the complex coefficient F(u,v)=R(u,v)+i·I(u,v) contains amplitude and phase information, reflecting the intensity and spatial position of the image at this specific frequency component. This conversion operation is repeated for all frames in the sequence to finally generate a set of frequency domain components for each frame.

[0050] formula: The formula is useful in that it calculates the frequency domain representation F of the current frame j j The normalized difference between (u, v) and the frequency domain template G(u, v) of the ideal clear image can quantitatively evaluate the degree of motion blur of the image; motion blur usually leads to the attenuation of high-frequency components. This formula can effectively capture this attenuation characteristic by comparing the spectrum differences at each frequency point and averaging them, thereby obtaining a comprehensive blur index B j This indicator is crucial for subsequent screening of clear images; F is used in the denominator j Normalizing (u,v)+∈ can reduce the impact of overall image brightness changes on blur evaluation and enhance the robustness of the evaluation. Parameter B jThe steps for obtaining are as follows: This parameter is the final calculation result of the formula, representing the motion blur value of the j-th frame image. It is calculated by substituting other parameters and does not need to be obtained separately. Its numerical value directly reflects the degree of blur of the image. The smaller the value, the clearer the image and the closer it is to a state without motion blur.

[0051] Parameter F j The acquisition steps of (u, v) are as follows: This parameter represents the complex frequency domain component amplitude of the j-th frame face image at the frequency coordinate (u, v), which is obtained by performing a two-dimensional discrete Fourier transform (usually using the FFT algorithm) on the grayscale matrix data of the j-th frame. In the previous step, the complete frequency domain component set of each frame j of the image sequence has been calculated, that is, a series of F j (u,v) values, where u and v are the horizontal and vertical coordinate indices of the frequency domain, respectively. For example, for a 640×480 grayscale image, after performing a two-dimensional FFT, a 640×480 complex matrix will be obtained. j (u, v) is the complex value of the matrix at position (u, v). Acquisition example: Perform FFT on the 640×480 grayscale matrix of the fifth frame (j=5) image to obtain a complex matrix. The value at the frequency coordinate (u=10,v=20) is F5(10,20)=15.2+8.1i.

[0052] The acquisition steps of the parameter G(u,v) are as follows: This parameter represents a reference frequency domain component template without motion blur, which is used to compare and evaluate the blur degree of the current frame. The template is obtained by pre-collecting a set of high-quality and clear face images (for example, 100) under the same lighting and shooting conditions, ensuring that the person is still. These clear images are gray-scaled and two-dimensionally discrete Fourier transformed to obtain their respective frequency domain components F sharp_i (u,v), and then calculate the average amplitude of the frequency domain components of these clear images at each frequency coordinate (u,v). Acquisition example: Collect 100 clear face images, calculate their respective FFTs, and for the frequency coordinates (u=10,v=20), calculate the complex FFT of these 100 images at that point sharp_i The average value of (10,20) is G(10,20)=18.5+6.5i.

[0053] Parameter N u and N v The steps to obtain are as follows: These two parameters represent the horizontal and vertical resolutions of the frequency domain components, respectively. Their values ​​directly correspond to the width and height (in pixels) of the grayscale image input for Fourier transform. These two values ​​are determined by the sensor resolution during the image acquisition phase or the image size during the preprocessing phase. Acquisition example: If the image sensor resolution used is 640×480 and the image size is not changed before Fourier transform, then Nu =640, N v =480.

[0054] The steps to obtain the parameter ∈ are as follows: This is a very small positive number added to prevent the denominator from being zero. Its value is usually set to a very small positive number based on experience to ensure the stability of the calculation. At the same time, its value is small enough not to significantly affect the calculation result. In this application, it is set to 10 -6 Acquisition example: directly set ∈ = 0.000001.

[0055] Calculation process: Take the calculation of the motion blur B5 of the j=5th frame image as an example, the image resolution used is N u =640, N v =480, minimum quantity∈=10 -6 The reference frequency domain template G(u,v) has been pre-calculated. First, obtain the frequency domain component set F5(u,v) of the 5th frame. Then, for each frequency coordinate (u,v) (where u ranges from 0 to 639 and v ranges from 0 to 479), perform the following calculations:

[0056] 1. Calculate the difference between F5(u,v) and G(u,v): Diff(u,v) = F5(u,v) - G(u,v).

[0057] 2. Calculate the denominator: Denom(u,v)=F5(u,v)+∈.

[0058] 3. Calculate the absolute value of the ratio at this point: Calculates the division and absolute value (modulus) of complex numbers. For example, for the point (u=10,v=20), using the values ​​obtained in the previous example:

[0059] Then, for all N u ×N v =640×480=307200 frequency coordinates (u, v) and the sum of the Ratio(u, v) values ​​calculated:

[0060]

[0061] For example, by calculating the Ratio(u,v) of all points and summing them, we get SumRatio = 78643.2. Finally, we calculate the average value to get B5:

[0062]

[0063] The results show that the calculated motion blur value B5 of the 5th frame is 0.256, which represents the average normalized difference between the frequency domain characteristics of the frame image and the clear image reference template.

[0064] Based on the motion blur value B of all frames calculated in the previous step j (For example, a frames =10 frames, the ambiguity value sequence is {B1,B2,...,B 10}, the specific value is

[0065] {0.15,0.42,0.21,0.55,0.26,0.18,0.39,0.60,0.28,0.36}), first calculate the arithmetic mean μ of this sequence B and standard deviation σ B The average value is calculated by dividing the sum of all frame blur values ​​by the number of frames, and the standard deviation is calculated by the square root of the average value of the sum of the squares of the differences between each blur value and the mean. Then, the calculated mean minus the standard deviation is used as the dynamic screening threshold T for this sequence screening. blur , that is, T blur =μ B -σ B = 0.34-0.1457 = 0.1943. This threshold is set based on statistical principles and aims to select frames whose blur is significantly lower than the average level of the sequence, so as to dynamically adapt to the changes in the overall quality of the image sequence under different lighting and motion speeds. Subsequently, the motion blur value B of each frame in the sequence is traversed. j , and compare it with the dynamic screening threshold T blur For comparison, if B j ≤T blur (i.e. B j ≤0.1943), the frame is determined to be a qualified frame and is selected. According to the sample data, B1=0.15≤0.1943, B6=0.18≤0.1943, and B of other frames j The values ​​are all greater than 0.1943, so the 1st and 6th frames are selected, and all the selected frames (in this case, the image data of the 1st and 6th frames) are collected to generate a qualified face image set for subsequent face feature extraction and recognition.

[0066] The steps to obtain the current face feature vector are:

[0067] Based on qualified face images, the face area is geometrically corrected, facial key points are aligned to preset positions in the standard coordinate system, rotation and scaling differences are eliminated, and histogram equalization is applied to eliminate lighting interference, generating full-face image data after geometric and lighting normalization.

[0068] Based on the full-face image data after geometric and illumination normalization, a multi-scale pyramid structure is used to perform hierarchical feature extraction on the full-face image data. Three convolution kernels, 3×3, 5×5, and 7×7, are used to scan the full-face image data respectively to locate the facial key points. A high-dimensional texture descriptor fused with the histogram of directional gradient and the local binary pattern is extracted based on the 16×16 pixel area around each key point to generate a high-dimensional texture descriptor set of the full-face key points.

[0069] Based on the high-dimensional texture descriptor set of all facial key points, principal component analysis is used to independently reduce the dimensionality of the high-dimensional texture descriptor of each key point. Subsequently, the descriptors of all key points are spliced ​​through a fully connected network layer to establish the current facial feature vector.

[0070] Specifically, based on each image in the qualified face image set obtained in the previous step, a pre-trained facial key point detection model (for example, a method based on cascade regression or deep learning, such as MTCNN) is first used to preliminarily locate several core key points in the face area, including at least 5 points such as the left and right eye corners, the nose tip, and the mouth corners. Subsequently, according to the current coordinates of these detected key points and the corresponding point target positions in a pre-defined set of standard facial coordinate systems (for example, in a target image size of 112 pixels × 112 pixels, the left eye center coordinates are set to (30, 40), the right eye center coordinates are set to (82, 40), and the nose tip coordinates are set to (56, 70)), an optimal two-dimensional affine transformation matrix (including rotation, scaling, and translation parameters) is calculated. The calculation process is performed by solving the least squares The problem is to find a transformation that minimizes the sum of the squares of the distances between the source key points and the target standard positions after transformation. The calculated affine transformation matrix is ​​then applied to the entire original qualified face image. A geometrically corrected image is generated by pixel resampling (for example, using bilinear interpolation). At this time, the facial posture is basically aligned and the size is unified to the preset standard. Based on this geometrically corrected image, the adaptive histogram equalization (CLAHE) technology is further applied. This technology divides the image into several small rectangular areas (for example, 8x8 pixel blocks), calculates the grayscale histogram of each area independently and equalizes it, and then smoothes the boundaries between blocks through interpolation, effectively enhancing the local contrast of the image and suppressing the influence of global illumination unevenness, generating the final full-face image data after geometric and illumination normalization.

[0071] According to the full face image data after geometric and illumination normalization generated in the previous step (for example, a grayscale image with a uniform size of 112x112 pixels), an image pyramid is constructed. By performing continuous Gaussian smoothing and downsampling (for example, halving the resolution each time) operations on the original normalized image, a series of image levels with different resolutions (for example, 112x112, 56x56, 28x28) are generated to form a multi-scale representation. Then, on each layer of the pyramid image, three different sizes of convolution kernels (3×3, 5×5, 7×7) are used for convolution operations. The unit stride (stride = 1) and edge padding are used during the product to maintain the size of the feature map. These convolution kernels of different sizes are designed to capture facial details and structural information of different scales and receptive fields, generate multiple sets of feature maps, and use these multi-scale feature map information to accurately locate all the preset key points of the face (for example, the 68 facial key points that meet the dlib standard, including eyebrows, eyes, nose, mouth and cheek contour points) through the trained key point positioning network (which can be a cascaded convolutional network or a heat map regression model). The precise pixel coordinates of the key points are obtained. Then, for each located key point, a fixed-size neighborhood image block (e.g., a 16×16 pixel area) is extracted around it, and the directional gradient histogram (HOG) features and local binary pattern (LBP) features are calculated in this image block. When extracting the HOG feature, the 16×16 area is divided into smaller units (cell, such as 4×4 pixels), and the directional histogram of the pixel gradient in each unit is calculated (e.g., divided into 9 directional intervals), and then several units are grouped into blocks (e.g., 2×2 cells). ll), the histogram of all units in the block is normalized (for example, L2-norm); the LBP feature calculates the grayscale comparison result of each pixel and its neighboring pixels (for example, 8 neighbors with a radius of 1) to form a binary code, and the histogram of the coding mode is statistically analyzed (for example, using a rotation-invariant uniform mode). Finally, the calculated HOG descriptor vector and the LBP descriptor vector are concatenated to form a high-dimensional texture descriptor that combines shape and texture information. This operation is repeated for all key points to generate a set of high-dimensional texture descriptors for the key points of the entire face.

[0072] Based on the high-dimensional texture descriptor set of full-face key points generated in the previous step (which contains, for example, 68 key points, each key point corresponds to a high-dimensional descriptor, for example, its original dimension is D orig=1024), first apply principal component analysis (PCA) to the descriptors of each key point type (for example, the descriptor set of all left eye corner key points, the descriptor set of all nose tip key points, etc.) for independent dimensionality reduction processing. This process requires pre-calculating and storing the PCA transformation basis (i.e., eigenvector) of each key point type based on a large-scale face dataset, or calculating the transformation basis of the current batch of data online, setting the target dimensionality reduction dimension N, and selecting the dimension N based on compressing the data as much as possible while retaining sufficient information (for example, retaining 95% of the original variance). It is set based on experience and experimental results. For example, the 1024-dimensional descriptor of each key point is reduced to N = 128 dimensions. The specific operation is to project the original high-dimensional descriptor vector of each key point onto its corresponding The first N principal component feature vectors of the type are used to obtain the reduced dimensionality descriptors. This independent dimensionality reduction operation is performed on the descriptors of all 68 key points to generate 68 128-dimensional vectors. Subsequently, these 68 reduced dimensionality descriptor vectors are spliced ​​in a predetermined order (for example, the order of key points from left to right and from top to bottom) to form a single long vector with a total dimension of 68×128=8704. Finally, this spliced ​​long vector is input into a pre-trained fully connected network layer. The function of this layer is to further integrate the information of all key points and map it to a more compact and more discriminative facial feature space. The number of input nodes of the fully connected layer is 8704, and the number of output nodes is set to the target dimension of the final facial feature vector, for example, D final =512, a linear transformation is performed within the layer (multiplied by the weight matrix) and a bias term may be added, and then an activation function is passed (for example, no activation function is used, a linear result is directly output, or ReLU is used). The weight parameters of this fully connected layer are obtained by supervised learning training on a large dataset of face recognition tasks. Finally, the output of this layer is the 512-dimensional feature vector of the current face, which is used to establish the current face feature vector.

[0073] The steps to obtain the personnel exclusion status are:

[0074] Based on the current facial feature vector, perform spectral decomposition on the covariance matrix of the preset unauthorized feature space, extract eigenvalues ​​and orthogonal basis vectors, perform normalized projection on the current facial feature vector along the orthogonal basis direction, and generate a set of normalized projection feature components;

[0075] Based on the standardized projected feature component set, the dynamic distance value to the boundary of the unauthorized feature space is calculated using the following formula:

[0076]

[0077] Among them, D p is the dynamic distance value from the current face feature vector to the boundary of the unauthorized feature space, vpi is the coordinate value of the standardized projection feature component in the i-th dimension, μ pi is the mean value of the unauthorized feature space samples in the i-th dimension, zσ pi is the standard deviation of the i-th dimension, m is the total number of feature space dimensions, and t is the integral variable;

[0078] Based on the dynamic distance value, the dynamic distance value is compared with the preset exclusion judgment threshold. If the logarithmic transformation value of the dynamic distance value is greater than or equal to 1.5 times the preset exclusion judgment threshold, the personnel exclusion status is judged to be true, otherwise it is false and the personnel exclusion status is generated.

[0079] Specifically, based on the current face feature vector v established in the previous step current First, we load the pre-built statistical information of the unauthorized feature space (e.g., a 512-dimensional vector). The data in this space is obtained by collecting and processing facial images of a group of people who are clearly not allowed to obtain access rights (e.g., unregistered visitors, interference samples for system testing, etc., on the order of 10,000 samples), and extracting their corresponding feature vectors v. non_auth_j And form, calculate the covariance matrix Σ of this set of unauthorized feature vectors non , then the covariance matrix Σ of the m×m (e.g. 512×512) dimension non Perform spectral decomposition (eigenvalue decomposition) to obtain m eigenvalues ​​λ i (arranged in descending order) and its corresponding m mutually orthogonal unit eigenvectors u i , these eigenvectors u i The principal axis direction (orthogonal basis) of the unauthorized feature space is formed. Then, the current face feature vector v current Project onto this orthogonal basis and calculate the projection components In order to standardize the subsequent distance calculation, it is necessary to use the distribution information of unauthorized samples in these main axis directions. Specifically, calculate the projection of unauthorized samples onto each feature vector u i The mean μ on pi and standard deviation σ pi , then normalize the projection component of the current face feature vector and calculate the standardized projection feature component Finally, a standardized projection feature component set v composed of these components is generated p =[v p1 ,v p2 ,...,v pm ].

[0080] formula: Used to assess the degree to which the current facial feature vector deviates from the center of the preset unauthorized feature space. The first term effectively accounts for the variability (variance) of unauthorized samples in different directions by calculating the sum of squares of the standardized coordinate differences along the principal axis of the unauthorized space, making the distance metric adaptive to the data distribution. The second term introduces an integral term that calculates the average absolute value of the standardized coordinate deviation vector along the path from the origin to the target point. This term may capture the overall amplitude information of the deviation vector and complements the simple sum of squares term.

[0081] Parameter v pi The acquisition steps are as follows: This parameter represents the coordinate value of the current facial feature vector after being standardized and projected onto the i-th dimension principal axis direction of the unauthorized space. This value is obtained by current First project to the i-th eigenvector u i Get Then, according to the mean μ′ of the unauthorized samples in this direction pi and standard deviation σ pi (Here μ′ pi Refers to unauthorized sample projection p i Get example: calculate the standardized projection feature component set v p =[v p1 ,v p2 ,...,v pm ], then v pi It is the value of the i-th element in the set. For example, for the first dimension, v p1 =1.2.

[0082] Parameter μ pi The steps to obtain is: This parameter represents the mean value of the unauthorized feature space samples on the i-th dimension coordinate (after projection basis transformation). This value is a statistic calculated in advance when building the unauthorized space model. The specific calculation method is: collect all unauthorized sample feature vectors v non_auth_j , projecting them onto the i-th eigenvector u i The projection value p is obtained ji , and then calculate the mean μ of these projection values pi ; In practical applications, if v pi The centering (mean subtracted) has been completed when obtaining, so μ here is pi Should be 0; if v pi Only the projection value p i , then μ here pi It is the pre-calculated projection mean. Obtain example: When building the model, the projection mean of the unauthorized sample on the first axis is calculated as μ p1 =0.05.

[0083] Parameter zσ pi The steps to obtain are: This parameter represents the standard deviation σ of the unauthorized feature space sample in the i-th dimension coordinate pi Multiply by a scaling factor z, standard deviation σ pi It is also a pre-calculated statistic, which is equal to the i-th eigenvalue λ i The square root of Or by calculating the unauthorized sample projection value p ji The scaling factor z is a hyperparameter used to adjust the role of the standard deviation in distance calculations. It can be set according to the system's requirements for exclusion strictness. For example, setting z = 1 means using the original standard deviation, and setting z = 2 or z = 3 is equivalent to requiring the current vector to be 2σ or 3σ outside the range of the unauthorized distribution to be considered far enough. The basis for setting z is usually based on a trade-off analysis of the system's false alarm rate (FAR) and false negative rate (FRR) on the validation set, and a z value that can achieve the required security level is selected. Acquisition example: The first eigenvalue obtained by spectral decomposition is λ1 = 0.81, then If the scaling factor z is set to 1.5, then zσ p1 =1.5×0.9=1.35.

[0084] The steps to obtain the parameter m are as follows: This parameter is the total dimension of the feature space, which is equal to the current face feature vector v current The dimension is also equal to the covariance matrix Σ non The dimension of is determined by the design of the facial feature extraction model. Example: If the facial feature vector obtained in the previous step is 512-dimensional, then m = 512.

[0085] The steps to obtain the parameter t are: the parameter is the definite integral ∫0 1 ...The integration variable in dt ranges from 0 to 1 and is used to parameterize the path from the origin to the destination.

[0086] Substitute the parameters into the formula to calculate the dynamic distance value D p It is approximately 1.365. This value quantifies the degree of deviation between the current facial feature vector (in its standardized projection coordinates) and the preset unauthorized feature space center, taking into account the sum of the squares of the standardized distances in each main axis direction and the integral measure of the overall deviation vector size.

[0087] Based on the dynamic distance value D calculated in the previous step p (For example, D p =1.365), first perform natural logarithm transformation and calculate ln(D p )=ln(1.365)≈0.311, and then compare this logarithmic value with an adjusted preset exclusion decision threshold, which is Texclude The setup process is as follows: collect a large number of face feature vectors of people with known identities, including authorized personnel (such as N auth = 5000 people) and unauthorized personnel (e.g. N non_auth =10000 people), calculate the dynamic distance value D for each sample p And take the logarithm ln(D p ), respectively draw the ln(D p ) value distribution histogram, analyze the overlap of the two distributions, and select a threshold T according to the security requirements of the application scenario (for example, the probability of unauthorized personnel being mistakenly accepted is required to be less than 0.1%). compare So that ln(D p )≥T compare The decision rule can meet the security requirements while minimizing the error rejection of authorized personnel. Based on this principle, T is selected. compare For example, the analysis found that when T compare =5.0 can meet the safety requirements, and then according to the comparison logic of this step (compared with 1.5 times the threshold), the basic threshold T is calculated inversely. exclude =T compare / 1.5=5.0 / 1.5≈3.33, so the preset exclusion threshold is set to T exclude =3.33, when making a judgment, calculate the comparison reference value 1.5×T exclude =

[0088] 1.5×3.33=4.995, the ln(D p ) value (0.311) is compared with the comparison benchmark value (4.995). Since 0.311<4.995, that is, the logarithmic transformation value of the dynamic distance value is less than 1.5 times the preset exclusion judgment threshold, the personnel exclusion status is judged to be false (False), and the generated personnel exclusion status Boolean value is False.

[0089] The steps to obtain the dual feature vector groups of the whole face and the eye area are as follows:

[0090] Based on a qualified face image, a convolutional neural network is used to identify the coordinates of key points in the full face area. The minimum bounding rectangle area image covering the entire face is cropped based on the key point coordinates. The feature vector of the full face area image is extracted using a pre-trained ResNet-50 model to generate a full face area feature vector set.

[0091] Based on the qualified face image, the coordinates of six key points, including the inner and outer corners of the eyes and the pupil center, are located. The image blocks of the periocular area are expanded 10 pixels outward according to the coordinates of the key points. The features of the periocular area image blocks are extracted to generate a periocular area feature vector set. The full-face area feature vector set is combined with the full-face area feature vector set to obtain a full-face and periocular area dual feature vector set.

[0092] Specifically, based on the qualified face images screened in the previous step, a pre-trained convolutional neural network (CNN) is first used to detect facial key points. The CNN model is based on, for example, MobileNetV2, with an output layer for coordinate regression or generating key point heat maps connected to the top. The network input is a single qualified face image (for example, resized to 256×256 pixels), and the output is the pixel coordinates (x i ,y i ), these key points cover the eyes, eyebrows, nose, mouth and facial contours. The CNN model has been fully supervised and trained on a dataset containing a large number of diverse face images and corresponding accurately labeled key points (such as WFLW, 300W or COFW), and has learned the mapping relationship from facial appearance to key point positions. In the inference stage, the key point coordinate set can be obtained by directly forward propagating the input qualified face image. The coordinates of all key points {(x i ,y i )}, calculate the minimum bounding rectangle of these coordinates, that is, determine all x i The minimum value x in min and the maximum value x max , and all y i The minimum value y in min and the maximum value y max , a small border extension can be appropriately added (for example, the width and height are expanded by 15% each) to include a more complete facial area, and the upper left corner coordinate (x′) of the cropped area is obtained. min ,y′ min ) and the lower right corner coordinate (x′ max ,y′ max), use these coordinates to crop a rectangular area image from the original qualified face image, then preprocess the cropped full-face area image block, including adjusting its size to a standard size that meets the input requirements of a specific feature extraction model (for example, ResNet-50 typically uses 224×224 pixels) and normalizing it (for example, subtracting the mean of the ImageNet dataset and dividing it by its standard deviation). Then, input the preprocessed image into a pre-trained ResNet-50 model, which is trained on large-scale face recognition benchmark datasets (such as VGGFace2 and MS-Celeb-1M). Its original top classification layer (for identity recognition) is removed, and its convolutional base part is retained. A forward propagation calculation is performed to extract the output of the last global average pooling layer, which is a fixed-length vector. This vector is regarded as the deep feature representation of the full-face area image. Repeat this process for each image in the qualified face image set to generate the corresponding full-face area feature vector set.

[0093] Based on the same qualified face image and the full-face key point coordinate set identified by CNN, the coordinates of 6 specific key points related to both eyes are accurately extracted, namely the inner corner of the left eye, the outer corner of the left eye, the center of the left pupil, the inner corner of the right eye, the outer corner of the right eye, and the center of the right pupil. If the initial detection key point set (such as the 68-point model) already contains these points, they are directly selected. If not, or the accuracy is insufficient, a high-precision key point detector specially trained for the eye area can be called to re-locate and obtain the precise coordinates of these 6 points (x eye_j ,y eye_j ) where j = 1...6. Then, based on these coordinates, define the image blocks of the left and right eye areas. For the left eye area, find the minimum and maximum horizontal and vertical coordinate values ​​of the three key points (inner corner, outer corner, pupil) corresponding to it, and record them as (x Lmin ,y Lmin ) and (x Lmax ,y Lmax ), and then expand outwards by a fixed pixel value. The expansion value is set based on the face image resolution and eye size experience. Here it is set to 10 pixels, that is, the coordinate of the upper left corner of the left eye area image block is (x Lmin -10,y Lmin -10), the coordinate of the lower right corner is (x Lmax +10,y Lmax+10), perform the same operation on the right eye area to obtain the coordinate range of the right eye area image block, use the calculated coordinate range to crop the left and right eye area image blocks from the original qualified face image, and then perform feature extraction on these two image blocks. A different strategy from the full face feature extraction can be used, such as using a lighter pre-trained CNN model (such as MobileNetV3-Small or a specially designed eye feature extraction network), or reusing the used ResNet-50 model, and convert each eye area image block (which also needs pre-processing, such as resizing to 112×112 Pixels are normalized) are independently input into the model, and the output of the middle layer or the last pooling layer is extracted as the feature vector of the periocular area (for example, a 512-dimensional feature vector is extracted for each periocular area). In this way, feature vectors are generated for the left and right periocular areas respectively, forming a periocular area feature vector set (including two vectors for the left eye and the right eye). Finally, this periocular area feature vector set (for example, 2×512 dimensions) is logically combined with the generated full-face area feature vector (for example, 2048 dimensions) and recorded as a data structure to obtain a full-face and periocular dual feature vector group, which contains feature information from the overall face and the key eye area.

[0094] The steps to obtain the fusion face recognition features are:

[0095] Based on the full-face and peri-eye double feature vector groups, the vector of the full-face region feature vector and the vector of the peri-eye region feature vector are concatenated into a composite vector in dimensional order to generate a fused face recognition feature.

[0096] Specifically, based on the full face and eye area dual feature vector group generated in the previous step, the group includes a feature vector representing the full face area (for example, denoted as V full , dimension is 2048) and a set of feature vectors of the left and right eye regions (e.g., V lefteye and V righteye , each with a dimension of 512), perform a vector concatenation operation. Specifically, concatenate these three vectors in a predetermined dimensional order to form a single, higher-dimensional composite vector. The concatenation order is defined as placing the full-face region feature vector V first. full , and then place the left eye region feature vector V lefteye , and finally place the right eye region feature vector V righteye , indicating that the upper side is V fused =[V full ; V lefteye ; V righteye], where the semicolon represents the vertical or horizontal concatenation of vectors, depending on whether the vector is a row vector or a column vector. The dimension of the final composite vector is the sum of the dimensions of each part of the vector, that is, 2048+512+512=3072. This composite vector is the fused face recognition feature.

[0097] The steps to obtain the initial access instruction are:

[0098] Based on the personnel exclusion status, the truth state of the Boolean judgment result is parsed, and if the truth state is true, the original access denial instruction data including the timestamp, device number and personnel exclusion status code is generated;

[0099] Based on the original access denial instruction data, the predefined instruction template library is called to match the text description rules of the corresponding status code, the timestamp is converted to a standard format, the device number is converted to a hexadecimal code, and then combined with the status code into a structured string to generate a standardized access denial instruction text;

[0100] Based on the standardized access denial instruction text, the protocol header, checksum and terminator are added in accordance with the access control system communication protocol requirements, and the text is encapsulated into a binary data frame that meets the RS-485 bus transmission requirements to generate a preliminary access instruction.

[0101] Specifically, based on the personnel exclusion status determined in the previous step (a Boolean value, denoted as S exclude , whose value is True or False), first check the truth state of the Boolean value, and only if S exclude When it is true, it means that the current person is judged as an unauthorized person, and the system will continue to execute the operation of generating the access denial instruction. exclude If it is true, the current system timestamp is obtained immediately, and the time information accurate to milliseconds is obtained by calling the time service at the operating system level (for example, the format is "2025-04-28T18:44:15.123+08:00"). At the same time, the unique identifier (device number) of this access control terminal device is read. This number is set when the device is initialized or configured. For example, the string "Terminal_WardA_001" is read from the configuration file. In addition, according to the judgment result (the personnel exclusion status is true), a predefined personnel exclusion status code is specified. This code is used to internally identify the rejection reason. For example, the code is set to "REJ_NPA".

[0102] (Reject-Non-PermittedArea / Person), the obtained timestamp, device number and specified status code are combined and stored in a temporary internal data structure, such as a record containing a key-value pair {timestamp: "2025-04-28T18:44:15.123+08:00", deviceID: "Terminal_WardA_001", statusCode: "REJ_NPA"}, to generate the original access rejection instruction data containing the timestamp, device number and personnel exclusion status code.

[0103] According to the original access denial instruction data generated in the previous link (for example, record {timestamp: "2025-04-28T18:44:15.123+08:00", deviceID: "Terminal_WardA_001", statusCode: "REJ_NPA"}), access a pre-defined instruction template library in the system. The library uses the personnel exclusion status code ("REJ_NPA") as the index key to find the corresponding instruction formatting rules or text description templates. For example, the template may define the output word The field order and separator of the string are determined. Then, the contents of the original data are formatted. The timestamp ("2025-04-28T18:44:15.123+08:00") is converted to the standard format required by the communication protocol, for example, it is converted to the more compact "YYYYMMDDHHMMSS" format to obtain "20250428184415". The device number ("Terminal_WardA_001") is converted to hexadecimal encoding according to the protocol requirements, for example, its ASCII code sequence is converted to a hexadecimal string.

[0104] 5465726d696e616c5f57617264415f303031", or if the device number is a numeric ID (such as 1001), convert it to hexadecimal representation "03E9", select one of the methods such as the latter to obtain "03E9", and then combine the formatted timestamp, hexadecimal device number and original status code ("REJ_NPA") into a structured string according to the rules defined by the template library, for example, use comma separators to combine as: "20250428184415, 03E9, REJ_NPA", to generate a standardized access denial instruction text.

[0105] Based on the standardized access denial instruction text generated in the previous link (for example, the string "20250428184415, 03E9, REJ_NPA"), it is encapsulated to construct a complete binary data frame according to the specific communication protocol specification based on the RS-485 bus adopted by the access control system. First, the protocol header information is added before the standardized text data. The protocol header usually contains a fixed frame start identifier (for example, 1 byte 0xAA), the target device address (for example, 1 byte, indicating the address of the access control controller receiving the instruction, set to 0x01), the source device address (for example, 1 byte, indicating the terminal address, set to 0xFE), and a command code that identifies the instruction type (for example, 1 byte, indicating that this is an access denial instruction). , set to 0xD1), then, convert the standardized access denial instruction text string into a byte sequence according to the encoding method specified by the protocol (such as UTF-8 or ASCII) as the data payload part, then calculate the check code, select the check algorithm according to the protocol, such as calculating the CRC-16-CCITT checksum of all bytes from the target address to the end of the data payload, and obtain a 2-byte check code, which is appended to the data payload, and finally, add a fixed frame end marker (such as 1-byte 0x55) after the check code, and combine the protocol header, the encoded data payload, the calculated check code and the frame end marker in sequence into a continuous byte stream, such as [0xAA, 0x01, 0xFE, 0xD1,

[0106] <UTF8_bytes_of_"20250428184415,03E9,REJ_NPA"> ,<CRC_high_byte> ,<CRC_low_byte> , 0x55], this complete byte stream is the encapsulated binary data frame that meets the transmission requirements of the RS-485 bus, generating a preliminary access instruction as an access denial indication.

[0107] The steps to obtain access control pass instructions are as follows:

[0108] Based on the preliminary access instructions, the feature similarity score is calculated by integrating the face recognition features and the authorized personnel feature database. The calculation formula is:

[0109]

[0110] Among them, S a is the feature similarity score, X k To fuse the face recognition features in the kth dimension, Y k is the value of the target feature in the authorized personnel feature library in the kth dimension, m is the total dimension of the feature vector, δ = 10 -5 is a very small quantity, σ is the Gaussian kernel width parameter;

[0111] Based on the feature similarity score, the feature similarity score is compared with the preset authorization threshold in a two-way interval. If the feature similarity score falls within the interval, a pass instruction is generated, otherwise a secondary verification request instruction is generated. When encapsulating the instruction, a CRC-32 check code and a timestamp suffix are added to generate an access control pass instruction.

[0112] Specifically, it should be noted that this step is only executed when the previous "personnel exclusion status" is judged to be false (False). If the status is true, the access denial instruction has been generated and no similarity calculation is required. Based on the preliminary access instruction that does not trigger the access denial (that is, the personnel exclusion status is false), and based on the fused face recognition feature X (an m-dimensional vector, for example, m = 3072) obtained in the previous step, and the pre-loaded authorized personnel feature library (containing the reference feature vectors Y of all registered authorized personnel in the system), j ), perform feature comparison to calculate feature similarity score S a This process typically involves comparing the current feature X to each reference feature Y in the library. j is compared (1:N matching), or if the system includes an identity assertion step, is compared with a specific reference feature Y corresponding to the asserted identity (1:1 matching), calculated as:

[0113]

[0114] The usefulness of the formula is that it combines the relative differences in the dimensions of the eigenvectors (via max(|X k |,|Y k |) normalized) and absolute differences (passed through the Gaussian kernel e - ...weighted) to evaluate similarity; the normalization term makes the score insensitive to changes in the amplitude of the feature vector itself, focusing more on differences in feature structure; the Gaussian kernel term introduces consideration of absolute distance, causing the contribution of dimensions with large differences to the final score to be rapidly reduced; the parameter σ controls the width of the Gaussian kernel, determining the score's sensitivity to absolute differences. This composite metric provides a unique feature comparison strategy, the specific effectiveness of which depends on the choice of parameter σ and the distribution characteristics of the features themselves.

[0115] Parameter X k The steps for obtaining : This parameter represents the value of the fusion face recognition feature vector X of the current person to be verified in the kth dimension. The vector X is generated through the above steps (full face and eye feature extraction and splicing) and its dimension is m. Example: If the fusion face recognition feature X is a 3072-dimensional vector, X k It is the kth element value of the vector. For example, when k=1, X1=0.75.

[0116] Parameter Yk The steps for obtaining are as follows: This parameter represents the value of the kth dimension of the reference feature vector Y of the authorized person selected as the comparison target in the authorized person feature database. This vector Y is pre-extracted and stored in the feature database during the person registration phase, and its dimension is also m. Acquisition example: The feature vector Y of the target person is retrieved from the authorized person feature database, and its first element value is Y1 = 0.71.

[0117] The steps for obtaining parameter m are as follows: This parameter is the total dimension of the fused face recognition feature vector, determined by the feature extraction and fusion strategy, and is equal to the length of vectors X and Y. Acquisition example: Based on the aforementioned fusion steps, the full face feature has 2048 dimensions, and the left and right eye features have 512 dimensions each. After splicing, the total dimension m = 2048 + 512 + 512 = 3072.

[0118] The steps to obtain the parameter δ are: This is a step to prevent the denominator (max(|X k |,|Y k |)) in X k and Y k When the absolute values ​​are close to zero, a small positive constant is set to increase the numerical stability of the calculation, such as 10 -5 Acquisition example: Directly set δ = 0.00001.

[0119] The steps to obtain the parameter σ are as follows: This parameter is the width parameter of the Gaussian kernel, which controls the sensitivity of the similarity score to the absolute difference between the feature vectors. A smaller σ makes only dimensions with small differences contribute significantly (the Gaussian kernel decays quickly), while a larger σ allows dimensions with slightly larger differences to also have an impact (the Gaussian kernel decays slowly). The value of σ usually needs to be determined through experiments based on the specific data set and feature distribution of the application. The setting process is generally as follows: prepare a verification data set containing a large number of known "matching pairs" (different feature vectors from the same person) and "non-matching pairs" (feature vectors from different people), select a series of candidate σ values ​​(for example, gradually increasing from 0.1 to 5.0), and use each candidate σ value to calculate S for all pairs in the verification set. a The system then evaluates the recognition performance for each σ value (e.g., plotting a receiver operating characteristic (ROC) curve, calculating the equal error rate (EER), or the FRR at a specific FAR), and selects the σ value that provides the best performance (e.g., the lowest EER) as the final parameter. For example, when testing on an internal validation set, it was found that when σ = 1.2, the system performs best in distinguishing between matching and non-matching pairs, achieving the lowest equal error rate, so σ = 1.2 was selected.

[0120] Substitute the parameters into the formula to calculate the feature similarity score S aIt is approximately 0.02306. This value indicates the degree of match between the currently captured fusion face recognition features and the target reference features in the authorized personnel database. According to the characteristics of this formula, the closer the score is to 0, the more similar the two feature vectors are.

[0121] Based on the feature similarity score S calculated in the previous step a (For example, S a =0.02306), and compare it with a pre-set authorization threshold interval [T low ,T high ] for comparison. The setting of this interval is based on the system's requirements for security and convenience. By comparing a large number of known authorized personnel (calculating the feature vectors of the same person at different times S a ) and unauthorized personnel comparison (feature vector calculation S of different people) a The goal is to select an interval that can contain the S of the authorized personnel to the greatest extent. a Scores (these scores are expected to be low, indicating high similarity), while excluding scores of unauthorized persons as much as possible. For example, the setting process is: collect 10,000 pairs of same-person comparison scores and 1 million pairs of different-person comparison scores. Analysis shows that the scores of same-person comparisons are mainly concentrated in the range of [0.01, 0.18], while the scores of different-person comparisons are mostly greater than 0.25. In order to maximize the correct acceptance rate (TAR) while ensuring a low false acceptance rate (FAR), the lower limit of the authorization threshold T is set. low =0.01, upper limit T high =0.18, (this interval setting needs to ensure that it covers most of the real matching situations, and the selection of the boundary value needs to be based on ROC curve analysis, balancing FAR and FRR), and the currently calculated S a =0.02306 is compared with the interval, and the judgment condition is T low ≤S a ≤T high , that is, 0.01≤0.02306≤0.18, this condition is true, so the system determines that the match is successful and generates the instruction content of allowing passage. The instruction content may contain a status code, such as "ALLOW_PASS". If S aIf the value falls outside this range (less than 0.01 or greater than 0.18), it is determined that the match is uncertain or failed, and a secondary verification request instruction is generated, such as the status code "REQ_SEC_VERIF". Subsequently, the generated instruction status code (such as "ALLOW_PASS") and the current timestamp (the system time is obtained and formatted, such as "20250428184646") are combined into the instruction core data, and the CRC-32 checksum of the core data is calculated (using the standard CRC-32 algorithm, such as CRC-32-IEEE). The instruction status code, timestamp suffix and CRC-32 checksum are encapsulated into a binary data frame according to the predefined communication protocol format to generate the final access control pass instruction, which is ready to be sent to the access control controller via the RS-485 bus.

Claims

1. The facial recognition ward access control system is characterized by: The system comprises: The facial image acquisition module activates the image sensor of the ward access control to capture facial image sequences based on the people in the field of view, evaluates the image motion blur value, and screens single-frame images to obtain qualified facial images; The forbidden feature comparison module extracts the corresponding full-face feature vector value based on the qualified face image, establishes the current face feature vector, calculates the distance value to the preset unauthorized feature space boundary based on the current face feature vector, and compares it with the preset exclusion judgment threshold to obtain the person exclusion status; A multi-region feature fusion module extracts a full-face region feature vector set based on the qualified face image and simultaneously locates and crops the periocular region image block to extract the periocular region feature vector set, establishes a full-face and periocular region dual feature vector set, and based on the full-face and periocular region dual feature vector set, splices the two vector sets within the group to obtain a fused face recognition feature; The access control authority judgment module, based on the personnel exclusion status, determines that if the Boolean value is true, it outputs the access denial instruction text and establishes a preliminary access instruction. Based on the preliminary access instruction, if the text content is not rejection, the fused facial recognition feature is compared and retrieved with the authorized personnel feature library and an access control pass instruction is generated based on the relationship between the comparison score and the authorization threshold.

2. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the qualified face image are: Based on the facial image sequence captured by the image sensor, the grayscale matrix data of each frame is extracted, and a two-dimensional discrete Fourier transform is performed on the grayscale matrix of each frame to generate a set of frequency domain components of each frame; Calculate the motion blur value based on the frequency domain component set of each frame; Based on the motion blur values ​​of all frames, the mean and standard deviation are calculated, and the mean minus the standard deviation is used as the dynamic screening threshold. The motion blur values ​​of all frames are traversed, and frames with motion blur values ​​less than or equal to the dynamic screening threshold are selected to generate qualified face images.

3. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the current face feature vector are: Based on the qualified face image, geometric correction is performed on the face area, facial key points are aligned to preset positions in a standard coordinate system, rotation and scaling differences are eliminated, and histogram equalization is applied to eliminate lighting interference, thereby generating full-face image data after geometric and lighting normalization; Based on the geometrically and illumination-normalized full-face image data, hierarchical feature extraction is performed on the full-face image data using a multi-scale pyramid structure. Three convolution kernels, 3×3, 5×5, and 7×7, are used to scan the full-face image data, locate facial key points, and extract a high-dimensional texture descriptor fused with a histogram of oriented gradients and a local binary pattern based on a 16×16 pixel region around each key point to generate a set of high-dimensional texture descriptors for the full-face key points. Based on the high-dimensional texture descriptor set of all facial key points, principal component analysis is used to independently reduce the dimensionality of the high-dimensional texture descriptor of each key point, and then the descriptors of all key points are spliced ​​through a fully connected network layer to establish the current facial feature vector.

4. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the personnel exclusion status are: Based on the current facial feature vector, performing spectral decomposition on the covariance matrix of the preset unauthorized feature space, extracting eigenvalues ​​and orthogonal basis vectors, and performing normalized projection on the current facial feature vector along the orthogonal basis direction to generate a normalized projection feature component set; Calculating a dynamic distance value to the boundary of the unauthorized feature space based on the set of standardized projected feature components; Based on the dynamic distance value, the dynamic distance value is compared with the preset exclusion judgment threshold. If the logarithmic transformation value of the dynamic distance value is greater than or equal to 1.5 times the preset exclusion judgment threshold, the personnel exclusion status is determined to be true, otherwise it is false, and the personnel exclusion status is generated.

5. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the dual feature vector group of the whole face and the eye area are as follows: Based on the qualified face image, identifying the coordinates of key points of the full face area through a convolutional neural network, cropping a minimum circumscribed rectangular area image covering the full face according to the key point coordinates, extracting feature vectors of the full face area image through a pre-trained ResNet-50 model, and generating a full face area feature vector set; Based on the qualified face image, the coordinates of six key points, including the inner corners of the eyes, the outer corners of the eyes, and the center of the pupil, are located. An image block of the periocular area is generated by expanding 10 pixels outward according to the coordinates of the key points. Feature extraction is performed on the image block of the periocular area to generate a feature vector set of the periocular area. Combined with the feature vector set of the full face area, a dual feature vector group of the full face and periocular area is obtained.

6. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the fused face recognition features are as follows: Based on the full-face and peri-eye dual feature vector group, the vector of the full-face region feature vector and the vector of the peri-eye region feature vector are spliced ​​into a composite vector in dimensional order to generate a fused face recognition feature.

7. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the preliminary access instruction are: Based on the personnel exclusion status, analyzing the truth state of the Boolean determination result, and if the truth state is true, generating original access denial instruction data including a timestamp, a device number and a personnel exclusion status code; Based on the original access denial instruction data, a predefined instruction template library is called to match the text description rules of the corresponding status code, the timestamp is converted to a standard format, the device number is converted to a hexadecimal code, and the result is combined with the status code into a structured string to generate a standardized access denial instruction text; Based on the standardized access denial instruction text, a protocol header, a checksum and a terminator are added according to the access control system communication protocol requirements, and the text is encapsulated into a binary data frame that meets the RS-485 bus transmission requirements to generate a preliminary access instruction.

8. The face recognition ward access control system according to claim 1 is characterized in that: The steps for obtaining the access control instruction are as follows: Based on the preliminary access instruction, calculating a feature similarity score according to the fused facial recognition features and the authorized personnel feature library; Based on the feature similarity score, a two-way interval comparison is performed between the feature similarity score and the preset authorization threshold. If the feature similarity score falls within the interval, a pass permission instruction is generated; otherwise, a secondary verification request instruction is generated. When encapsulating the instruction, a CRC-32 checksum and a timestamp suffix are added to generate an access control pass instruction.

Citation Information

Cited By

  • Anti-tracking access control system based on face recognition

    CN121075022A