Multi-mode biological recognition intelligent access control security system and method
By using a multimodal biometric intelligent access control system, combined with channel and spatial attention enhancement units, a standard feature vector set is constructed, which solves the problems of incomplete feature extraction and recognition failure in face recognition in coal mine environments, improves recognition accuracy and stability, and ensures the safety management of coal mine production areas.
Patent Information
- Application Number
- CN202511679155.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-03
AI Technical Summary
In complex environments such as coal mines, facial recognition is easily affected by factors such as coal dust, oil stains, obstruction by protective equipment, and complex lighting, resulting in incomplete or failed feature extraction and recognition failure.
The intelligent access control system employing multimodal recognition, including the multimodal recognition intelligent access control system, integrates modules such as a face recognition matching module, a research and development module, a recognition failure handling module, and a secondary acquisition triggering module. Combined with channel attention enhancement units and spatial attention enhancement units, it constructs a standard feature vector set, performs feature channel filtering and missing feature channel splicing, and improves recognition accuracy.
In complex environments, it improves the accuracy and stability of facial recognition, ensures the success rate and security of access control and security systems, reduces false alarms and recognition failures, and ensures the management and security of production areas.
Smart Images

Figure CN121600626A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of facial recognition technology, and more specifically, to a multimodal biometric intelligent access control security system and method. Background Technology
[0002] In industries such as coal mining, to ensure production safety and prevent unauthorized personnel from entering work areas, existing intelligent access control systems are often inadequate due to the complex working environment. This is because workers' hands may be contaminated with coal dust or oil, leading to uncleanliness, or they may experience hand fatigue and reduced dexterity after prolonged work. Manual methods such as account password verification are prone to contamination and damage to the access control equipment and may also affect passage efficiency due to operational errors. Therefore, multimodal identity verification using account password verification and facial recognition is employed. However, the complex working environment in industries like coal mining presents several reasons why facial recognition is prone to failure.
[0003] First, the high concentration of coal dust in the working environment makes it easy for coal dust to adhere to workers' faces, obscuring or altering facial texture features and covering up effective information during facial feature extraction. Second, workers may come into contact with oil during the work process, which can blur facial details and interfere with the accuracy of feature extraction. Third, for safety reasons, workers need to wear safety helmets, protective masks, and other protective equipment, which can obscure key areas of the face such as the forehead, cheeks, and even eyes, reducing the effective facial features available for recognition. Fourth, the lighting conditions in the working area are complex, with situations such as excessively dim lighting, direct strong light, or local shadows affecting the clarity and contrast of facial images, making it difficult for feature extraction algorithms to accurately capture key facial features. As a result, incomplete facial feature extraction, obscuring or interfering with key features, and ultimately, recognition failure may occur during face recognition. In view of this, we propose a multimodal biometric intelligent access control security system and method. Summary of the Invention
[0004] The purpose of this invention is to solve the problems that cause facial recognition to fail easily, and feature extraction to be incomplete or interfered with due to factors such as coal dust adhesion, facial oil stains, protective equipment obstruction, and complex lighting.
[0005] To achieve the above objectives, this invention provides a multimodal biometric intelligent access control and security system, comprising a multimodal recognition integration module, a face recognition matching module, a recognition failure handling module, and a secondary acquisition triggering module, wherein:
[0006] The multimodal recognition integration module acquires facial images and converts them into preliminary feature maps with multiple feature channels. The preliminary feature map is then enhanced into a spatial feature map, which is further enhanced to obtain an enhanced feature map. Facial feature vectors are extracted from the enhanced feature map, including recognition feature vectors and standard feature vectors, and a set of standard feature vectors is constructed. The face recognition matching module is used to identify the user's identity. If user identification fails, a secondary acquisition signal is output to the multimodal recognition integration module for re-identification. An upper limit for the number of comparisons is set; if the number of comparisons exceeds this limit, an alarm signal is output.
[0007] The recognition failure processing module includes a feature channel filtering and matching unit and a missing feature channel splicing unit. The feature channel filtering and matching unit receives the face recognition failure signal, analyzes whether the facial image is missing a feature channel, and then determines whether a complete facial feature vector can be constructed after splicing the feature channels of different facial images. The missing feature channel splicing unit uses different feature channels to construct a complete facial feature vector. When the multimodal recognition integration module constructs the standard feature vector, the secondary acquisition triggering module receives the missing feature channels from the feature channel filtering and matching unit and establishes a mapping relationship between the failure signal and the secondary acquisition signal.
[0008] As a further improvement to this technical solution, the multimodal recognition integration module includes a channel attention enhancement unit and a facial feature vector construction unit. The channel attention enhancement unit converts the facial image into a preliminary feature map with multiple feature channels, and the preliminary feature map is then transformed into a spatial feature map F through channel attention enhancement. c Then, spatial feature map F is enhanced through spatial attention. c An enhanced feature map is obtained; the facial feature vector construction unit is used to extract the enhanced feature map F. att The facial feature vector v for each feature channel c And sort each facial feature vector v in a fixed order. c This yields the final facial feature vector.
[0009] As a further improvement to this technical solution, the channel attention enhancement unit receives a facial image and obtains multiple preliminary feature maps for each feature channel through a basic feature extraction network. Each feature channel corresponds to a response distribution of a facial feature. The spatial dimension of the preliminary feature map F is processed by taking the maximum value of the spatial dimension for each feature channel and taking the spatial average value for each feature channel to obtain two channel descriptors. The two channel descriptors are input into a dimensionality-reducing fully connected layer, and the ReLU activation function is applied to the dimensionality-reduced result to obtain the dimensionality-reduced result: values greater than zero in the dimensionality-reduced result are taken, and values less than or equal to zero are set to zero. Then, the result after ReLU activation is input into a dimensionality-increasing fully connected layer to obtain the dimensionality-increasing result. The two results after the fully connected layer are then added together and passed through the Sigmoid activation function to generate channel attention weights. The dimensionality-reducing fully connected layer reduces the number of feature channels, while the dimensionality-increasing fully connected layer increases the number of channels back to the original number.
[0010] The channel attention enhancement unit multiplies the preliminary feature map channel by channel attention weight to obtain the spatial feature map after channel attention enhancement.
[0011] The beneficial effects of the above-mentioned further scheme are as follows: by obtaining the channel descriptor by taking the maximum and average values of the spatial dimension for each feature channel, significant facial features and global response intensity within each feature channel can be accurately captured; dimensionality reduction, ReLU activation, dimensionality increase, and subsequent operations can effectively highlight the importance of key feature channels and suppress interference from irrelevant channels while reducing computational load; finally, the channel attention weights are multiplied with the preliminary feature map channel by channel, which can specifically enhance channels containing effective facial features and weaken invalid channels affected by coal dust, oil stains, protective equipment occlusion, etc., thereby improving the accuracy and completeness of facial feature extraction. This provides a more reliable feature foundation for subsequent multimodal recognition such as face recognition in complex coal mine environments, and improves the recognition success rate and stability of access control and security systems.
[0012] Based on the above technical solution, the present invention can also be improved as follows: the channel attention enhancement unit receives the spatial feature map and performs max pooling and average pooling on the feature channel dimension of the spatial feature map to obtain two single feature channel spatial descriptors: max pooling is to take the maximum value on all feature channels at each spatial location, and average pooling is to take the average value on all feature channels at each spatial location.
[0013] Two spatial descriptors are concatenated along the channel dimension and merged into a feature map of a specific size. The concatenated feature map is then input into a convolutional layer to capture long-range spatial dependencies. The result after processing by the convolutional layer is input into a sigmoid activation function to generate a spatial attention weight map. The spatial attention weight map is then multiplied element-wise with the spatial feature map to obtain an enhanced feature map.
[0014] As a further improvement to this technical solution, the facial feature vector construction unit enhances each feature channel of the feature map, sums the pixel values at all row and column positions in the feature channel, and divides the summation result by the product of the number of rows and columns of the feature channel to obtain the feature vector of the feature channel.
[0015] The facial feature vector construction unit receives a standard feature vector template construction signal and a face recognition signal. When constructing the standard feature vector template signal, the facial feature vector v extracted by the facial feature vector construction unit is a standard feature vector. Then, multiple standard feature vectors extracted from multiple facial images are used to construct a standard feature vector template, and a mapping relationship between user identity and standard features is established. During face recognition, the facial feature vector extracted by the facial feature vector construction unit is the standard feature vector recognition feature vector, and the recognition feature vector is output to the face recognition matching module.
[0016] The beneficial effects of the above-mentioned further scheme are as follows: by summing and averaging each feature channel of the enhanced feature map to obtain the feature vector, the facial feature information contained in that channel can be extracted efficiently and accurately, providing a concise and crucial feature expression for subsequent recognition; when constructing the standard feature vector template, multiple standard feature vectors are extracted from multiple facial images to jointly construct the template, which can fully cover facial features under different conditions, reduce feature deviations caused by a single image or accidental factors, establish a mapping relationship between user identity and standard features, and lay a reliable foundation for accurate user identity recognition; when extracting and transmitting the recognition feature vector during face recognition, the actual facial features to be recognized can be matched with the standard template. Combined with the more accurate features extracted from the enhanced feature map, it can effectively cope with problems such as facial obstruction by coal dust, oil stains, protective equipment, and poor lighting in complex environments such as coal mines, significantly improving the accuracy and reliability of face recognition and ensuring the effective operation of access control and security systems in special operating scenarios.
[0017] Based on the above technical solution, the present invention can be further improved as follows: The face recognition matching module receives the recognition feature vector, and then calculates the similarity between the recognition feature vector and multiple standard feature vectors in each standard feature vector template in turn. If the similarity between the recognition feature and the standard feature corresponding to a certain user is greater than the similarity threshold, the user identity recognition is determined to be successful; otherwise, the user identity recognition fails. A second acquisition of the user's facial image is performed to extract facial feature vectors for comparison. An upper limit value for the number of comparisons is set. If the number of comparisons exceeds the upper limit value, an alarm signal is output to remind staff.
[0018] As a further improvement to this technical solution, the feature channel filtering and matching unit is used to receive the face recognition failure signal in the face recognition matching module, and retrieve the channel attention weight and the brightness of each position in the spatial feature map of each facial image when recognition fails; set the weight threshold and the pixel threshold, select the feature channels in each facial image whose channel attention weight is greater than the weight threshold and whose brightness at each position in the spatial feature map is greater than the pixel threshold, and compare whether the attributes of the selected feature channels are the same as the number of feature channels used to construct the complete facial feature vector. If the number of feature channels is the same, it means that different facial images can construct a complete facial feature vector after being stitched together.
[0019] The beneficial effects of the above-mentioned further solutions are that, when facial recognition fails in complex environments such as coal mines due to coal dust, oil stains, or obstruction by protective equipment, effective feature channels can be accurately selected by setting weight thresholds and pixel thresholds, while invalid channels that are greatly affected by interference can be eliminated. By comparing the number of selected feature channels with the number of feature channels required to construct a complete facial feature vector, it is possible to scientifically determine whether different facial images, after being stitched together, meet the conditions for constructing a complete feature vector. If so, multiple images can be stitched together to compensate for the feature loss caused by environmental interference in a single image, thereby providing the possibility for subsequent successful facial recognition. This improves the success rate and adaptability of facial recognition in complex operating scenarios for access control and security systems, and ensures the accuracy and security of personnel entry and exit management in production areas.
[0020] Based on the above technical solution, the present invention can also be improved as follows: the missing feature channel splicing unit receives a signal that can construct a complete facial feature vector, selects the feature channels with channel attention weight ≤ weight threshold and pixel value ≤ pixel threshold at each position in the spatial feature map as missing feature channels, records the number of missing feature channels in each facial image, selects the facial image with the fewest missing feature channels as the image to be spliced, and the remaining facial images as images to be selected.
[0021] Then, the similarity between channel attention weights and spatial feature maps is calculated sequentially. First, the similarity of channel attention weights is calculated as the absolute value of the difference between the two channel attention weights; the similarity of spatial feature maps is calculated as the absolute value of the difference between the pixel values at corresponding positions in the two spatial feature maps. The difference within the feature channels is obtained by multiplying the fusion coefficient of the channel attention weight difference and the spatial feature map difference by the attention weight difference and the spatial feature map difference. Then, the differences of each feature channel are weighted and fused again to obtain the similarity between the image to be stitched and the image to be selected.
[0022] The missing feature channels in the image with the highest similarity to the image to be stitched are retrieved, and the missing feature channels are stitched to the image to be stitched in a fixed order to form a complete set of feature channels.
[0023] The beneficial effects of the above-mentioned further solutions are as follows: In complex environments such as coal mines, missing facial feature channels can be accurately located. By selecting the image with the fewest missing feature channels as the image to be stitched, a better foundation is provided for subsequent stitching. When calculating channel attention weights and spatial feature map similarity, the differences between the two are comprehensively considered and weighted by the fusion coefficient, which can more comprehensively and accurately measure the similarity between images, ensuring that the image most similar to the image to be stitched is selected. Stitching the missing feature channels of the most similar image to the image to be stitched in a fixed order can effectively supplement the missing features of the image to be stitched, thereby constructing a complete set of feature channels. This provides sufficient and effective facial features for subsequent accurate face recognition, greatly improving the success rate of face recognition in harsh environments such as coal dust, complex lighting, and protective equipment obstruction. This ensures the reliability and security of access control and security systems in industries such as coal mines, and avoids affecting production efficiency and personnel management order due to recognition failure.
[0024] Based on the above technical solution, the present invention can be further improved as follows: when the face recognition matching module constructs the standard feature vector, the secondary acquisition triggering module receives the missing feature channel signal in the feature channel filtering and matching unit, and outputs the face feature extraction failure signal, establishes the mapping relationship between the failure signal and the secondary acquisition signal, and outputs the secondary acquisition signal synchronously after outputting the failure signal.
[0025] A multimodal biometric intelligent access control security method includes the following steps:
[0026] Step 1: Acquire facial images and convert them into preliminary feature maps with multiple feature channels; extract facial feature vectors from the facial images and construct a standard feature vector set.
[0027] Step 2: Identify the user's identity. If the user identification fails, output a secondary acquisition signal for secondary identification and set an upper limit for the number of comparisons. If the number of comparisons exceeds the upper limit, output an alarm signal.
[0028] Step 3: Analyze whether the facial image is missing feature channels. If there are missing feature channels, construct a complete facial feature vector from different feature channels.
[0029] Step 4: Receive information from missing feature channels and establish a mapping relationship between failure signals and secondary acquisition signals.
[0030] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the figures. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the overall module of the present invention;
[0032] Figure 2 This is a flowchart illustrating the working principle of the multimodal recognition and integration module of the present invention.
[0033] Figure 3 This is a flowchart illustrating the working principle of the failure handling module of the present invention.
[0034] Figure 4 This is a schematic diagram illustrating the working steps of the present invention.
[0035] The meanings of the labels in the diagram are as follows:
[0036] 1. Multimodal recognition integration module; 11. Channel attention enhancement unit; 12. Facial feature vector construction unit; 2. Face recognition matching module; 3. Recognition failure handling module; 31. Feature channel filtering and matching unit; 32. Missing feature channel splicing unit; 4. Secondary acquisition triggering module. Detailed Implementation
[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] refer to Figures 1-4 As shown, a multimodal biometric intelligent access control and security system includes a multimodal recognition integration module 1, a face recognition matching module 2, a recognition failure handling module 3, and a secondary acquisition triggering module 4, wherein:
[0039] To improve the efficiency of face recognition, the multimodal recognition integration module 1 includes a channel attention enhancement unit 11 and a facial feature vector construction unit 12. The channel attention enhancement unit 11 uses camera technology to acquire facial images I and converts them into preliminary feature maps F for the i-th feature channel. The preliminary feature map F is then transformed into a spatial feature map F through channel attention enhancement. c Then, spatial feature map F is enhanced through spatial attention. c Obtain the enhanced feature map F att Facial feature vector construction unit 12 is used to extract enhanced feature map F. att The facial feature vector v for each feature channel c And sort each facial feature vector v in a fixed order. c This yields the final facial feature vector v:
[0040]
[0041] v cLet be the feature vector of the c-th feature channel, representing the average response intensity of the c-th feature channel. To enhance feature map F att The pixel value of the c-th feature channel, the h-th row, and the w-th column; v c Let v be the c-th element of the feature vector, representing the average response intensity of the c-th feature channel; v represents the average response intensity of all feature channels. c Arranged in a fixed order, the final facial feature vector v is obtained;
[0042] The facial feature vector construction unit 12 receives a standard feature vector template construction signal and a face recognition signal. When constructing the standard feature vector template signal, the facial feature vector v extracted by the facial feature vector construction unit 12 is a standard feature vector. Then, multiple standard feature vectors extracted from multiple facial images are used to construct a standard feature vector template and establish a mapping relationship between user identity and standard features. During face recognition, the facial feature vector v extracted by the facial feature vector construction unit 12 is the standard feature vector recognition feature vector, and the recognition feature vector is output to the face recognition matching module 2.
[0043] Channel attention enhancement unit 11 progressively enhances the initial feature map F into feature map Fi. att The working principle is as follows: Receive facial image I∈R H×W×3 H and W represent the height and width of the facial image I, and 3 represents the RGB channels; preliminary feature maps F∈R of multiple feature channels c are obtained through a basic feature extraction network (the basic feature extraction network is a convolutional neural network, CNN, etc.). h×w×c h = H / 3, w = W / s, s is the downsampling rate, c is the number of feature channels, and each feature channel c corresponds to the response distribution of a facial feature (such as eyebrow edge, eye contour, etc.).
[0044] Global max pooling and global average pooling are performed on the spatial dimensions of the initial feature map F to obtain two channel descriptors F. gmax =max h∈{1,2,…,H},w∈{1,2,…,W} F(h,w,:) and F gavg =avg h∈{1,2,…,H},w∈{1,2,…,W} F(h,w,:) is the feature channel vector at spatial location (h,w) of the initial feature map F. max is the maximum value of the spatial dimension for each feature channel, highlighting the most significant facial features within the feature channel (e.g., if a channel has the strongest response in the "eye region", then the maximum value of that region is retained). avg is the average value of the spatial dimension for each feature channel, taking the average value at all spatial locations to reflect the global response intensity of the feature channel.
[0045] Channel descriptor F gmax and channel descriptor F gavgChannel attention weights M are generated by passing each weight through a shared fully connected layer (first reducing dimensionality, then increasing it to reduce computation), summing the results, and then passing them through a sigmoid activation function. c ∈R 1×1×C :
[0046] M c =σ(FC) up (ReLU(FC down (F gmax ))))+FC up (ReLU(FC down (F gavg )));
[0047] FC down For dimensionality reduction, a fully connected layer (reducing the number of feature channels from C to C / r, where r is the reduction rate, e.g., r = 16) is called an FC layer. up For a fully connected layer of increased dimensionality (increasing the number of channels from C / r back to C), ReLU is the activation function and σ is the Sigmoid activation function;
[0048] Then, the channel attention weight M is used. c Multiplying the initial feature map F channel by channel yields the spatial feature map F after channel attention enhancement. c =F⊙M c , where ⊙ represents the element-wise multiplication of all spatial locations of each feature channel by its corresponding channel attention weight M. c Scaling, through channel attention weights M c It can highlight the feature channels that are more important for facial recognition, such as the feature channels of key parts like the eyes and nose, while suppressing those feature channels that are greatly affected by coal dust and oil pollution and contain interfering information. This allows for effective focus on effective feature channels in complex environments, reducing interference from irrelevant information, making the features extracted subsequently more targeted and reliable, and improving the recognition accuracy.
[0049] Channel attention enhancement unit 11 received spatial feature map F c And the spatial feature map F c Max pooling and average pooling are performed on the feature channel dimension to obtain two single-feature-channel spatial descriptors F. c,max =max c∈{1,2,…,C} F(:,:,c) and F avg =avg c∈{1,2,…,C} F(:,:,c), where F(:,:,c) is the spatial feature map F cThe c-th feature channel is a two-dimensional matrix (size H×W). max is the maximum value of all feature channels for each spatial location, highlighting the most significant feature channel response within the location (e.g., if a location has the strongest response in the "eye texture" feature channel, then the maximum value of that feature channel is retained). avg is the element-wise average value, which is the average value of all feature channels for each spatial location, reflecting the global feature channel response intensity of the spatial location.
[0050] Concatenate the space descriptor F in the channel dimension max and space descriptor F avg Furthermore, a spatial attention weight map M is generated through convolutional layers (7×7 convolutional kernels) and a sigmoid activation function. s ∈R H×W×1 :
[0051] M s =σ(Conv(Concat(F) max ,F avg )));
[0052] Where σ is the Sigmoid activation function, used to compress the output to the [0,1] interval, and Concat is the feature channel dimension concatenation operation, which concatenates the spatial descriptor F. max and space descriptor F avg The feature maps are merged into H×W×2 maps. Conv is a convolution operation used to capture long-range spatial dependencies (covering a wider region and determining "what is important"). The spatial attention weight map is M. s By enhancing the features of key facial areas (such as the unobstructed facial portion) and suppressing interference from background and obstructed areas, such as enhancing the feature expression of non-shadowed facial areas in the presence of shadows and highlighting the features of the unobstructed portion when protective equipment partially obscures the face, facial features can be extracted more accurately in complex spatial environments, thereby improving recognition accuracy.
[0053] Spatial attention weight map M s Spatial Feature Map F c Element-wise multiplication yields the enhanced feature map F. att =F c ⊙M s Where ⊙ represents element-wise multiplication, features of regions with high weights are preserved or even enhanced, while features of regions with low weights are suppressed, through the spatial attention weight map M. s Spatial Feature Map F c Element-wise multiplication yields the enhanced feature map F. attThis reduces a large amount of invalid and interfering information; when comparing with the facial feature database in the future, the amount of computation is reduced and the comparison speed is accelerated, thereby improving the efficiency of facial recognition. For example, it is no longer necessary to perform complex calculations on the feature channels corresponding to areas severely polluted by coal dust, and can directly focus on the effective feature channels, thus speeding up the recognition speed.
[0054] The face recognition matching module 2 receives the recognition feature vector and then calculates the similarity between the recognition feature vector and multiple standard feature vectors in each standard feature vector template. If the similarity between the recognition feature and the standard feature corresponding to a user is greater than the similarity threshold, the user's identity is successfully recognized; otherwise, the user's identity is recognized but not recognized. A secondary acquisition signal is then output to the multimodal recognition integration module 1 for secondary acquisition of the user's facial image and extraction of facial feature vectors for secondary recognition by the face recognition matching module 2. An upper limit for the number of comparisons is set. If the number of comparisons exceeds the upper limit, an alarm signal is output to alert staff, allowing them to quickly become aware of any identity recognition anomalies and promptly intervene and verify the information. This effectively prevents security risks such as unauthorized personnel entering the premises due to system malfunctions, severe facial contamination, or other special circumstances. It ensures the safety management of entrances and exits in places such as coal mines, prevents unauthorized personnel from entering the work area, and maintains production safety order.
[0055] This invention further considers that before face recognition, users usually tidy themselves up and try to present a clear and complete face. However, during the acquisition process, due to unstable lighting in working environments such as coal mines, the easy fall and adhesion of coal dust to the face, and the possibility of fatigue, changes in facial muscle state, or even slight deviations in the wearing angle of protective equipment after long-term work, some facial features in the user's facial image I may be obscured by coal dust, blurred due to shadows caused by lighting issues, or the unstable and inconsistent presentation of features due to subtle changes in facial muscles. This can lead to multiple face recognition failures by the face recognition matching module 2 and the false output of alarm signals. Therefore, the recognition failure processing module 3 includes a feature channel filtering and matching unit 31 and a missing feature channel splicing unit 32. The feature channel filtering and matching unit 31 is used to receive face recognition failure signals from the face recognition matching module 2 (the number of failures is less than or equal to the comparison number threshold) and retrieve the channel attention weight M of each facial image when recognition fails. c Spatial feature map F c The pixel value at each position; setting the weight threshold T. w Pixel threshold T w Select the channel attention weight M in each facial image c > Weight threshold T w And spatial feature map F cThe pixel value at each location is greater than the pixel threshold T. w The feature channels are compared to the number of feature channels used to construct the complete facial feature vector v. The attributes of the selected feature channels (such as the attributes of feature channels corresponding to facial features like eyebrow edges and eye contours) are compared to the number of feature channels used to construct the complete facial feature vector v. If the number of feature channels is the same, it means that the feature channels of different facial images can be concatenated to construct the complete facial feature vector v. The expression is: A selected ∩A target =A target A selected Let A be the set of feature channel attributes in a facial image. target The set of feature channel attributes required to construct a complete facial feature vector v; if the number of feature channels is not the same, it is determined that a complete facial feature vector v cannot be constructed.
[0056] Missing feature channel splicing unit 32 receives a signal capable of constructing a complete facial feature vector v, and retrieves the channel attention weights M. c ≤ weight threshold T w Spatial feature map F c The pixel value at each position is ≤ pixel threshold T w The feature channel is the missing feature channel C. missing Record the missing feature channels C in each facial image missing The number of missing feature channels C missing The facial image with the fewest faces is selected as the image to be stitched together, and the remaining facial images are selected as images to be chosen.
[0057] Calculate the C of each feature channel with the same attribute in the image to be stitched and each image to be selected in turn. target Channel attention weights M c Spatial feature map F c Similarity between them: the difference in channel attention weights is ΔM = |M c '-M c |, where M c 'For the feature channel C of the image to be stitched target Channel attention weights, M c Image feature channel C to be selected target Channel attention weights; spatial feature map difference is ΔF = |F c '-F| c , where F c 'For the feature channel C of the image to be stitched target Spatial feature map, F c Image feature channel C to be selected target Spatial feature map;
[0058] Fusion of multiple feature channels C targetThe channel attention weight difference ΔM and spatial feature map difference ΔF are: Δi = αΔM + βΔF, where Δi is the feature channel C. target The difference after fusion, α and β are fusion coefficients, representing the contribution of the channel attention weight difference ΔM and spatial feature map difference ΔF to the fusion process; then, each feature channel C... target The difference Δi is then fused again to obtain the similarity between the image to be stitched and the image to be selected: Where w i The weights for each feature channel are...
[0059] Missing feature channel stitching unit 32 and retrieves the missing feature channel C from the image to be stitched with the highest similarity to the image to be stitched. target The missing feature channel C missing The fixed order of feature channels (such as the spatial logical order of facial features: forehead → eyes → nose → chin, etc.) is stitched to the image to be stitched to form a complete set of feature channels. This effectively supplements missing facial features, improves the accuracy of face recognition, reduces false alarms, and ensures the reliable operation of the coal mine access control system in complex environments. It avoids interference with normal workers due to false alarms and can more accurately identify authorized personnel, maintaining the safety of the work area. Then, the complete facial feature vector is extracted and output to the face recognition matching module 2 for face recognition.
[0060] When the multimodal recognition integration module 1 constructs the standard feature vector, the secondary acquisition trigger module 4 receives the missing feature channel C from the feature channel filtering and matching unit 31. missing The system outputs a face feature extraction failure signal and establishes a mapping relationship between the failure signal and the secondary acquisition signal. After outputting the failure signal, the secondary acquisition signal is output synchronously, which can promptly detect face feature extraction failures. By synchronously triggering secondary acquisition, it provides an opportunity to supplement missing features and construct a complete standard feature vector, thereby improving the integrity and accuracy of the standard feature vector. This lays the foundation for high accuracy in subsequent face recognition, reduces recognition errors caused by feature extraction failures, and ensures the reliability of the access control system for identity recognition in complex environments.
[0061] refer to Figure 4 As shown, a multimodal biometric intelligent access control security method includes the following steps:
[0062] Step 1: Acquire facial images and convert them into preliminary feature maps with multiple feature channels; extract facial feature vectors from the facial images and construct a standard feature vector set.
[0063] Step 2: Identify the user's identity. If the user identification fails, output a secondary acquisition signal for secondary identification and set an upper limit for the number of comparisons. If the number of comparisons exceeds the upper limit, output an alarm signal.
[0064] Step 3: Analyze whether the facial image is missing feature channels. If there are missing feature channels, construct a complete facial feature vector from different feature channels.
[0065] Step 4: Receive information from missing feature channels and establish a mapping relationship between failure signals and secondary acquisition signals.
[0066] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multimodal biometric intelligent access control and security system, characterized in that, It includes a multimodal recognition integration module (1), a face recognition matching module (2), a recognition failure handling module (3), and a secondary acquisition triggering module (4), wherein: The multimodal recognition integration module (1) acquires facial images and converts them into preliminary feature maps of multiple feature channels. The preliminary feature map is enhanced into a spatial feature map, and then the spatial feature map is enhanced to obtain an enhanced feature map. Facial feature vectors are extracted from the enhanced feature map. The facial feature vectors are the recognition feature vector and the standard feature vector, and a set of standard feature vectors is constructed. The face recognition matching module (2) is used to identify the user's identity. If the user's identity recognition fails, a secondary acquisition signal is output to the multimodal recognition integration module (1), and then the recognition is performed again. An upper limit value for the number of comparisons is set. If the number of comparisons exceeds the upper limit value, an alarm signal is output. The recognition failure processing module (3) includes a feature channel filtering and matching unit (31) and a missing feature channel splicing unit (32). The feature channel filtering and matching unit (31) receives the face recognition failure signal, analyzes whether the facial image is missing a feature channel, and then judges whether a complete facial feature vector can be constructed after splicing the feature channels of different facial images. The missing feature channel splicing unit (32) constructs a complete facial feature vector using different feature channels. When the multimodal recognition integration module (1) constructs the standard feature vector, the secondary acquisition triggering module (4) receives the missing feature channel in the feature channel filtering and matching unit (31) and establishes a mapping relationship between the failure signal and the secondary acquisition signal.
2. The multimodal biometric intelligent access control and security system according to claim 1, characterized in that: The multimodal recognition integration module (1) includes a channel attention enhancement unit (11) and a facial feature vector construction unit (12). The channel attention enhancement unit (11) converts the facial image into a preliminary feature map of multiple feature channels. The preliminary feature map F is enhanced by channel attention to obtain a spatial feature map, and then the spatial feature map is enhanced by spatial attention to obtain an enhanced feature map. The facial feature vector construction unit (12) is used to extract the facial feature vector of each feature channel in the enhanced feature map and sort each facial feature vector in a fixed order to obtain the final facial feature vector.
3. The multimodal biometric intelligent access control and security system according to claim 2, characterized in that: The channel attention enhancement unit (11) receives a facial image and obtains multiple preliminary feature maps of feature channels through a basic feature extraction network. Each feature channel corresponds to a response distribution of a facial feature. The spatial dimension of the preliminary feature map F is processed by taking the maximum value of the spatial dimension for each feature channel and taking the average value of the spatial dimension for each feature channel to obtain two channel descriptors. The two channel descriptors are input into a dimension-reducing fully connected layer. The ReLU activation function is applied to the dimension-reduced result to obtain the dimension-reduced result: the value greater than zero in the dimension-reduced result is taken, and the value less than or equal to zero is set to zero. Then the result after ReLU activation is input into a dimension-increasing fully connected layer to obtain the dimension-increasing result. Then the two results after the fully connected layer are added together and then processed by the Sigmoid activation function to generate channel attention weights. The dimension-reducing fully connected layer will reduce the number of feature channels, and the dimension-increasing fully connected layer will increase the number of channels back to the original number. The channel attention enhancement unit (11) multiplies the preliminary feature map by feature channel through the channel attention weight to obtain the spatial feature map after channel attention enhancement.
4. The multimodal biometric intelligent access control and security system according to claim 3, characterized in that: The channel attention enhancement unit (11) receives the spatial feature map and performs max pooling and average pooling on the feature channel dimension of the spatial feature map to obtain two single feature channel spatial descriptors: max pooling is to take the maximum value on all feature channels at each spatial location, and average pooling is to take the average value on all feature channels at each spatial location. Two spatial descriptors are concatenated along the channel dimension and merged into a feature map of a specific size. The concatenated feature map is then input into a convolutional layer to capture long-range spatial dependencies. The result after processing by the convolutional layer is then input into a Sigmoid activation function to generate a spatial attention weight map. Then, the spatial attention weight map and the spatial feature map are multiplied element-wise to obtain the enhanced feature map.
5. The multimodal biometric intelligent access control and security system according to claim 4, characterized in that: The facial feature vector construction unit (12) enhances each feature channel of the feature map, sums the pixel values at all row and column positions in the feature channel, divides the summation result by the product of the number of rows and columns of the feature channel, and generates the feature vector of the feature channel. The facial feature vector construction unit (12) receives a standard feature vector template construction signal and a face recognition signal. When constructing the standard feature vector template signal, the facial feature vector extracted by the facial feature vector construction unit (12) is a standard feature vector. Then, multiple standard feature vectors extracted from multiple facial images are used to construct a standard feature vector template and establish a mapping relationship between user identity and standard features. During face recognition, the facial feature vector extracted by the facial feature vector construction unit (12) is a standard feature vector recognition feature vector, and the recognition feature vector is output to the face recognition matching module (2).
6. The multimodal biometric intelligent access control and security system according to claim 5, characterized in that: The face recognition matching module (2) receives the recognition feature vector and then calculates the similarity between the recognition feature vector and multiple standard feature vectors in each standard feature vector template. If the similarity between the recognition feature and the standard feature corresponding to a user is greater than the similarity threshold, the user identity recognition is successful. Otherwise, the user identity recognition fails, and a second collection of the user's facial image is performed to extract facial feature vectors for comparison. An upper limit for the number of comparisons is set. If the number of comparisons is greater than the upper limit, an alarm signal is output to remind the staff.
7. The multimodal biometric intelligent access control and security system according to claim 6, characterized in that: The feature channel filtering and matching unit (31) is used to receive the face recognition failure signal in the face recognition matching module (2), and retrieve the channel attention weight and the brightness of each position in the spatial feature map of each face image when recognition fails; set the weight threshold and the pixel threshold, select the feature channels in each face image where the channel attention weight is greater than the weight threshold and the brightness of each position in the spatial feature map is greater than the pixel threshold, and compare whether the attributes of the selected feature channels are the same as the number of feature channels used to construct the complete face feature vector. If the number of feature channels is the same, it means that different face images can construct a complete face feature vector after being stitched together.
8. The multimodal biometric intelligent access control and security system according to claim 7, characterized in that: The missing feature channel splicing unit (32) receives a signal that can construct a complete facial feature vector, selects the feature channels with channel attention weight ≤ weight threshold and pixel value ≤ pixel threshold at each position in the spatial feature map as missing feature channels, records the number of missing feature channels in each facial image, selects the facial image with the fewest missing feature channels as the image to be spliced, and the remaining facial images as the images to be selected. Then, the similarity between channel attention weights and spatial feature maps is calculated sequentially. First, the similarity of channel attention weights is calculated as the absolute value of the difference between the two channel attention weights; the similarity of spatial feature maps is calculated as the absolute value of the difference between the pixel values at corresponding positions in the two spatial feature maps. The difference within the feature channels is obtained by multiplying the fusion coefficient of the channel attention weight difference and the spatial feature map difference by the attention weight difference and the spatial feature map difference. Then, the differences of each feature channel are weighted and fused again to obtain the similarity between the image to be stitched and the image to be selected. The missing feature channels in the image with the highest similarity to the image to be stitched are retrieved, and the missing feature channels are stitched to the image to be stitched in a fixed order to form a complete set of feature channels.
9. The multimodal biometric intelligent access control and security system according to claim 8, characterized in that: When the face recognition matching module (2) constructs the standard feature vector, the secondary acquisition trigger module (4) receives the missing feature channel signal in the feature channel filtering and matching unit (31), and outputs the face feature extraction failure signal, establishes the mapping relationship between the failure signal and the secondary acquisition signal, and outputs the secondary acquisition signal synchronously after outputting the failure signal.
10. A method for a multimodal biometric intelligent access control and security system, applied to the multimodal biometric intelligent access control and security system according to any one of claims 1-9, characterized in that, Includes the following steps: Step 1: Acquire facial images and convert them into preliminary feature maps with multiple feature channels; extract facial feature vectors from the facial images and construct a standard feature vector set. Step 2: Identify the user's identity. If the user identification fails, output a secondary acquisition signal for secondary identification and set an upper limit for the number of comparisons. If the number of comparisons exceeds the upper limit, output an alarm signal. Step 3: Analyze whether the facial image is missing feature channels. If there are missing feature channels, construct a complete facial feature vector from different feature channels. Step 4: Receive information from missing feature channels and establish a mapping relationship between failure signals and secondary acquisition signals.