A one-way gate control method, system, and device based on image recognition

By combining feature fusion technology of polarized light images and depth point cloud data in a one-way door control system, the problem of interference from material reflection characteristics in highly reflective scenes is solved, achieving high security and reliability of identity verification and reducing the false recognition rate and the risk of accidental door opening.

CN120375503BActive Publication Date: 2026-03-06TIANJIN XINLIJIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510440313.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2026-03-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Existing one-way door control solutions cannot effectively eliminate interference from material reflection characteristics in highly reflective scenarios, leading to authentication failure and easily causing false rejection or false opening of the access control system, thus affecting security and reliability.

Method used

By acquiring polarized light images and binocular depth point cloud data in highly reflective scenes, spatial alignment and feature stitching are performed to generate joint polarization features of the point cloud. Feature fusion is performed using a multi-scale attention mechanism and dynamic mask to generate multimodal feature vectors. Identity verification is performed by combining a classification network, and the consistency between polarization features and depth features is corrected by a material geometry correlation model. The fusion weights are dynamically adjusted to suppress noise interference in reflective areas.

Benefits of technology

It significantly improves the robustness of facial verification in highly reflective scenarios, reduces the false rejection rate and the risk of accidental door opening, and ensures the security and reliability of the one-way access control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375503B_ABST
    Figure CN120375503B_ABST
Patent Text Reader

Abstract

This application discloses a one-way door control method, system, and device based on image recognition, belonging to the field of image processing. This application simultaneously acquires polarized light images, binocular depth point cloud data, and face images of the target area; it fuses the polarized light images with the three-dimensional features of the depth point cloud through spatial registration to construct joint polarization features of the point cloud; it aligns the joint features with the face image pixels and generates a dynamic mask using the polarization degree data of the polarized light image; under the constraint of the dynamic mask, it employs a multi-scale attention mechanism to perform cross-modal fusion of the face image and spatial feature map, extracting multimodal feature vectors containing facial spatial structure, material properties, and texture features; it compares the similarity of this vector with a preset database and outputs verification information through a classification network; based on the mapping relationship between the verification information and the control strategy, it triggers adaptive control commands for the one-way door access control system. This enables accurate and effective control of the one-way door in complex reflective scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing, and in particular relates to a one-way gate control method, system and device based on image recognition. Background Technology

[0002] In highly reflective environments such as laboratories, cleanrooms, and medical isolation areas, reliable identity verification of one-way door access control systems faces significant challenges. In these scenarios, metal door frames, glass viewing windows, and high-intensity ambient lighting cause strong specular reflections between the background and the person's image, resulting in overexposure of facial images and identity verification failures. This can easily lead to false rejections or false openings, threatening area security. Therefore, there is an urgent need for a one-way door control method that can adapt to high reflectivity interference to ensure the accuracy and effectiveness of one-way door control.

[0003] Existing one-way door control solutions typically reduce environmental reflection interference through infrared illumination or polarizing filters. However, the filter parameters are fixed and cannot dynamically adapt to complex reflective scenarios. The fixed filtering strategies of existing one-way door control solutions cannot effectively eliminate interference from the material reflection characteristics of highly reflective areas, making it difficult to distinguish between reflective artifacts and the highlight areas of a real face. This can lead to authentication failures and false rejections or openings. Therefore, existing methods suffer from reduced accuracy and effectiveness of one-way door control in complex reflective scenarios. Summary of the Invention

[0004] This application provides a one-way door control method, system, device, and computer-readable storage medium based on image recognition, which can accurately and effectively control one-way doors in complex reflective scenarios.

[0005] In a first aspect, embodiments of this application provide a one-way gate control method based on image recognition, the method comprising:

[0006] Acquire polarized light images of the target area under highly reflective scenes, as well as face images and binocular depth point cloud data of the human figure area within the target area. The target area includes the human figure area and the background area.

[0007] Spatial alignment and feature stitching are performed between polarized light images and binocular depth point cloud data to generate joint polarization features of the point cloud.

[0008] The point cloud polarization joint features are spatially aligned with the face image to obtain the aligned spatial feature map. Based on the polarization degree of the polarized light image, the spatial feature map is partitioned by reflective intensity to generate a dynamic mask.

[0009] By utilizing a multi-scale attention mechanism and using a dynamic mask as a constraint, feature fusion is performed on spatial feature maps and face images to generate multimodal feature vectors that include spatial structure, material properties, and texture information.

[0010] The similarity between the multimodal feature vector and the identity features in the pre-set face identity database is calculated using a classification network to obtain target verification information;

[0011] Based on the pre-defined correspondence between verification information and control information, the target control information corresponding to the target verification information is determined, and the one-way gate is controlled through the target control information.

[0012] In one feasible implementation, a multi-scale attention mechanism is used, constrained by a dynamic mask, to fuse spatial feature maps and face images, generating a multimodal feature vector that includes spatial structure, material properties, and texture information, including:

[0013] Using a multi-scale attention mechanism and a dynamic mask as a constraint, and based on a material geometry association model, feature fusion is performed on spatial feature maps and face images to generate multimodal feature vectors that include spatial structure, material properties, and texture information. The material geometry association model is trained based on the polarization angle of historical polarized light images and the face geometric parameters of historical spatial feature maps through a predefined correspondence between skin polarization angle and curvature.

[0014] In one feasible implementation, before utilizing a multi-scale attention mechanism, constrained by a dynamic mask, to perform feature fusion on spatial feature maps and face images to generate multimodal feature vectors including spatial structure, material properties, and texture information, the method includes:

[0015] Consistency verification is performed on polarization features and depth features in the spatial feature map. If a conflict region is detected in the human figure region where the polarization degree is lower than the preset polarization threshold and the depth value difference between adjacent regions exceeds the preset depth threshold, the conflict region in the spatial feature map is determined.

[0016] The depth values ​​of conflict regions in the spatial feature map are repaired by using neighborhood feature interpolation, resulting in a repaired spatial feature map.

[0017] In one feasible implementation, before spatially aligning the point cloud polarization joint features with a face image to obtain an aligned spatial feature map, and before generating a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image to reflect light intensity, the method includes:

[0018] The polarized light image is decomposed into high-frequency detail components and low-frequency structural components, which are then fused across scales with the geometric information in the point cloud polarization joint features to generate point cloud polarization layered features.

[0019] The point cloud polarization joint features are spatially aligned with the face image to obtain an aligned spatial feature map. Based on the polarization degree of the polarized light image, the spatial feature map is partitioned by reflective intensity to generate a dynamic mask, including:

[0020] The point cloud polarization layer features are spatially aligned with the face image to obtain an aligned spatial feature map. Based on the polarization degree of the polarized light image, the spatial feature map is partitioned by reflective intensity to generate a dynamic mask.

[0021] In one feasible implementation, a multi-scale attention mechanism is used, constrained by a dynamic mask, and feature fusion is performed on the spatial feature map and face image based on a material geometry correlation model to generate a multimodal feature vector including spatial structure, material properties, and texture information, including:

[0022] Based on the predefined polarization angle and curvature correspondence in the material geometry correlation model, the target curvature range of each region in the spatial feature map is determined;

[0023] Based on the dynamic mask, the high reflectivity suppression region and normal region in the spatial feature map are divided, and an initial fusion weight is assigned to each region.

[0024] For each level of features in the spatial feature map, the polarization angle distribution and actual curvature value are extracted, the difference between the curvature value and the target curvature range is calculated, and the initial fusion weight of the corresponding region is adjusted according to the difference.

[0025] The adjusted initial fusion weights are superimposed with the face image features of the corresponding level to generate weighted fusion features layer by layer, and the resulting multimodal feature vector is output after merging.

[0026] In one feasible implementation, the polarized light image is spatially aligned and feature-stitched with the binocular depth point cloud data to generate joint polarization features of the point cloud, including:

[0027] For each pixel in the polarized light image, the corresponding three-dimensional spatial point in the binocular depth point cloud is determined through the physical position mapping relationship, and an association index table between the pixel and the three-dimensional spatial point is established.

[0028] Extract the polarization parameters of each pixel in the polarized light image. The polarization parameters include the direction of polarization intensity change and the difference in polarization intensity.

[0029] Based on the association index table, the polarization parameters are added to the corresponding three-dimensional spatial points of the association index to obtain joint data points including spatial coordinates, color information and polarization parameters;

[0030] Based on the geometric distribution density of the binocular depth point cloud, the polarization parameters of the joint data points are regionally smoothed to obtain smoothed joint data points.

[0031] The smoothed joint data point set is used as the joint polarization feature of the point cloud.

[0032] In one feasible implementation, a classification network is used to calculate the similarity between the multimodal feature vector and the identity features in a pre-defined face identity database to obtain target verification information, including:

[0033] The multimodal feature vector is divided into multiple feature blocks according to spatial structure, material properties, and texture information, and the feature expression value of each feature block is extracted.

[0034] Obtain the feature template of the target identity from the facial identity database. The feature template includes the historical average and fluctuation range of the block features for each type.

[0035] Calculate the feature expression value of each block feature and the feature difference ratio of the corresponding historical mean, and determine the weight coefficient of each block feature based on the fluctuation range;

[0036] The target matching degree is obtained by weighting and summing the feature difference ratios based on the weight coefficients of each feature block.

[0037] The target verification information is obtained by comparing the target matching degree with the preset matching degree threshold.

[0038] Secondly, embodiments of this application provide a one-way door control system based on image recognition, the system comprising:

[0039] The acquisition module is used to acquire polarized light images of the target area under high reflectivity scenes, as well as face images and binocular depth point cloud data of the human figure area in the target area. The target area includes the human figure area and the background area.

[0040] The generation module is used to spatially align and feature-stitch the polarized light image with the binocular depth point cloud data to generate joint polarization features of the point cloud.

[0041] The generation module is also used to spatially align the point cloud polarization joint features with the face image to obtain the aligned spatial feature map, and generate a dynamic mask by partitioning the spatial feature map by reflective intensity based on the polarization degree of the polarized light image.

[0042] The generation module is also used to utilize a multi-scale attention mechanism, with dynamic masks as constraints, to perform feature fusion on spatial feature maps and face images, generating multimodal feature vectors that include spatial structure, material properties, and texture information;

[0043] The verification module is used to calculate the similarity between the multimodal feature vector and the identity features in the preset face identity database using a classification network, so as to obtain the target verification information.

[0044] The control module is used to determine the target control information corresponding to the target verification information based on the preset correspondence between verification information and control information, and to complete the control of the one-way door through the target control information.

[0045] Thirdly, embodiments of this application provide an electronic device, the device including: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the image recognition-based one-way gate control method as described in any embodiment of the first aspect.

[0046] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the image recognition-based one-way gate control method as described in any embodiment of the first aspect.

[0047] This application discloses a one-way door control method, system, device, and computer-readable storage medium based on image recognition. First, it quantifies the reflectivity of polarized light images using polarization measurements and combines this with geometric information from depth point clouds to dynamically divide reflectivity suppression regions, eliminating specular reflection artifacts on metal and glass surfaces. Second, it utilizes multimodal feature fusion to differentiate real faces using polarization characteristics, such as low-polarization skin and high-reflectivity materials like high-polarization metals or glass. Furthermore, it strengthens facial texture and structural features through a multi-scale attention mechanism, accurately extracting identity information even under strong reflectivity interference. Finally, it employs a feature weighting strategy based on dynamic mask constraints to suppress invalid noise in reflective areas, improving the robustness of face verification. This enables reliable one-way door access control in high-security scenarios such as laboratories and medical isolation areas, significantly reducing false rejection rates and the risk of accidental door opening.

[0048] Furthermore, by introducing a material geometry correlation model, the interference of material reflection characteristics on normal estimation in highly reflective scenes is effectively solved. Specifically, this model, based on the correspondence between polarization angle and curvature of historical polarized light images, utilizes predefined facial skin polarization characteristics, such as the physical correlation between polarization angle and curvature, to dynamically correlate polarization parameters with 3D geometric features. It can distinguish between metal or glass reflection artifacts and the diffuse reflection characteristics of a real face, for example, by correcting specular reflection distortion in depth point clouds through curvature differences, thereby improving the accuracy of material property discrimination. Simultaneously, by combining a multi-scale attention mechanism to perform layered fusion of spatial feature maps and face images, the fusion weights can be adaptively adjusted to suppress noise interference in highly reflective areas and enhance the complementarity between texture details and 3D structures. Ultimately, this solution significantly improves the spatial consistency of facial features and reduces the false recognition rate caused by material reflection in highly reflective scenes such as laboratories and medical isolation areas, ensuring the security and reliability of one-way access control systems. Attached Figure Description

[0049] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating a one-way gate control method based on image recognition provided in one embodiment of this application;

[0051] Figure 2 This is a flowchart illustrating a method for generating multimodal feature vectors according to an embodiment of this application;

[0052] Figure 3 This is a flowchart illustrating a method for generating point cloud polarization joint features according to an embodiment of this application;

[0053] Figure 4 This is a schematic diagram of the structure of a one-way door control system based on image recognition provided in one embodiment of this application;

[0054] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0055] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0056] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0057] In highly reflective environments such as laboratories, cleanrooms, and medical isolation areas, reliable identity verification of one-way door access control systems faces significant challenges. In these scenarios, metal door frames, glass viewing windows, and high-intensity ambient lighting cause strong specular reflections between the background and the person's image, resulting in overexposure of facial images and identity verification failures. This can easily lead to false rejections or false openings, threatening area security. Therefore, there is an urgent need for a one-way door control method that can adapt to high reflectivity interference to ensure the accuracy and effectiveness of one-way door control.

[0058] Existing one-way door control solutions typically reduce environmental reflection interference through infrared illumination or polarizing filters. However, the filter parameters are fixed and cannot dynamically adapt to complex reflective scenarios. The fixed filtering strategies of existing one-way door control solutions cannot effectively eliminate interference from the material reflection characteristics of highly reflective areas, making it difficult to distinguish between reflective artifacts and the highlight areas of a real face. This can lead to authentication failures and false rejections or openings. Therefore, existing methods suffer from reduced accuracy and effectiveness of one-way door control in complex reflective scenarios.

[0059] To address the problems of the prior art, embodiments of this application provide a one-way door control method, system, device, and computer storage medium based on image recognition. The image recognition-based one-way door control method provided in this application embodiment will be described first below.

[0060] Figure 1 A flowchart illustrating a one-way gate control method based on image recognition according to an embodiment of this application is shown. Figure 1 As shown, the method includes steps S110 to S160.

[0061] S110: Acquire polarized light images of the target area under high reflectivity scenes, as well as face images and binocular depth point cloud data of the human figure area in the target area. The target area includes the human figure area and the background area.

[0062] The target area refers to the physical space that needs to be identified within the monitoring range of a one-way access control system. This includes the human face area (the area where the face to be identified is located) and the background area, such as metal door frames, glass windows, and other objects that may produce reflections. Polarized light images are two-dimensional images acquired through a polarization sensor. Each pixel contains optical parameters such as polarization angle and degree of polarization, used to quantify the reflective properties of an object's surface, such as specular or diffuse reflection. Face images are two-dimensional color images captured by an RGB camera, containing facial texture and color information, used to extract facial features such as facial contours and skin texture. Binocular depth point cloud data refers to three-dimensional point cloud data generated by a binocular stereo vision system. Each point contains three-dimensional spatial coordinates (X / Y / Z) and color information, used to describe the geometric structure of the human face area, such as facial curvature and depth differences.

[0063] Multi-sensor synchronous triggering involves deploying a polarization camera, an RGB camera, and a binocular stereo vision device in the target area. Hardware synchronization signals ensure consistency in the timestamps of polarized light images, face images, and depth point cloud data. The polarization camera uses multi-angle polarizers (e.g., 0°, 45°, 90°, 135°) to capture images in a time-division manner, calculating the polarization angle and degree of polarization for each pixel using Stokes vectors to generate a polarized light image. The binocular camera, based on the parallax principle, uses feature matching algorithms such as SIFT or ORB to calculate the parallax of corresponding points in the left and right images, combining this with camera calibration parameters to convert it into 3D point cloud data, generating a binocular depth point cloud.

[0064] For example, polarization cameras and binocular cameras are installed on both sides of the one-way door frame at the laboratory entrance. When a person approaches, the system synchronously triggers the sensors to collect data. The polarization camera captures the polarization parameters of the reflective area of ​​the door frame glass (background) and the person's face (image), while the binocular camera generates a 3D point cloud of the face, such as the height of the bridge of the nose and the depth of the eye sockets.

[0065] S120: Spatially align and feature-stitch the polarized light image with the binocular depth point cloud data to generate joint polarization features of the point cloud.

[0066] Spatial alignment refers to establishing a mapping relationship between pixels in a 2D polarized image and corresponding coordinate points in a 3D point cloud, ensuring that the data are in the same coordinate system. Feature stitching refers to fusing polarization parameters with the geometric and color information of the point cloud to form joint data containing multi-dimensional attributes. Point cloud polarization joint features: a multi-dimensional dataset that fuses the spatial coordinates and color information of a binocular depth point cloud with the polarization parameters (polarization angle, degree of polarization) of a polarized light image, characterizing the 3D geometry and material reflection properties of the target region.

[0067] First, using the calibration parameters of the polarization camera and the binocular camera, a physical position mapping relationship between pixels and 3D points is established. Projection calculations determine the corresponding 3D coordinates of each pixel in the polarization image within the point cloud, generating an association index table. Second, the polarization direction and intensity difference values ​​of each pixel are extracted from the polarized light image. Based on the index table, these parameters are appended to the corresponding 3D point cloud data, ensuring that each 3D point contains coordinate, color, and polarization information. Finally, all processed 3D point data are integrated to form a point cloud polarization joint feature, which simultaneously includes the spatial structure, surface reflection characteristics, and color information of the target region.

[0068] S130: Spatial alignment of the point cloud polarization joint features with the face image is performed to obtain the aligned spatial feature map. Based on the polarization degree of the polarized light image, the spatial feature map is partitioned by reflective intensity to generate a dynamic mask.

[0069] Spatial alignment refers to registering the point cloud polarization joint features (3D spatial data) with a face image (2D image) in a unified coordinate system to ensure their consistent position in physical space. The spatial feature map is the aligned multidimensional data matrix containing the geometric structure, polarization characteristics, and RGB texture information of the portrait region. Polarization degree refers to the quantized value of the difference in polarization intensity of each pixel in a polarized light image, used to distinguish between specular reflection (high polarization degree) and diffuse reflection (low polarization degree) regions. A dynamic mask is a binary mask that divides the spatial feature map based on polarization degree, used to mark highly reflective areas and normal areas.

[0070] First, using multi-sensor calibration parameters, such as camera intrinsic and extrinsic parameters, a mapping relationship is established between the 3D coordinates of the point cloud polarization joint features and the pixel coordinates of the face image. For example, the 3D points are projected onto the 2D image plane using the projection matrix of the binocular camera, generating a one-to-one coordinate index table. Next, for sparse areas of the point cloud, bilinear interpolation or nearest-neighbor interpolation is used to complete the missing polarization parameters and face texture information, ensuring that each pixel has complete geometric and polarization attributes in the spatial feature map. Then, the polarization degree channel of the polarized light image is extracted, and an adaptive thresholding algorithm, such as the Otsu method, is used to divide high-reflectivity regions (i.e., polarization degrees higher than the threshold) and normal regions. Morphological operations, such as dilation and erosion, are combined to eliminate noise interference, mark the boundaries of continuous high-reflectivity regions, and generate a binary mask. Finally, the transition weights of the mask edges are adjusted according to the polarization degree gradient to avoid artifacts caused by hard segmentation, ensuring that illumination interference in high-reflectivity regions is effectively suppressed during subsequent feature fusion.

[0071] For example, a stereo camera acquires a depth point cloud of the face, such as the height of the bridge of the nose and the depth of the eye sockets, while a polarization camera captures the polarization parameters of the face and background (such as the specular reflection polarization angle of the glass door frame). The 3D point cloud is mapped to a 2D face image using calibration parameters, ensuring that the tip of the nose is aligned with the image center. The polarization degree image shows that the polarization degree of the glass area of ​​the door frame is as high as 0.8 (strong specular reflection), while the polarization degree of the facial skin area is less than 0.2 (diffuse reflection). The system uses adaptive threshold segmentation to mark the glass area as a high-reflectivity suppression area and generate a mask; the face area is retained as a normal area for subsequent identity feature extraction. When changes in ambient light cause fluctuations in the reflectivity of the door frame, the mask is dynamically updated according to the real-time polarization degree.

[0072] S140: Utilizing a multi-scale attention mechanism and a dynamic mask as a constraint, feature fusion is performed on spatial feature maps and face images to generate multimodal feature vectors that include spatial structure, material properties, and texture information.

[0073] Multi-scale attention mechanisms are feature fusion techniques that extract features at different resolutions or receptive fields in parallel and dynamically allocate the importance of features at each scale using attention weights to enhance the joint modeling capability of local details and global structure. This mechanism can simultaneously capture the microscopic texture (e.g., pores) and macroscopic geometric structure (e.g., facial contours) of a face. Dynamic mask constraints apply spatial restrictions to attention weights during feature fusion, forcing the model to reduce feature responses in highly reflective suppression regions and preserve or enhance feature expressions in normal regions. This constraint ensures that the fusion process is not affected by reflective artifacts through weight adjustment or region feature suppression. Multimodal feature vectors are high-dimensional vectors obtained by fusing spatial feature maps (including polarization parameters and geometric structure) and face images (including RGB textures), containing three key identity features: spatial structure, material properties (e.g., skin reflection characteristics), and texture information (e.g., facial feature details).

[0074] First, pyramid downsampling is performed on both the spatial feature map and the face image to generate multi-scale feature maps (e.g., original resolution, 1 / 2 scale, 1 / 4 scale). For example, dilated convolution or pooling operations are used to extract local features (e.g., eye contour) and global features (e.g., head pose) under different receptive fields. At each scale level, the geometric-polarization features of the spatial feature map and the RGB texture features of the face image are calculated to form hierarchical feature pairs. Next, a dynamic mask is used as a spatial attention template, applying suppression weights (e.g., weights approaching 0) to the feature channels of highly reflective areas, while retaining the original weights for normal areas. For example, element-wise multiplication of the mask matrix with the feature map eliminates the interference of door frame reflections on the face region. A channel attention mechanism (e.g., the SENet module) is introduced to adaptively adjust channel weights based on the response intensity of the feature channels, enhancing the expression of key material properties (e.g., skin polarization angle). Finally, a cross-attention mechanism is used to interact the geometric-polarization features of the spatial feature map with the texture features of the face image. For example, by using a query-key-value structure, a similarity matrix between different modal features is calculated to generate weighted fusion features. The fusion results at multiple scale levels are then upsampled and concatenated level by level, and the original information is preserved through residual connections. Finally, a multimodal feature vector containing multi-dimensional identity features is output.

[0075] S150: Calculate the similarity between the multimodal feature vector and the identity features in the preset face identity database using a classification network to obtain target verification information.

[0076] A classification network is a deep learning-based discriminative model that calculates the similarity between input features and a pre-defined database to output an authentication result. It typically employs cosine similarity or cross-entropy loss functions, combined with a multilayer perceptron (MLP) or support vector machine (SVM). Similarity refers to the degree of matching between multimodal feature vectors and feature templates in the database, derived by quantifying the difference ratio and weighted summation, and used to determine whether an identity is legitimate. Target verification information refers to the verification results output by the classification network, including successful and failed match indicators and confidence levels, serving as the basis for controlling the opening or closing of one-way doors.

[0077] First, the multimodal feature vectors are aligned with feature templates in a pre-defined face identity database to ensure consistency in feature dimensions and semantics. For example, normalization is used to eliminate dimensional differences between data from different sensors. Next, a classification network, such as a neural network or support vector machine, is used to calculate the similarity between the multimodal feature vectors and each feature template in the database. By comparing differences in key features, such as geometric structure, material properties, and texture information, a matching score is output. Finally, based on the comparison between the similarity score and a pre-defined threshold, it is determined whether the identity is legitimate. If the matching degree exceeds the threshold, a "verification passed" message is generated; otherwise, it is marked as "verification failed."

[0078] S160: Based on the preset correspondence between verification information and control information, determine the target control information corresponding to the target verification information, and complete the control of the one-way door through the target control information.

[0079] The predefined correspondence between verification and control information refers to a predefined rule base that clearly defines the mapping relationship between different verification results (e.g., pass / fail / abnormal) and one-way door access control actions (e.g., open / alarm / retry). Target control information refers to the instruction signals generated based on the verification results, including door lock drive commands, alarm signals, or log recording commands, which directly control the one-way door actuator. The control instruction distribution module refers to the hardware interface or communication protocol that converts the target control information into physical signals (e.g., relay on / off, motor drive pulses) to drive the access control actuator.

[0080] First, the target verification information output by the classification network, such as "verification passed, confidence level 0.92," is matched against conditions in a preset rule base. The rule base is defined based on business requirements; for example, if the verification status is "passed" and the confidence level is higher than a threshold, an opening command is triggered; if the status is "failed," an alarm is activated and a log is recorded. Next, target control information is generated based on the matching results. For example, "verification passed" generates a control command containing parameters such as door lock opening time and motor speed; "verification failed" generates an audible and visual alarm signal and an identity anomaly log. The command format must be compatible with the access control hardware interface protocol, such as Modbus or CAN bus. Finally, the target control information is sent to the access control actuator, such as an electromagnetic lock or motor drive board, through the control command distribution module. The system monitors the execution status in real time, such as whether the door lock has successfully unlocked. If no feedback is received within a timeout, an abnormal retry mechanism is triggered to ensure control reliability.

[0081] For example, when the target verification information is "verification passed, confidence level 0.95", the system matches the "high confidence level passed" condition in the rule base and generates a "open door for 5 seconds" command. The control command distribution module sends the command to the motor controller via the RS485 bus to drive the door lock to open. The door status sensor provides real-time feedback of an "open" signal, and the system records the passage log. If the verification result is "failed", the system triggers a buzzer alarm and pushes an exception information to the management terminal. When a network failure causes the command distribution to time out, the system enables a local cache command retransmission mechanism to ensure that the one-way door access control status is controllable.

[0082] This embodiment first uses polarized light images to quantify reflectivity intensity and combines it with geometric information from depth point clouds to dynamically divide reflectivity suppression regions, eliminating specular reflection artifacts on metal and glass surfaces. Second, multimodal feature fusion utilizes polarization characteristics to distinguish real faces, such as low-polarization skin and high-reflectivity materials, such as high-polarization metals or glass. A multi-scale attention mechanism is used to enhance facial texture and structural features, enabling accurate extraction of identity information even under strong reflectivity interference. Finally, a feature weighting strategy based on dynamic mask constraints suppresses invalid noise in reflective areas, improving the robustness of face verification. This enables reliable one-way access control in high-security scenarios such as laboratories and medical isolation areas, significantly reducing false rejection rates and the risk of accidental door opening.

[0083] In highly reflective scenes, strong specular reflection not only interferes with the background area but also creates dynamic specular artifacts on facial surfaces, such as those caused by reflections from glasses, highly confusing with the texture features of a real face. Furthermore, the material properties of polarized light images, such as the polarization angle, have a cross-modal correlation with the spatial geometric features of facial images. However, in dynamically highly reflective scenes, such as reflections from glass curtain walls or specular reflections from metal surfaces, strong background reflections distort the physical properties of polarized light, disrupting the inherent mapping relationship between material properties and geometric structures. For example, the pseudo-polarization angle generated by glass reflection may be incorrectly mapped to the low-curvature cheek area, while the high-curvature nose area of ​​real skin may lose its polarization features due to reflection interference. This cross-modal feature misalignment can cause semantic conflicts in multimodal feature fusion, such as incorrectly associating the material properties of reflective artifacts with the geometric structure of a real face. This application solves the above-mentioned technical problems through the following technical solution.

[0084] In one feasible implementation, step S140: Utilizing a multi-scale attention mechanism, with a dynamic mask as a constraint, feature fusion is performed on the spatial feature map and the face image to generate a multimodal feature vector including spatial structure, material properties, and texture information, including:

[0085] Using a multi-scale attention mechanism and a dynamic mask as a constraint, and based on a material geometry association model, feature fusion is performed on spatial feature maps and face images to generate multimodal feature vectors that include spatial structure, material properties, and texture information. The material geometry association model is trained based on the polarization angle of historical polarized light images and the face geometric parameters of historical spatial feature maps through a predefined correspondence between skin polarization angle and curvature.

[0086] Material geometry correlation model is a physical constraint model that establishes a mapping between material properties (polarization angle) and geometric structure (curvature) by using a predefined correspondence between skin polarization angle and curvature (e.g., the high curvature of the nose tip corresponds to a specific polarization angle distribution), thereby solving the problem of cross-modal feature misalignment.

[0087] First, pyramid downsampling is performed on the spatial feature map (including polarization parameters and geometric structure) and the face image to generate multi-scale feature maps (e.g., original, 1 / 2, and 1 / 4 resolution). At each level, a spatial attention module, such as SENet, is used to calculate feature channel weights, and a dynamic mask is introduced to apply suppression weights to the feature channels of highly reflective areas. For example, the mask matrix is ​​multiplied element-wise with the feature map to make the weights of reflective areas approach 0. Next, based on a material geometry correlation model, the nonlinear relationship between the polarization angle and curvature of the skin region is learned from historical data. For example, the greater the curvature, the more significant the change in polarization angle. The consistency of polarization angle and curvature of each region in the current spatial feature map is verified in real time. For inconsistent regions, such as those with abnormal polarization angles but curvature consistent with facial features, feature correction is performed. A cross-attention mechanism interacts the geometric-polarization features with the facial texture features to generate weighted fused features. Finally, the fused features at multiple scales are upsampled and concatenated, and the original information is preserved through residual connections. The final output is a multimodal feature vector containing spatial structure (such as nose bridge height), material properties (such as skin reflection characteristics), and texture information (such as pore details).

[0088] For example, in highly reflective scenes, strong specular reflection from metal door frames and glass viewing windows causes overexposure of facial images. Feature fusion is achieved through the following process: First, dynamic mask generation is performed: polarization analysis shows that the polarization degree of the door frame area is as high as 0.8 (specular reflection), while the polarization degree of the facial skin area is less than 0.2 (diffuse reflection). The system generates a mask to suppress the reflective areas of the door frame. Second, multi-scale feature fusion is performed: at the 1 / 4 resolution level, the spatial feature map extracts the geometric curvature of the face, such as a nasal tip curvature of 0.5. At the same time, the polarization angle distribution shows that the polarization angle of the nasal tip area is 45°, which conforms to the preset relationship in the material geometry association model, such as a curvature of 0.5 corresponding to a polarization angle of 45°±5°, and this area is given high weight. In the reflective areas of the door frame, the mask suppresses the features of this area to avoid the polarization angle of the glass reflection, for example, a polarization angle of 80°, interfering with facial feature fusion. Finally, the multimodal features are output: the final feature vector integrates the three-dimensional structure of the bridge of the nose (depth point cloud), the polarization characteristics of the skin material (low polarization degree) and facial texture (such as wrinkles around the eyes), providing robust input for subsequent identity verification.

[0089] This embodiment effectively solves the problem of interference between material reflection characteristics and normal estimation in highly reflective scenes by introducing a material geometry correlation model. Specifically, the model is based on the correspondence between polarization angle and curvature of historical polarized light images, and utilizes predefined facial skin polarization characteristics, such as the physical correlation between polarization angle and curvature, to dynamically correlate polarization parameters with 3D geometric features. It can distinguish between metal or glass reflection artifacts and the diffuse reflection characteristics of real faces, for example, by correcting specular reflection distortion in depth point clouds through curvature differences, thereby improving the accuracy of material property discrimination. Simultaneously, by combining a multi-scale attention mechanism to perform layered fusion of spatial feature maps and face images, the fusion weights can be adaptively adjusted to suppress noise interference in highly reflective areas and enhance the complementarity between texture details and 3D structures. Ultimately, this solution significantly improves the spatial consistency of facial features and reduces the false recognition rate caused by material reflection in highly reflective scenes such as laboratories and medical isolation areas, ensuring the security and reliability of one-way access control systems.

[0090] Because of the physical inconsistency between polarization and depth features in multimodal fusion of polarized light images and face images, feature conflicts occur in areas with low polarization and abrupt depth changes, such as reflective materials on the face surface and edge contours. Specifically, when the polarization degree of the portrait area is lower than a preset threshold due to specular reflection or environmental interference, traditional depth perception algorithms will produce abnormal depth jumps due to polarization information distortion, such as the depth difference between adjacent pixels exceeding the threshold. This contradictory feature directly leads to structural misalignment and material misjudgment in the fused multimodal feature vector. Polarized imaging is sensitive to surface materials, while depth perception depends on geometric structure. Feature conflicts can disrupt the complementarity of multimodal information. Especially in scenarios such as financial-grade face recognition and liveness detection, subtle material reflections, such as those from eyeglasses, oily skin, and depth jumps in the real face contour, may be misjudged as fraud attacks, threatening system security. This application solves the above technical problems through the following technical solution.

[0091] In one feasible implementation, before step S140: using a multi-scale attention mechanism and a dynamic mask as a constraint to perform feature fusion on the spatial feature map and the face image to generate a multimodal feature vector including spatial structure, material properties, and texture information, the method includes:

[0092] The consistency of polarization features and depth features in the spatial feature map is checked. If a polarization degree is lower than the preset polarization threshold and the depth value difference between adjacent regions exceeds the preset depth threshold in the human image region, the conflict region in the spatial feature map is identified. The depth value of the conflict region in the spatial feature map is repaired by interpolation of neighborhood features to obtain the repaired spatial feature map.

[0093] Polarization features refer to the degree of linear polarization (DoLP) and angle of linear polarization (AoLP) of each pixel in a polarized light image, used to quantify the reflective properties of an object's surface. For example, high polarization corresponds to specular reflection, such as a metal door frame, while low polarization corresponds to diffuse reflection, such as human skin. Depth features refer to 3D coordinate data derived from binocular depth point clouds, describing the geometric structure of the target region, such as facial curvature and depth differences. For example, the depth value at the bridge of the nose is higher than that at the cheek, while the depth value at the eye socket is recessed. Consistency verification refers to verifying the matching degree between polarization features and depth features through physical correlation. For example, skin regions should have low polarization and continuously changing curvature. If a region has an abnormally low polarization but adjacent depths change abruptly, it indicates a feature contradiction caused by reflective interference. Conflict regions refer to abnormal regions within a portrait area where polarization features and depth features show significant contradictions. For example, if a skin region's polarization is below a threshold due to glass reflection, it is inconsistent with skin characteristics, and the adjacent depth values ​​differ too much, it is inconsistent with smooth facial geometry. Neighborhood feature interpolation refers to the process of repairing abnormal depth values ​​based on the depth value distribution of valid pixels around the conflict region, using the assumption of spatial continuity, such as using bilinear interpolation or the K-nearest neighbor algorithm to fill in missing data.

[0094] First, the polarization degree of each pixel in the extracted spatial feature map is compared with a preset threshold, such as the typical upper limit of polarization degree for diffuse skin reflection. Regions below the threshold are marked as potential anomalies. The depth difference between adjacent pixels is then calculated; if it exceeds a preset depth threshold, such as a reasonable range for changes in facial curvature, it is identified as a geometrically discontinuous region. Next, pixels that simultaneously meet the criteria of "polarization degree below the threshold" and "depth difference exceeding the limit" are identified as conflict regions. For example, reflections from a metal door frame projected onto a portrait area cause abnormal polarization degree and abrupt depth changes in that region. Finally, neighborhood feature interpolation is performed for repair. A sample set is constructed using the depth values ​​of surrounding valid pixels centered on the conflict region, excluding the influence of other conflict regions. A bilinear interpolation algorithm is used to generate repaired values ​​based on the depth value distribution of the neighborhood samples. For example, for a conflict pixel, the weighted average depth of its four effective neighboring pixels (up, down, left, and right) is taken as the new depth value. Gaussian filtering is then used to regionally smooth the repaired depth map, eliminating local noise caused by interpolation and ensuring geometric continuity.

[0095] For example, in a one-way door scene at a laboratory entrance, the high reflectivity of the metal door frame causes anomalies in the polarization of the human face: due to the mirror reflection from the door frame, the polarization of a person's right cheek drops to 0.1 (below the typical skin threshold of 0.2), while the binocular depth point cloud shows a depth difference of 15mm between adjacent pixels in this area (exceeding the preset threshold of 10mm). The system marks this area as a conflict region through consistency verification and performs bilinear interpolation based on the depth values ​​(continuously distributed) of the surrounding normal cheek area. After repair, the depth difference is reduced to 3mm, consistent with the curvature of a real human face. The repaired spatial feature map eliminates reflection artifacts, ensuring that the subsequent multi-scale attention mechanism can accurately fuse facial geometry and texture features.

[0096] This embodiment effectively identifies and locates depth anomaly conflict regions in portrait areas by verifying the consistency of polarization and depth features, significantly improving the reliability of depth data in 3D face reconstruction under complex lighting and occlusion scenarios. By combining neighborhood feature interpolation algorithms for adaptive repair of conflict regions, it can restore true depth information while maintaining the continuity of facial topology, providing high-precision spatial geometric constraints for subsequent multimodal feature fusion, thereby enhancing the physical rationality and detail fidelity of material texture reconstruction.

[0097] Under complex lighting conditions, 3D face reconstruction faces a technical barrier due to the mismatch between the polarization mode and the cross-domain feature scale of RGB images. Traditional methods directly align polarization point cloud features with 2D images, but the high-frequency details (such as skin micro-textures) and low-frequency geometric structures (such as facial contours) are coupled, causing high-frequency information to be submerged by geometric features during cross-modal fusion. Simultaneously, low-frequency reflective areas (such as highlights on oily skin) experience decoupling between depth features and material reflection characteristics due to abrupt changes in polarization, resulting in both texture blurring and material distortion in the reconstructed model. This application addresses these technical problems through the following solution.

[0098] In one feasible implementation, in step S130: spatially aligning the point cloud polarization joint features with the face image to obtain an aligned spatial feature map, before generating a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image to reflective intensity, the method includes: performing multi-frequency domain decomposition on the polarized light image to obtain high-frequency detail components and low-frequency structural components, and fusing them with the geometric information in the point cloud polarization joint features across scales to generate point cloud polarization layered features.

[0099] Step S130: Spatially align the point cloud polarization joint features with the face image to obtain an aligned spatial feature map. Based on the polarization degree of the polarized light image, partition the spatial feature map by reflective intensity to generate a dynamic mask. This includes: spatially aligning the point cloud polarization layer features with the face image to obtain an aligned spatial feature map, and partitioning the spatial feature map by reflective intensity based on the polarization degree of the polarized light image to generate a dynamic mask.

[0100] Multi-frequency decomposition refers to the process of decomposing polarized light images into different frequency components using frequency domain analysis techniques, such as wavelet transform. High-frequency detail components correspond to rapidly changing features in the image, such as edges and textures (e.g., high-frequency noise in glass reflections), while low-frequency structural components correspond to the overall contour and gradually changing regions of the image (e.g., the approximate shape of a face). High-frequency detail components are high-frequency signals composed of local abrupt changes in the image, such as reflection artifacts and skin textures, reflecting changes in the microscopic reflective properties of the surface. Low-frequency structural components are low-frequency signals composed of smooth regions in the image (e.g., facial contours, background planes), reflecting macroscopic geometric structural features. Cross-scale fusion refers to associating high-frequency detail and low-frequency structural information with point cloud geometric features (e.g., curvature, normal vectors) to form a hierarchical feature representation. For example, high-frequency details are associated with the local curvature of the point cloud to capture reflection artifacts, while low-frequency structure is associated with the overall normal vector of the point cloud to describe facial pose. Point cloud polarization hierarchical features refer to a hierarchical dataset generated through cross-scale fusion, containing a high-frequency layer (detail texture and local geometry), a low-frequency layer (overall structure and material properties), and cross-layer association indexes.

[0101] First, the polarized light image is decomposed into three levels using Discrete Wavelet Transform (DWT) to obtain high-frequency components (LH, HL, HH) and low-frequency components (LL). The high-frequency components capture edge oscillations in reflective areas, such as stripes in glass reflections, while the low-frequency components preserve the macroscopic structure of the face and background. The balance between high-frequency noise suppression and low-frequency structure preservation is optimized by adjusting the filter parameters of the wavelet basis (e.g., Daubechies wavelet). Next, high-frequency details and geometry are correlated: the polarization angle gradient of the high-frequency components is extracted and matched with the local curvature in the joint features of the point cloud. For example, a high-frequency detail layer is generated by calculating the mapping relationship between high-frequency polarization angle changes and point cloud curvature using a Convolutional Neural Network (CNN), such as marking the abnormal curvature distribution of reflective glass areas. Then, low-frequency structure and geometry are correlated: the polarization degree distribution of the low-frequency components is spatially aligned with the point cloud normal vectors, and a weighted average method is used to fuse the angle between the low-frequency polarization degree and the point cloud normal vectors, generating a low-frequency structure layer, such as the consistency between the overall curvature of the face and the normal vector of the background plane. Then, a hierarchical index is constructed: a cross-layer association table is established between high-frequency and low-frequency layers to ensure that features at different scales can be synchronously invoked during subsequent dynamic mask generation, such as the material properties corresponding to high-frequency reflective areas in the low-frequency layer. Finally, the point cloud polarization layer features are projected onto the 2D face image plane, and coordinate registration is completed using the calibration parameters of the binocular camera to ensure that the high-frequency detail layer (such as nasal wing texture) and the low-frequency structure layer (such as facial contour) are aligned at the pixel level. The reflection intensity of each pixel is calculated based on the polarization degree channel, and adaptive threshold segmentation is used to divide the high reflective area (polarization degree ≥ threshold) and the normal area. Morphological closing operations are combined to eliminate segmentation noise and generate a binary dynamic mask, in which the weight of the high reflective area is 0 (suppressed) and the weight of the normal area is 1 (preserved).

[0102] For example, in a laboratory entrance scene, the glass area of ​​the door frame captured by the polarization camera exhibits high-frequency specular reflection (polarization degree 0.85), while the facial skin area shows low-frequency diffuse reflection (polarization degree 0.15). First, Haar wavelet decomposition is performed on the polarized image. The high-frequency components highlight the striped artifacts of the glass reflection, while the low-frequency components preserve the overall contour of the face. Next, for the high-frequency detail layer: the high-frequency polarization angle gradient of the glass reflection is associated with abnormally flat areas (curvature close to 0) in the point cloud, and these are marked as background interference. For the low-frequency structure layer: the low-frequency polarization degree of the face is fused with the high-curvature area of ​​the nose bridge in the point cloud (curvature > 0.2) to enhance realistic facial features. Finally, based on a polarization degree threshold of 0.6, the glass area is marked as a suppression region, and the face area is marked as a normal region, generating a dynamic mask. The mask edges are smoothed using Gaussian transitions to avoid fusion artifacts caused by hard segmentation.

[0103] Figure 2 A flowchart illustrating a method for generating multimodal feature vectors according to an embodiment of this application is shown. Figure 2As shown, the method includes steps S210 to S240.

[0104] In one feasible implementation, step S140: Utilizing a multi-scale attention mechanism, constrained by a dynamic mask, and based on a material geometry correlation model, feature fusion is performed on the spatial feature map and the face image to generate a multimodal feature vector including spatial structure, material properties, and texture information, including:

[0105] S210: Based on the predefined polarization angle-curvature correspondence in the material geometry correlation model, determine the target curvature range for each region in the spatial feature map. Using the polarization angle-curvature mapping table in the material geometry correlation model, convert the polarization angle of each region in the spatial feature map into a target curvature range. For example, if the polarization angle of a region is 45°, the corresponding curvature range is 0.05-0.1 (high curvature nose bridge) or 0.01-0.03 (low curvature cheek).

[0106] S220: Based on a dynamic mask, the spatial feature map is divided into high-reflectivity suppression regions and normal regions, and initial fusion weights are assigned to each region. The spatial feature map is divided into high-reflectivity suppression regions (mask value 0) and normal regions (mask value 1) using a dynamic mask. Initial fusion weights are assigned according to region type: the weight for suppression regions approaches 0, and the weight for normal regions is 1.

[0107] S230: For each level of feature in the spatial feature map, extract the polarization angle distribution and actual curvature value, calculate the difference between the curvature value and the target curvature range, and adjust the initial fusion weights of the corresponding region based on the difference. For each level of feature (e.g., original, downsampling scale), extract the actual curvature value and calculate the difference between it and the target curvature range. For example, if the actual curvature is 0.12 (exceeding the target range of 0.05-0.1), the difference is 0.12-0.1=0.02. The weights are dynamically adjusted based on the difference: the greater the difference, the higher the weight decay coefficient, suppressing the feature response of abnormal regions.

[0108] S240: The adjusted initial fusion weights are superimposed with the corresponding level of face image features to generate weighted fusion features layer by layer, and the merged features are output as a multimodal feature vector. The adjusted weights are multiplied element-wise with the corresponding level of face image features (such as SIFT texture descriptors) to generate a weighted feature map. For example, the texture features of the suppressed region are weakened, while the geometric-polarization features of the normal region are strengthened. Finally, the multi-scale feature maps are merged into a unified multimodal feature vector through channel concatenation and global pooling.

[0109] Figure 3 A flowchart illustrating a method for generating joint polarization features of point clouds according to an embodiment of this application is shown. Figure 3 As shown, the method includes steps S310 to S350.

[0110] In one feasible implementation, step S120: spatially aligning and feature stitching the polarized light image with the binocular depth point cloud data to generate joint polarization features of the point cloud, including:

[0111] S310: For each pixel in the polarized light image, determine the corresponding three-dimensional spatial point in the binocular depth point cloud through physical position mapping relationship, and establish an association index table between the pixel and the three-dimensional spatial point.

[0112] The physical location mapping relationship refers to the geometric correspondence between two-dimensional polarized light image pixels and three-dimensional stereo depth point cloud coordinates established through sensor calibration parameters (such as camera intrinsic and extrinsic parameters), ensuring that the same physical location can be accurately aligned in data from different sensors. The association index table is a lookup table that records the corresponding three-dimensional point cloud spatial coordinates for each polarized image pixel, used for data association.

[0113] Based on the calibration parameters (such as focal length and baseline distance) of the polarization camera and the binocular camera, the three-dimensional point cloud coordinates corresponding to each pixel in the polarized light image are calculated through projection transformation, the physical position mapping relationship between the pixel and the three-dimensional spatial point is established, and an associated index table is generated.

[0114] S320: Extract the polarization parameters of each pixel in the polarized light image. The polarization parameters include the direction of polarization intensity change and the difference in polarization intensity.

[0115] Polarization parameters include the difference in polarization intensity along the direction of polarization, which represents the polarization angle indicating the direction of light wave vibration, and the degree of polarization quantifying the proportion of polarized light. The polarization angle and degree of polarization parameters for each pixel are extracted from the polarized light image.

[0116] S330: Based on the association index table, add the polarization parameters to the three-dimensional spatial points of the corresponding association index to obtain joint data points including spatial coordinates, color information and polarization parameters.

[0117] Joint data points refer to multi-dimensional data units that integrate three-dimensional coordinates, RGB color, and polarization parameters, used to describe the geometric and optical properties of an object's surface. Polarization parameters are appended to the corresponding three-dimensional point cloud data according to an association index table, forming joint data points that contain spatial coordinates, color, and polarization parameters.

[0118] S340: Based on the geometric distribution density of the binocular depth point cloud, the polarization parameters of the joint data points are regionally smoothed to obtain smoothed joint data points.

[0119] Geometric distribution density refers to the number of data points per unit volume in a 3D point cloud, reflecting the spatial sampling density. Based on the geometric distribution density of the point cloud (e.g., high-density areas represent rich details, and low-density areas represent smooth surfaces), regional smoothing is performed on the polarization parameters of the joint data points: in high-density areas, a small neighborhood window is used for local averaging to preserve details; in low-density areas, the neighborhood window is expanded to enhance parameter continuity.

[0120] S350: Use the smoothed joint set of data points as the joint polarization feature of the point cloud.

[0121] Regional smoothing refers to performing a neighborhood-weighted average of polarization parameters based on point cloud density to suppress noise interference. The smoothed joint data point set is output as a joint polarization feature of the point cloud, which simultaneously includes the spatial structure, surface reflection properties, and color information of the target region.

[0122] For example, a polarization camera and a binocular camera simultaneously acquire polarized light images and depth point clouds of the door frame area (background) and the person's face (portrait). By calibrating parameters, the system maps the pixel at the tip of the nose in the polarization image to the corresponding 3D coordinates in the depth point cloud, establishing an association index table. For example, a pixel with coordinates (x=320, y=240) in the polarization image corresponds to a 3D point (X=1.2m, Y=0.5m, Z=2.0m) in the point cloud. The polarization angle of 45° and polarization degree of 0.15 of this pixel are extracted and appended to the 3D point data. For the glass area of ​​the door frame (high reflectivity, low point cloud density), the system expands the smoothing window and performs regional averaging of the polarization degree to eliminate parameter jumps caused by glass reflection; while the facial area (high point cloud density) uses a small window for smoothing to preserve skin texture details. The final generated point cloud polarization joint features contain the geometry of the face, skin reflection characteristics, and background material information, providing a data foundation for subsequent reflection suppression and feature fusion.

[0123] In one feasible implementation, step S150: Calculate the similarity between the multimodal feature vector and the identity features in a preset face identity database using a classification network to obtain target verification information, including:

[0124] The multimodal feature vector is divided into multiple feature blocks according to spatial structure, material properties, and texture information, and the feature expression value of each feature block is extracted. The feature template of the target identity is obtained from the face identity database. The feature template includes the historical mean and fluctuation range of each type of feature block. The feature difference ratio between the feature expression value of each feature block and the corresponding historical mean is calculated, and the weight coefficient of each feature block is determined according to the fluctuation range. The feature difference ratio is weighted and summed according to the weight coefficient of each feature block to obtain the target matching degree. The target verification information is obtained based on the comparison result between the target matching degree and the preset matching degree threshold.

[0125] Block features refer to dividing a multimodal feature vector into multiple sub-feature segments according to semantic dimensions, including spatial structure blocks (such as facial curvature distribution), material attribute blocks (such as skin polarization angle distribution), and texture information blocks (such as facial contour details). Feature representation values ​​are quantified values ​​of block features after normalization, used to describe the statistical characteristics of the block, such as the average curvature of the spatial structure block and the polarization angle variance of the material attribute block. Feature templates refer to feature reference data for legitimate identities pre-stored in the face identity database. Each template contains the mean and fluctuation range of multiple historically collected block features; for example, the mean curvature of a user's spatial structure block is 0.5, with a fluctuation range of ±0.1. Feature difference ratio refers to the degree of difference between the current block feature value and the template mean, quantified by a ratio. The weight coefficient is dynamically adjusted based on the fluctuation range of the block feature; the smaller the fluctuation range, the higher the contribution weight of the block to identity determination. Target matching degree is the weighted sum of the feature difference ratios of each block, used to comprehensively evaluate the similarity between the current feature and the template.

[0126] First, the multimodal feature vectors are divided into three categories of features based on a preset semantic dimension: spatial structure, material attributes, and texture information. For example, fixed-length segmentation or dynamic segmentation based on feature importance ranking can be used to ensure that each category contains the same type of semantic information. For each category, feature representation values ​​are extracted using statistical methods (such as mean, variance) or deep learning encoders (such as fully connected layers). For example, the material attribute category uses the peak value of the polarization angle distribution histogram as its representation value. Second, feature templates for the target identity are retrieved from the face identity database to obtain the historical mean and fluctuation range for each category. For example, for a user's texture information category, the mean vector and standard deviation range of texture features from 100 historically collected face images are loaded. Then, the difference ratio between the current category's feature representation value and the template mean is calculated. For example, cosine similarity or Euclidean distance is used to quantify the difference, and normalization is applied to convert the difference value into a proportional value. Based on the fluctuation range of the blocks in the template, weight coefficients are dynamically assigned: if a block has a small historical fluctuation range (e.g., the polarization angle of a material property block is stable), it is given a higher weight; conversely, if a block has a large fluctuation range (e.g., the texture block is greatly affected by lighting), its weight is reduced. Finally, the weighted sum of the difference ratios of all blocks is calculated to obtain the target matching degree. If the target matching degree exceeds a preset threshold (e.g., 0.9), the identity verification is considered successful; otherwise, it is marked as a failure.

[0127] For example, in a one-way access control scenario in a medical isolation area, a multimodal feature vector of a medical worker is collected and divided into three blocks: spatial structure, material properties, and texture information. The spatial structure block extracts the average curvature (0.48) of the facial depth point cloud; the material properties block calculates the skin polarization angle variance (0.12); and the texture block extracts the SIFT feature descriptor of the eye corner texture details. The feature template of this medical worker is retrieved from the database. Its historical spatial structure mean is 0.5 (fluctuation ±0.05), material property mean variance is 0.1 (fluctuation ±0.02), and texture block similarity threshold is 0.85. The current spatial structure difference ratio is calculated as (0.48-0.5) / 0.5 = 0.04, the material property difference ratio is (0.12-0.1) / 0.1 = 0.2, and the texture block cosine similarity is 0.88. Weights are assigned based on the fluctuation range: spatial structure weight 0.6 (small fluctuation), material attribute weight 0.3 (medium fluctuation), and texture weight 0.1 (large fluctuation). The weighted matching degree is 0.04×0.6+0.2×0.3+(1-0.88)×0.1=0.084, which is lower than the threshold of 0.1, so the verification is considered successful, and an opening command is generated.

[0128] Based on the same concept, this application provides a one-way door control system based on image recognition, which will be described below in conjunction with... Figure 4 The image recognition-based one-way door control system provided in the embodiments of this application will be described in detail.

[0129] Figure 4 This is a structural block diagram of a one-way door control system based on image recognition, as shown in an embodiment of this application.

[0130] like Figure 4 As shown, the image recognition-based one-way door control system may include:

[0131] The acquisition module 410 is used to acquire polarized light images of the target area under high reflectivity scenes, as well as face images and binocular depth point cloud data of the human image area in the target area. The target area includes the human image area and the background area.

[0132] The generation module 420 is used to spatially align and feature-stitch the polarized light image with the binocular depth point cloud data to generate point cloud polarization joint features.

[0133] The generation module 420 is also used to spatially align the point cloud polarization joint features with the face image to obtain the aligned spatial feature map, and generate a dynamic mask by partitioning the spatial feature map by reflective intensity based on the polarization degree of the polarized light image.

[0134] The generation module 420 is also used to utilize a multi-scale attention mechanism, with a dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image, and generate a multimodal feature vector including spatial structure, material properties, and texture information.

[0135] The verification module 430 is used to calculate the similarity between the multimodal feature vector and the identity features in the preset face identity database using a classification network, so as to obtain the target verification information.

[0136] The control module 440 is used to determine the target control information corresponding to the target verification information based on the preset correspondence between verification information and control information, and to complete the control of the one-way door through the target control information.

[0137] In one embodiment, the generation module 420 is specifically used to utilize a multi-scale attention mechanism, with a dynamic mask as a constraint, and to perform feature fusion on the spatial feature map and the face image based on the material geometry association model, to generate a multimodal feature vector including spatial structure, material properties, and texture information. The material geometry association model is trained based on the polarization angle of the historical polarized light image and the face geometric parameters of the historical spatial feature map through a predefined correspondence between skin polarization angle and curvature.

[0138] In one embodiment, before using a multi-scale attention mechanism with a dynamic mask as a constraint to perform feature fusion on the spatial feature map and the face image to generate a multimodal feature vector including spatial structure, material properties, and texture information, the generation module 420 is also used to perform consistency verification on the polarization features and depth features in the spatial feature map. If it is detected that there is a polarization degree lower than a preset polarization threshold and the depth value difference between adjacent regions exceeds a preset depth threshold in the portrait region, the conflict region in the spatial feature map is determined; the depth value of the conflict region in the spatial feature map is repaired by using neighborhood feature interpolation to obtain the repaired spatial feature map.

[0139] In one embodiment, before spatially aligning the point cloud polarization joint features with the face image to obtain an aligned spatial feature map, and before generating a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image to divide it into reflective intensity, the generation module 420 is further configured to perform multi-frequency domain decomposition on the polarized light image to obtain high-frequency detail components and low-frequency structural components, and then perform cross-scale fusion with the geometric information in the point cloud polarization joint features to generate point cloud polarization layered features; spatially aligning the point cloud polarization layered features with the face image to obtain an aligned spatial feature map, and then generating a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image to divide it into reflective intensity.

[0140] In one embodiment, the generation module 420 is specifically used to determine the target curvature range of each region in the spatial feature map according to the predefined polarization angle and curvature correspondence in the material geometry correlation model; divide the high reflectivity suppression region and normal region in the spatial feature map based on the dynamic mask, and assign initial fusion weights to each region; extract the polarization angle distribution and actual curvature value for each level of features in the spatial feature map, calculate the difference between the curvature value and the target curvature range, and adjust the initial fusion weights of the corresponding region according to the difference; superimpose the adjusted initial fusion weights with the face image features of the corresponding level, generate weighted fusion features layer by layer, and output a multimodal feature vector after merging.

[0141] In one embodiment, the generation module 420 is specifically used to determine the corresponding three-dimensional spatial point in the binocular depth point cloud for each pixel in the polarized light image through physical position mapping relationship, and establish an association index table between the pixel and the three-dimensional spatial point; extract the polarization parameters of each pixel in the polarized light image, the polarization parameters including the direction of polarization intensity change and the polarization intensity difference value; add the polarization parameters to the three-dimensional spatial point of the corresponding association index according to the association index table, to obtain joint data points including spatial coordinates, color information and polarization parameters; perform regional smoothing processing on the polarization parameters of the joint data points based on the geometric distribution density of the binocular depth point cloud, to obtain smoothed joint data points; and use the set of smoothed joint data points as the point cloud polarization joint feature.

[0142] In one embodiment, the verification module 430 is specifically used to divide the multimodal feature vector into multiple block features according to spatial structure, material properties, and texture information, and extract the feature expression value of each block feature; obtain the feature template of the target identity from the face identity database, the feature template including the historical mean and fluctuation range of each type of block feature; calculate the feature difference ratio between the feature expression value of each block feature and the corresponding historical mean, and determine the weight coefficient of each block feature according to the fluctuation range; perform a weighted summation of the feature difference ratio according to the weight coefficient of each block feature to obtain the target matching degree; and obtain the target verification information based on the comparison result between the target matching degree and the preset matching degree threshold.

[0143] Figure 4 Each module in the system shown has an implementation Figure 1 Of Figure 3 The functions of each step in the process and their corresponding technical effects are described in detail here for the sake of brevity.

[0144] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application is shown.

[0145] The electronic device may include a processor 510 and a memory 520 storing computer program instructions.

[0146] Specifically, the processor 510 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0147] Memory 520 may include mass storage for data or instructions. For example, and not limitingly, memory 520 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 520 may include removable or non-removable (or fixed) media. Where appropriate, memory 520 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 520 is non-volatile solid-state memory.

[0148] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.

[0149] The processor 510 reads and executes computer program instructions stored in the memory 520 to implement any of the image recognition-based one-way gate control methods in the above embodiments.

[0150] In one example, the electronic device may also include a communication interface 530 and a bus 540. Wherein, such as Figure 5 As shown, the processor 510, memory 520, and communication interface 530 are connected through bus 540 and complete communication with each other.

[0151] The communication interface 530 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0152] Bus 540 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 540 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0153] This electronic device can execute the image recognition-based one-way gate control method in the embodiments of this application, thereby achieving a combination of Figures 1 to 3 The image recognition-based one-way gate control method is described.

[0154] Furthermore, in conjunction with the image recognition-based one-way gate control method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the image recognition-based one-way gate control methods in the above embodiments.

[0155] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0156] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0157] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0158] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0159] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. An image recognition-based one-way door control method, characterized by, The method comprises the following steps: acquiring a polarized light image of a target region in a high-reflective scene, and a face image and a binocular depth point cloud data of a human image region in the target region, the target region comprising the human image region and a background region; spatially aligning and feature splicing the polarized light image and the binocular depth point cloud data to generate point cloud polarization joint features; spatially aligning the point cloud polarization joint features and the face image to obtain an aligned spatial feature map, and generating a dynamic mask based on a polarized degree of the polarized light image and a reflective light intensity partition of the spatial feature map; using a multi-scale attention mechanism, and performing feature fusion on the spatial feature map and the face image by taking the dynamic mask as a constraint to generate a multi-modal feature vector comprising spatial structure, material attribute and texture information; calculating a similarity between the multi-modal feature vector and an identity feature in a preset face identity database by using a classification network to obtain target verification information; determining target control information corresponding to the target verification information according to a preset corresponding relationship between verification information and control information, and completing control on the one-way door by using the target control information.

2. The method of claim 1, wherein, The method of using a multi-scale attention mechanism, taking the dynamic mask as a constraint, and performing feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector comprising spatial structure, material attribute and texture information comprises: using a multi-scale attention mechanism, taking the dynamic mask as a constraint, and performing feature fusion on the spatial feature map and the face image based on a material geometry correlation model to generate a multi-modal feature vector comprising spatial structure, material attribute and texture information, wherein the material geometry correlation model is trained based on a polarized angle of a historical polarized light image and a face geometry parameter of a historical spatial feature map by using a predefined corresponding relationship between a skin polarized angle and a curvature.

3. The method of claim 2, wherein, Before the method of using a multi-scale attention mechanism, taking the dynamic mask as a constraint, and performing feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector comprising spatial structure, material attribute and texture information, the method comprises: performing consistency verification on polarization features and depth features in the spatial feature map, and determining a conflict region in the spatial feature map when it is detected that there is a polarization degree lower than a preset polarization threshold and a depth value difference of adjacent regions exceeding a preset depth threshold in the human image region; using a neighborhood feature interpolation to repair the depth value of the conflict region in the spatial feature map to obtain a repaired spatial feature map.

4. The method of claim 1, wherein, Before the method of spatially aligning the point cloud polarization joint features and the face image to obtain an aligned spatial feature map, and generating a dynamic mask based on a polarized degree of the polarized light image and a reflective light intensity partition of the spatial feature map, the method comprises: performing multi-frequency domain decomposition on the polarized light image to obtain a high-frequency detail component and a low-frequency structure component, and performing cross-scale fusion on the geometric information in the point cloud polarization joint features, respectively, to generate point cloud polarization layered features; The point cloud polarization joint feature is spatially aligned with the face image to obtain an aligned spatial feature map, and a dynamic mask is generated by performing inverse light intensity zoning on the spatial feature map based on the polarization degree of the polarized light image. The point cloud polarization joint feature is spatially aligned with the face image to obtain an aligned spatial feature map, and a dynamic mask is generated by performing inverse light intensity zoning on the spatial feature map based on the polarization degree of the polarized light image.

5. The method of claim 2, wherein, The spatial feature map and the face image are fused based on a material-geometry correlation model, and a multi-modal feature vector including spatial structure, material attribute, and texture information is generated by using a multi-scale attention mechanism and taking the dynamic mask as a constraint. A target curvature range of each region in the spatial feature map is determined according to a predefined correspondence between a polarization angle and a curvature in the material-geometry correlation model. High-light-suppression regions and normal regions in the spatial feature map are divided based on the dynamic mask, and initial fusion weights are assigned to each region. For each level of feature of the spatial feature map, a polarization angle distribution and an actual curvature value are extracted, a difference degree of the curvature value and the target curvature range is calculated, and the initial fusion weight of the corresponding region is adjusted according to the difference degree. The adjusted initial fusion weight is superimposed on the face image feature of the corresponding level to generate a weighted fusion feature layer by layer, and the multi-modal feature vector is output after merging.

6. The method of claim 1, wherein, The polarized light image and the binocular depth point cloud data are spatially aligned and feature-spliced to generate a point cloud polarization joint feature, including: For each pixel point in the polarized light image, a corresponding three-dimensional space point in the binocular depth point cloud is determined through a physical position mapping relationship, and an association index table of pixel points and three-dimensional space points is established; Polarization parameters of each pixel point in the polarized light image are extracted, including a polarization light intensity change direction and a polarization light intensity difference value; According to the association index table, the polarization parameters are added to the three-dimensional space points corresponding to the association indexes to obtain joint data points including spatial coordinates, color information, and polarization parameters; Based on the geometric distribution density of the binocular depth point cloud, the polarization parameters of the joint data points are regionally smoothed to obtain smoothed joint data points; The set of smoothed joint data points is taken as the point cloud polarization joint feature.

7. The method of claim 1, wherein, The similarity between the multi-modal feature vector and the identity feature in a preset face identity database is calculated by using a classification network to obtain target verification information, including: The multi-modal feature vector is divided into multiple block features according to spatial structure, material attribute, and texture information, and a feature expression value of each block feature is extracted; A feature template of a target identity is obtained from the face identity database, and the feature template includes a historical mean value and a fluctuation range of each type of block feature; A feature difference proportion of the feature expression value of each block feature and the corresponding historical mean value is calculated, and a weight coefficient of each block feature is determined according to the fluctuation range. The feature difference proportions are weighted and summed according to the weight coefficients of each sub-block feature to obtain a target matching degree; According to a comparison result of the target matching degree and a preset matching degree threshold, the target verification information is obtained.

8. An image recognition based one-way door control system, characterized in that, The system comprises: An acquisition module is configured to acquire a polarized light image of a target region in a high-reflective scene, a face image of a portrait region in the target region, and binocular depth point cloud data of the target region, the target region comprising the portrait region and a background region; A generation module is configured to perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate point cloud polarization joint features; The generation module is further configured to perform spatial alignment on the point cloud polarization joint features and the face image to obtain an aligned spatial feature map, perform reflective intensity partitioning on the spatial feature map based on a polarization degree of the polarized light image to generate a dynamic mask, and perform feature fusion on the spatial feature map and the face image by using a multi-scale attention mechanism with the dynamic mask as a constraint to generate a multi-modal feature vector comprising spatial structure, material attribute, and texture information. A verification module is configured to calculate a similarity between the multi-modal feature vector and an identity feature in a preset face identity database by using a classification network to obtain target verification information. A control module is configured to determine target control information corresponding to the target verification information according to a preset corresponding relationship between verification information and control information, and complete control of the one-way door by using the target control information. The device comprises a processor and a memory storing computer program instructions; 9. An electronic device, comprising: The processor executes the computer program instructions to implement the one-way door control method based on image recognition according to any one of claims 1-7. The computer readable storage medium stores computer program instructions, and the computer program instructions are executed by the processor to implement the one-way door control method based on image recognition according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Image recognition method and device, computer equipment, readable storage medium and program product

    CN119314013A

  • Access control system based on face recognition and recognition method

    CN119672782A