One-way door control method, system and equipment based on image recognition

By combining the feature fusion technology of polarized light images and depth point cloud data in the one-way gate control system, the problem of material reflection interference in high-reflection scenarios is solved, and accurate identity verification and reliable access control are achieved in complex environments.

CN120375503AActive Publication Date: 2025-07-25TIANJIN XINLIJIA TECH CO LTD

Patent Information

Application Number
CN202510440313.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art one-way door control scheme cannot effectively eliminate material reflection characteristics interference in high-reflection scenarios, resulting in invalid identity verification, which can easily cause false rejection or misopening of access control, reducing the accuracy and effectiveness of one-way door control.

Method used

By obtaining polarized light images and binocular depth point cloud data in high-reflection scenes, spatial alignment and feature splicing are performed, point cloud polarization joint features are generated, feature fusion is used for multi-scale attention mechanism and dynamic mask, multi-modal feature vectors are generated, and classification network is used for identity verification, and the opening or closing of the one-way door is finally controlled.

Benefits of technology

In high reflective scenarios, the robustness of face verification is significantly improved, the false rejection rate and the risk of door opening are reduced, and the safety and reliability of the one-way access control system is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375503A_ABST
    Figure CN120375503A_ABST
Patent Text Reader

Abstract

The invention discloses a one-way door control method, system and equipment based on image recognition, and belongs to the field of images. The method comprises the following steps: synchronously obtaining a polarized light image, binocular depth point cloud data and a face image of a target area; the polarized light image and the depth point cloud three-dimensional feature are fused through spatial registration, and a point cloud polarization joint feature is constructed; performing pixel alignment on the joint feature and the face image, and generating a dynamic mask by using polarization degree data of the polarized light image; under the constraint of a dynamic mask, performing cross-modal fusion on the face image and the spatial feature map by adopting a multi-scale attention mechanism, and extracting a multi-modal feature vector containing a face spatial structure, a material attribute and a texture feature; performing similarity comparison on the vector and a preset database, and outputting verification information through a classification network; and triggering a self-adaptive control instruction of the one-way access control system according to the mapping relationship between the verification information and the control strategy. The one-way door can be accurately and effectively controlled in a complex reflective scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing, and particularly relates to a one-way door control method, system, and device based on image recognition. Background Art

[0002] In high-reflective scenarios such as laboratories, clean rooms, and medical isolation areas, reliable identity verification of one-way door access control systems faces severe challenges. In such scenarios, metal door frames, glass observation windows, and high-intensity ambient lighting cause strong specular reflections between the background area and the human portrait area, resulting in problems such as overexposure of face images, causing identity verification to fail, easily leading to false rejections or unauthorized door openings, and threatening regional security control. Therefore, there is an urgent need for a one-way door control method that can adapt to high-reflective interference to ensure the accuracy and effectiveness of one-way door control.

[0003] In the existing one-way door control solutions, generally, infrared supplementary lighting or polarization filters are used to reduce ambient light reflection interference. However, the filtering parameters are fixed and cannot dynamically adapt to complex reflective scenarios. The fixed filtering strategy of the existing one-way door control solutions cannot effectively eliminate the interference of the material reflection characteristics in the high-reflective area, and it is difficult to distinguish the reflective artifacts from the highlight area of the real face, resulting in identity verification failure and easily leading to false rejections or unauthorized door openings. Therefore, there are technical problems in the existing methods that the accuracy and effectiveness of one-way door control in complex reflective scenarios are reduced. Summary of the Invention

[0004] Embodiments of this application provide a one-way door control method, system, device, and computer-readable storage medium based on image recognition, which can accurately and effectively control one-way doors in complex reflective scenarios.

[0005] In a first aspect, embodiments of this application provide a one-way door control method based on image recognition, and the method includes:

[0006] Obtain the polarized light image of the target area, as well as the face image and binocular depth point cloud data of the human portrait area in the target area, where the target area includes the human portrait area and the background area;

[0007] Perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate a point cloud polarization joint feature;

[0008] Perform spatial alignment on the point cloud polarization joint feature and the face image to obtain an aligned spatial feature map, and generate a dynamic mask by dividing the spatial feature map according to the polarization degree of the polarized light image;

[0009] Use a multi-scale attention mechanism, with the dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information;

[0010] Calculate the similarity between the multi-modal feature vector and the identity features in the preset face identity database by using a classification network to obtain target verification information;

[0011] Determine the target control information corresponding to the target verification information according to the corresponding relationship between the preset verification information and the control information, and complete the control of the one-way door through the target control information.

[0012] In an implementable embodiment, use a multi-scale attention mechanism, with a dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, including:

[0013] Use a multi-scale attention mechanism, with a dynamic mask as a constraint, and perform feature fusion on the spatial feature map and the face image based on a material geometry correlation model to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, where the material geometry correlation model is trained based on the polarization angle of the historical polarized light image for the face geometry parameters of the historical spatial feature map through a predefined correspondence between the skin polarization angle and the curvature.

[0014] In an implementable embodiment, before using a multi-scale attention mechanism, with a dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, the method includes:

[0015] Perform consistency verification on the polarization feature and the depth feature in the spatial feature map, and determine the conflict area in the spatial feature map when it is detected that there is a polarization degree lower than the preset polarization threshold and the depth value difference in the adjacent area exceeds the preset depth threshold in the portrait area;

[0016] Use neighborhood feature interpolation to repair the depth value of the conflict area in the spatial feature map to obtain a repaired spatial feature map.

[0017] In an implementable embodiment, before spatially aligning the point cloud polarization joint feature with the face image to obtain an aligned spatial feature map and generating a dynamic mask by partitioning the reflective intensity of the spatial feature map based on the polarization degree of the polarized light image, the method includes:

[0018] Perform multi-frequency domain decomposition on the polarized light image to obtain a high-frequency detail component and a low-frequency structure component, and perform cross-scale fusion with the geometric information in the point cloud polarization joint feature respectively to generate a point cloud polarization hierarchical feature;

[0019] Spatially align the point cloud polarization joint feature with the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the reflective intensity of the spatial feature map based on the polarization degree of the polarized light image, including:

[0020] Spatially align the polarization layer features of the point cloud with the face image to obtain the spatially aligned feature map, and generate a dynamic mask by partitioning the reflective intensity of the spatially aligned feature map based on the degree of polarization of the polarized light image.

[0021] In an implementable embodiment, using a multi-scale attention mechanism, with the dynamic mask as a constraint, and based on the material geometry correlation model, perform feature fusion on the spatially aligned feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, including:

[0022] According to the predefined correspondence between the polarization angle and curvature in the material geometry correlation model, determine the target curvature range of each region in the spatially aligned feature map;

[0023] Based on the dynamic mask, divide the highly reflective suppression region and the normal region in the spatially aligned feature map, and assign initial fusion weights to each region;

[0024] For each level of features in the spatially aligned feature map, extract the polarization angle distribution and the actual curvature value, calculate the difference degree between the curvature value and the target curvature range, and adjust the initial fusion weight of the corresponding region according to the difference degree;

[0025] Overlay the adjusted initial fusion weight with the face image features of the corresponding level, generate weighted fusion features layer by layer, and output the multi-modal feature vector after merging.

[0026] In an implementable embodiment, perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate point cloud polarization joint features, including:

[0027] For each pixel point in the polarized light image, determine the corresponding three-dimensional spatial point in the binocular depth point cloud through the physical position mapping relationship, and establish an association index table between the pixel point and the three-dimensional spatial point;

[0028] Extract the polarization parameters of each pixel point in the polarized light image, where the polarization parameters include the change direction of the polarized light intensity and the polarized light intensity difference value;

[0029] According to the association index table, add the polarization parameters to the three-dimensional spatial points corresponding to the associated index to obtain joint data points including spatial coordinates, color information, and polarization parameters;

[0030] Based on the geometric distribution density of the binocular depth point cloud, perform regional smoothing processing on the polarization parameters of the joint data points to obtain smoothed joint data points;

[0031] Use the set of smoothed joint data points as the point cloud polarization joint features.

[0032] In an implementable embodiment, a classification network is used to calculate the similarity between the multi-modal feature vector and the identity features in a preset face identity database, and target verification information is obtained, including:

[0033] The multi-modal feature vector is divided into multiple block features according to the spatial structure, material attributes, and texture information, and the feature expression value of each block feature is extracted;

[0034] The feature template of the target identity is obtained from the face identity database, and the feature template includes the historical mean and fluctuation range of each type of block feature;

[0035] The feature difference ratio between the feature expression value of each block feature and the corresponding historical mean is calculated, and the weight coefficient of each block feature is determined according to the fluctuation range;

[0036] The weighted sum of the feature difference ratios is calculated according to the weight coefficient of each block feature to obtain the target matching degree;

[0037] The target verification information is obtained according to the comparison result between the target matching degree and the preset matching degree threshold.

[0038] In a second aspect, an embodiment of the present application provides a one-way door control system based on image recognition. The system includes:

[0039] An acquisition module, configured to acquire a polarized light image of a target area, a face image of a human portrait area in the target area, and binocular depth point cloud data in a high-reflection scene. The target area includes a human portrait area and a background area;

[0040] A generation module, configured to perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate a point cloud polarization joint feature;

[0041] The generation module is further configured to perform spatial alignment on the point cloud polarization joint feature and the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the reflection intensity of the spatial feature map based on the polarization degree of the polarized light image;

[0042] The generation module is further configured to use a multi-scale attention mechanism to perform feature fusion on the spatial feature map and the face image with the dynamic mask as a constraint, and generate a multi-modal feature vector including spatial structure, material attributes, and texture information;

[0043] A verification module, configured to use a classification network to calculate the similarity between the multi-modal feature vector and the identity features in a preset face identity database, and obtain target verification information;

[0044] A control module, configured to determine the target control information corresponding to the target verification information according to the corresponding relationship between the preset verification information and the control information, and complete the control of the one-way door through the target control information.

[0045] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the image recognition-based one-way door control method in any one of the implementation manners of the first aspect.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the image recognition-based one-way door control method in any one of the implementation manners of the first aspect is implemented.

[0047] For the image recognition-based one-way door control method, system, device, and computer-readable storage medium according to the embodiments of the present application, in this embodiment, first, the polarized light image is used to quantify the specular reflection intensity through the degree of polarization, and combined with the geometric information of the depth point cloud, the specular reflection suppression area is dynamically divided to eliminate the specular reflection artifacts on the metal and glass surfaces; secondly, the multi-modal feature fusion uses the polarization characteristics to distinguish real human faces, such as low-polarization skin and high-specular reflection materials, such as high-polarization metal or glass, and strengthens the face texture and structural features through a multi-scale attention mechanism, so as to accurately extract identity discrimination information even under strong specular reflection interference; finally, based on the feature weighting strategy constrained by the dynamic mask, the invalid noise in the specular reflection area is suppressed, and the robustness of face verification is improved, so as to realize reliable one-way door access control in high-security scenarios such as laboratories and medical isolation areas, and significantly reduce the false rejection rate and the risk of unauthorized door opening.

[0048] Furthermore, by introducing a material geometry correlation model, the interference problem of the material reflection characteristics on the normal estimation in high-specular reflection scenarios is effectively solved. Specifically, based on the corresponding relationship between the polarization angle and curvature of historical polarized light images, the model dynamically correlates the polarization parameters with the three-dimensional geometric features by using the predefined polarization characteristics of human skin, such as the physical correlation between the polarization angle and curvature. It can distinguish the specular reflection artifacts of metal or glass from the diffuse reflection characteristics of real human faces. For example, the specular reflection distortion in the depth point cloud is corrected through the curvature difference, so as to improve the discrimination accuracy of material properties. At the same time, by combining the multi-scale attention mechanism to perform hierarchical fusion on the spatial feature map and the face image, the fusion weight can be adaptively adjusted to suppress the noise interference in the high-specular reflection area and enhance the complementarity between texture details and three-dimensional structures. Finally, in high-specular reflection scenarios such as laboratories and medical isolation areas, this solution significantly improves the spatial consistency of face features, reduces the false recognition rate caused by material reflection, and ensures the security and reliability of the one-way access control system. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 is a schematic flowchart of a one-way door control method based on image recognition provided by an embodiment of the present application;

[0051] Figure 2 is a schematic flowchart of a method for generating multi-modal feature vectors provided by an embodiment of the present application;

[0052] Figure 3 is a schematic flowchart of a method for generating point cloud polarization joint features provided by an embodiment of the present application;

[0053] Figure 4 is a schematic structural diagram of a one-way door control system based on image recognition provided by an embodiment of the present application;

[0054] Figure 5 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0055] The following will describe in detail the features and exemplary embodiments of various aspects of the present application. To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.

[0056] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article, or device including the said elements.

[0057] In high-reflective scenarios such as laboratories, clean rooms, and medical isolation areas, reliable authentication of one-way door access control systems faces severe challenges. In such scenarios, metal door frames, glass observation windows, and high-intensity ambient lighting cause strong specular reflections between the background area and the portrait area, resulting in problems such as overexposure of face images, leading to invalid identity verification, easy occurrence of false rejections or unauthorized door openings, and threatening the security control of the area. Therefore, there is an urgent need for a one-way door control method that can adapt to high-reflective interference to ensure the accuracy and effectiveness of one-way door control.

[0058] In the existing one-way door control solutions, generally, infrared supplementary lighting or polarization filters are used to reduce ambient reflective interference. However, the filter parameters are fixed and cannot dynamically adapt to complex reflective scenarios. The fixed filter strategy of the existing one-way door control solutions cannot effectively eliminate the interference of the material reflection characteristics in the high-reflective area, making it difficult to distinguish the reflective artifacts from the highlight areas of real faces, resulting in invalid identity verification, easy occurrence of false rejections or unauthorized door openings. Therefore, there are technical problems in the existing methods, such as the reduction of the accuracy and effectiveness of one-way door control in complex reflective scenarios.

[0059] To solve the problems of the existing technology, the embodiments of the present application provide a one-way door control method, system, device, and computer storage medium based on image recognition. First, the one-way door control method based on image recognition provided by the embodiments of the present application will be introduced below.

[0060] Figure 1 The flowchart of the one-way door control method based on image recognition provided by an embodiment of the present application is shown. As Figure 1 shown, the method includes steps S110 to S160.

[0061] S110: Obtain a polarized light image of a target area in a high-reflective scenario, as well as a face image and binocular depth point cloud data of the portrait area in the target area. The target area includes a portrait area and a background area.

[0062] The target area refers to the physical space that needs to be recognized within the monitoring range of the one-way access control system, including the portrait area, that is, the area where the face to be recognized is located, and the background area, such as areas of objects that may cause reflections, such as metal door frames and glass windows. The polarized light image is a two-dimensional image collected by a polarization sensor, and each pixel point contains optical parameters such as polarization angle and polarization degree, which are used to quantify the reflection characteristics of the object surface, such as specular reflection or diffuse reflection. The face image is a two-dimensional color image captured by an RGB camera, which contains face texture and color information and is used to extract facial features, such as facial contours and skin texture. The binocular depth point cloud data is three-dimensional point cloud data generated by a binocular stereo vision system, and each point contains three-dimensional spatial coordinates (X / Y / Z) and color information, which are used to describe the geometric structure of the portrait area, such as face curvature and depth difference.

[0063] Multi-sensor synchronous triggering. A polarization camera, an RGB camera, and a binocular stereo vision device are deployed in the target area. Through a hardware synchronization signal, the timestamps of the polarized light image, the face image, and the depth point cloud data are ensured to be consistent. The polarization camera uses multi-angle polarizers, such as 0°, 45°, 90°, and 135°, to capture images in a time-sharing manner. The polarization angle and degree of polarization of each pixel are calculated through the Stokes vector to generate a polarized light image. The binocular camera, based on the parallax principle, uses feature matching algorithms, such as SIFT or ORB, to calculate the parallax of corresponding points in the left and right images, and converts them into three-dimensional point cloud data through camera calibration parameters to generate a binocular depth point cloud.

[0064] Exemplarily, a polarization camera and a binocular camera are installed on both sides of the door frame of the one-way door at the laboratory entrance. When a person approaches, the system synchronously triggers the sensors to collect data. The polarization camera captures the polarization parameters of the reflective area of the door frame glass (background) and the person's face (portrait), and the binocular camera generates a three-dimensional point cloud of the face, such as the height of the nose bridge and the depth of the eye socket depression.

[0065] S120: Align the polarized light image and the binocular depth point cloud data spatially and splice the features to generate a point cloud polarization joint feature.

[0066] Spatial alignment means establishing a mapping relationship between the pixel points of the two-dimensional polarized image and the corresponding coordinate points in the three-dimensional point cloud to ensure that the data is in the same coordinate system. Feature splicing means fusing the polarization parameters with the geometric and color information of the point cloud to form a joint data containing multi-dimensional attributes. Point cloud polarization joint feature: A multi-dimensional data set that fuses the spatial coordinates, color information of the binocular depth point cloud, and the polarization parameters (polarization angle, degree of polarization) of the polarized light image, characterizing the three-dimensional geometry and material reflection characteristics of the target area.

[0067] First, using the calibration parameters of the polarization camera and the binocular camera, establish a physical position mapping relationship between the pixel points and the three-dimensional points. Determine the corresponding three-dimensional coordinates of each pixel in the polarized image in the point cloud through projection calculation to generate an association index table. Secondly, extract the polarization direction and intensity difference value of each pixel from the polarized light image. According to the index table, attach these parameters to the corresponding three-dimensional point cloud data so that each three-dimensional point contains coordinate, color, and polarization information. Finally, integrate all the processed three-dimensional point data to form a point cloud polarization joint feature, which contains the spatial structure, surface reflection characteristics, and color information of the target area at the same time.

[0068] S130: Align the point cloud polarization joint feature with the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the reflective intensity of the spatial feature map based on the degree of polarization of the polarized light image.

[0069] Spatial alignment refers to registering the combined point cloud polarization features (3D spatial data) with the face image (2D image) in a unified coordinate system to ensure the position consistency of the two in the physical space. The spatial feature map refers to the multi-dimensional data matrix after alignment, which contains the geometric structure, polarization characteristics, and RGB texture information of the portrait area. The degree of polarization refers to the quantization value of the polarization light intensity difference of each pixel in the polarization light image, which is used to distinguish the specular reflection (high degree of polarization) and diffuse reflection (low degree of polarization) areas. The dynamic mask refers to a binary mask that divides the spatial feature map based on the degree of polarization and is used to mark the highly reflective area and the normal area.

[0070] First, using multi-sensor calibration parameters, such as camera internal parameters and external parameters, establish the mapping relationship between the 3D coordinates of the combined point cloud polarization features and the pixel coordinates of the face image. For example, project the 3D points onto the 2D image plane through the projection matrix of the binocular camera to generate a one-to-one corresponding coordinate index table. Then, for the sparse area of the point cloud, use bilinear interpolation or nearest neighbor interpolation method to complement the missing polarization parameters and face texture information to ensure that each pixel has complete geometric and polarization attributes in the spatial feature map. Subsequently, extract the polarization degree channel of the polarization light image, and through an adaptive threshold algorithm, such as the Otsu method, divide the highly reflective area, that is, the area where the polarization degree is higher than the threshold, and the normal area. Combine morphological operations, such as dilation and erosion, to eliminate noise interference, mark the boundary of the continuous highly reflective area, and generate a binary mask. Finally, adjust the transition weight of the mask edge according to the polarization degree gradient to avoid artifacts caused by hard segmentation and ensure that the light interference in the highly reflective area is effectively suppressed during subsequent feature fusion.

[0071] Exemplarily, obtain the depth point cloud of the face, such as the height of the nose bridge and the depth of the eye sockets, through a binocular camera, and at the same time, the polarization camera captures the polarization parameters of the face and the background (such as the polarization angle of the specular reflection of the glass door frame). Map the 3D point cloud to the 2D face image through the calibration parameters to ensure that the tip of the nose is aligned with the center of the image. The polarization degree image shows that the polarization degree of the glass area of the door frame is as high as 0.8 (strong specular reflection), while the polarization degree of the human face skin area is lower than 0.2 (diffuse reflection). The system uses adaptive threshold segmentation to mark the glass area as a highly reflective suppression area and generate a mask; the human face area is reserved as the normal area for subsequent identity feature extraction. When the ambient light changes cause the reflection intensity of the door frame to fluctuate, the mask is dynamically updated according to the real-time polarization degree.

[0072] S140: Use a multi-scale attention mechanism, with the dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information.

[0073] The multi-scale attention mechanism refers to a feature fusion technique that extracts features at different resolutions or receptive fields in parallel and combines attention weights to dynamically allocate the importance of features at each scale, aiming to enhance the joint modeling ability of local details and global structures. This mechanism can capture both the micro-texture of a face (such as pores) and the macro-geometric structure (such as facial contours) simultaneously. The dynamic mask constraint means that during the feature fusion process, a dynamic mask is used to impose spatial restrictions on the attention weights, forcing the model to reduce the feature response in the high-specular reflection suppression area and retain or enhance the feature expression in the normal area. This constraint ensures that the fusion process is not interfered by specular artifacts through weight adjustment or regional feature suppression. The multi-modal feature vector refers to the high-dimensional vector obtained by fusing the spatial feature map (including polarization parameters and geometric structures) and the face image (including RGB textures), which contains three types of key identity features: spatial structure, material attributes (such as skin reflection characteristics), and texture information (such as facial feature details).

[0074] First, perform pyramid downsampling on the spatial feature map and the face image respectively to generate multi-scale feature maps (such as the original resolution, 1 / 2, and 1 / 4 scales). For example, use dilated convolution or pooling operations to extract local features (such as eye contours) and global features (such as head poses) at different receptive fields. At each scale level, calculate the geometric-polarization features of the spatial feature map and the RGB texture features of the face image to form hierarchical feature pairs. Then, use the dynamic mask as a spatial attention template to impose suppression weights (such as weights approaching 0) on the feature channels in the high-specular reflection area, and retain the original weights in the normal area. For example, eliminate the interference of doorframe specular reflection on the face area by element-wise multiplication of the mask matrix and the feature map. Introduce a channel attention mechanism (such as the SENet module) to adaptively adjust the channel weights according to the response intensity of the feature channels and strengthen the expression of key material attributes (such as skin polarization angle). Finally, adopt a cross-attention mechanism to interact the geometric-polarization features of the spatial feature map and the texture features of the face image. For example, calculate the similarity matrix between different modal features through the Query-Key-Value structure to generate weighted fusion features. Upsample the fusion results at the multi-scale levels and splice them gradually, and retain the original information through residual connections, and finally output a multi-modal feature vector containing multi-dimensional identity features.

[0075] S150: Use a classification network to calculate the similarity between the multi-modal feature vector and the identity features in the preset face identity database to obtain the target verification information.

[0076] A classification network refers to a discriminative model based on deep learning that calculates the similarity between input features and a preset database and outputs an authentication result. Cosine similarity or cross-entropy loss functions are usually adopted and combined with a multi-layer perceptron (MLP) or a support vector machine (SVM) to achieve this. Similarity refers to the degree of matching between multi-modal feature vectors and feature templates in the database, which is obtained by quantifying the difference ratio and weighted summation and is used to determine whether it is a legitimate identity. The target verification information refers to the verification result output by the classification network, including success and failure indicators and confidence levels, which serves as the basis for controlling the opening or closing of a one-way door.

[0077] First, align the multi-modal feature vectors with the feature templates in the preset face identity database to ensure consistent feature dimensions and semantics. For example, eliminate the dimensional differences of data from different sensors through normalization. Then, use a classification network, such as a neural network or a support vector machine, to calculate the similarity between the multi-modal feature vectors and each feature template in the database. By comparing the differences in key features, such as geometric structures, material properties, and texture information, a matching score is output. Finally, based on the comparison result between the similarity score and a preset threshold, determine whether it is a legitimate identity. If the matching degree exceeds the threshold, generate a "verification passed" message; otherwise, mark it as "verification failed".

[0078] S160: Determine the target control information corresponding to the target verification information according to the corresponding relationship between the preset verification information and the control information, and complete the control of the one-way door through the target control information.

[0079] The corresponding relationship between the preset verification information and the control information refers to a predefined rule library that clarifies the mapping relationship between different verification results (such as passed / failed / abnormal) and the access control actions of the one-way door (such as opening the door / alarming / retrying). The target control information refers to the command signal generated according to the verification result, including door lock drive instructions, alarm signals, or log recording commands, which directly controls the one-way door actuator. The control instruction distribution module refers to the hardware interface or communication protocol that converts the target control information into physical signals (such as relay on / off, motor drive pulses) to drive the access control actuator to act.

[0080] First, match the target verification information output by the classification network, such as "Verification passed, confidence level 0.92", with the conditions in the preset rule library. The rule library is defined based on business requirements. For example: If the verification status is "passed" and the confidence level is higher than the threshold, trigger the door opening instruction; if the status is "failed", initiate an alarm and record the log. Then, generate the target control information according to the matching result. For example, when "verification passed", generate a control instruction containing the door lock opening duration and motor speed parameters; when "verification failed", generate an audible and visual alarm signal and an identity anomaly log. The instruction format needs to be compatible with the access control hardware interface protocol, such as Modbus or CAN bus. Finally, send the target control information to the access control actuator, such as an electromagnetic lock or a motor drive board, through the control instruction distribution module. The system monitors the execution status in real time, such as whether the door lock is successfully unlocked. If no feedback is received after a timeout, trigger the exception retry mechanism to ensure control reliability.

[0081] Exemplarily, when the target verification information is "Verification passed, confidence level 0.95", the system matches the "High confidence level passed" condition in the rule library and generates an instruction of "Open the door for 5 seconds". The control instruction distribution module sends the instruction to the motor controller through the RS485 bus to drive the door lock to open. The door status sensor gives a real-time feedback of the "Opened" signal, and the system records the access log. If the verification result is "failed", the system triggers the buzzer alarm and pushes the abnormal information to the management terminal. When the instruction distribution times out due to a network failure, the system enables the local cache instruction retransmission mechanism to ensure the controllability of the one-way door access status.

[0082] In this embodiment, first, the polarized light image is used to quantify the specular reflection intensity through the degree of polarization, and combined with the geometric information of the depth point cloud, the specular reflection suppression area is dynamically divided to eliminate the specular reflection artifacts on the metal and glass surfaces. Second, the multi-modal feature fusion uses the polarization characteristics to distinguish real human faces, such as low polarization skin and high specular reflection materials, such as high polarization metal or glass, and strengthens the face texture and structure features through the multi-scale attention mechanism, so that the identity discrimination information can be accurately extracted even under strong specular reflection interference. Finally, based on the feature weighting strategy constrained by the dynamic mask, the invalid noise in the specular reflection area is suppressed, and the robustness of face verification is improved, so as to achieve reliable one-way door access control in high-security scenarios such as laboratories and medical isolation areas, and significantly reduce the false rejection rate and the risk of unauthorized door opening.

[0083] In a high-reflectivity scenario, strong specular reflection not only interferes with the background area but also forms dynamic specular artifacts on the human face surface, such as glare on glasses, which highly confuses with the real human face texture features. Moreover, there is a cross-modal correlation between the material properties of polarized light images, such as the polarization angle, and the spatial geometric features of human face images. However, in dynamic high-reflectivity scenarios, such as specular reflection on glass curtain walls and metal surfaces, strong background reflection will distort the physical properties of polarized light, resulting in the destruction of the inherent mapping relationship between material properties and geometric structures. For example, the pseudo-polarization angle generated by glass reflection may be wrongly mapped to the low-curvature cheek area, while the high-curvature nasal bridge area of real skin may lose polarization features due to glare interference. This cross-modal feature misalignment will trigger semantic conflicts in multi-modal feature fusion, such as wrongly associating the material properties of specular artifacts with the geometric structure of the real human face. The present application solves the above technical problems through the following technical solutions.

[0084] In an implementable embodiment, step S140: Using a multi-scale attention mechanism and constrained by a dynamic mask, perform feature fusion on the spatial feature map and the human face image to generate a multi-modal feature vector including spatial structure, material properties, and texture information, including:

[0085] Using a multi-scale attention mechanism, constrained by a dynamic mask, and based on a material-geometry correlation model, perform feature fusion on the spatial feature map and the human face image to generate a multi-modal feature vector including spatial structure, material properties, and texture information, where the material-geometry correlation model is trained based on the polarization angle of historical polarized light images for the human face geometric parameters of the historical spatial feature map through a predefined correspondence between skin polarization angle and curvature.

[0086] The material-geometry correlation model refers to a physical constraint model that establishes a mapping between material properties (polarization angle) and geometric structures (curvature) through a predefined correspondence between skin polarization angle and curvature (for example, a specific polarization angle distribution corresponding to the high curvature of the nose tip), so as to solve the problem of cross-modal feature misalignment.

[0087] First, perform pyramid downsampling on the spatial feature map (including polarization parameters and geometric structures) and the face image respectively to generate multi-scale feature maps (such as original, 1 / 2, 1 / 4 resolution). In each level, a spatial attention module, such as SENet, is used to calculate the feature channel weights, and a dynamic mask is introduced to apply inhibitory weights to the feature channels in the highly reflective area. For example, the mask matrix is multiplied element-wise with the feature map to make the weights in the reflective area approach 0. Then, based on the material geometry correlation model, learn the non-linear relationship between the polarization angle and curvature of the skin area from historical data. For example, the larger the curvature, the more significant the change in the polarization angle, and real-time verify the consistency of the polarization angle and curvature in each area of the current spatial feature map. For inconsistent areas, such as abnormal polarization angles but curvatures that conform to facial features, perform feature correction. Through the cross-attention mechanism, interact the geometry-polarization features with the facial texture features to generate weighted fusion features. Finally, upsample and splice the fusion features at multiple scales, retain the original information through residual connections, and finally output a multi-modal feature vector containing spatial structures (such as the height of the nose bridge), material attributes (such as skin reflection characteristics), and texture information (such as pore details).

[0088] Exemplarily, in a highly reflective scene, the strong specular reflection of the metal door frame and the glass observation window causes overexposure of the face image. Feature fusion is achieved through the following process. First, perform dynamic mask generation: Polarization analysis shows that the polarization degree in the door frame area is as high as 0.8 (specular reflection), while the polarization degree in the human face skin area is lower than 0.2 (diffuse reflection). The system generates a mask to suppress the reflective area of the door frame. Second, perform multi-scale feature fusion: At the 1 / 4 resolution level, the spatial feature map extracts the geometric curvature of the human face. For example, the curvature of the tip of the nose is 0.5, and at the same time, the polarization angle distribution shows that the polarization angle in the tip of the nose area is 45°, which conforms to the preset relationship in the material geometry correlation model. For example, a curvature of 0.5 corresponds to a polarization angle of 45° ± 5°, and a high weight is assigned to this area. In the reflective area of the door frame, the mask suppresses the features in this area to avoid the polarization angle of the glass reflection, such as a polarization angle of 80°, from interfering with the facial feature fusion. Finally, output multi-modal features: The final feature vector integrates the three-dimensional structure of the nose bridge (depth point cloud), the polarization characteristics of the skin material (low polarization degree), and the facial texture (such as crow's feet at the corners of the eyes), providing a robust input for subsequent identity verification.

[0089] In this embodiment, by introducing a material geometry correlation model, the problem of interference of material reflection characteristics on normal estimation in high-reflection scenes is effectively solved. Specifically, based on the correspondence between the polarization angle and curvature of historical polarized light images, the model dynamically correlates polarization parameters with three-dimensional geometric features by using predefined polarization characteristics of human face skin, such as the physical correlation between the polarization angle and curvature. It can distinguish the specular reflection artifacts of metal or glass from the diffuse reflection characteristics of real human faces. For example, it corrects the specular reflection distortion in the depth point cloud through curvature differences, thereby improving the discrimination accuracy of material properties. At the same time, by combining the multi-scale attention mechanism to perform hierarchical fusion on the spatial feature map and the face image, the fusion weights can be adaptively adjusted to suppress the noise interference in the high-reflection area and enhance the complementarity of texture details and three-dimensional structures. Finally, in high-reflection scenes such as laboratories and medical isolation areas, this solution significantly improves the spatial consistency of face features, reduces the misrecognition rate caused by material reflection, and ensures the security and reliability of the one-way access control system.

[0090] Due to the inconsistency of the physical properties of polarization features and depth features in the multi-modal fusion of polarized light images and face images, feature conflicts occur in areas with low polarization and depth mutations, such as reflective materials on the human face surface and edge contours. Specifically, when the polarization degree of the portrait area is lower than the preset threshold due to specular reflection or environmental interference, traditional depth perception algorithms will have abnormal depth jumps due to distorted polarization information. For example, the depth difference between adjacent pixels exceeds the threshold. Such contradictory features directly lead to structural misalignment and material misjudgment in the fused multi-modal feature vector. Polarization imaging is sensitive to surface materials while depth perception depends on geometric structures. Feature conflicts will destroy the complementarity of multi-modal information. Especially in scenarios such as financial-grade face recognition and live detection, subtle material reflections, such as glasses and oily skin, and depth jumps in the contours of real human faces may be misjudged as fraud attacks, threatening the security of the system. This application solves the above technical problems through the following technical solutions.

[0091] In an implementable embodiment, before step S140: using the multi-scale attention mechanism, with a dynamic mask as a constraint, performing feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, the method includes:

[0092] Performing consistency verification on the polarization features and depth features in the spatial feature map. When it is detected that there is a conflict area in the portrait area where the polarization degree is lower than the preset polarization threshold and the depth value difference between adjacent areas exceeds the preset depth threshold, determining the conflict area in the spatial feature map; using neighborhood feature interpolation to repair the depth value of the conflict area in the spatial feature map to obtain a repaired spatial feature map.

[0093] The polarization feature refers to the degree of linear polarization (DoLP) and the angle of linear polarization (AoLP) of each pixel in the polarized light image, which are used to quantify the reflection characteristics of the object surface. For example, a high degree of polarization corresponds to specular reflection such as a metal door frame, and a low degree of polarization corresponds to diffuse reflection such as human face skin. The depth feature refers to the three-dimensional coordinate data derived from the binocular depth point cloud, which describes the geometric structure of the target area, such as the curvature of the human face and depth differences. For example, the depth value at the bridge of the nose is higher than that of the cheeks, and the depth value is concave at the eye sockets. The consistency check refers to verifying the matching degree between the polarization feature and the depth feature through physical relevance. For example, the skin area should have a low degree of polarization and a continuous change in curvature. If a certain area has an abnormally low degree of polarization but a sudden change in adjacent depth, it indicates a feature contradiction caused by specular reflection interference. The conflict area refers to the abnormal area where there are significant contradictions between the polarization feature and the depth feature within the portrait area. For example, due to glass reflection, the degree of polarization of a certain skin area is lower than the threshold, which does not conform to the skin characteristics, and at the same time, the difference in adjacent depth values is too large, which does not conform to the smooth human face geometry. The neighborhood feature interpolation refers to repairing abnormal depth values based on the depth value distribution of valid pixels around the conflict area through the assumption of spatial continuity. For example, bilinear interpolation or the K-nearest neighbor algorithm is used to fill in missing data.

[0094] First, extract the degree of polarization of each pixel in the spatial feature map and compare it with a preset threshold, such as the upper limit of the typical degree of polarization of skin diffuse reflection. The area below the threshold is marked as potentially abnormal. Calculate the depth difference between adjacent pixels. If it exceeds the preset depth threshold, such as the reasonable range of human face curvature change, it is determined as a geometric discontinuity area. Then, for the pixels that simultaneously satisfy "degree of polarization below the threshold" and "depth difference exceeding the limit", they are determined as conflict areas. For example, the specular reflection of a metal door frame projects onto the portrait area, resulting in abnormal polarization and depth jump in this area. Finally, perform neighborhood feature interpolation repair. Taking the conflict area as the center, select the depth values of the surrounding valid pixels to form a sample set, excluding the influence of other conflict areas. Use the bilinear interpolation algorithm to generate repair values according to the depth value distribution of the neighborhood samples. For example, for a conflict pixel, take the weighted average of the depth values of its four valid neighboring pixels above, below, left, and right as the new depth value. Perform regional smoothing on the repaired depth map through Gaussian filtering to eliminate local noise caused by interpolation and ensure geometric continuity.

[0095] Exemplarily, in the scenario of a one-way door at the laboratory entrance, the high reflectivity of the metal door frame causes abnormal polarization in the portrait area: due to the interference of the mirror reflection of the door frame, the polarization degree of a certain person's right cheek area drops to 0.1 (lower than the typical skin threshold of 0.2), and at the same time, the binocular depth point cloud shows that the depth difference between adjacent pixels in this area reaches 15 mm (exceeding the preset threshold of 10 mm). The system marks this area as a conflict area through consistency verification, and performs bilinear interpolation based on the depth values of the surrounding normal cheek areas (continuously distributed). After repair, the depth difference drops to 3 mm, which is consistent with the true face curvature. The repaired spatial feature map eliminates the reflection artifacts, ensuring that the subsequent multi-scale attention mechanism can accurately fuse the face geometry and texture features.

[0096] Through the consistency verification of polarization features and depth features, this embodiment effectively identifies and locates the depth abnormal conflict area in the portrait area, significantly improving the reliability of depth data for 3D face reconstruction in complex lighting and occlusion scenarios. Combining the neighborhood feature interpolation algorithm to adaptively repair the conflict area can restore the true depth information while maintaining the continuity of the facial topology structure, providing high-precision spatial geometric constraints for subsequent multi-modal feature fusion, and further enhancing the physical rationality and detail fidelity of material texture reconstruction.

[0097] Under complex lighting conditions, 3D face reconstruction faces the technical barrier of the mismatch of cross-domain feature scales between the polarization light modality and RGB images: when traditional methods directly align the polarization point cloud features with 2D images, due to the feature coupling between the polarization high-frequency details (such as skin micro-textures) and the low-frequency geometric structures (such as facial contours), the high-frequency information is submerged by the geometric features during the cross-modal fusion process. At the same time, the low-frequency reflection areas (such as the highlights of oily skin) cause the decoupling of depth features and material reflection characteristics due to the sudden change of polarization degree, resulting in the dual defects of texture blur and material distortion in the reconstructed model. This application solves the above technical problems through the following technical solutions.

[0098] In an implementable embodiment, before step S130: spatially align the point cloud polarization joint feature with the face image to obtain the aligned spatial feature map, and generate a dynamic mask by partitioning the reflection intensity of the spatial feature map based on the polarization degree of the polarization light image, the method includes: decomposing the polarization light image into high-frequency detail components and low-frequency structure components in multiple frequency domains, and performing cross-scale fusion with the geometric information in the point cloud polarization joint feature respectively to generate the point cloud polarization hierarchical feature.

[0099] Step S130: Spatially align the joint point cloud polarization features with the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image, including: Spatially align the hierarchical point cloud polarization features with the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image.

[0100] Multi-frequency domain decomposition refers to the process of decomposing a polarized light image into different frequency components through frequency domain analysis techniques such as wavelet transform. Among them, the high-frequency detail component corresponds to the rapidly changing features in the image, such as edges and textures, like the high-frequency noise of glass reflection, and the low-frequency structure component corresponds to the overall contour and slowly changing regions of the image, such as the general shape of the face). The high-frequency detail component refers to the high-frequency signal composed of local mutation features in the image, such as reflection artifacts and skin textures, reflecting the changes in surface micro-reflection characteristics. The low-frequency structure component refers to the low-frequency signal composed of smooth regions in the image (such as facial contours and background planes), reflecting the macroscopic geometric structure features. Cross-scale fusion refers to associating the high-frequency detail and low-frequency structure information with the point cloud geometric features (such as curvature and normal vector) respectively to form a hierarchical feature representation. For example, the high-frequency detail is associated with the local curvature of the point cloud to capture reflection artifacts, and the low-frequency structure is associated with the overall normal vector of the point cloud to describe the facial pose. The hierarchical point cloud polarization feature refers to the hierarchical data set generated through cross-scale fusion, including a high-frequency layer (detail texture and local geometry), a low-frequency layer (overall structure and material properties), and a cross-layer correlation index.

[0101] First, perform a three-level decomposition of the polarized light image using discrete wavelet transform (DWT) to obtain high-frequency components (LH, HL, HH) and low-frequency components (LL). The high-frequency components capture the edge oscillations in the specular reflection areas, such as the stripes reflected by the glass mirror surface, and the low-frequency components retain the macroscopic structures of the face and the background. By adjusting the filter parameters of the wavelet basis (such as Daubechies wavelet), the balance between high-frequency noise suppression and low-frequency structure retention is optimized. Then, perform high-frequency detail and geometric correlation: extract the polarization angle gradient of the high-frequency components and match it with the local curvature in the joint features of the point cloud. For example, calculate the mapping relationship between the high-frequency polarization angle change and the point cloud curvature through a convolutional neural network (CNN) to generate a high-frequency detail layer, such as marking the abnormal curvature distribution in the glass specular reflection area). And perform low-frequency structure and geometric correlation: spatially align the polarization degree distribution of the low-frequency components with the point cloud normal vector, and use the weighted average method to fuse the included angle between the low-frequency polarization degree and the point cloud normal vector to generate a low-frequency structure layer, such as the consistency between the overall curvature of the face and the normal vector of the background plane. Then perform hierarchical index construction: establish a cross-layer correlation table between the high-frequency layer and the low-frequency layer to ensure that features at different scales can be synchronously called during subsequent dynamic mask generation, such as the material attributes corresponding to the high-frequency specular reflection area in the low-frequency layer. Finally, project the point cloud polarization hierarchical features onto the two-dimensional face image plane, and complete coordinate registration through the calibration parameters of the binocular camera to ensure pixel-level alignment between the high-frequency detail layer (such as nose wing texture) and the low-frequency structure layer (such as facial contour). And calculate the reflection intensity of each pixel based on the polarization degree channel, and use adaptive threshold segmentation to divide the high-specular reflection area (polarization degree ≥ threshold) and the normal area. Combine morphological closing operation to eliminate segmentation noise and generate a binary dynamic mask, where the weight of the high-specular reflection area is 0 (suppression) and the weight of the normal area is 1 (retention).

[0102] Exemplarily, in the laboratory entrance scene, the door frame glass area captured by the polarization camera shows high-frequency specular reflection (polarization degree 0.85), while the human face skin area is low-frequency diffuse reflection (polarization degree 0.15). First, perform Haar wavelet decomposition on the polarized light image. The high-frequency components highlight the striped artifacts of the glass specular reflection, and the low-frequency components retain the overall contour of the face. Then, for the high-frequency detail layer: correlate the high-frequency polarization angle gradient of the glass specular reflection with the abnormally flat area (curvature close to 0) in the point cloud and mark it as background interference. For the low-frequency structure layer: fuse the low-frequency polarization degree of the face with the high-curvature area (curvature > 0.2) of the point cloud nose bridge to enhance the real facial features. Finally, according to the polarization degree threshold of 0.6, mark the glass area as the suppression area and the face area as the normal area to generate a dynamic mask. The mask edge is smoothly transitioned through Gaussian smoothing to avoid fusion artifacts caused by hard segmentation.

[0103] Figure 2 The flowchart of the method for generating a multi-modal feature vector provided by an embodiment of the present application is shown. As Figure 2As shown, the method includes steps S210 to S240.

[0104] In an implementable embodiment, step S140: Using a multi-scale attention mechanism, with a dynamic mask as a constraint, and based on a material geometry correlation model, perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, including:

[0105] S210: According to the predefined polarization angle - curvature correspondence in the material geometry correlation model, determine the target curvature range of each region in the spatial feature map. Based on the polarization angle - curvature mapping table in the material geometry correlation model, convert the polarization angle of each region in the spatial feature map into the target curvature range. For example, if the polarization angle of a certain region is 45°, the corresponding curvature range is 0.05 - 0.1 (high curvature nose bridge) or 0.01 - 0.03 (low curvature cheek).

[0106] S220: Based on the dynamic mask, divide the high - specular reflection suppression region and the normal region in the spatial feature map, and assign initial fusion weights to each region. Use the dynamic mask to divide the spatial feature map into a high - specular reflection suppression region (mask value is 0) and a normal region (mask value is 1). The initial fusion weights are assigned according to the region type: the weight of the suppression region approaches 0, and the weight of the normal region is 1.

[0107] S230: For each level of features in the spatial feature map, extract the polarization angle distribution and the actual curvature value, calculate the difference degree between the curvature value and the target curvature range, and adjust the initial fusion weight of the corresponding region according to the difference degree. For each level of features (such as the original, down - sampled scale), extract the actual curvature value and calculate the difference degree with the target curvature range. For example, if the actual curvature is 0.12 (exceeding the target range of 0.05 - 0.1), the difference degree is 0.12 - 0.1 = 0.02. Dynamically adjust the weight according to the difference degree: the greater the difference degree, the higher the weight attenuation coefficient, suppressing the feature response of the abnormal region.

[0108] S240: Superimpose the adjusted initial fusion weight on the face image features of the corresponding level, generate weighted fusion features layer by layer, and merge them to output a multi - modal feature vector. Multiply the adjusted weight element - by - element with the face image features of the corresponding level (such as the SIFT texture descriptor) to generate a weighted feature map. For example, the texture features in the suppression region are weakened, while the geometric - polarization features in the normal region are strengthened. Finally, merge the multi - scale feature maps through channel concatenation and global pooling into a unified multi - modal feature vector.

[0109] Figure 3 The flowchart of the method for generating point cloud polarization joint features provided by an embodiment of the present application is shown. As Figure 3 shown, the method includes steps S310 to S350.

[0110] In an implementable embodiment, step S120: spatially align and feature splice the polarized light image and the binocular depth point cloud data to generate a point cloud polarization joint feature, including:

[0111] S310: For each pixel point in the polarized light image, determine the corresponding three-dimensional space point in the binocular depth point cloud through the physical position mapping relationship, and establish an association index table between the pixel point and the three-dimensional space point.

[0112] The physical position mapping relationship refers to the geometric correspondence between the two-dimensional polarized light image pixel points and the three-dimensional binocular depth point cloud coordinates established through sensor calibration parameters (such as camera internal parameters and external parameters), ensuring that the same physical position can be accurately aligned in different sensor data. The association index table refers to a lookup table that records the spatial coordinates of the three-dimensional point cloud corresponding to each polarized image pixel point and is used for data association.

[0113] Based on the calibration parameters of the polarized camera and the binocular camera (such as focal length and baseline distance), calculate the three-dimensional point cloud coordinates corresponding to each pixel point in the polarized light image through projection transformation, establish the physical position mapping relationship between the pixel point and the three-dimensional space point, and generate an association index table.

[0114] S320: Extract the polarization parameters of each pixel point in the polarized light image. The polarization parameters include the polarization light intensity change direction and the polarization light intensity difference value.

[0115] The polarization parameters include the polarization light intensity change direction and the polarization light intensity difference value, that is, the polarization angle representing the vibration direction of the light wave and the polarization degree quantifying the proportion of polarized light. Extract the polarization angle and polarization degree parameters of each pixel from the polarized light image.

[0116] S330: According to the association index table, add the polarization parameters to the three-dimensional space points corresponding to the associated index to obtain joint data points including spatial coordinates, color information, and polarization parameters.

[0117] The joint data point refers to a multi-dimensional data unit that fuses three-dimensional coordinates, RGB colors, and polarization parameters and is used to describe the geometric and optical characteristics of the object surface. Attach the polarization parameters to the corresponding three-dimensional point cloud data according to the association index table to form joint data points including spatial coordinates, colors, and polarization parameters.

[0118] S340: Based on the geometric distribution density of the binocular depth point cloud, perform regional smoothing processing on the polarization parameters of the joint data points to obtain smoothed joint data points.

[0119] The geometric distribution density refers to the number of data points per unit volume in a three-dimensional point cloud, reflecting the density of spatial sampling. According to the geometric distribution density of the point cloud (for example, a high-density area indicates rich details, and a low-density area indicates a smooth surface), regional smoothing processing is performed on the polarization parameters of the combined data points: a small neighborhood window is used for local averaging in the high-density area to retain details; in the low-density area, the neighborhood window is enlarged to enhance parameter continuity.

[0120] S350: Use the set of combined data points after smoothing processing as the joint feature of point cloud polarization.

[0121] Regional smoothing processing refers to performing neighborhood weighted averaging on the polarization parameters according to the point cloud density to suppress noise interference. The set of combined data points after smoothing is output as the joint feature of point cloud polarization, and this feature includes the spatial structure, surface reflection attributes, and color information of the target area at the same time.

[0122] Exemplarily, a polarization camera and a binocular camera synchronously collect polarization light images and depth point clouds of a door frame area (background) and a person's face (portrait). Through calibration parameters, the system maps the pixel at the tip of the nose position in the polarization image to the corresponding three-dimensional coordinates in the depth point cloud to establish an association index table. For example, the pixel with coordinates (x = 320, y = 240) in the polarization image corresponds to the three-dimensional point (X = 1.2m, Y = 0.5m, Z = 2.0m) in the point cloud. The polarization angle of 45° and the polarization degree of 0.15 of this pixel are extracted and attached to the three-dimensional point data. For the door frame glass area (high reflectivity, low point cloud density), the system enlarges the smoothing window to perform regional averaging on the polarization degree to eliminate parameter jumps caused by glass reflection; while for the face area (high point cloud density), a small window is used for smoothing to retain skin texture details. The finally generated joint feature of point cloud polarization includes the geometric shape of the human face, skin reflection characteristics, and background material information, providing a data basis for subsequent reflection suppression and feature fusion.

[0123] In an implementable embodiment, step S150: Use a classification network to calculate the similarity between the multi-modal feature vector and the identity features in a preset face identity database to obtain target verification information, including:

[0124] Divide the multi-modal feature vector into multiple block features according to the spatial structure, material attributes, and texture information, and extract the feature expression values of each block feature; obtain the feature template of the target identity from the face identity database, and the feature template includes the historical mean and fluctuation range of each type of block feature; calculate the feature difference ratio between the feature expression value of each block feature and the corresponding historical mean, and determine the weight coefficient of each block feature according to the fluctuation range; perform weighted summation on the feature difference ratios according to the weight coefficients of each block feature to obtain the target matching degree; obtain the target verification information according to the comparison result between the target matching degree and the preset matching degree threshold.

[0125] The block features refer to dividing the multi-modal feature vector into multiple sub-feature segments according to semantic dimensions, including spatial structure blocks (such as face curvature distribution), material property blocks (such as skin polarization angle distribution), and texture information blocks (such as facial feature contour details). The feature expression value refers to the quantized value of the block feature after normalization, which is used to describe the statistical characteristics of the block. For example, the average curvature of the spatial structure block and the polarization angle variance of the material property block. The feature template refers to the feature reference data of legal identities pre-stored in the face identity database. Each template contains the mean and fluctuation range of multiple block features collected historically. For example, the mean curvature of the spatial structure block of a certain user is 0.5, and the fluctuation range is ±0.1. The feature difference ratio refers to the degree of difference between the current block feature value and the template mean, which is quantified by a ratio. The weight coefficient is dynamically adjusted based on the fluctuation range of the block feature. The smaller the fluctuation range, the higher the contribution weight of the block to identity discrimination. The target matching degree is the weighted sum of the feature difference ratios of each block, which is used to comprehensively evaluate the similarity between the current feature and the template.

[0126] First, divide the multi-modal feature vector into three types of block features: spatial structure, material property, and texture information according to the preset semantic dimensions. For example, use fixed-length cutting or dynamic division based on feature importance ranking to ensure that each type of block contains the same type of semantic information. For each block, extract the feature expression value through statistical methods (such as mean, variance) or deep learning encoders (such as fully connected layers). For example, the material property block calculates the peak value of the polarization angle distribution histogram as the expression value. Secondly, retrieve the feature template of the target identity from the face identity database to obtain the historical mean and fluctuation range of each block. For example, for the texture information block of a certain user, load the mean vector and standard deviation range of the texture features of 100 face images collected historically. Subsequently, calculate the difference ratio between the current block feature expression value and the template mean. For example, use cosine similarity or Euclidean distance to quantify the difference, and convert the difference value into a ratio value through normalization. According to the fluctuation range of the block in the template, dynamically assign weight coefficients: if the historical fluctuation range of a certain block is small (such as the stable polarization angle of the material property block), a higher weight is given; otherwise (such as the large fluctuation of the texture block affected by light), the weight is reduced. Finally, perform a weighted sum of the difference ratios of all blocks to obtain the target matching degree. If the target matching degree exceeds the preset threshold (such as 0.9), it is determined that the identity verification is passed; otherwise, it is marked as failed.

[0127] Exemplarily, in the one-way access control scenario of the medical isolation area, multi-modal feature vectors of a certain medical staff are collected and divided into three blocks: spatial structure, material property, and texture information. The average curvature (0.48) of the facial depth point cloud is extracted from the spatial structure block, the variance of the skin polarization angle (0.12) is calculated from the material property block, and the SIFT feature descriptor of the texture details at the corners of the eyes is extracted from the texture block. The feature template of the medical staff is retrieved from the database, with the historical average value of the spatial structure being 0.5 (fluctuating ±0.05), the mean square deviation of the material property being 0.1 (fluctuating ±0.02), and the similarity threshold of the texture block being 0.85. The current spatial structure difference ratio is calculated as (0.48 - 0.5) / 0.5 = 0.04, the material property difference ratio is (0.12 - 0.1) / 0.1 = 0.2, and the cosine similarity of the texture block is 0.88. Weights are assigned according to the fluctuation range: the spatial structure weight is 0.6 (small fluctuation), the material property weight is 0.3 (medium fluctuation), and the texture weight is 0.1 (large fluctuation). The weighted matching degree is 0.04×0.6 + 0.2×0.3 + (1 - 0.88)×0.1 = 0.084, which is lower than the threshold of 0.1, so it is determined that the verification is passed and an opening instruction is generated.

[0128] Based on the same concept, an embodiment of the present application provides a one-way door control system based on image recognition. The following will be combined with Figure 4 to elaborate on the one-way door control system based on image recognition provided by the embodiment of the present application.

[0129] Figure 4 is a structural block diagram of a one-way door control system based on image recognition shown in an embodiment of the present application.

[0130] As Figure 4 shown, the one-way door control system based on image recognition may include:

[0131] An acquisition module 410, configured to acquire a polarized light image of a target area and a face image and binocular depth point cloud data of a human portrait area in the target area, where the target area includes a human portrait area and a background area;

[0132] A generation module 420, configured to perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate a point cloud polarization joint feature;

[0133] The generation module 420 is further configured to perform spatial alignment on the point cloud polarization joint feature and the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the spatial feature map according to the polarization degree of the polarized light image;

[0134] The generation module 420 is further configured to use a multi-scale attention mechanism, with the dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image, and generate a multi-modal feature vector including spatial structure, material attributes, and texture information;

[0135] The verification module 430 is configured to calculate the similarity between the multi-modal feature vector and the identity features in the preset face identity database by using a classification network, and obtain target verification information;

[0136] The control module 440 is configured to determine the target control information corresponding to the target verification information according to the corresponding relationship between the preset verification information and the control information, and complete the control of the one-way door through the target control information.

[0137] In one embodiment, the generation module 420 is specifically configured to use a multi-scale attention mechanism, with the dynamic mask as a constraint, and based on the material geometry correlation model, perform feature fusion on the spatial feature map and the face image, and generate a multi-modal feature vector including spatial structure, material attributes, and texture information, where the material geometry correlation model is trained based on the polarization angle of the historical polarized light image and the face geometry parameters of the historical spatial feature map through a predefined correspondence between the skin polarization angle and the curvature.

[0138] In one embodiment, before using the multi-scale attention mechanism, with the dynamic mask as a constraint, to perform feature fusion on the spatial feature map and the face image and generate a multi-modal feature vector including spatial structure, material attributes, and texture information, the generation module 420 is further configured to perform consistency verification on the polarization feature and the depth feature in the spatial feature map, and determine the conflict area in the spatial feature map when it is detected that the polarization degree in the portrait area is lower than the preset polarization threshold and the depth value difference in the adjacent area exceeds the preset depth threshold; use neighborhood feature interpolation to repair the depth value of the conflict area in the spatial feature map to obtain a repaired spatial feature map.

[0139] In one embodiment, before spatially aligning the point cloud polarization joint feature with the face image to obtain an aligned spatial feature map and generating a dynamic mask by partitioning the reflection intensity of the spatial feature map based on the polarization degree of the polarized light image, the generation module 420 is further configured to perform multi-frequency domain decomposition on the polarized light image to obtain a high-frequency detail component and a low-frequency structure component, and perform cross-scale fusion with the geometric information in the point cloud polarization joint feature respectively to generate a point cloud polarization hierarchical feature; spatially align the point cloud polarization hierarchical feature with the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the reflection intensity of the spatial feature map based on the polarization degree of the polarized light image.

[0140] In one embodiment, the generation module 420 is specifically configured to determine the target curvature range of each region in the spatial feature map according to the predefined correspondence between the polarization angle and the curvature in the material geometry association model; divide the high-reflectivity suppression region and the normal region in the spatial feature map based on the dynamic mask, and assign an initial fusion weight to each region; for each hierarchical feature of the spatial feature map, extract the polarization angle distribution and the actual curvature value, calculate the difference degree between the curvature value and the target curvature range, and adjust the initial fusion weight of the corresponding region according to the difference degree; superimpose the adjusted initial fusion weight on the face image feature of the corresponding level, generate the weighted fusion feature layer by layer, and output the multi-modal feature vector after merging.

[0141] In one embodiment, the generation module 420 is specifically configured to, for each pixel point in the polarized light image, determine the corresponding three-dimensional spatial point in the binocular depth point cloud through the physical position mapping relationship, and establish an association index table between the pixel point and the three-dimensional spatial point; extract the polarization parameters of each pixel point in the polarized light image, where the polarization parameters include the polarization light intensity change direction and the polarization light intensity difference value; according to the association index table, add the polarization parameters to the three-dimensional spatial point corresponding to the associated index to obtain a joint data point including spatial coordinates, color information, and polarization parameters; based on the geometric distribution density of the binocular depth point cloud, perform regional smoothing processing on the polarization parameters of the joint data points to obtain the smoothed joint data points; use the set of smoothed joint data points as the point cloud polarization joint feature.

[0142] In one embodiment, the verification module 430 is specifically configured to divide the multi-modal feature vector into multiple block features according to the spatial structure, material attributes, and texture information, and extract the feature expression values of each block feature; obtain the feature template of the target identity from the face identity database, where the feature template includes the historical mean and the fluctuation range of each type of block feature; calculate the feature difference ratio between the feature expression value of each block feature and the corresponding historical mean, and determine the weight coefficient of each block feature according to the fluctuation range; perform weighted summation on the feature difference ratio according to the weight coefficient of each block feature to obtain the target matching degree; obtain the target verification information according to the comparison result between the target matching degree and the preset matching degree threshold.

[0143] Figure 4 Each module in the system shown has the function of implementing Figure 1 the Figure 3 functions of each step therein, and can achieve the corresponding technical effects. For the sake of brevity, the description is not repeated here.

[0144] Figure 5 FIG. shows a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application.

[0145] The electronic device may include a processor 510 and a memory 520 storing computer program instructions.

[0146] Specifically, the above-mentioned processor 510 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0147] The memory 520 may include a mass storage for data or instructions. By way of example and not limitation, the memory 520 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 520 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 520 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, the memory 520 is a non-volatile solid state memory.

[0148] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present disclosure.

[0149] The processor 510 reads and executes the computer program instructions stored in the memory 520 to implement any one of the one-way door control methods based on image recognition in the above embodiments.

[0150] In one example, the electronic device may further include a communication interface 530 and a bus 540. Among them, as Figure 5 shown, the processor 510, the memory 520, and the communication interface 530 are connected through the bus 540 and complete communication with each other.

[0151] The communication interface 530 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application.

[0152] The bus 540 includes hardware, software, or both, and couples the components of the online data flow metering device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, the bus 540 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0153] The electronic device can execute the one-way door control method based on image recognition in the embodiments of the present application, so as to implement the combination Figures 1 to 3 of the one-way door control method based on image recognition described above.

[0154] In addition, in combination with the one-way door control method based on image recognition in the above embodiments, the embodiments of the present application can be implemented by providing a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the one-way door control methods based on image recognition in the above embodiments is implemented.

[0155] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0156] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, and so on. When implemented in software, the elements of the present application are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via a data signal carried in a carrier wave. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0157] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0158] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware for performing the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0159] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A one-way door control method based on image recognition, characterized in that, Including: Obtain a polarized light image of a target area in a highly reflective scene, as well as a face image and binocular depth point cloud data of a portrait area in the target area, where the target area includes the portrait area and a background area; Perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate a point cloud polarization joint feature; Perform spatial alignment on the point cloud polarization joint feature and the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the spatial feature map according to the polarization degree of the polarized light image; Using a multi-scale attention mechanism, with the dynamic mask as a constraint, perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information; Use a classification network to calculate the similarity between the multi-modal feature vector and the identity features in a preset face identity database to obtain target verification information; According to the corresponding relationship between the preset verification information and the control information, determine the target control information corresponding to the target verification information, and complete the control of the one-way door through the target control information.

2. The method according to claim 1, wherein The step of using a multi-scale attention mechanism, with the dynamic mask as a constraint, performing feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information includes: Using a multi-scale attention mechanism, with the dynamic mask as a constraint, and based on a material geometry correlation model, perform feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, where the material geometry correlation model is trained based on the polarization angle of historical polarized light images and the face geometry parameters of historical spatial feature maps through a predefined correspondence between skin polarization angle and curvature.

3. The method according to claim 2, wherein Before the step of using a multi-scale attention mechanism, with the dynamic mask as a constraint, performing feature fusion on the spatial feature map and the face image to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, the method includes: Perform consistency verification on the polarization feature and the depth feature in the spatial feature map. When it is detected that there is a region in the portrait area where the polarization degree is lower than a preset polarization threshold and the depth value difference in the adjacent region exceeds a preset depth threshold, determine the conflict region in the spatial feature map; Use neighborhood feature interpolation to repair the depth value of the conflict region in the spatial feature map to obtain the repaired spatial feature map.

4. The method according to claim 1, wherein Before the step of performing spatial alignment on the point cloud polarization joint feature and the face image to obtain an aligned spatial feature map, and generating a dynamic mask by partitioning the spatial feature map according to the polarization degree of the polarized light image, the method includes: Perform multi-frequency domain decomposition on the polarized light image to obtain a high-frequency detail component and a low-frequency structure component, and perform cross-scale fusion with the geometric information in the point cloud polarization joint feature respectively to generate a point cloud polarization hierarchical feature; Spatially aligning the joint point cloud polarization features with the face image to obtain an aligned spatial feature map, and generating a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image, including: Spatially aligning the hierarchical point cloud polarization features with the face image to obtain an aligned spatial feature map, and generating a dynamic mask by partitioning the spatial feature map based on the polarization degree of the polarized light image.

5. The method according to claim 2, wherein Using a multi-scale attention mechanism, with the dynamic mask as a constraint, and performing feature fusion on the spatial feature map and the face image based on a material geometry correlation model to generate a multi-modal feature vector including spatial structure, material attributes, and texture information, including: Determining the target curvature range of each region in the spatial feature map according to the predefined correspondence between the polarization angle and the curvature in the material geometry correlation model; Dividing the high-reflection suppression region and the normal region in the spatial feature map based on the dynamic mask, and assigning initial fusion weights to each region; For each level of features in the spatial feature map, extracting the polarization angle distribution and the actual curvature value, calculating the difference degree between the curvature value and the target curvature range, and adjusting the initial fusion weight of the corresponding region according to the difference degree; Overlaying the adjusted initial fusion weight with the face image features of the corresponding level, generating weighted fusion features layer by layer, and outputting the multi-modal feature vector after merging.

6. The method according to claim 1, characterized in that, Spatially aligning and feature stitching the polarized light image and the binocular depth point cloud data to generate joint point cloud polarization features, including: For each pixel point in the polarized light image, determining the corresponding three-dimensional spatial point in the binocular depth point cloud through a physical position mapping relationship, and establishing an association index table between the pixel point and the three-dimensional spatial point; Extracting the polarization parameters of each pixel point in the polarized light image, where the polarization parameters include the polarization light intensity change direction and the polarization light intensity difference value; According to the association index table, adding the polarization parameters to the three-dimensional spatial points with the corresponding association index to obtain joint data points including spatial coordinates, color information, and polarization parameters; Based on the geometric distribution density of the binocular depth point cloud, performing regional smoothing processing on the polarization parameters of the joint data points to obtain the smoothed joint data points; Taking the set of the smoothed joint data points as the joint point cloud polarization features.

7. The method according to claim 1, characterized in that, Using a classification network to calculate the similarity between the multi-modal feature vector and the identity features in a preset face identity database to obtain target verification information, including: Dividing the multi-modal feature vector into multiple block features according to spatial structure, material attributes, and texture information, and extracting the feature expression values of each block feature; Obtaining the feature template of the target identity from the face identity database, where the feature template includes the historical mean and the fluctuation range of each type of block feature; Calculating the feature difference ratio between the feature expression value of each block feature and the corresponding historical mean, and determining the weight coefficient of each block feature according to the fluctuation range; The weighted sum of the feature difference ratios is calculated according to the weight coefficients of each sub-block feature to obtain the target matching degree; The target verification information is obtained according to the comparison result between the target matching degree and the preset matching degree threshold.

8. A one-way door control system based on image recognition, characterized in that, The system includes: An acquisition module, configured to acquire a polarized light image of a target area in a highly reflective scene, as well as a face image and binocular depth point cloud data of a portrait area in the target area, where the target area includes the portrait area and a background area; A generation module, configured to perform spatial alignment and feature stitching on the polarized light image and the binocular depth point cloud data to generate a point cloud polarization joint feature; The generation module is further configured to perform spatial alignment on the point cloud polarization joint feature and the face image to obtain an aligned spatial feature map, and generate a dynamic mask by partitioning the reflective intensity of the spatial feature map based on the polarization degree of the polarized light image; The generation module is further configured to use a multi-scale attention mechanism to perform feature fusion on the spatial feature map and the face image with the dynamic mask as a constraint to generate a multi-modal feature vector including spatial structure, material attributes, and texture information; A verification module, configured to calculate the similarity between the multi-modal feature vector and the identity features in a preset face identity database by using a classification network to obtain target verification information; A control module, configured to determine target control information corresponding to the target verification information according to the corresponding relationship between the preset verification information and the control information, and complete the control of the one-way door through the target control information.

9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the one-way door control method based on image recognition according to any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by the processor, the one-way door control method based on image recognition according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Face recognition system and method for access control system

    CN118711237A

  • Image recognition method and device, computer equipment, readable storage medium and program product

    CN119314013A

  • Access control system based on face recognition and recognition method

    CN119672782A

  • Method and apparatus to complement depth image

    US20220067950A1

  • Integrating model reuse with model retraining for video analytics

    US20240096063A1

Cited By

  • Method and system for face image recognition

    CN120726684A

  • Identification processing method and system for key area in aerial power line surveying and mapping image

    CN121074345A

  • Face depth detection method and system based on multi-mode double shooting

    CN121096003A

  • A multi-modal dual-camera face depth detection method and system

    CN121096003B

  • Multi-reflection suppression method and device based on super depth-of-field fusion algorithm

    CN121414632A