Intelligent door lock face recognition unlocking method

By using time-domain sparse points and frequency-domain spatial textures in smart door locks to extract image features, and combining multi-mode mechanisms and flat mechanisms to generate target images, the problem of low facial recognition accuracy in facial occlusion or multi-user dynamic scenarios is solved, and higher recognition accuracy and robustness are achieved.

CN120014740AActive Publication Date: 2025-05-16SHENZHEN ISURPASS TECH CO LTD

Patent Information

Application Number
CN202510479916.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing smart door lock face recognition unlocking method cannot accurately obtain facial features in facial occlusion or multi-user dynamic scenarios, resulting in a decrease in recognition accuracy.

Method used

Image features are extracted based on time-domain sparse points and frequency-domain spatial textures, combined with multi-mode mechanism and flat state mechanism, the net area and dirty area are determined, the target image is generated, and the user's height is estimated through the head key points, a three-dimensional mapping table for height-domain-gradient parameters is constructed, and the gradient calculation area is adjusted to identify face features.

Benefits of technology

Accurate facial features acquisition in facial occlusion or multi-user dynamic scenarios is realized, and the accuracy and robustness of facial recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014740A_ABST
    Figure CN120014740A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent door lock face recognition unlocking method, and relates to the technical field of intelligent door lock control. Determining the number of users and a corresponding time-space curve based on the plurality of pre-acquired images, and obtaining a primary image; if the moving speed value is not greater than the first threshold value, marking the pre-acquired image as a primary image; obtaining the time-frequency characteristics of the primary image, determining a plurality of local areas to obtain an interference value, and if the interference value is greater than a second threshold value, determining a clean area and a dirty area based on a multimode mechanism, and generating a target image; if the interference value is not greater than the second threshold value, obtaining a target image according to a flat state mechanism; estimating the height of a user, constructing a height-distance-gradient parameter three-dimensional mapping table, adjusting a gradient calculation region based on a target image to recognize face features, and comparing the face features with a pre-stored image; based on the time domain sparse points and the frequency domain spatial texture, the feature information in the image is accurately extracted, a multimode mechanism is set, and the effect of accurately recognizing the user with the shielded face is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent door lock control, and in particular to an intelligent door lock face recognition unlocking method. Background Art

[0002] Smart door locks unlocked by face recognition use cameras to capture face images and compare the features of face images with pre-stored face images to verify the identity of the user, which can effectively reduce the time it takes for the user to open the door. At the same time, the user does not need to carry keys with him, which improves the convenience of opening the door. However, the face image is captured by a camera at a fixed angle, and the viewing angle range is fixed. For users of different heights to be identified, the complete face image cannot be captured, resulting in misidentification and failure to successfully open the door lock.

[0003] The Chinese invention patent with application number 202310247026.5 provides a smart door lock face recognition unlocking method, device, medium and smart door lock. By collecting multiple pre-collected images with different collection angles, the target size of the mask is determined from the preset mask size according to the collection distance between the face of the user to be identified and the image collection device, and the mask is placed on the pre-collected image with the target size. The mask is moved on the pre-collected image to obtain the corresponding pixel feature information; the overlapping pixels of each pre-collected image are determined according to the pixel feature information, and the relative position relationship of multiple pre-collected images is predicted according to the overlapping pixels; and the multiple pre-collected images are spliced ​​according to the relative position relationship to obtain the target image.

[0004] However, in the prior art, facial features from multiple angles are selected using masks and then stitched together, which is limited to static single-user scenarios. When there is facial occlusion or dynamic multi-user scenarios, facial features cannot be accurately obtained, resulting in misjudgment and reducing the accuracy of face recognition. Summary of the invention

[0005] This application solves the problem in the prior art that facial features cannot be accurately acquired, thereby reducing the accuracy of facial recognition, by providing a smart door lock facial recognition unlocking method, and achieves the technical effect of accurately acquiring facial features and improving facial recognition accuracy.

[0006] This application provides a smart door lock face recognition unlocking method, including:

[0007] S100: determining the number of users and corresponding space-time curves based on a plurality of pre-collected images; and obtaining a moving speed value based on the space-time curve;

[0008] S200: If the moving speed value is greater than the first threshold, the blur degree of the pre-collected image is calculated, and a deconvolution operation and layered fusion are performed to obtain a preliminary image;

[0009] If the moving speed value is not greater than the first threshold, the pre-captured image is marked as a preliminary selected image;

[0010] S300: obtaining time-frequency features of the preliminary image, and then determining multiple local areas to obtain interference values. If the interference value is greater than a second threshold, determining a clean area and a dirty area based on a multi-mode mechanism to generate a target image; if the interference value is not greater than the second threshold, obtaining the target image according to a flat state mechanism;

[0011] S400: estimating the height of the user through key points of the head, constructing a three-dimensional mapping table of height-distance-gradient parameters, adjusting the gradient calculation area based on the target image to recognize facial features, and comparing with the pre-stored image;

[0012] The moving speed value is used to measure the change range of the user's posture.

[0013] Furthermore, the method further includes: S110: identifying the number of users, if the number of users is 1, executing S200; if the number of users is greater than 1, executing S120;

[0014] S120: obtaining a preliminary user based on the splitting mechanism, performing facial segmentation on the preliminary user, identifying facial feature points and comparing them with pre-stored images, determining a target user, and obtaining a moving speed value based on the time-space curve, and executing S200;

[0015] Among them, the splitting mechanism includes: determining the deviation from the target path based on the space-time curve, marking the users corresponding to the deviation less than a preset third threshold as preliminary users; the target path refers to the straight path directly in front of the door lock; the deviation is used to measure the difference between each user's moving path and the preset target path.

[0016] Furthermore, the height of the user is estimated through the key points of the head, a three-dimensional mapping table of height-distance-gradient parameters is constructed, and the gradient calculation area is adjusted based on the target image to identify facial features, including: establishing a head posture model, using the three-dimensional head model to match the facial key points, establishing a mapping relationship between the head posture and the key point coordinates, and estimating the height of the user according to the relationship between the user's head posture, the position of the head key points and the position of the ground; constructing a three-dimensional mapping table of height-distance-gradient parameters, and the estimated height of the target image to obtain the corresponding gradient parameters; based on the gradient parameters, determining the gradient calculation area to highlight the facial features.

[0017] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0018] By accurately extracting feature information from images based on sparse points in the time domain and spatial textures in the frequency domain, a more precise local division effect is achieved. A multi-mode mechanism is set up to determine the boundary curve based on infrared imaging and depth differences, which can adapt to users with different occlusions, lighting conditions and facial features, and achieves the effect of accurately obtaining facial features for users with facial occlusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The present invention is a flowchart of a smart door lock face recognition unlocking method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] To facilitate the understanding of the present invention, the present application will be described more comprehensively below with reference to the relevant drawings; the drawings show preferred embodiments of the present invention, but the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to enable a more thorough and comprehensive understanding of the disclosed content of the present invention.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technical users in the technical field of the invention; the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more related listed items.

[0022] Embodiment 1: Figure 1 As shown, a smart door lock face recognition unlocking method includes:

[0023] S100: Determine the number of users and the corresponding space-time curve based on a plurality of pre-collected images; and obtain a moving speed value based on the space-time curve.

[0024] S200: If the moving speed value is greater than the first threshold, the blur degree of the pre-captured image is calculated and the blur kernel is estimated, and a deconvolution operation and hierarchical fusion are performed to obtain a preliminary image; if the moving speed value is not greater than the first threshold, the pre-captured image is marked as a preliminary image.

[0025] In some embodiments, in response to the unlocking operation of the smart door lock, the present application collects multiple pre-captured images at different collection angles and distances through an image acquisition device configured on the smart door lock, wherein the multiple pre-captured images are obtained by the image acquisition device adjusting the collection angle multiple times according to a preset frequency and a preset angle; the image acquisition device can be one or more cameras, and also includes an infrared camera. Each time the collection angle is adjusted, one or more pre-captured images can be captured. After adjusting the collection angle once, multiple pre-captured images are captured, and the one with the highest clarity can be used as the pre-captured image after adjusting the collection angle this time to ensure that there is only one pre-captured image for each collection angle. The pre-captured image is the captured user image, including a face image and a body part image.

[0026] In some embodiments, the first threshold is pre-set based on historical experimental data. First, a large number of face images with different movement speeds are obtained, and then recognition is performed to find the critical point for successful recognition. The movement speed corresponding to the critical point is the corresponding optimal first threshold. The first threshold can also be dynamically adjusted according to actual conditions to meet different experimental requirements.

[0027] S300: Obtain the time-frequency characteristics of the preliminary image, determine multiple local areas based on the time-frequency characteristics and calculate the interference value. If the interference value is greater than the second threshold, determine the clean area and the dirty area based on the multi-mode mechanism and generate a target image; if the interference value is not greater than the second threshold, obtain the target image according to the flat state mechanism.

[0028] In this embodiment, pixel analysis is performed based on the obtained multiple pre-collected images to obtain time-frequency features and space-time curves; a number of local areas are determined based on the time-frequency features, and the local areas are key points on the pre-collected images determined based on the changes in pixels and texture features. The face key points are constructed based on face organs, and the local areas include the vector values ​​of each pixel, the number of face key points corresponding to the local pixel set, and the vector values ​​of each pixel in the pixel set to be determined, and the number of face key points corresponding to the pixel set to be determined, and the feature vector values ​​of multiple pixels in the local pixel set and the feature vector values ​​of multiple pixels in the pixel set to be determined are recalculated until there is no local pixel set with the same feature vector value in each of the pre-collected images. In this way, the existence of local pixel sets with the same feature vector value is avoided, and the same local pixel set is avoided when calculating overlapping pixels, resulting in the failure of overlapping pixel determination.

[0029] In some embodiments, a facial key point detection algorithm is used to perform key point detection on multiple pre-collected images collected continuously, extract the coordinates of facial key points in each image, including key positions such as eyes, nose, and mouth, and obtain the displacement characteristics of each key point based on the space-time curve, wherein the displacement characteristics include: key point coordinates, displacement direction, speed value, and time interval. Specifically, for the same key point between consecutive frames, its displacement vector (i.e., coordinate difference) is calculated. According to the time interval between the displacement vector and the image acquisition, the speed value of the key point is calculated (the displacement size divided by the time interval). The displacement direction and speed value of each key point are plotted as a curve over time to form a space-time curve. The space-time curve reflects the motion trajectory and speed change of facial key points in three-dimensional space.

[0030] The space-time curve is used to reflect the user's motion trajectory and speed changes. The average speed value is obtained by dividing the modulus of the displacement vector by the time interval. The displacement features include: key point coordinates, displacement direction, speed value and time interval.

[0031] Based on the spatiotemporal curve, the displacement characteristics of facial key points between consecutive frames in the pre-collected image are calculated. For each key point, the average speed value between consecutive frames is calculated, and the standard deviation of the speed value of each key point is calculated as the change amplitude quantification value, which is used to reflect the stability of the key point speed; the average speed value of all key points and the change amplitude quantification value are weighted summed to obtain the moving speed value. The calculation formula is as follows:

[0032] ;

[0033] Among them, M is the moving speed value, is the weight coefficient of the i-th key point, and its value range is [0, 1]; is the average velocity value of the i-th key point; It is the quantitative value of the change amplitude of the i-th key point.

[0034] In some embodiments, the time-frequency features include time-domain sparse points and frequency-domain spatial textures, and the time-domain sparse points are: extracting pixel feature information of each pre-collected image, monitoring the grayscale value change of each pixel in real time, and marking the pixel points greater than the grayscale change threshold based on a preset grayscale change threshold; recording the spatiotemporal coordinates of each time-domain sparse point, including a timestamp and a pixel position;

[0035] The frequency domain spatial texture is: obtain the pre-collected image, perform wavelet transform, decompose the frequency domain components at different scales, extract the low-frequency contour and high-frequency detail features from the frequency domain components, and form the frequency domain spatial texture. Wavelet transform is a new transform analysis method, which performs multi-scale refinement analysis on the image through calculation functions such as scaling and translation; by convolving a scale-adjustable and position-variable wavelet function with the signal, the components of the signal at different scales and positions are obtained. These wavelet components can reflect the local characteristics of the signal at different frequencies and times, and this application will not go into details here.

[0036] In some embodiments, the temporal sparse points are determined by extracting pixel feature information of each pre-collected image and monitoring the grayscale value change of each pixel in real time. When the grayscale value change of a pixel is greater than a preset grayscale change threshold, the pixel is marked as a temporal sparse point.

[0037] Time-domain sparse points can capture pixels with significant grayscale changes in the image, which often correspond to key features or dynamically changing areas in the image. By recording the spatiotemporal coordinates of each time-domain sparse point (including timestamp and pixel position), it is possible to accurately track and locate dynamically changing images. Since only focusing on pixels with significant grayscale changes, the time-domain sparse point method can greatly reduce the amount of calculation in subsequent processing and improve processing efficiency. Through time-domain sparse points, dynamic features in the image can be more effectively extracted, providing an accurate data basis for subsequent local determination and interference value calculation. The spatiotemporal positioning capability of time-domain sparse points helps to establish a stable feature correspondence between consecutive frames and improve the robustness of image processing.

[0038] Frequency domain spatial texture is obtained by acquiring pre-acquired images and performing wavelet transform. Wavelet transform can decompose frequency domain components at different scales, from which low-frequency contours and high-frequency detail features can be extracted to form frequency domain spatial texture. By extracting low-frequency contours and high-frequency detail features, frequency domain spatial texture can enhance useful information in the image, suppress noise and irrelevant details, and provide an effective way to describe texture features in the image, which is helpful for texture analysis and recognition in subsequent processing. It makes local determination more accurate and comprehensive, and can capture feature information at different scales; it can more effectively describe texture features in the image, and provide strong support for interference value calculation and image quality assessment.

[0039] Specifically, the acquisition device is initialized, multiple face images are continuously acquired, time domain sparse points and frequency domain spatial textures are extracted, and the local area is determined and the interference value is calculated. According to the interference value, it is judged whether the acquired face image is blocked, and a second threshold is pre-set. The second threshold is pre-set according to historical experimental data and is used to measure the degree of interference factors existing in the face image. Specifically, a large number of face images with different degrees of occlusion can be acquired, and the critical point corresponding to the degree of occlusion is found according to the success rate of recognition. The degree of occlusion corresponding to the critical point is set as the second threshold. Specifically, it can be adjusted above and below the critical point according to the actual situation; if the interference value is greater than the second threshold, it means that there is occlusion, such as sunglasses, masks, hats, etc. For the case of occlusion, the demarcation curve is determined based on the multi-mode mechanism.

[0040] In some embodiments, the multi-mode mechanism is: determining a demarcation curve based on the difference in depth information between infrared image data and pixel points, and obtaining a clean area and a dirty area based on the demarcation curve; wherein, infrared image data is obtained and image boundaries are identified, boundary points are randomly selected and the temperature gradient change rate is calculated one by one, a temperature change threshold is pre-set, and an area where the temperature gradient change rate is greater than the temperature change threshold is identified as a demarcation area; edge points are determined based on the depth information difference of the pixel points in the demarcation area; and a demarcation curve is obtained based on the edge points. Specifically, an infrared camera is used to collect thermal imaging images of the user's face to ensure that the collection environment is stable and to avoid interference from external heat sources. The collected infrared image is denoised, such as using Gaussian filtering to remove noise points in the image, and image enhancement is performed to improve the image contrast, so as to facilitate subsequent temperature gradient analysis. In the pre-processed infrared image, multiple points on the facial boundary are randomly selected as initial analysis points. For each selected boundary point, its temperature gradient change rate along the normal direction is calculated, and the temperature gradient change rate can be obtained by calculating the temperature difference and distance ratio of adjacent pixel points. A temperature change threshold is preset, which is determined based on experimental data or empirical values, and regions where the temperature gradient change rate is greater than the temperature change threshold are identified, and these regions are regarded as boundary regions.

[0041] Collect the depth data of the user's face, and analyze the depth information difference of the pixels in the identified boundary area. The depth information difference can be obtained by calculating the difference in the depth values ​​of adjacent pixels. Based on the depth information difference, determine the edge points in the boundary area. The edge points are usually located in areas with significant depth changes, such as the junction of the occluder and the face. Connect the identified edge points to form an initial boundary curve. Preferably, a curve fitting algorithm (such as B-spline curve fitting) is used to optimize the initial boundary curve to make it smoother and consistent with the actual boundary shape; considering the prior knowledge of the facial contour, the boundary curve is further adjusted to improve its accuracy.

[0042] In some embodiments, the clean area refers to a clean and unobstructed face image; the dirty area refers to a face image with occlusion; and the dividing curve refers to the dividing line between the clean area and the dirty area.

[0043] In some embodiments, a clean area and a dirty area are obtained based on a multi-mode mechanism, and a target image is generated, including: obtaining boundary pixel points of the clean area, extracting visible light texture features, and generating a mixed feature vector in combination with temperature distribution features in infrared thermal images; obtaining contour boundaries and temperature distribution features of the dirty area, and performing prediction and completion in combination with the mixed feature vector of the clean area to obtain a target image.

[0044] In this embodiment, based on infrared imaging and depth difference, the accuracy and robustness of boundary curve recognition are improved, and it can adapt to different obstructions, different lighting conditions and users with different facial features. It has a wide range of applicability and meets application scenarios with high real-time requirements such as smart door locks. It solves the problem that the user's face image cannot be accurately recognized when it is blocked.

[0045] In some embodiments, if the interference value is not greater than the second threshold, it indicates that the obtained pre-captured image is a clear and stable face image, and the target image is obtained according to the flat state mechanism; the flat state mechanism is: determining the overlapping pixels and the relative position relationship according to a number of local areas, and splicing the multiple pre-captured images according to the relative position relationship to obtain the target image.

[0046] Specifically, the overlapping pixels of each of the pre-captured images are determined based on several local areas, and based on the overlapping pixels, the relative position relationship of the multiple pre-captured images on the face of the user to be identified is predicted. It can be understood that if the pixel feature information in different pre-captured images is the same, then they are overlapping pixels. The pixel feature information can be the vector value of each pixel. If there are multiple pixels under local coverage in different pre-captured images, and the vector values ​​of the pixels at each position are the same, then the pixel feature information in the pre-captured images is the same. According to the acquisition distance between the face of the user to be identified and the image acquisition device, a local target size is determined from preset local sizes, wherein there is a one-to-one correspondence between the preset local size and the acquisition distance, and the pixel feature information of each pre-acquired image is obtained based on the extraction of the target size on each pre-acquired image, a number of local areas are determined according to time-frequency characteristics, and the local areas are moved on each pre-acquired image with the same step length; the one-to-one correspondence between the local size and the acquisition distance can be made into a table, and by looking up the table, according to the acquisition distance between the face of the user to be identified and the image acquisition device, the local target size is determined from the time domain characteristics; the multiple pre-acquired images are spliced ​​according to the relative position relationship to obtain the target image.

[0047] S400: Compare the target image with the pre-stored image; obtain the target image and compare it with the pre-stored image to obtain a comparison result; if the comparison result indicates a successful match (i.e., the similarity is greater than a preset similarity threshold), unlock the smart door lock; otherwise, re-compare or issue an early warning. The similarity threshold needs to be set according to the actual situation. To ensure safety, the optimal setting range is between 80% and 100%.

[0048] In this embodiment, based on the time domain sparse points and frequency domain spatial texture, multiple local areas are directly clustered and divided in the image. By comprehensively considering the information in the time domain and frequency domain, the feature distribution and texture structure in the image can be more accurately reflected. Combined with the texture information of the image, the feature information in the image can be more accurately extracted, making the division of the local area more accurate and reasonable; avoiding the problem of being unable to accurately identify facial features due to directly dividing the area to obtain features.

[0049] In this embodiment, a multi-mode mechanism is set to determine the demarcation curve based on infrared imaging and depth difference, which improves the accuracy and robustness of demarcation curve recognition, solves the problem that the user's face image cannot be accurately recognized when it is blocked, and can adapt to users with different obstructions, lighting conditions and facial features; a flat mechanism is set to determine the overlapping pixels and relative position relationship based on the local area, and the local target size is determined in combination with the acquisition distance, which can accurately splice multiple pre-collected images to obtain the target image. It solves the problem that the face features cannot be accurately obtained when the user's face is blocked, and then misjudgment occurs, reducing the accuracy of face recognition.

[0050] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0051] This application is based on time domain sparse points and frequency domain spatial textures to accurately extract feature information from images, comprehensively consider time domain and frequency domain information, accurately reflect feature distribution and texture structure in images, and achieve more precise local division. A multi-mode mechanism is set up to determine the boundary curve based on infrared imaging and depth difference, which can adapt to users with different occlusions, lighting conditions and facial features, and achieves the effect of accurately obtaining facial features for users with facial occlusion. The accuracy of face recognition is further improved.

[0052] Embodiment 2: Embodiment 1 achieves the effect of improving the accuracy of face recognition for users with obstructed faces. However, when the user moves quickly, the recognition accuracy for users with obstructed faces is reduced.

[0053] Based on the original boundary curve judgment mechanism, this implementation further introduces solutions such as spatiotemporal curve analysis, head posture estimation, and motion blur processing to improve the accuracy and robustness of face recognition in complex occlusion and motion scenes. By continuously collecting multiple pre-collected images, calculating the spatiotemporal curve of facial key points, outputting the user's head posture in real time, and judging the user's movement speed and change range based on the spatiotemporal curve, a clear face image is finally obtained for recognition.

[0054] In some embodiments, if the moving speed value is greater than a first threshold, the blur degree of the pre-captured image is calculated and the blur kernel is estimated, and a deconvolution operation and hierarchical fusion are performed to obtain a preliminary image.

[0055] Specifically, multiple consecutive pre-collected images are obtained, and the degree of dynamic blur in the image is analyzed by a blind deconvolution algorithm to estimate the blur kernel. The blind deconvolution algorithm is a technology that restores a blurred image by estimating the blur kernel and the original image without knowing the specific form of the blur kernel. It is mainly used in the field of image deblurring, especially when dealing with image blur problems caused by camera shake, motion blur, etc., it can play an important role. Specifically, an initial blur kernel estimate and an original image estimate are randomly generated, the current estimated blur kernel is fixed, and the original image is estimated, that is, a problem of minimizing an energy function is solved. The energy function includes a data fidelity term (that is, the difference between the estimated image and the observed image) and a regularization term (used to constrain the prior information such as the smoothness and sparsity of the image). According to the estimated values ​​of the original image and blur kernel obtained in each iteration, their estimated values ​​are updated. Repeat the above iterative optimization process until a certain stopping condition is met (such as reaching the maximum number of iterations, the energy function value no longer changes significantly, etc.). The final estimated value of the original image is the restored clear image. Preprocess the face image in fast motion, and analyze and determine the size and shape of the blur kernel. The estimated blur kernel is used to perform a deconvolution operation to restore the clear features of the image and obtain a preliminary clear image. Specifically, each of the multiple pre-acquired images input has undergone a different blurring process, so it is necessary to perform a deconvolution operation on each image separately. For each image, a blur kernel is estimated based on its own blur characteristics, and the blur kernel is used to perform a deconvolution operation to restore the clear image. Therefore, the number of restored images obtained in the end is the same as the number of blurred images input, that is, multiple preliminary clear images are obtained after the deconvolution operation.

[0056] The preliminary clear image is layered and fused to obtain a preliminary image, including: the layered fusion is based on short-term features and long-term features, extracting local feature points of a single frame image in the preliminary clear image, calculating local feature descriptors; obtaining local feature descriptors of multiple consecutive frames of images, obtaining aggregated long-term feature vectors; splicing and fusing the local feature descriptors and long-term feature vectors respectively, obtaining a fused feature vector as the preliminary image. Multiple preliminary clear images are fused to form a preliminary image.

[0057] In some embodiments, short-term feature extraction focuses on local occlusion-invariant features in a single-frame image. These features can maintain a certain stability when the face is partially occluded (such as wearing a mask or sunglasses), which helps to perform face recognition under occlusion. Typical local occlusion-invariant features include forehead texture, auricle shape, etc. Local feature points in a single-frame image are extracted using a local feature extraction algorithm (such as SIFT, SURF, etc.), and key points in the image (such as corner points, edge points, etc.) are detected. For each detected key point, its feature descriptor is calculated. The feature descriptor is a vector that describes the local image pattern around the key point. The local feature extraction algorithm is a commonly used technology, and the corresponding extraction algorithm can be selected according to the actual situation. This application does not make specific restrictions here.

[0058] In some embodiments, long-term feature aggregation aims to capture the dynamic changes of the user's face by aggregating multi-frame temporal features. This helps to maintain the stability of face recognition in dynamic scenes (such as rapid movement of the user and changes in expression). Collect multiple consecutive frames of images, extract the local feature descriptors of each frame, and build an LSTM network to aggregate the local feature descriptors of multiple frames of images. LSTM (Long Short-Term Memory) network is a special recurrent neural network that can process sequence data and capture long-term dependencies. The local feature descriptors of multiple consecutive frames of images are input into the LSTM network for training. During the training process, the LSTM network learns how to aggregate these feature descriptors to capture the dynamic changes of the user's face. After the training is completed, the local feature descriptors of the new multiple consecutive frames of images are input into the LSTM network to obtain the aggregated long-term feature vector.

[0059] Feature fusion and comparison is to fuse short-term features with long-term features and compare them with the feature vectors of pre-stored images to determine the degree of match. This helps to improve the robustness of face recognition in dynamic occlusion scenes. The short-term features and long-term features are concatenated or weighted to obtain the final fused feature vector, i.e., the preliminary image.

[0060] In this embodiment, the blur kernel is estimated by a blind deconvolution algorithm and a deconvolution operation is performed, which effectively restores the clear features of the face image in fast motion, obtains a preliminary clear image, and provides a clearer image basis for subsequent processing; hierarchical fusion is based on short-term features and long-term features, and short-term feature extraction focuses on local occlusion invariant features, and can still maintain stability when the face is partially occluded; long-term feature aggregation captures the dynamic changes of the user's face through the LSTM network, and maintains the stability of face recognition in dynamic scenes. The two are fused to obtain the final fusion feature vector, which further ensures that when the user's face is occluded and moves quickly, the effect of rapid recognition and improved recognition accuracy is achieved. The problem of reduced recognition accuracy caused by image blur when the user moves quickly is solved, and the face recognition accuracy in complex motion scenes is improved.

[0061] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0062] This application obtains a preliminary clear image by setting a deconvolution operation, maintains the stability of face recognition in dynamic scenes based on short-term features and long-term features, and obtains the final fused feature vector, which further ensures the effect of rapid recognition and improved recognition accuracy when the user's face is occluded and moves quickly; and further improves the face recognition accuracy in complex motion scenes.

[0063] Embodiment 3: In the above embodiment, the face recognition accuracy of users with face occlusion and fast movement is improved. However, in actual use, when the number of users increases, the recognition rate and accuracy are reduced. This embodiment makes further improvements on the basis of the above content.

[0064] The method further includes: S110: identifying the number of users, if the number of users is 1, executing S200; if the number of users is greater than 1, executing S120.

[0065] S120: Obtain preliminary users based on the splitting mechanism, perform facial segmentation on the preliminary users, identify facial feature points and compare them with pre-stored images, determine the target user, and obtain movement speed values ​​based on the space-time curves, and execute S200.

[0066] In some embodiments, the splitting mechanism includes: determining the deviation from the target path based on the space-time curve, marking the users corresponding to the deviation less than a preset third threshold as preliminary users; the target path refers to the straight path directly in front of the door lock.

[0067] Perform facial segmentation on the preliminary selected users, identify facial feature points and compare them with the pre-stored images to determine the target users, including: extracting facial regions and identifying facial feature points based on the pre-collected images corresponding to the preliminary selected users, and comparing the feature vectors of the facial feature points with the feature vectors of the pre-stored images to obtain the matching degree and then determine the target users. A matching degree threshold is pre-set, for example, to 70%, that is, if the matching degree is greater than 70%, it is determined to be a target user. When setting the matching degree threshold, the best setting range is between 50% and 80%, because the number of users is now preliminarily determined, which belongs to rough screening, and because the distance between the users is far, and the user's face may be blocked, so the matching degree threshold cannot be set too high. Specifically, the images of the preliminary selected users are pre-processed, such as denoising, contrast enhancement, etc., to improve the accuracy of facial segmentation. Use a facial segmentation algorithm (such as a segmentation network based on deep learning) to perform facial segmentation on the pre-processed image and extract the facial region. Use a facial feature point detection algorithm to identify feature points in the segmented facial region and extract feature points at key positions such as eyes, nose, and mouth. For the extracted facial feature points, their feature vectors are calculated, and the extracted feature vectors are compared with the feature vectors of the pre-stored images using a comparison algorithm (such as cosine similarity, Euclidean distance, etc.). The matching degree is calculated based on the comparison results. If the matching degree is greater than the preset threshold, it is determined that the match is successful and the primary selected user is the target user.

[0068] In some embodiments, the moving path of each user, such as direction, speed, etc., is analyzed based on the space-time curve, and then the moving path of each user in the field of view of the acquisition device is obtained, and the moving path of each person is compared with the preset target path to calculate the deviation. The deviation is used to measure the difference between the moving path of each user and the preset target path. In this embodiment, the deviation is defined as a comprehensive measure of the angle difference and distance difference between the moving path and the target path. Users corresponding to the deviation greater than a preset third threshold are marked as interfering users and eliminated, such as neighbors or passers-by.

[0069] The angle difference refers to the angle between the direction vector of the moving path and the direction vector of the target path. The direction vectors of the moving path and the target path are determined, and the cosine value of the angle between the two direction vectors is calculated using the vector dot product formula. The distance difference refers to the vertical distance from a point on the moving path to the target path. Specifically, a point on the moving path is determined, and the distance from the point to the target path is calculated using the point-to-straight line distance formula. The deviation is defined as the weighted sum of the angle difference and the distance difference. The weight coefficients are set in advance to adjust the relative importance of the angle difference and the distance difference in the comprehensive deviation.

[0070] Based on the above embodiments, this embodiment further introduces a user number identification and splitting mechanism to deal with the problem of face recognition unlocking of smart door locks in multi-user scenarios. By identifying the number of users, when the number of users is greater than 1, a preliminary user is obtained based on the splitting mechanism, and then the face of the preliminary user is segmented, the facial feature points are identified, and the comparison with the pre-stored image is performed to finally determine the target user and execute the subsequent unlocking process. This solves the problem that the target user cannot be quickly locked in the case of multiple users with facial occlusion and rapid movement, thereby reducing the efficiency and accuracy of door lock face recognition.

[0071] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0072] This application distinguishes between single-user and multi-user scenarios through user number identification, and sets up a splitting mechanism to screen preliminary users based on the deviation between the target path and the user's movement trajectory, and then performs facial segmentation, feature point identification and comparison on the preliminary users. It can accurately screen out target users from multiple users, and quickly lock the target user in the case of multiple users with facial occlusion and rapid movement, further improving the efficiency and accuracy of door lock face recognition.

[0073] Embodiment 4: In the above embodiment, user face recognition in a complex multi-user scenario is achieved, but the efficiency of user face recognition is affected. This embodiment makes further improvements based on the above content.

[0074] In some embodiments, the height of the user is estimated by using key points of the head, a three-dimensional mapping table of height-distance-gradient parameters is constructed, and the gradient calculation area is adjusted based on the target image to identify facial features, including: establishing a head posture model, using the three-dimensional head model to match the facial key points, establishing a mapping relationship between the head posture and the key point coordinates, estimating the height of the user based on the relationship between the user's head posture, the position of the head key points and the position of the ground; constructing a three-dimensional mapping table of height-distance-gradient parameters, and the estimated height of the target image to obtain the corresponding gradient parameters; determining the gradient calculation area to highlight the facial features based on the gradient parameters.

[0075] In some embodiments, a head posture model is established, and a three-dimensional head model is used to match the facial key points to establish a mapping relationship between the head posture and the key point coordinates. The pitch angle and yaw angle of the head are obtained by solving the rotation matrix or quaternion in the head posture model. The pitch angle reflects the degree of up and down tilt of the head, and the yaw angle reflects the degree of left and right rotation of the head. The calculated pitch angle and yaw angle are output in real time for subsequent face recognition and posture adjustment. Assume that the pitch angle of a user is 10 degrees and the yaw angle is -5 degrees calculated by the head posture model. The head posture output in real time is "pitch angle 10 degrees, yaw angle -5 degrees". The user's motion trajectory and motion amplitude are further determined according to the user's head posture and the space-time curve. At the same time, the user's height is estimated according to the relationship between the user's head posture, the position of the head key points and the ground. In the face recognition process, due to the user's movement and environmental influences, it is often impossible to accurately obtain the user's full-body image, so the height is estimated based on the face image.

[0076] Specifically, a large number of face images with different head postures are collected, and the three-dimensional coordinates of facial key points are marked. A pre-built three-dimensional head model is used to match the key points in the collected image with the model by minimizing the error between the key points to obtain the pitch angle and yaw angle of the head; the intrinsic parameters (such as focal length, optical center coordinates) and extrinsic parameters (the position and posture of the collection device in the world coordinate system) of the collection device are determined, and the key points of the head, such as the head vertex and chin point, are accurately located in the collected image. According to the imaging principle, the coordinates of the key points in the image are converted into actual space coordinates, and the user's height is estimated based on the actual space coordinates and pitch angle of the head key points.

[0077] In some embodiments, facial images of different heights (such as 1.5m, 1.6m, 1.7m, etc.) and corresponding full-body images are obtained at different distances (such as 0.5m, 1m, 1.5m, etc.). The gradient features of each collected image are calculated, and the gradient features can be calculated by common gradient operators, including the size and direction of the gradient. Key gradient parameters, such as the average gradient size, gradient direction distribution, etc., are extracted from the gradient features of each image. The variation law of the gradient parameters at different heights and different distances is analyzed. For example, as the distance increases, the gradient size may gradually decrease; the gradient distribution of facial images of different heights may be different at the same distance. The height, distance and extracted gradient parameters are used as three dimensions to construct a three-dimensional mapping table, and the height, distance and extracted gradient parameters are organized into a data table, and each collection point corresponds to a set of parameter values. For example, with the height as the x-axis, the distance as the y-axis, and the average gradient size as the z-axis, the parameter value corresponding to each collection point is filled in the mapping table. The mapping table can be interpolated using an interpolation algorithm so that the corresponding gradient parameters can be quickly obtained through the mapping table at any given height and distance. Height: affects the vertical distribution of the face in the image (the head of a tall user is close to the top of the image); distance: determines the resolution of the face (close-range images are rich in details, and long-range images are blurred); gradient parameters: include gradient sensitive regions (ROI), sampling density, and high and low frequency weights. For example, close distance (0.5m): height ≤ 160cm: focus on the lower half of the face (chin to nose tip), step length = 2 pixels, low frequency weight = 0.7; height ≥ 180cm: focus on the upper half of the face (forehead to eyes), step length = 1 pixel, high frequency weight = 0.6. Long distance (2m): uniformly focus on the eye-nose triangle, step length = 3 pixels, high frequency weight = 0.8 (compensate for loss of details). Using common gradient operators such as the Sobel operator and the Scharr operator to calculate the gradient features of an image is a known technical method, and this application will not go into details. Extract key parameters such as the average gradient size and gradient direction distribution from the gradient features.

[0078] During the face recognition process, for the current face image collected, firstly, the corresponding gradient parameters are obtained from the three-dimensional mapping table according to the user's height (obtained through dynamic estimation of the head posture model) and distance. The obtained gradient parameters are fused with the gradient features of the current image. For example, the average gradient size obtained from the mapping table can be used as a weight to adjust the gradient features of the current image, enhance the important gradient information in the image, and suppress noise and irrelevant information. This application does not make specific restrictions.

[0079] In some embodiments, by using gradient information to enhance the facial feature extraction effect, the extracted features are more distinguishable and can better characterize the facial features of different users. In the face recognition comparison stage, the enhanced feature vector is compared with the feature vector of the pre-stored face image to improve the accuracy of the match and reduce the false recognition rate and rejection rate. The gradient calculation range is limited by the mapping table to reduce the number of processed pixels. For example, only 30% of the image area is calculated in a long-distance scene, and the time consumption is reduced by 50%. The time for feature extraction is reduced, thereby improving the response speed of the entire door lock system, allowing users to complete the unlocking operation faster.

[0080] For example, in complex environments (such as lighting changes, facial occlusion, etc.), gradient information can provide additional feature information to help the system identify the user more accurately.

[0081] In some embodiments, in the face recognition scenario of smart door locks, the optimal gradient refers to the gradient calculation strategy that can most effectively characterize facial features (such as contours and textures) under specific height, distance and environmental conditions. Its core goals are: maximize feature differentiation: highlight the gradient information of key areas of the face (such as eye and nose contours); minimize noise interference: suppress invalid gradients caused by lighting changes, occlusions, etc.; balance computational efficiency: find the optimal solution between recognition accuracy and processing speed.

[0082] In some embodiments, image gradient is the rate of change of pixel grayscale value in the spatial direction. Areas with high gradient amplitude correspond to image edges (such as facial contours and boundaries of facial features), which are the core basis for feature extraction. The combination of high-frequency gradient (details) and low-frequency gradient (contours) can comprehensively describe the facial structure, and the gradient direction distribution can distinguish between real features and noise (such as light spots).

[0083] In some embodiments, facial images are collected under different heights (150-190cm), distances (0.5-3m), and lighting conditions; the sampling step size (1-5) and weight ratio (low frequency: high frequency = 0.2:0.8 to 0.8:0.2) are traversed. With recognition accuracy (%) and single frame processing time (ms) as indicators, analysis charts are drawn to select the optimal parameter combination. Parameters are dynamically fine-tuned according to the current recognition result (success / failure). For example, if there are three consecutive failures, the gradient threshold is automatically lowered. Preferably, the high-frequency weight is adjusted using light sensor data (low-frequency weight is increased in strong light, and high frequency is enhanced in weak light).

[0084] After acquiring the target image, the user's height is determined based on the distance between the head key points of the image and the ground. Based on the three-dimensional mapping of height-distance-gradient parameters, the optimal gradient calculation strategy is selected to accurately identify the user's facial features, and then compared with the pre-stored image, that is, the pre-stored face image, to ensure recognition accuracy and speed.

[0085] In some embodiments, the gradient features (such as edges and textures) of facial images show significant differences at different heights and distances. By constructing a mapping table, the optimal gradient calculation parameters for different scenarios can be pre-stored to avoid the complexity and resource consumption of real-time calculations. In a dynamic environment, the stability of the gradient feature directly affects the recognition accuracy. The mapping table provides adaptive parameter configuration through historical data analysis to enhance the system's adaptability to complex scenarios. For different user groups and interaction distances, key gradient areas are automatically selected to avoid feature extraction of invalid areas, significantly reducing the amount of calculation.

[0086] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0087] This application uses gradient information to enhance the facial feature extraction effect. The extracted features are more distinguishable and can better characterize the facial features of different users. The enhanced feature vector is compared with the feature vector of the pre-stored face image to improve the matching accuracy and reduce the false recognition rate and rejection rate. The gradient calculation range is limited by the mapping table, the number of processed pixels is reduced, and the door lock response speed is improved.

[0088] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For users skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A smart door lock face recognition unlocking method, characterized in that: include: S100: determining the number of users and corresponding space-time curves based on a plurality of pre-collected images; Get the moving speed value based on the time-space curve; S200: If the moving speed value is greater than the first threshold, the blur degree of the pre-collected image is calculated, and a deconvolution operation and layered fusion are performed to obtain a preliminary image; If the moving speed value is not greater than the first threshold, the pre-captured image is marked as a preliminary selected image; S300: obtaining time-frequency features of the preliminary image, and then determining multiple local areas to obtain interference values. If the interference value is greater than a second threshold, determining a clean area and a dirty area based on a multi-mode mechanism to generate a target image; If the interference value is not greater than the second threshold, the target image is obtained according to the flat state mechanism; S400: estimating the height of the user through key points of the head, constructing a three-dimensional mapping table of height-distance-gradient parameters, adjusting the gradient calculation area based on the target image to recognize facial features, and comparing with the pre-stored image; The moving speed value is used to measure the change range of the user's posture.

2. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The multi-mode mechanism is: determining a demarcation curve based on the difference between the infrared image data and the depth information of the pixel points, and obtaining a clean area and a dirty area based on the demarcation curve; Among them, infrared image data is acquired and image boundaries are identified, boundary points are randomly selected and the temperature gradient change rate is calculated successively, a temperature change threshold is set in advance, and an area where the temperature gradient change rate is greater than the temperature change threshold is identified as a boundary area; based on the depth information difference of pixel points in the boundary area, edge points are determined; and a boundary curve is obtained based on the edge points.

3. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: Based on the multi-mode mechanism, the clean area and the dirty area are obtained, and the target image is generated, including: The clean area is an image without occlusion; the dirty area is an image with occlusion; Obtain boundary pixel points of the clean area, extract visible light texture features, and generate a mixed feature vector by combining the temperature distribution features in the infrared thermal image; The contour boundary and temperature distribution characteristics of the dirty area are obtained, and the mixed feature vector of the clean area is combined for prediction and completion to obtain the target image.

4. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The time-frequency features include time-domain sparse points and frequency-domain spatial textures; the time-domain sparse points are: extracting pixel feature information of each pre-collected image, monitoring the grayscale value change of each pixel in real time, and based on a preset grayscale change threshold, marking the pixel points greater than the grayscale change threshold as time-domain sparse points; Record the spatiotemporal coordinates of each time-domain sparse point, including timestamp and pixel position; The frequency domain spatial texture is: obtaining a pre-collected image, performing wavelet transform, decomposing frequency domain components at different scales, extracting low-frequency contours and high-frequency detail features from the frequency domain components, and forming a frequency domain spatial texture.

5. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The flat state mechanism is: determining overlapping pixel points and relative position relationships according to a number of local areas, and splicing the plurality of pre-collected images according to the relative position relationships to obtain a target image; The overlapping pixel points of the pre-captured images are obtained, and the relative position relationship of the multiple pre-captured images on the face of the user to be identified is predicted based on the overlapping pixel points.

6. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The moving speed value is obtained based on the time-space curve, including: Based on the spatiotemporal curve, the displacement characteristics, average speed value and change amplitude quantization value of the facial key points between consecutive frames in the pre-collected image are calculated, and the average speed value and the change amplitude quantization value of all key points are weighted summed to obtain the moving speed value; among them, for each key point, the average speed value between consecutive frames is calculated; the standard deviation of the speed value of each key point is calculated as the change amplitude quantization value, which is used to reflect the stability of the key point speed; the calculation formula of the moving speed value is as follows: ; Among them, M is the moving speed value, is the weight coefficient of the i-th key point, and its value range is [0, 1]; is the average velocity value of the i-th key point; is the quantitative value of the change amplitude of the i-th key point; n is the total number of key points; The average speed value is obtained by dividing the modulus of the displacement vector by the time interval; the displacement features include: key point coordinates, displacement direction, speed value and time interval.

7. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: Calculate the blur degree of the pre-collected image and estimate the blur kernel, perform deconvolution and layered fusion to obtain the preliminary image, including: The estimated blur kernel is used to perform deconvolution operation to restore the clear features of the image and obtain a preliminary clear image; Perform layered fusion on the preliminary clear images to obtain the preliminary selected images; Among them, the hierarchical fusion is based on the fusion of short-term features and long-term features, extracting local feature points of a single frame image in a preliminary clear image, calculating local feature descriptors; obtaining local feature descriptors of multiple consecutive frames of images, and obtaining aggregated long-term feature vectors; splicing and fusing the local feature descriptors and long-term feature vectors respectively, and obtaining a fused feature vector as a preliminary image.

8. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The method further includes: S110: identifying the number of users, if the number of users is 1, executing S200; if the number of users is greater than 1, executing S120; S120: obtaining a preliminary user based on the splitting mechanism, performing facial segmentation on the preliminary user, identifying facial feature points and comparing them with pre-stored images, determining a target user, and obtaining a moving speed value based on the time-space curve, and executing S200; Among them, the splitting mechanism includes: determining the deviation from the target path based on the space-time curve, marking the users corresponding to the deviation less than a preset third threshold as preliminary users; the target path refers to the straight path directly in front of the door lock; the deviation is used to measure the difference between each user's moving path and the preset target path.

9. The face recognition unlocking method for smart door locks according to claim 8, characterized in that: Based on the pre-collected images corresponding to the primary selected users, the facial area is extracted to identify the facial feature points, and the feature vectors based on the facial feature points are compared with the feature vectors of the pre-stored images to obtain the matching degree and thus determine the target user.

10. The smart door lock face recognition unlocking method according to claim 1, characterized in that: The user's height is estimated through the key points of the head, a three-dimensional mapping table of height-distance-gradient parameters is constructed, and the gradient calculation area is adjusted based on the target image to recognize the facial features, including: establishing a head posture model, using the three-dimensional head model to match the facial key points, establishing a mapping relationship between the head posture and the key point coordinates, and estimating the user's height based on the user's head posture, the position of the head key points and the position of the ground; constructing a three-dimensional mapping table of height-distance-gradient parameters, and the estimated height of the target image to obtain the corresponding gradient parameters; determining the gradient calculation area to highlight the facial features based on the gradient parameters.

Citation Information

Patent Citations

  • Smart door lock facial recognition unlocking methods, devices, media, and smart door locks

    CN115938023B

  • Method and device for training face recognition model, equipment and storage medium

    CN117612223A

  • Face recognition method and system based on AI intelligence

    CN118430054A

  • Access control system based on face recognition and recognition method

    CN119672782A

  • Door lock control method, system and intelligent door lock

    CN119741775A

Cited By

  • Digital image enhancement method and system based on endoscope

    CN120339112A

  • Bridge dynamic displacement high-precision identification method based on digital image processing

    CN120833573A

  • Multi-level permission access control identification method and system based on biometric feature weight distribution

    CN122598282A