An intelligent door lock face recognition unlocking method
Through the stitching and space-time curve analysis of multiple pre-acquisition images in intelligent door lock face recognition technology, combined with infrared imaging and boundary curve recognition of depth differences, the problem of low facial recognition accuracy under different heights and facial occlusion is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510479916.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing smart door lock face recognition technology cannot accurately obtain the facial features of users of different heights and facial obstruction users, resulting in a reduced recognition accuracy.
Through the stitching and spatiotemporal curve analysis of multiple pre-acquisition images, the target image is generated, and the boundary curve is determined based on infrared imaging and depth differences, which can adapt to different occlusions and lighting conditions, and improve the accuracy of facial features acquisition.
Accurate face recognition under different heights and facial occlusion is achieved, improving the accuracy and robustness of recognition.
Smart Images

Figure CN120014740B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent door lock control, and particularly to a face recognition unlocking method for an intelligent door lock. Background Art
[0002] The intelligent door lock with face recognition unlocking uses a camera to collect face images, and compares the face images with the pre-stored face images for feature comparison, so as to identify and verify the user's identity. It can effectively reduce the time for the user to open the door, and at the same time, the user does not need to carry the key with them, improving the convenience of opening the door. However, by collecting face images through a camera with a fixed angle, the viewing angle range is fixed, and for users to be identified with different heights, a complete face image cannot be collected, resulting in misidentification and failure to successfully open the door lock.
[0003] Chinese Patent Application No. 202310247026.5 provides a face recognition unlocking method, device, medium and intelligent door lock for an intelligent door lock. By collecting multiple pre-captured images from different capture perspectives, according to the capture distance between the face of the user to be identified and the image capture device, the target size of the mask is determined from the preset mask sizes, the mask is placed on the pre-captured image with the target size, and the mask is moved on the pre-captured image to obtain the corresponding pixel point feature information; according to the pixel point feature information, the overlapping pixel points of each pre-captured image are determined, and the relative position relationship of the multiple pre-captured images is predicted according to the overlapping pixel points; the multiple pre-captured images are stitched together according to the relative position relationship to obtain a target image.
[0004] However, in the prior art, the method of using a mask to select face features from multiple angles and then stitching them is only limited to static single-user scenarios. In the case of scenarios with facial occlusion or multi-user dynamics, the face features cannot be accurately obtained, resulting in misjudgment and reducing the accuracy of face recognition. Summary of the Invention
[0005] This application provides a face recognition unlocking method for an intelligent door lock, which solves the problem in the prior art that face features cannot be accurately obtained, thereby reducing the accuracy of face recognition, and realizes the technical effect of accurately obtaining face features and improving the accuracy of face recognition.
[0006] This application provides a face recognition unlocking method for an intelligent door lock, including:
[0007] S100: Based on multiple pre-captured images, determine the number of users and the corresponding spatio-temporal curve; obtain the moving speed value based on the spatio-temporal curve;
[0008] S200: If the moving speed value is greater than the first threshold, calculate the blur degree of the pre-captured image, perform deconvolution operation and hierarchical fusion to obtain a primary selected image;
[0009] If the movement speed value is not greater than the first threshold, mark the pre-captured image as the preliminary selected image;
[0010] S300: Obtain the time-frequency characteristics of the preliminary selected image, and then determine the interference values of multiple local areas. If the interference value is greater than the second threshold, determine the clean area and the dirty area based on the multi-mode mechanism, and generate the target image; if the interference value is not greater than the second threshold, obtain the target image according to the flat state mechanism;
[0011] S400: Estimate the user's height through the head key points, construct a three-dimensional mapping table of height-distance-gradient parameters, adjust the gradient calculation area based on the target image to identify the facial features, and compare with the pre-stored image;
[0012] Among them, the movement speed value is used to measure the change range of the user's posture.
[0013] Furthermore, the method further includes: S110: Identify the number of users. If the number of users is 1, execute S200; if the number of users is greater than 1, execute S120;
[0014] S120: Obtain the preliminary selected users based on the splitting mechanism, perform face segmentation on the preliminary selected users, identify the facial feature points and compare with the pre-stored image to determine the target user, and respectively obtain the movement speed value based on the spatio-temporal curve, and execute S200;
[0015] Among them, the splitting mechanism includes: determining the deviation degree from the target path based on the spatio-temporal curve, and marking the users corresponding to the deviation degree less than the preset third threshold as the preliminary selected users; the target path refers to the straight line path directly in front of the door lock; the deviation degree is used to measure the difference between the movement path of each user and the preset target path.
[0016] Furthermore, estimating the user's height through the head key points, constructing a three-dimensional mapping table of height-distance-gradient parameters, and adjusting the gradient calculation area based on the target image to identify the facial features includes: establishing a head pose model, using a three-dimensional head model to match with the facial key points, establishing a mapping relationship between the head pose and the key point coordinates, and estimating the user's height according to the relationship between the user's head pose, the position of the head key points and the position of the ground; constructing a three-dimensional mapping table of height-distance-gradient parameters and the estimated height of the target image to obtain the corresponding gradient parameters; determining the gradient calculation area based on the gradient parameters to highlight the facial features.
[0017] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0018] By accurately extracting the feature information in the image based on the time-domain sparse points and the frequency-domain spatial texture, a more accurate local division effect is achieved; a multi-mode mechanism is set to determine the boundary curve based on infrared imaging and depth difference, which can adapt to users with different occlusions, lighting conditions, and facial features, and achieves the effect of accurately obtaining the facial features of users with facial occlusions. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. is a schematic flow chart of a face recognition and unlocking method for an intelligent door lock in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant drawings; the drawings show preferred embodiments of the present invention, however, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs; the terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0022] Embodiment 1: As Figure 1 shown, a face recognition and unlocking method for an intelligent door lock includes:
[0023] S100: Based on multiple pre-acquired images, determine the number of users and the corresponding spatio-temporal curve; obtain the moving speed value based on the spatio-temporal curve.
[0024] S200: If the moving speed value is greater than the first threshold, calculate the blur degree of the pre-acquired image and estimate the blur kernel, perform deconvolution operation and hierarchical fusion to obtain a preliminary selected image; if the moving speed value is not greater than the first threshold, mark the pre-acquired image as the preliminary selected image.
[0025] In some embodiments, in response to an unlocking operation of the intelligent door lock, the present application collects multiple pre-captured images at different acquisition perspectives and different distances through an image acquisition device configured on the intelligent door lock. Among them, the multiple pre-captured images are obtained by the image acquisition device adjusting the acquisition perspective multiple times according to a preset frequency and a preset angle; the image acquisition device can be one or more cameras, and also includes an infrared camera. Each time the acquisition perspective is adjusted, one or more pre-captured images can be collected. After collecting multiple pre-captured images in one adjustment of the acquisition perspective, the one with the highest clarity can be used as the pre-captured image after this adjustment of the acquisition perspective to ensure that there is only one pre-captured image for each acquisition perspective. The pre-captured images are the collected user images, including face images and body part images.
[0026] In some embodiments, the first threshold is pre-set based on historical experimental data. First, a large number of face images with different moving speeds are obtained, and then recognition is performed to find the critical point of successful recognition. The moving speed corresponding to the critical point is the corresponding optimal first threshold. The first threshold can also be dynamically adjusted according to actual situations to meet different experimental requirements.
[0027] S300: Obtain the time-frequency characteristics of the primary selected image, determine multiple regions based on the time-frequency characteristics and calculate the interference value. If the interference value is greater than the second threshold, determine the clean region and the dirty region based on the multi-mode mechanism and generate the target image; if the interference value is not greater than the second threshold, obtain the target image according to the flat state mechanism.
[0028] In this embodiment, based on the obtained multiple pre-captured images, pixel point analysis is performed to obtain time-frequency characteristics and spatio-temporal curves; several regions are determined based on the time-frequency characteristics. The region is the key point on the pre-captured image determined based on the change of pixel points and texture characteristics. The face key points are constructed according to facial organs. The region includes the vector values of each pixel point, the number of face key points corresponding to the set of local pixel points, as well as the vector values of each pixel point in the set of pixel points to be determined, the number of face key points corresponding to the set of pixel points to be determined, and recalculate the feature vector values of multiple pixel points in the set of local pixel points and the feature vector values of multiple pixel points in the set of pixel points to be determined until there is no set of local pixel points with the same feature vector value in each pre-captured image. In this way, it is avoided that there is a set of local pixel points with the same feature vector value, and it is avoided that when calculating overlapping pixel points, there is the same set of local pixel points, resulting in the failure of determining overlapping pixel points.
[0029] In some embodiments, a facial key-point detection algorithm is used to detect key points in multiple pre-captured images collected continuously, extract the facial key-point coordinates in each image, including key positions such as eyes, nose, mouth, etc., and obtain the displacement features of each key point based on the spatio-temporal curve. The displacement features include: key-point coordinates, displacement direction, speed value, and time interval. Specifically, for the same key point between consecutive frames, calculate its displacement vector (i.e., coordinate difference). According to the displacement vector and the time interval of image acquisition, calculate the speed value of the key point (displacement magnitude divided by the time interval). Plot the change of the displacement direction and speed value of each key point over time to form a spatio-temporal curve. The spatio-temporal curve reflects the movement trajectory and speed change of facial key points in three-dimensional space.
[0030] The spatio-temporal curve is used to reflect the user's movement trajectory and speed change. The average speed value is obtained by dividing the magnitude of the displacement vector by the time interval. The displacement features include: key-point coordinates, displacement direction, speed value, and time interval.
[0031] Calculate the displacement features of facial key points between consecutive frames in the pre-captured images based on the spatio-temporal curve. For each key point, calculate the average speed value between consecutive frames, calculate the standard deviation of the speed values of each key point as the change amplitude quantization value, which is used to reflect the stability of the key point speed; perform a weighted sum of the average speed values and the change amplitude quantization values of all key points to obtain the movement speed value. The calculation formula is as follows:
[0032] ;
[0033] where M is the movement speed value, is the weight coefficient of the i-th key point, and the value range is [0, 1]; is the average speed value of the i-th key point; is the change amplitude quantization value of the i-th key point.
[0034] In some embodiments, the time-frequency features include time-domain sparse points and frequency-domain spatial textures. The time-domain sparse points are: extract the pixel point feature information of each pre-captured image, monitor the gray value change of each pixel point in real time, and based on a pre-set gray value change threshold, mark the pixel points with gray values greater than the gray value change threshold as; record the spatio-temporal coordinates of each time-domain sparse point, including the time stamp and the pixel point position;
[0035] The frequency-domain spatial texture is as follows: Obtain a pre-acquired image, perform wavelet transform, decompose the frequency-domain components at different scales, extract the low-frequency contour and high-frequency detail features from the frequency-domain components, and form the frequency-domain spatial texture. Wavelet transform is a new transform analysis method that performs multi-scale refinement analysis on an image through operations such as dilation and translation; by convolving a wavelet function with adjustable scale and variable position with the signal, the components of the signal at different scales and positions can be obtained. These wavelet components can reflect the local characteristics of the signal at different frequencies and times, and this application will not elaborate too much here.
[0036] In some embodiments, the time-domain sparse points are determined by extracting the pixel feature information of each pre-acquired image and monitoring the change in the gray value of each pixel in real time. When the change in the gray value of a certain pixel is greater than a pre-set gray value change threshold, that pixel is marked as a time-domain sparse point.
[0037] Time-domain sparse points can capture the pixels with significant gray value changes in the image, and these points often correspond to the key features or dynamic change regions in the image. By recording the spatio-temporal coordinates (including the timestamp and pixel position) of each time-domain sparse point, precise tracking and positioning of the dynamic image can be achieved. Since only the pixels with significant gray value changes are concerned, the time-domain sparse point method can greatly reduce the computational amount in subsequent processing and improve the processing efficiency. Through time-domain sparse points, the dynamic features in the image can be more effectively extracted, providing an accurate data basis for subsequent local determination and interference value calculation. The spatio-temporal positioning ability of time-domain sparse points helps to establish a stable feature correspondence relationship between consecutive frames and improve the robustness of image processing.
[0038] The frequency-domain spatial texture is obtained by acquiring a pre-acquired image and performing wavelet transform. Wavelet transform can decompose the frequency-domain components at different scales, and the low-frequency contour and high-frequency detail features can be extracted from these components, thereby forming the frequency-domain spatial texture. By extracting the low-frequency contour and high-frequency detail features, the frequency-domain spatial texture can enhance the useful information in the image, suppress noise and irrelevant details, provide an effective description method for the texture features in the image, and help with texture analysis and recognition in subsequent processing. It makes the local determination more accurate and comprehensive, and can capture the feature information at different scales; it can more effectively describe the texture features in the image and provide strong support for interference value calculation and image quality assessment.
[0039] Specifically, initialize the acquisition device, continuously acquire multiple face images, extract time-domain sparse points and frequency-domain spatial textures, determine the local area and calculate the interference value. Determine whether the acquired face image has occlusion based on the interference value. A second threshold is preset. The second threshold is preset according to historical experimental data and is used to measure the degree of interference factors existing on the face image. Specifically, a large number of face images with different occlusion degrees can be obtained, and the critical point corresponding to the occlusion degree can be found according to the recognition success rate, and the occlusion degree corresponding to the critical point is set as the second threshold. Specifically, it can be adjusted up and down at the critical point according to the actual situation; if the interference value is greater than the second threshold, it indicates that there is an occlusion situation, such as sunglasses, masks, hats, etc. For the occlusion situation, judge the demarcation curve based on the multi-modal mechanism.
[0040] In some embodiments, the multi-modal mechanism is as follows: determine the demarcation curve based on the difference between the infrared image data and the depth information of the pixel points, and obtain the clean area and the dirty area based on the demarcation curve; among them, acquire the infrared image data and identify the image boundary, randomly select boundary points and calculate the temperature gradient change rate successively. Preset the temperature change threshold, and identify the area where the temperature gradient change rate is greater than the temperature change threshold as the demarcation area; based on the difference in the depth information of the pixel points in the demarcation area, determine the edge points; obtain the demarcation curve based on the edge points. Specifically, use an infrared camera to collect the thermal imaging image of the user's face, ensure that the acquisition environment is stable, and avoid interference from external heat sources. Denoise the collected infrared image, such as using Gaussian filtering to remove the noise points in the image, perform image enhancement, and improve the image contrast for subsequent temperature gradient analysis. In the preprocessed infrared image, randomly select multiple points on the face boundary as the initial analysis points. For each selected boundary point, calculate the temperature gradient change rate along the normal direction. The temperature gradient change rate can be obtained by calculating the ratio of the temperature difference between adjacent pixel points to the distance. Preset the temperature change threshold, which is determined based on experimental data or empirical values, and identify the area where the temperature gradient change rate is greater than the temperature change threshold. These areas are regarded as the demarcation areas.
[0041] Collect the depth data of the user's face. In the identified demarcation area, analyze the difference in the depth information of the pixel points. The difference in the depth information can be obtained by calculating the difference value of the depth values of adjacent pixel points. Based on the difference in the depth information, determine the edge points in the demarcation area. The edge points are usually located in the area where the depth changes significantly, such as the junction between the occluder and the face. Connect the identified edge points to form an initial demarcation curve. Preferably, use a curve fitting algorithm (such as B-spline curve fitting) to optimize the initial demarcation curve to make it smoother and conform to the actual boundary shape; consider the prior knowledge of the face contour and further adjust the demarcation curve to improve its accuracy.
[0042] In some embodiments, the clean region refers to a clean and unobstructed face image; the dirty region refers to a face image with an obstruction; and the demarcation curve refers to the boundary line between the clean region and the dirty region.
[0043] In some embodiments, obtaining the clean region and the dirty region based on a multi-modal mechanism and generating a target image includes: acquiring boundary pixel points of the clean region, extracting visible light texture features, and generating a mixed feature vector by combining the temperature distribution features in the infrared thermal image; acquiring the contour boundary and temperature distribution features of the dirty region, and performing prediction and completion in combination with the mixed feature vector of the clean region to obtain the target image.
[0044] In this embodiment, based on infrared imaging and depth difference, the accuracy and robustness of demarcation curve recognition are improved, which can adapt to users with different occlusions, different lighting conditions, and different facial features, has wide applicability, and meets application scenarios with high real-time requirements such as intelligent door locks. The problem that the user's face image cannot be accurately recognized when it is occluded is solved.
[0045] In some embodiments, if the interference value is not greater than the second threshold, it indicates that the pre-acquired image obtained is a clear and stable face image, and then the target image is obtained according to the flat state mechanism; the flat state mechanism is: determining overlapping pixel points and relative position relationships according to several local regions, and splicing the multiple pre-acquired images according to the relative position relationships to obtain the target image.
[0046] Specifically, determining the overlapping pixel points of each of the pre-acquired images according to several local regions, and predicting the relative position relationships of the multiple pre-acquired images on the face of the user to be recognized based on the overlapping pixel points. It can be understood that if the pixel point feature information in different pre-acquired images is the same, they are overlapping pixel points. The pixel point feature information can be the vector value of each pixel point. If the vector values of multiple pixel points under the local coverage in different pre-acquired images are the same at each position, the pixel point feature information in the pre-acquired images is the same. According to the acquisition distance between the face of the user to be recognized and the image acquisition device, determining the target size of the local region from the preset local region sizes, where there is a one-to-one correspondence between the preset local region sizes and the acquisition distance. Based on the extracted pixel point feature information of each of the pre-acquired images with the target size, determining several local regions according to time-frequency features, and moving the local regions on each of the pre-acquired images with the same step size; the one-to-one correspondence between the local region size and the acquisition distance can be made into a table, and by looking up the table, according to the acquisition distance between the face of the user to be recognized and the image acquisition device, determining the target size of the local region from the time domain features; splicing the multiple pre-acquired images according to the relative position relationships to obtain the target image.
[0047] S400: Compare the target image with the pre-stored image; obtain the target image and compare it with the pre-stored image to get a comparison result; if the comparison result indicates a successful match (i.e., the similarity is greater than the preset similarity threshold), then unlock the smart door lock; otherwise, conduct a re-comparison or issue a warning. The similarity threshold needs to be set according to the actual situation. For the sake of security, the best setting range is between 80% and 100%.
[0048] In this embodiment, based on time-domain sparse points and frequency-domain spatial texture, multiple local regions are directly clustered and divided in the image. By comprehensively considering the information in the time domain and the frequency domain, it can more accurately reflect the feature distribution and texture structure in the image. Combining the texture information of the image, it can more accurately extract the feature information in the image, making the division of local regions more accurate and reasonable; avoiding the problem that it is impossible to accurately identify facial features when directly dividing regions to obtain features.
[0049] In this embodiment, a multi-mode mechanism is set to determine the boundary curve based on infrared imaging and depth difference, improving the accuracy and robustness of boundary curve recognition, solving the problem that accurate recognition cannot be performed when the user's face image is blocked, and being able to adapt to users with different occluders, lighting conditions, and facial features; setting a flat state mechanism to determine the overlapping pixel points and relative position relationships according to the local region, and combining the acquisition distance to determine the size of the local region target, can accurately splice multiple pre-acquired images to obtain the target image. Solving the problem that it is impossible to accurately obtain facial features when the user's face is blocked, resulting in misjudgment and reducing the accuracy of face recognition.
[0050] The technical solutions in the above embodiments of the present application at least have the following technical effects or advantages:
[0051] Based on time-domain sparse points and frequency-domain spatial texture, the present application accurately extracts the feature information in the image, comprehensively considers the information in the time domain and the frequency domain, accurately reflects the feature distribution and texture structure in the image, and realizes the effect of more accurate local region division; setting a multi-mode mechanism to determine the boundary curve based on infrared imaging and depth difference can adapt to users with different occluders, lighting conditions, and facial features, and realizes the effect of accurately obtaining the facial features of users with facial occlusion; further improving the accuracy of face recognition.
[0052] Embodiment 2: In Embodiment 1, the effect of improving the accuracy of face recognition for users with facial occlusion is achieved. However, when the user moves quickly, the recognition accuracy for users with facial occlusion is reduced.
[0053] Based on the original boundary curve judgment mechanism, this implementation further introduces solutions such as spatio-temporal curve analysis, head pose estimation, and motion blur processing to improve the accuracy and robustness of face recognition in complex occlusion and motion scenarios. By continuously collecting multiple pre-captured images, calculating the spatio-temporal curves of facial key points, real-time outputting the user's head pose, and judging the user's moving speed and change amplitude based on the spatio-temporal curves, a clear face image is finally obtained for recognition.
[0054] In some embodiments, if the moving speed value is greater than the first threshold, calculate the blur degree of the pre-captured image and estimate the blur kernel, perform deconvolution operation and hierarchical fusion to obtain a preliminary selected image.
[0055] Specifically, obtain multiple consecutive pre-captured images, use the blind deconvolution algorithm to analyze the dynamic blur degree in the images and estimate the blur kernel; the blind deconvolution algorithm is a technique that restores a blurred image by estimating the blur kernel and the original image without knowing the specific form of the blur kernel. It is mainly applied in the field of image deblurring. Especially when dealing with image blurring problems caused by camera shake, motion blur, etc., it can play an important role. Specifically, randomly generate an initial estimated value of the blur kernel and the original image. Fix the currently estimated blur kernel and estimate the original image, that is, solve a problem of minimizing an energy function. The energy function includes a data fidelity term (i.e., the difference between the estimated image and the observed image) and a regularization term (used to constrain prior information such as the smoothness and sparsity of the image). According to the estimated values of the original image and the blur kernel obtained in each iteration, update their estimated values. Repeat the above iterative optimization process until a certain stopping condition is met (such as reaching the maximum number of iterations, the value of the energy function no longer changes significantly, etc.). The finally obtained estimated value of the original image is the restored clear image. Preprocess the face image in fast motion, analyze and determine the size and shape of the blur kernel. Use the estimated blur kernel to perform deconvolution operation to restore the clear features of the image and obtain a preliminary clear image. Specifically, for the multiple input pre-captured images, each pre-captured image has experienced a different blur process, so the deconvolution operation needs to be performed on each image separately. For each image, a blur kernel is estimated according to its own blur characteristics, and the deconvolution operation is performed using this blur kernel to restore the clear image. Therefore, the number of restored images finally obtained is the same as the number of input blurred images, that is, multiple preliminary clear images are obtained after the deconvolution operation.
[0056] Perform hierarchical fusion on the initially clear images to obtain the preliminary selection images, including: the hierarchical fusion is based on short-term features and long-term features for fusion, extract the local feature points of a single-frame image in the initially clear image, and calculate the local feature descriptors; obtain the local feature descriptors of consecutive multiple frames of images to get the aggregated long-term feature vectors; respectively splice and fuse the local feature descriptors and the long-term feature vectors to obtain the fused feature vectors as the preliminary selection images. Fuse multiple initially clear images to form one preliminary selection image.
[0057] In some embodiments, short-term feature extraction mainly focuses on local occlusion-invariant features in a single-frame image. These features can still maintain a certain degree of stability when there is partial occlusion of the face (such as wearing a mask, sunglasses), which helps in face recognition in occlusion situations. Typical local occlusion-invariant features include forehead texture, auricle shape, etc. Use local feature extraction algorithms (such as SIFT, SURF, etc.) to extract the local feature points in a single-frame image, detect the key points (such as corner points, edge points, etc.) in the image, and for each detected key point, calculate its feature descriptor. The feature descriptor is a vector used to describe the local image pattern around the key point. Local feature extraction algorithms are commonly used techniques, and the corresponding extraction algorithm can be selected according to the actual situation. This application does not make specific limitations here.
[0058] In some embodiments, long-term feature aggregation aims to capture the dynamic change rules of the user's face by aggregating multi-frame temporal features. This helps to maintain the stability of face recognition in dynamic scenarios (such as the user moving quickly, changing expressions). Collect consecutive multiple frames of images, and extract the local feature descriptors of each frame of image, and construct an LSTM network for aggregating the local feature descriptors of multi-frame images. The LSTM (Long Short-Term Memory) network is a special recurrent neural network that can process sequence data and capture the long-term dependencies therein. Input the local feature descriptors of consecutive multiple frames of images into the LSTM network for training. During the training process, the LSTM network will learn how to aggregate these feature descriptors to capture the dynamic change rules of the user's face. After training, input the local feature descriptors of the new consecutive multiple frames of images into the LSTM network to obtain the aggregated long-term feature vectors.
[0059] Feature fusion and comparison is to fuse short-term features and long-term features and compare them with the feature vectors of the pre-stored images to judge the matching degree. This helps to improve the robustness of face recognition in dynamic occlusion scenarios. Splice or weight-fuse the short-term features and long-term features to obtain the final fused feature vectors, that is, the preliminary selection images.
[0060] In this embodiment, the blur kernel is estimated through a blind deconvolution algorithm and deconvolution operation is performed, effectively restoring the clear features of the face image during fast movement, obtaining a preliminary clear image, and providing a clearer image basis for subsequent processing. The hierarchical fusion is based on short-term features and long-term features. The short-term feature extraction focuses on local occlusion-invariant features and can maintain stability even when the face is partially occluded. The long-term feature aggregation captures the dynamic change law of the user's face through an LSTM network and maintains the stability of face recognition in a dynamic scenario. The fusion of the two obtains the final fused feature vector, further ensuring the effect of fast recognition and improved recognition accuracy when the user's face is occluded and moving quickly. It solves the problem of reduced recognition accuracy caused by image blur when the user is moving quickly and improves the face recognition accuracy in complex motion scenarios.
[0061] The technical solutions in the above embodiments of the present application at least have the following technical effects or advantages:
[0062] The present application obtains a preliminary clear image by setting deconvolution operation, maintains the stability of face recognition in a dynamic scenario based on short-term features and long-term features, obtains the final fused feature vector, and further ensures the effect of fast recognition and improved recognition accuracy when the user's face is occluded and moving quickly; further improves the face recognition accuracy in complex motion scenarios.
[0063] Embodiment 3: In the above embodiment, the face recognition accuracy of users with occluded faces and moving quickly is improved. However, in actual use, when the number of users increases, the recognition rate and accuracy are reduced. This embodiment makes further improvements on the basis of the above content.
[0064] The method further includes: S110: Identify the number of users. If the number of users is 1, execute S200; if the number of users is greater than 1, execute S120.
[0065] S120: Obtain preliminary selected users based on the splitting mechanism, perform face segmentation on the preliminary selected users, identify facial feature points and compare them with pre-stored images to determine the target users, and respectively obtain the moving speed values based on the spatio-temporal curves, and then execute S200.
[0066] In some embodiments, the splitting mechanism includes: determining the deviation degree from the target path based on the spatio-temporal curve, and marking the users with a deviation degree less than a pre-set third threshold as preliminary selected users; the target path refers to the straight path directly in front of the door lock.
[0067] Perform face segmentation on the initially selected users, identify facial feature points, and compare them with pre-stored images to determine the target users, including: based on the pre-acquired images corresponding to the initially selected users, extract the facial regions to identify facial feature points, and compare the feature vectors of the facial feature points with the feature vectors of the pre-stored images to obtain the matching degree and then determine the target users. Set a matching degree threshold in advance, for example, set it to 70%, that is, if the matching degree is greater than 70%, it is determined as the target user. When setting the matching degree threshold, the best setting range is between 50% - 80% because currently it is a preliminary judgment of the number of users, belonging to a rough screening, and due to the relatively long distance of the users and the possible occlusion of the users' faces, the matching degree threshold cannot be set too high. Specifically, preprocess the images of the initially selected users, such as denoising, enhancing contrast, etc., to improve the accuracy of face segmentation. Use a face segmentation algorithm (such as a segmentation network based on deep learning) to perform face segmentation on the preprocessed images and extract the facial regions. Use a facial feature point detection algorithm to identify the feature points in the segmented facial regions and extract the feature points at key positions such as eyes, nose, mouth, etc. For the extracted facial feature points, calculate their feature vectors, and use a comparison algorithm (such as cosine similarity, Euclidean distance, etc.) to compare the extracted feature vectors with the feature vectors of the pre-stored images. Calculate the matching degree according to the comparison result. If the matching degree is greater than the preset threshold, it is determined that the match is successful, and the initially selected user is the target user.
[0068] In some embodiments, analyze the movement paths of each user based on spatio-temporal curves, such as direction, speed, etc., and then obtain the movement paths of each user within the field of view of the acquisition device. Compare the movement path of each person with the preset target path and calculate the deviation degree. The deviation degree is used to measure the difference between the movement path of each user and the preset target path. In this implementation scheme, the deviation degree is defined as a comprehensive measure of the angle difference and distance difference between the movement path and the target path. Mark the users corresponding to the deviation degree greater than the preset third threshold as interfering users and eliminate them, such as neighbors or passers-by.
[0069] The angle difference refers to the included angle between the direction vector of the movement path and the direction vector of the target path. Determine the direction vectors of the movement path and the target path, and use the vector dot product formula to calculate the cosine value of the included angle between the two direction vectors; the distance difference refers to the perpendicular distance from a certain point on the movement path to the target path, specifically: determine a point on the movement path and use the point-to-line distance formula to calculate the distance from the point to the target path; define the deviation degree as the weighted sum of the angle difference and the distance difference, and set the weight coefficients in advance to adjust the relative importance of the angle difference and the distance difference in the comprehensive deviation degree.
[0070] Based on the above embodiment, this embodiment further introduces a user quantity recognition and splitting mechanism to address the issue of face recognition and unlocking of intelligent door locks in multi-user scenarios. By recognizing the number of users, when the number of users is greater than 1, the primary selected users are obtained based on the splitting mechanism. Then, face segmentation, facial feature point recognition, and comparison with pre-stored images are performed on the primary selected users. Finally, the target user is determined and the subsequent unlocking process is executed. This solves the problem that in the case of multiple users with facial occlusion and rapid movement, the target user cannot be quickly locked, thereby reducing the efficiency and accuracy of face recognition of the door lock.
[0071] The technical solutions in the above embodiments of the present application at least have the following technical effects or advantages:
[0072] Through user quantity recognition, this application distinguishes between single-user and multi-user scenarios, and sets up a splitting mechanism. The primary selected users are screened based on the deviation degree obtained from the target path and the movement trajectory of the users. Then, face segmentation, feature point recognition, and comparison are performed on the primary selected users, enabling accurate screening of the target user from multiple users. It realizes quickly locking the target user in the case of multiple users with facial occlusion and rapid movement, further improving the efficiency and accuracy of face recognition of the door lock.
[0073] Embodiment 4: In the above embodiment, face recognition of users in complex multi-user scenarios is achieved, but the efficiency of user face recognition is affected. This embodiment makes further improvements based on the above content.
[0074] In some embodiments, the user's height is estimated through head key points, a three-dimensional mapping table of height - distance - gradient parameters is constructed, and face features are recognized based on the target image by adjusting the gradient calculation area, including: establishing a head pose model, using a three-dimensional head model to match with facial key points, establishing a mapping relationship between the head pose and the coordinates of the key points, and estimating the user's height according to the relationship between the user's head pose, the position of the head key points, and the position of the ground; constructing a three-dimensional mapping table of height - distance - gradient parameters and the estimated height of the target image to obtain the corresponding gradient parameters; determining the gradient calculation area based on the gradient parameters to highlight the face features.
[0075] In some embodiments, a head pose model is established. A three-dimensional head model is used to match facial key points, and a mapping relationship between the head pose and the key point coordinates is established. By solving the rotation matrix or quaternion in the head pose model, the pitch angle and yaw angle of the head are obtained. The pitch angle reflects the degree of the head tilting up and down, and the yaw angle reflects the degree of the head rotating left and right. The calculated pitch angle and yaw angle are output in real time for subsequent face recognition and pose adjustment. Suppose the pitch angle of a certain user calculated by the head pose model is 10 degrees and the yaw angle is -5 degrees. Then the head pose output in real time is "pitch angle 10 degrees, yaw angle -5 degrees". The movement trajectory and movement amplitude of the user are further determined according to the user's head pose and spatio-temporal curve. At the same time, the height of the user is estimated according to the relationship between the user's head pose, the position of the head key points and the position of the ground. During the face recognition process, due to the movement of the user and environmental influence, it is often impossible to accurately obtain the full-body image of the user. Therefore, the height of the user needs to be estimated based on the face image.
[0076] Specifically, a large number of face images containing different head poses are collected, and the three-dimensional coordinates of the facial key points are labeled. Using the pre-constructed three-dimensional head model, the key points in the collected images are matched with the model by minimizing the error between the key points, and the pitch angle and yaw angle of the head are obtained; the internal parameters (such as focal length, optical center coordinates) and external parameters (the position and pose of the acquisition device in the world coordinate system) of the acquisition device are determined, and the head key points are accurately located in the collected images, such as the top vertex and the chin point. According to the imaging principle, the key point coordinates in the image are converted into actual space coordinates, and the user's height is estimated according to the actual space coordinates of the head key points and the pitch angle.
[0077] In some embodiments, face images and corresponding full-body images of different heights (such as 1.5 m, 1.6 m, 1.7 m, etc.) collected at different distances (such as 0.5 m, 1 m, 1.5 m, etc.) are obtained. For each collected image, its gradient features are calculated. The gradient features can be calculated by common gradient operators, including the magnitude and direction of the gradient. Key gradient parameters, such as the average gradient magnitude and the gradient direction distribution, are extracted from the gradient features of each image. Analyze the variation rules of the gradient parameters at different heights and different distances. For example, as the distance increases, the gradient magnitude may gradually decrease; at the same distance, the gradient distributions of face images of different heights may vary. Taking height, distance, and the extracted gradient parameters as three dimensions, a three-dimensional mapping table is constructed, and the height, distance, and the extracted gradient parameters are organized into a data table, with each collection point corresponding to a set of parameter values. For example, with height as the x-axis, distance as the y-axis, and the average gradient magnitude as the z-axis, the parameter values corresponding to each collection point are filled into the mapping table. An interpolation algorithm can be used to perform interpolation processing on the mapping table, so that at any given height and distance, the corresponding gradient parameters can be quickly obtained through the mapping table. Height: Affects the vertical distribution of the face in the image (the heads of tall users are closer to the top of the image); Distance: Determines the face resolution (the details of the image are rich at close range, and the image is blurred at long range); Gradient parameters: Include the gradient-sensitive region (ROI), sampling density, and high-low frequency weights. For example, at a close distance (0.5 m): When height ≤ 160 cm, focus on the lower half of the face (from the chin to the tip of the nose), step size = 2 pixels, low-frequency weight = 0.7; When height ≥ 180 cm, focus on the upper half of the face (from the forehead to the eyes), step size = 1 pixel, high-frequency weight = 0.6. At a long distance (2 m): Uniformly focus on the eye-nose triangle area, step size = 3 pixels, high-frequency weight = 0.8 (to compensate for the loss of details). Using common gradient operators such as the Sobel operator and the Scharr operator to calculate the gradient features of the image is a known technical method, and this application will not elaborate on it too much. Key parameters such as the average gradient magnitude and the gradient direction distribution are extracted from the gradient features.
[0078] During the face recognition process, for the currently collected face image, first, according to the user's height (dynamically estimated through the head pose model) and distance, the corresponding gradient parameters are obtained from the three-dimensional mapping table. The obtained gradient parameters are fused with the gradient features of the current image. For example, the average gradient magnitude obtained from the mapping table can be used as a weight to adjust the gradient features of the current image, enhancing the important gradient information in the image and suppressing noise and irrelevant information. This application does not make specific limitations.
[0079] In some embodiments, by utilizing gradient information to enhance the effect of face feature extraction, the extracted features are more discriminative and can better represent the facial features of different users. In the face recognition comparison stage, comparing the enhanced feature vector with the feature vector of the pre-stored face image can improve the matching accuracy and reduce the false recognition rate and rejection rate. By limiting the gradient calculation range through a mapping table, the number of processed pixels is reduced. For example, in a long-distance scenario, only 30% of the image area is calculated, and the time consumption is reduced by 50%. The time for feature extraction is reduced, thereby improving the response speed of the entire door lock system and enabling users to complete the unlocking operation faster.
[0080] For example, in complex environments (such as light changes, face occlusion, etc.), gradient information can provide additional feature information to help the system more accurately identify the user's identity.
[0081] In some embodiments, in the face recognition scenario of an intelligent door lock, the optimal gradient refers to the gradient calculation strategy that can most effectively represent face features (such as contours, textures) under specific height, distance, and environmental conditions. Its core objectives are: maximizing feature discrimination: highlighting the gradient information of key face areas (such as eye and nose contours); minimizing noise interference: suppressing invalid gradients caused by light changes, occlusions, etc.; balancing calculation efficiency: finding the optimal solution between recognition accuracy and processing speed.
[0082] In some embodiments, the image gradient is the rate of change of pixel gray values in the spatial direction. Areas with high gradient magnitudes correspond to image edges (such as face contours, facial feature boundaries), which are the core basis for feature extraction; the combination of high-frequency gradients (details) and low-frequency gradients (contours) can comprehensively describe the face structure, and the gradient direction distribution can distinguish real features from noise (such as light spots).
[0083] In some embodiments, face images are collected under different heights (150 - 190 cm), distances (0.5 - 3 m), and lighting conditions; the sampling step size (1 - 5) and weight ratio (low frequency: high frequency = 0.2:0.8 to 0.8:0.2) are traversed. Using the recognition accuracy (%) and single-frame processing time (ms) as indicators, an analysis chart is drawn to select the optimal parameter combination. The parameters are dynamically fine-tuned according to the current recognition result (success / failure). For example, if the recognition fails continuously 3 times, the gradient threshold is automatically reduced. Preferably, the high-frequency weight is adjusted using the light sensor data (increasing the low-frequency weight in strong light and enhancing the high-frequency in weak light).
[0084] After obtaining the target image, the user's height is determined based on the distance between the head key points of the image and the ground. Based on the three-dimensional mapping of height-distance-gradient parameters, the optimal gradient calculation strategy is selected to accurately identify the user's facial features, and then compared with the pre-stored image, that is, the pre-stored face image, ensuring the recognition accuracy and recognition speed.
[0085] In some embodiments, the gradient features (such as edges and textures) of face images exhibit significant differences at different heights and distances. By constructing a mapping table, the optimal gradient calculation parameters in different scenarios can be pre-stored to avoid the complexity and resource consumption of real-time calculation. In a dynamic environment, the stability of gradient features directly affects the recognition accuracy. Through historical data analysis, the mapping table provides adaptive parameter configurations to enhance the system's adaptability to complex scenarios. For different user groups and interaction distances, key gradient regions are automatically selected to avoid feature extraction from invalid regions, significantly reducing the computational amount.
[0086] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:
[0087] The present application enhances the face feature extraction effect by using gradient information. The extracted features are more discriminative and can better represent the facial features of different users. By comparing the enhanced feature vectors with the feature vectors of pre-stored face images, the matching accuracy is improved, and the false recognition rate and rejection rate are reduced. By limiting the gradient calculation range through the mapping table, the number of processed pixels is reduced, and the door lock response speed is increased.
[0088] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A smart door lock face recognition unlocking method, characterized in that: include: S100: determining the number of users and corresponding space-time curves based on a plurality of pre-collected images; Get the moving speed value based on the time-space curve; S200: If the moving speed value is greater than the first threshold, the blur degree of the pre-collected image is calculated and the blur kernel is estimated, and a deconvolution operation and hierarchical fusion are performed to obtain a preliminary image; If the moving speed value is not greater than the first threshold, the pre-captured image is marked as a preliminary selected image; S300: obtaining time-frequency features of the preliminary image, and then determining multiple local areas to obtain interference values. If the interference value is greater than a second threshold, determining a clean area and a dirty area based on a multi-mode mechanism to generate a target image; If the interference value is not greater than the second threshold, the target image is obtained according to the flat state mechanism; The multi-mode mechanism is as follows: determining a demarcation curve based on the difference in depth information between infrared image data and pixel points; acquiring infrared image data and identifying image boundaries, randomly selecting boundary points and calculating the temperature gradient change rate one by one, presetting a temperature change threshold, and identifying an area where the temperature gradient change rate is greater than the temperature change threshold as a demarcation area; determining edge points based on the difference in depth information of pixel points in the demarcation area; obtaining a demarcation curve based on the edge points; and obtaining a clean area and a dirty area based on the demarcation curve; The clean area is an image without occlusion; the dirty area is an image with occlusion; the boundary pixel points of the clean area are obtained, the visible light texture features are extracted, and a mixed feature vector is generated by combining the temperature distribution features in the infrared thermal image; the contour boundary and temperature distribution features of the dirty area are obtained, and the mixed feature vector of the clean area is combined for prediction and completion to obtain the target image; S400: estimating the height of the user through key points of the head, constructing a three-dimensional mapping table of height-distance-gradient parameters, adjusting the gradient calculation area based on the target image to recognize facial features, and comparing with the pre-stored image; The moving speed value is used to measure the change range of the user's posture.
2. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The time-frequency features include time-domain sparse points and frequency-domain spatial textures; the time-domain sparse points are: extracting pixel feature information of each pre-collected image, monitoring the grayscale value change of each pixel in real time, and based on a preset grayscale change threshold, marking the pixel points greater than the grayscale change threshold as time-domain sparse points; Record the spatiotemporal coordinates of each time-domain sparse point, including timestamp and pixel position; The frequency domain spatial texture is: obtaining a pre-collected image, performing wavelet transform, decomposing frequency domain components at different scales, extracting low-frequency contours and high-frequency detail features from the frequency domain components, and forming a frequency domain spatial texture.
3. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The flat state mechanism is: determining overlapping pixel points and relative position relationships according to a number of local areas, and splicing the plurality of pre-collected images according to the relative position relationships to obtain a target image; The overlapping pixel points of the pre-captured images are obtained, and the relative position relationship of the multiple pre-captured images on the face of the user to be identified is predicted based on the overlapping pixel points.
4. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The moving speed value is obtained based on the time-space curve, including: Based on the spatiotemporal curve, the displacement characteristics, average speed value and change amplitude quantization value of the facial key points between consecutive frames in the pre-collected image are calculated, and the average speed value and the change amplitude quantization value of all key points are weighted summed to obtain the moving speed value; among them, for each key point, the average speed value between consecutive frames is calculated; the standard deviation of the speed value of each key point is calculated as the change amplitude quantization value, which is used to reflect the stability of the key point speed; the calculation formula of the moving speed value is as follows: ; Among them, M is the moving speed value, is the weight coefficient of the i-th key point, and its value range is [0, 1]; is the average velocity value of the i-th key point; is the quantitative value of the change amplitude of the i-th key point; n is the total number of key points; The average speed value is obtained by dividing the modulus of the displacement vector by the time interval; the displacement features include: key point coordinates, displacement direction, speed value and time interval.
5. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: Calculate the blur degree of the pre-collected image and estimate the blur kernel, perform deconvolution and layered fusion to obtain the preliminary image, including: The estimated blur kernel is used to perform deconvolution operation to restore the clear features of the image and obtain a preliminary clear image; Perform layered fusion on the preliminary clear images to obtain the preliminary selected images; Among them, the hierarchical fusion is based on the fusion of short-term features and long-term features, extracting local feature points of a single frame image in a preliminary clear image, calculating local feature descriptors; obtaining local feature descriptors of multiple consecutive frames of images, and obtaining aggregated long-term feature vectors; splicing and fusing the local feature descriptors and long-term feature vectors respectively, and obtaining a fused feature vector as a preliminary image.
6. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The method further includes: S110: identifying the number of users, if the number of users is 1, executing S200; if the number of users is greater than 1, executing S120; S120: obtaining a preliminary user based on the splitting mechanism, performing facial segmentation on the preliminary user, identifying facial feature points and comparing them with pre-stored images, determining a target user, and obtaining a moving speed value based on the time-space curve, and executing S200; Among them, the splitting mechanism includes: determining the deviation from the target path based on the space-time curve, marking the users corresponding to the deviation less than a preset third threshold as preliminary users; the target path refers to the straight path directly in front of the door lock; the deviation is used to measure the difference between each user's moving path and the preset target path.
7. The face recognition unlocking method for smart door locks according to claim 6, characterized in that: Based on the pre-collected images corresponding to the primary selected users, the facial area is extracted to identify the facial feature points, and the feature vectors based on the facial feature points are compared with the feature vectors of the pre-stored images to obtain the matching degree and thus determine the target user.
8. The face recognition unlocking method for smart door locks according to claim 1, characterized in that: The user's height is estimated through the key points of the head, a three-dimensional mapping table of height-distance-gradient parameters is constructed, and the gradient calculation area is adjusted based on the target image to recognize the facial features, including: establishing a head posture model, using the three-dimensional head model to match the facial key points, establishing a mapping relationship between the head posture and the key point coordinates, and estimating the user's height based on the user's head posture, the position of the head key points and the position of the ground; constructing a three-dimensional mapping table of height-distance-gradient parameters, and the estimated height of the target image to obtain the corresponding gradient parameters; determining the gradient calculation area to highlight the facial features based on the gradient parameters.
Citation Information
Patent Citations
Smart door lock facial recognition unlocking methods, devices, media, and smart door locks
CN115938023B
Method and device for training face recognition model, equipment and storage medium
CN117612223A
Face recognition method and system based on AI intelligence
CN118430054A