Method, device, equipment and medium for intelligent door control

By combining gait and facial features for intelligent gate recognition, the problems of poor environmental adaptability and susceptibility to interference in single biometric recognition are solved, achieving higher recognition accuracy and security, and improving user experience.

CN121305658APending Publication Date: 2026-01-09WONLY SECURITY & PROTECTION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511265611.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing smart door systems rely on a single biometric feature for identity verification, which has problems such as limited functionality, poor environmental adaptability, susceptibility to changes in lighting and obstruction, vulnerability to deception, and poor user experience.

Method used

By combining gait and facial features for recognition, video information within the field of view of the smart gate is acquired, gait and facial features are extracted, confidence scores are fused for target recognition, and smart gate control is performed in conjunction with behavioral intent.

Benefits of technology

It improves the accuracy and security of recognition, reduces false triggers, enhances user experience and control precision, adapts to complex environments, and avoids the shortcomings of single features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305658A_ABST
    Figure CN121305658A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart home security, and discloses a method, a device, equipment and a medium for smart door control, the method comprises the following steps: acquiring video information in a smart door vision field, and extracting gait features and facial features in the video information; determining the gait feature confidence coefficient based on the matching degree of the gait features and the gait features in a gait feature library; determining facial feature confidence based on the similarity between the facial features and facial features in a facial feature library; fusing the gait feature confidence coefficient and the facial feature confidence coefficient to obtain a recognition result of the target in the video; according to the method, the gait features and the facial features are combined to perform target recognition, so that the accuracy and the safety of target recognition in the video can be improved, and the defect of single feature is avoided; meanwhile, the recognition result is combined with the behavior intention, corresponding intelligent door control is carried out, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart home security technology, and more specifically to methods, devices, equipment and media for smart door control. Background Technology

[0002] Smart door systems in related technologies rely on a single biometric feature, such as facial recognition or fingerprint recognition, to determine whether to open the door. This method has the drawback of being relatively limited in functionality, only capable of user authentication and unable to understand and respond to the user's actual behavior. Summary of the Invention

[0003] In view of this, the present invention provides a method, apparatus, device and medium for intelligent door control, to solve the problem that intelligent door systems in related technologies have relatively simple functions, can only perform user authentication, and cannot understand and respond to the user's actual behavior.

[0004] In a first aspect, the present invention provides a method for controlling a smart door, the method comprising: acquiring video information within the field of view of the smart door; extracting gait features and facial features from the video information; determining a gait feature confidence level based on the matching degree between the gait features and gait features in a gait feature library; determining a facial feature confidence level based on the similarity between the facial features and facial features in a facial feature library; fusing the gait feature confidence level and the facial feature confidence level to obtain a target recognition result in the video; recognizing the behavioral intent of the target; and controlling the smart door based on the recognition result and the behavioral intent.

[0005] In one optional implementation, the step of extracting gait features from the video information includes: determining the weight of each video frame based on the illumination intensity of each video frame in the video information; fusing the target video frame and the weight of the target video frame to extract the gait features from the target video frame.

[0006] In one optional implementation, determining the weight of each video frame based on the illumination intensity of each video frame in the video information includes: determining the weight of each video frame using the following formula:

[0007]

[0008] Where, α t The weights L used to characterize the image in frame t. t L is used to characterize the illumination intensity of the t-th frame image. opt The preset light intensity is used to characterize the light sensitivity adjustment parameter, and σ is used to characterize the light sensitivity adjustment parameter.

[0009] In one optional implementation, the step of fusing the target video frame and its weights to extract gait features from the target video frame includes: extracting the gait features using the following formula:

[0010]

[0011] Among them, GEI adj (x, y) represents the grayscale value of the gait energy map at pixel coordinates (x, y), N represents the number of image frames used to calculate the gait energy map, and α t The weights C used to characterize the image in frame t are: t (x,y) is used to characterize the pixel value at pixel coordinates (x,y) of the image in frame t.

[0012] In one optional implementation, determining the gait feature confidence level based on the matching degree between the gait features and gait features in the gait feature database, and determining the facial feature confidence level based on the similarity between the facial features and facial features in the facial feature database, includes: determining gait feature weights and facial feature weights based on the stability of the gait features and facial features; determining the gait feature confidence level based on the product of the matching degree and the gait feature weights; and determining the facial feature confidence level based on the product of the similarity and the facial feature weights.

[0013] In one optional implementation, determining the gait feature weights and facial feature weights based on the stability of the gait features and facial features includes: determining the gait feature weights using the following formula:

[0014]

[0015] Among them, w g Used to characterize gait feature weights Error variance estimates used to characterize facial features in video frames. Error variance estimates used to characterize gait features in the video frame; facial feature weights determined using the following formula:

[0016]

[0017] Among them, w f Used to characterize facial feature weights.

[0018] In one optional implementation, fusing the gait feature confidence and the facial feature confidence to obtain the target recognition result in the video includes: obtaining the recognition result using the following formula:

[0019]

[0020] Wherein, m(A) is used to characterize the confidence function value of the target belonging to category A in the recognition result, A is used to characterize the category of the target, B is used to characterize the gait recognition result, C is used to characterize the facial recognition result, m1(B) is used to characterize the confidence of the gait feature, m2(C) is used to characterize the confidence of the facial feature, and K is used to characterize the conflict coefficient.

[0021] In one optional implementation, identifying the target's behavioral intent includes obtaining the behavioral intent using the following formula:

[0022]

[0023] Among them, V * (s) represents the target's behavioral intention in state s, s represents the target's state vector, a represents the target's behavior in state s, R(s,a) represents the reward value of the target performing behavior a in state s, γ represents the discount factor, s′ represents the next state the target transitions to after performing behavior a in state s, P(s′|s,a) represents the probability of transitioning to state s′ after performing behavior a in state s, and V * (s′) is used to characterize the behavioral intention of the target in state s′.

[0024] In one optional implementation, controlling the smart door based on the identification result and the behavioral intent includes: if the identification result indicates that the target is an authorized user, and the behavioral intent indicates that the target's behavior is returning home, controlling the smart door to unlock; if the identification result indicates that the target is a visitor, and the behavioral intent indicates that the target's behavior is visiting, sending a message to the authorized user asking whether to unlock the smart door, and controlling the smart door to unlock or remain locked based on the authorized user's feedback on the message; if the identification result indicates that the target is a courier, and the behavioral intent indicates that the target's behavior is delivering a package, controlling the package delivery slot to open based on the smart door.

[0025] Secondly, the present invention provides a device for controlling a smart door, the device comprising: a data acquisition and feature extraction module, configured to acquire video information and extract gait features and facial features from the video information; determine gait feature confidence based on the matching degree between the gait features and gait features in a gait feature library; determine facial feature confidence based on the similarity between the facial features and facial features in a facial feature library; a feature fusion module, configured to fuse the gait feature confidence and the facial feature confidence to obtain a target recognition result in the video; and a control module, configured to recognize the behavioral intent of the target and control the smart door based on the recognition result and the behavioral intent.

[0026] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method for intelligent door control described in the first aspect or any corresponding embodiment thereof.

[0027] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method for intelligent door control described in the first aspect or any corresponding embodiment thereof.

[0028] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the method for intelligent door control described in the first aspect or any corresponding embodiment thereof.

[0029] The method for smart door control provided in this embodiment combines gait features and facial features for target recognition. Compared to single facial recognition, which is easily affected by light or occlusion, and single gait recognition, which is easily affected by clothing, this method can improve the accuracy and security of target recognition in video and avoid the defects of single features. At the same time, compared to solutions that can only perform identity recognition, this method combines the recognition result with behavioral intent. Smart door control is only performed when both the recognition result and the behavioral intent are met. This can reduce false triggering, improve the user experience of smart doors, and improve the control accuracy of smart doors. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of the present invention, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0031] Figure 1 A flowchart illustrating a method for smart door control according to an embodiment of this application is shown;

[0032] Figure 2 A schematic diagram of the structure of a system for intelligent door control provided in an embodiment of this application is shown;

[0033] Figure 3 A schematic diagram of the structure of a device for intelligent door control provided in an embodiment of this application is shown;

[0034] Figure 4 A schematic diagram of the hardware structure of a computer device according to an embodiment of this application is shown. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] Relying on a single biometric feature to determine whether a smart door is open has the following problems:

[0037] First, it has poor environmental adaptability; factors such as changes in lighting, occlusion, and camouflage can seriously affect the accuracy of recognition.

[0038] Secondly, it may require users to actively cooperate, such as pointing at the camera or touching the sensor, which lacks convenience and leads to a poor user experience.

[0039] Third, single-modal attacks are easily deceived, such as photo attacks or video attacks, which can easily lead to security risks.

[0040] Fourth, its functionality is limited; it can only determine whether to open the door and lacks an understanding of the user's behavioral intentions.

[0041] Fifth, gait recognition, as a long-distance, non-contact biometric identification technology, has advantages such as being difficult to fake and requiring no active cooperation from the user. However, when used alone, it suffers from problems such as high feature dimensionality and high computational complexity.

[0042] According to an embodiment of the present invention, a method embodiment for intelligent door control is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0043] This embodiment provides a method for controlling smart doors, which can be used in the aforementioned terminal devices, such as smart doors, mobile phones, tablets, desktop computers, laptops, or servers. Figure 1 A flowchart illustrating a method for smart door control according to an embodiment of this application is shown, as follows: Figure 1 As shown, the process includes the following steps:

[0044] Step S101: Obtain video information within the field of view of the smart gate, and extract gait features and facial features from the video information; determine the confidence level of gait features based on the matching degree between gait features and gait features in the gait feature library; determine the confidence level of facial features based on the similarity between facial features and facial features in the facial feature library.

[0045] In this step, the video information within the smart door's field of view can be a set of dynamic visual data continuously captured and generated by the smart door through its built-in or external image acquisition device within the physical field of view covered by the image acquisition device. This data can be acquired by an image acquisition device installed on the smart door, or by image acquisition devices located on the top, left, right, or bottom of the smart door. Specifically, the image acquisition device can be a camera, such as a 1920×1080 pixel high-definition camera, and can be configured on the smart door, to the left, or to the right of the smart door at a height of 1.8 meters to 2.2 meters above the ground.

[0046] Gait features characterize the unique movement patterns exhibited by the human body during walking, including stride frequency, stride length, gait cycle, and body swing amplitude. Gait features vary from person to person. Facial features characterize the distinctive biological features of the human face, including facial contours, key point coordinates, and skin texture, and also exhibit individual variations.

[0047] A gait feature library is a set of pre-stored gait feature data of authorized users (such as family members) in the smart gate system, which can be used for matching and comparison with extracted gait features. A facial feature library is a set of pre-stored facial feature data of authorized users (such as family members) in the smart gate system, which can be used for matching and comparison with extracted facial features.

[0048] Gait feature confidence score characterizes the degree of match between extracted gait features and the gait features of a specific user in a gait feature database. A higher score indicates stronger consistency between the two, meaning a higher confidence that the gait belongs to that user. Gait feature confidence score includes both the gait recognition result and the confidence level of the gait recognition result.

[0049] Facial feature confidence score characterizes the similarity between extracted facial features and the facial features of a user in a facial feature database. A higher score indicates stronger consistency between the two, meaning a higher confidence that the face belongs to that user. Facial feature confidence score includes both the facial recognition result and the confidence level of the facial recognition result.

[0050] Computer vision algorithms, such as skeletal keypoint extraction algorithms or improved ResNet models, can be used to extract gait and facial features from acquired video information. Gait feature confidence is obtained by matching the extracted gait features with stored authorized user gait features and calculating similarity.

[0051] Similarly, facial feature confidence can be achieved by matching extracted facial features with facial features in a facial feature database, and the calculated facial feature similarity score can be used as the facial feature confidence score. Specific similarity calculation algorithms can include cosine similarity algorithms, etc.

[0052] Step S102: Fusion of gait feature confidence and facial feature confidence to obtain the target recognition result in the video.

[0053] In this step, a multi-feature fusion algorithm, such as a weighted average or a voting method, can be used to fuse gait feature confidence scores and facial feature confidence scores to determine who the target in the video is. The target identification result in the video can be an authorized user (i.e., an authorized resident); a pre-recorded visitor; or an unauthorized user.

[0054] Step S103: Identify the target's behavioral intent, and control the smart door based on the identification result and behavioral intent.

[0055] In this step, behavioral intent is used to characterize the target's dynamic behavior in the video domain, such as the target's actions, the duration of its stay, and its interaction with the door. Behavioral analysis algorithms, such as Long Short-Term Memory (LSTM) network models, can be used to analyze the target's dynamic behavioral sequences in the video domain and predict the target's subjective intent. The control of the smart door can be the specific actions performed by the smart door, such as unlocking, triggering an alarm, or maintaining a locked state.

[0056] When the recognition result indicates that the target has the authority to open the door, and the behavioral intent indicates that the target intends to open the door, controlling the smart door to open can solve the problem of residents passing by the door but not needing to open it being mistakenly locked, and can avoid accidental triggering.

[0057] The method for smart door control provided in this embodiment combines gait features and facial features for target recognition. Compared to single facial recognition, which is easily affected by light or occlusion, and single gait recognition, which is easily affected by clothing, this method can improve the accuracy and security of target recognition in video and avoid the defects of single features. At the same time, compared to solutions that can only perform identity recognition, this method combines the recognition result with behavioral intent. Smart door control is only performed when both the recognition result and the behavioral intent are met. This can reduce false triggering, improve the user experience of smart doors, and improve the control accuracy of smart doors.

[0058] In some optional implementations, extracting gait features from video information includes: determining the weight of each video frame based on the illumination intensity of each video frame in the video information; fusing the target video frame and the weight of the target video frame to extract the gait features from the target video frame.

[0059] In this embodiment, the illumination intensity of each video frame in the video information is used to characterize the overall brightness level of a single frame image in the video, reflecting the degree of light illumination on that frame image. The weight of the video frame can be used to measure the contribution of that frame to gait feature extraction. The target video frame can represent any frame image in the video information. The video frames and their corresponding weights are fused, for example, by multiplying the image data of each video frame with its corresponding weight through weighted calculation. In the fusion result, gait features can be extracted using a skeleton detection algorithm.

[0060] Specifically, the image frames in the video information can be analyzed frame by frame to determine the illumination intensity of each frame, thus eliminating video frames without targets. Based on preset rules, the weights of each video frame are determined; for example, video frames with brightness between 50 and 150 are assigned a weight of 0.8 to 1, while those with brightness less than or equal to 50, or greater than or equal to 150, are assigned a weight of 0.1 to 0.7. A weighted average method is then used to fuse the video frames and their corresponding weights, extracting gait features from the fused data.

[0061] The lighting within the field of view of a smart door fluctuates greatly, such as during the day or at night. Excessive brightness or darkness can lead to the loss of limb details. By determining the weight of video frames based on the light intensity, illumination interference can be addressed in a targeted manner, thereby improving the accuracy of gait feature extraction.

[0062] In some optional implementations, illumination correction is performed based on video information to determine the weight of each video frame, including: determining the weight of each video frame using the following formula:

[0063]

[0064] Where, αt The weights L used to characterize the image in frame t. t L is used to characterize the illumination intensity of the t-th frame image. opt The preset light intensity is used to characterize the light sensitivity adjustment parameter, and σ is used to characterize the light sensitivity adjustment parameter.

[0065] In this embodiment, α t The value range of α is (0,1], and as an adaptive weighting coefficient, it can be based on α. t Illumination correction is applied to the gait energy map. When the actual illumination intensity L... t Approximately the preset light intensity L opt In the case of α t When L approaches 1, the adjustment range of the image is small; when L... t With L opt When the differences are large, α t Reducing the size of the image can increase the range of adjustments needed to counteract lighting interference.

[0066] L t The ambient light intensity, measured in lux (lx), can be collected in real time by an ambient light sensor and reflects the current lighting level in front of the smart door. It is a core input parameter for light compensation. The ambient light sensor can be configured at any location that can monitor the light intensity within the smart door's field of view. opt It can be used to characterize the preset optimal light intensity, and the unit is also lux (lx). It can be determined through experimental testing. Under this lighting condition, subsequent algorithms such as gait feature extraction and facial feature extraction can achieve the best recognition performance. It is a benchmark reference value for lighting compensation.

[0067] σ is a constant greater than zero, which can control the effect of changes in light intensity on α. t The degree of influence. The smaller σ is, the more a small change in light intensity will lead to α. t Significant fluctuations indicate that the system is more sensitive to changes in illumination; the larger σ is, the greater the α becomes, only when there are large changes in illumination. t Only then will adjustments be noticeable, making the system more adaptable to changes in lighting.

[0068] In this way, the weight allocation of video frames is more accurate and smooth, which can conform to the inherent laws of illumination and gait feature extraction and avoid the sudden error of threshold-based weight allocation. At the same time, the illumination within the field of view of the smart gate is dynamically and continuously changing. The continuity of the Gaussian function allows the weight of the video frame to be smoothly adjusted with the illumination, avoiding drastic weight jumps caused by small fluctuations in illumination, which can make subsequent gait feature extraction more stable.

[0069] In some optional implementations, the target video frame and its weights are fused to extract gait features from the target video frame, including: extracting the gait features using the following formula:

[0070]

[0071] Among them, GEI adj (x,y) represents the grayscale value of the gait energy map at pixel coordinates (x,y), N represents the number of image frames used to calculate the gait energy map, and α t The weights C used to characterize the image in frame t are: t (x,y) is used to characterize the pixel value at pixel coordinates (x,y) of the image in frame t.

[0072] In this embodiment, the Gait Energy Image (GEI) processes continuous video frames of the human walking process, converting dynamic gait motion information into a single static grayscale image, which simplifies the analysis and matching of gait features. The gait energy image in this embodiment incorporates illumination correction logic into its calculation process, fusing the illumination intensity weights of each video frame. This effectively eliminates illumination interference and provides stable input for subsequent gait feature analysis.

[0073] N is a positive integer. The value of N needs to balance computational efficiency and gait feature integrity. If the value is too small, the gait feature will not be captured completely. If the value is too large, the computational cost will be increased. This scheme determines through experiments that the optimal value range of N is 20 to 40 frames.

[0074] C t (x,y) is used to represent the pixel value of the binarized human body contour. The value is 0 or 1. 1 indicates that the pixel belongs to the human body contour region, and 0 indicates that the pixel is the background region. It is the basic data for constructing the gait energy map.

[0075] By integrating illumination correction logic into the calculation process of gait energy map, the gait energy map incorporates the illumination intensity weights of each video frame, which can effectively eliminate illumination interference and provide stable input for subsequent gait feature analysis.

[0076] In some optional implementations, gait feature confidence is determined based on the matching degree between gait features and gait features in a gait feature library, and facial feature confidence is determined based on the similarity between facial features and facial features in a facial feature library, including: determining gait feature weights and facial feature weights based on the stability of gait features and facial features; determining gait feature confidence based on the product of matching degree and gait feature weights; and determining facial feature confidence based on the product of similarity and facial feature weights.

[0077] In this embodiment, the degree to which gait and facial features are affected by environmental interference can be analyzed. Based on their stability, different feature weights are assigned to these two types of features; features with higher stability receive greater weights. In low-light conditions at night, facial feature stability decreases, so the weight of facial features can be reduced while the weight of gait features is increased to avoid misjudgments due to facial recognition errors. When the target's gait features are affected, but the face is clear, the weight of facial features can be increased to compensate for the deficiencies in gait features.

[0078] Specifically, algorithms such as similarity matching or feature vector distance can be used to obtain the matching degree between gait features and features in the gait feature database. Then, the matching degree is multiplied by the gait feature weights to obtain the gait feature confidence score. Similarly, the confidence score of facial features can be obtained.

[0079] In this way, adjusting feature weights based on feature stability can adapt to various complex scenarios and improve the robustness of this solution.

[0080] In some optional implementations, gait feature weights and facial feature weights are determined based on the stability of gait features and facial features, including: determining the gait feature weights using the following formula:

[0081]

[0082] Among them, w g Used to characterize gait feature weights Error variance estimates used to characterize facial features in video frames. Error variance estimates used to characterize gait features in video frames; facial feature weights are determined using the following formula:

[0083]

[0084] Among them, w f Used to characterize facial feature weights.

[0085] In this embodiment, w g The value range of w is [0,1]. g The size of w can be determined by comparing the stability of gait features and facial features. When facial features are occluded or interfered with by lighting, leading to a decrease in the reliability of facial features, w... g Increasing the weight of gait features can enhance the decision weight of gait characteristics.

[0086] w f The value range of w is [0,1], and it satisfies w g +w f =1, in well-lit scenes or scenes without facial obstruction, w f Enlarging the size of the face recognition system allows for full utilization of its high-precision advantages. The error variance estimate used to reflect gait characteristics in the current environment can reflect the stability of gait feature extraction. The smaller the value, the less the gait characteristics are affected by the environment and the higher the reliability. The larger the value, the greater the fluctuation in gait characteristics and the lower the reliability, so its fusion weight needs to be reduced. The error variance estimate used to reflect facial features in the current environment can reflect the stability of facial feature extraction. For example, under strong light, backlight, or when the face is covered by a mask, Enlarging the w area decreases the reliability of facial features, so the w area is automatically reduced. f To avoid misidentification.

[0087] In this way, by calculating the weights of gait features and facial features based on the statistical quantification index of error variance, the fluctuation of gait features or facial features in the current video frame can be accurately reflected. The smaller the error variance, the more stable and reliable the feature is, and the higher the weight will be. This data-driven weight allocation avoids the one-sidedness of artificially set rules, makes the fusion of dual features more in line with the statistical laws of actual data, reduces the impact of experience bias on the recognition results, and can improve the accuracy of the recognition results.

[0088] This embodiment provides a method for controlling a smart door, which can be used in the aforementioned terminal devices, such as smart doors, mobile phones, tablets, desktop computers, laptops, or servers. The process of the method for controlling a smart door according to this embodiment includes the following steps:

[0089] Step S201: Acquire video information within the field of view of the smart gate, and extract gait features and facial features from the video information; determine the confidence level of the gait features based on the matching degree between the gait features and gait features in the gait feature database; determine the confidence level of the facial features based on the similarity between the facial features and facial features in the facial feature database. For details, please refer to... Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0090] Step S202: Fusion of gait feature confidence and facial feature confidence to obtain the target recognition result in the video.

[0091] Specifically, step S202 includes:

[0092] Step S2021: Obtain the recognition result using the following formula:

[0093]

[0094] Where m(A) is used to characterize the confidence function value of the target belonging to category A in the recognition result, A is used to characterize the category of the target, B is used to characterize the gait recognition result, C is used to characterize the face recognition result, m1(B) is used to characterize the gait feature confidence, m2(C) is used to characterize the face feature confidence, and K is used to characterize the conflict coefficient.

[0095] In this step, decision fusion can be performed based on Dempster-Shafer evidence theory. The value of m(A) is [0,1]. The closer m(A) is to 1, the higher the confidence that the target in the video belongs to category A. m(A) is the core basis for the control decision of the smart gate. It should be noted that the numerator of m(A) is as follows:

[0096]

[0097] A includes: authorized users, visitors, delivery personnel, and unrecorded individuals. B includes the set of hypothetical categories in gait recognition mode, which can be the behavior categories of all identities that gait recognition may output, such as the gait of an authorized user, a delivery person, or an unrecorded individual. C includes the set of hypothetical categories in facial recognition mode, which can be all identities that facial recognition may output, such as the face of an authorized user or an unauthorized individual.

[0098] m1(B) can be used to determine the basic probability assignment value of the gait recognition result for the target in the video belonging to category B. The value of m1(B) is [0,1], which can reflect the degree of support of the gait recognition result for the target belonging to category B. m2(C) can be used to determine the basic probability assignment value of the face recognition result for the target in the video belonging to category C. The value of m2(C) is [0,1], which can reflect the degree of support of the face recognition result for the target belonging to category C. It can be calculated by the improved ResNet face feature matching algorithm.

[0099] K can be calculated using the following formula:

[0100]

[0101] The value of K is [0,1]. K is used to characterize the degree of conflict between the gait feature recognition result and the facial feature recognition result. The closer K is to 1, the greater the difference between the recognition results of the two modes. For example, the gait recognition result determines that the user is authorized, while the facial recognition result determines that the user is unauthorized. The closer K is to 0, the more consistent the recognition results of the two modes are.

[0102] Step S203: Identify the target's behavioral intent; based on the identification result and behavioral intent, control the smart door. For details, please refer to [link to relevant documentation]. Figure 1Step S103 of the illustrated embodiment will not be described again here.

[0103] In this way, by combining the DS evidence theory and quantifying the degree of conflict between gait and facial recognition results through the conflict coefficient, feature conflicts can be handled efficiently and the risk of misjudgment can be reduced. At the same time, the calculation method based on m(A) can not only describe the reliability of determining whether someone belongs to a certain type of person, but also distinguish between partially supported or completely uncertain reliability states. This is more flexible than simple probability judgment and can accurately express the uncertainty of reliability, thereby improving the accuracy of recognition.

[0104] In some alternative implementations, identifying the target's behavioral intent includes obtaining the behavioral intent using the following formula:

[0105]

[0106] Among them, V * (s) represents the target's behavioral intention in state s, s represents the target's state vector, a represents the target's behavior in state s, R(s,a) represents the reward value of the target performing behavior a in state s, γ represents the discount factor, s′ represents the next state the target transitions to after performing behavior a in state s, P(s′|s,a) represents the probability of transitioning to state s′ after performing behavior a in state s, and V * (s′) is used to characterize the target’s behavioral intention in state s′.

[0107] In this implementation, behavioral intent can be identified based on Markov decision processes. * (s) is the optimal value function under state s, representing the maximum cumulative reward that can be obtained by taking the optimal sequence of actions starting from the current state s. It can be used by the user to evaluate the value of the current state s. V * (s) can determine whether the target intends to go home or not.

[0108] `s` is used to characterize the target's current state vector, including parameters such as the target's distance from the smart door, movement speed, movement direction, and dwell time. It is the basic input for behavioral intent analysis. For example, a target 3 meters away from the door, moving towards the door, and moving at a speed of 0.8 m / s can constitute a specific state `s`. `a` can be continuing to move towards the door, staying, turning away, or approaching the express delivery area, etc., which can cover all possible behaviors of the target in front of the smart door.

[0109] R(s,a) can be determined by the system's preset reward rules. For example, if the target is in state s, 1 meter away from the door and moving towards the door, taking action a, which is to continue moving towards the door, will result in a positive reward, such as an increase of 5 in the reward value. If the target is in state s, 2 meters away from the door and staying, taking action a, which is to move closer to the door, will result in a negative reward, such as a decrease of 10 in the reward value.

[0110] γ takes a value of [0,1] and is used to balance the weight of immediate rewards and future rewards. The closer γ is to 1, the more emphasis is placed on long-term future rewards, such as the continuous behavior of a user moving towards the door; the closer γ is to 0, the more emphasis is placed on current immediate rewards, such as whether the user is currently moving towards the door.

[0111] s′ can be a state where the target is in a state where it is moving towards the door from a distance of 3 meters, and after taking action a to continue moving towards the door, it transitions to a state where it is moving towards the door from a distance of 2 meters.

[0112] The value of P(s′|s,a) is [0,1], which can be obtained by statistical training of a large amount of user behavior data. It can reflect the uncertainty of the target state transition. For example, if the target is 2 meters away from the door and takes the action of moving towards the door, there is a 90% probability that it will transfer to the state 1 meter away from the door, and a 10% probability that it will transfer to the state 2 meters away from the door and stay there due to staying.

[0113] In this way, by establishing the association between state, behavior and intent, the recognition accuracy can be improved; at the same time, for scenarios such as changes in lighting and background pedestrian interference, noise can be filtered through probabilistic state transitions to ensure the stability and anti-interference ability of target behavior and intent recognition in complex scenarios.

[0114] In some optional implementations, the smart door is controlled based on the identification result and behavioral intent, including: if the identification result indicates that the target is an authorized user and the behavioral intent indicates that the target's behavior is returning home, the smart door is unlocked; if the identification result indicates that the target is a visitor and the behavioral intent indicates that the target's behavior is visiting, the smart door sends a message to the authorized user asking whether to unlock, and based on the authorized user's feedback on the message, the smart door is controlled to unlock or remain locked; if the identification result indicates that the target is a courier and the behavioral intent indicates that the target's behavior is delivering a package, the smart door controls the package delivery slot to open.

[0115] In this implementation, different levels of smart door unlocking responses can be triggered based on the identification results and behavioral intent. The target's behavioral intent includes: returning home, passing by, delivering a package, or loitering abnormally. If the identification result indicates the target is a registered visitor, and the behavioral intent indicates the target's behavior is a visit, a notification message can be sent to the authorized user (i.e., a family member in the household). After confirmation by the authorized user, the smart door can be unlocked. If the identification result indicates the target is an unregistered person, and the behavioral intent indicates the target's behavior is abnormal loitering or unauthorized access, an alarm message can be sent to the authorized user based on the smart door.

[0116] In this way, for authorized users, when the user is identified as an authorized user and their intention is to return home, the smart door will automatically unlock without manual operation, which can improve the user experience and achieve seamless and convenient passage. At the same time, by judging the user's identity and intention, the security vulnerabilities of single identity recognition can be avoided. In addition, the smart door can be differentiated based on different identities and different behavioral purposes, which can realize the fine-grained scene adaptation of the smart door and improve its practicality.

[0117] In some alternative implementations, the target's movement trajectory can be tracked based on a particle filter algorithm. The target's trajectory can be obtained by detecting and tracking the target's position in each frame of a video frame sequence. The target's behavioral intent can be determined based on its movement trajectory.

[0118] Specifically, object detection algorithms, such as deep learning-based object detection models, can be used to locate the position of the target in each frame. Then, these positions are connected in chronological order to obtain the target's motion trajectory.

[0119] The behavioral characteristics of delivery personnel include lingering in specific areas in front of the smart door, such as near the package delivery area, for relatively fixed periods, and exhibiting a pattern of approaching the door or briefly stopping before leaving. Facial features of the delivery personnel can be recorded on the terminal device running the user smart door control method provided in this application, or they can be determined using common delivery personnel attire, etc. Due to the nature of their profession, delivery personnel may have gaits that differ from those of ordinary residents. Combined with the determination of their intention to deliver packages in the behavioral intent analysis, a comprehensive assessment is made to determine whether someone is a delivery personnel.

[0120] In some alternative implementations, if the identification result indicates that the target is an authorized user, a personalized greeting can be provided to the authorized user while controlling the unlocking of the smart door.

[0121] In some optional implementations, after acquiring video information within the smart door's field of view, noise reduction, size normalization, and illumination compensation can be performed on the video frames. Three-dimensional contour information of targets within the smart door's field of view can be acquired using a depth sensor positioned on or near the smart door.

[0122] The method for intelligent gate control provided in this application can significantly improve the recognition accuracy in complex environments by determining the weights of each video frame and dynamically determining the weights of gait features and facial features. At the same time, by using Markov decision processes to understand the target's behavioral intentions, a deep understanding of the target's behavioral intentions can be achieved. In addition, this seamless authentication of authorized users can enhance the user experience. Furthermore, the deep integration of gait recognition, facial recognition, and behavior analysis can provide comprehensive security.

[0123] Figure 2 A schematic diagram of the structure of a system for intelligent door control provided in an embodiment of this application is shown, such as... Figure 2 As shown, the system for intelligent door control includes:

[0124] The sensor data acquisition module 201 includes a high-definition camera 2011, an ambient light sensor 2012, and a depth sensor 2013. The sensor data acquired by the sensor data acquisition module 201 can be passed to the preprocessing and feature extraction module 202.

[0125] The preprocessing and feature extraction module 202 includes: an illumination compensation unit 2021, a gait feature extraction unit 2022, and a facial feature extraction unit 2023. The preprocessing and feature extraction module 202 inputs the extracted features into the multimodal fusion recognition module 203 and the behavior analysis module 204.

[0126] The multimodal fusion recognition module 203 includes an adaptive weight allocation unit 2031, a confidence evaluation unit 2032, and a decision fusion unit 2033. Through the multimodal fusion recognition module 203, weight allocation, confidence evaluation, and decision fusion are performed on the extracted facial features and gait features to obtain the recognition result. The obtained recognition result is then fed into the intelligent response module 205.

[0127] The behavior analysis module 204 includes a path planning and recognition unit 2041 and a behavior intent judgment unit 2042. Based on the behavior analysis module 204, the target behavior intent is identified. The obtained behavior intent is then passed to the intelligent response module 205.

[0128] The intelligent response module 205 includes a dynamic unlocking control unit 2051 and an abnormal situation handling unit 2052. The intelligent response module 205 controls the intelligent door based on the recognition results and behavioral intentions.

[0129] This embodiment also provides a device for intelligent door control, which implements the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0130] This embodiment provides a device for intelligent door control. Figure 3 A schematic diagram of the structure of a device for intelligent door control provided in an embodiment of this application is shown, as follows: Figure 3 As shown, it includes:

[0131] The data acquisition and feature extraction module 301 is used to acquire video information and extract gait features and facial features from the video information; determine the confidence level of gait features based on the matching degree between gait features and gait features in the gait feature library; and determine the confidence level of facial features based on the similarity between facial features and facial features in the facial feature library.

[0132] The feature fusion module 302 is used to fuse gait feature confidence and facial feature confidence to obtain the recognition result of the target in the video.

[0133] The control module 303 is used to identify the target's behavioral intent and control the smart door based on the identification result and behavioral intent.

[0134] In some optional implementations, the data acquisition and feature extraction module 301 includes:

[0135] The first feature extraction unit is used to determine the weight of each video frame based on the illumination intensity of each video frame in the video information; and to fuse the target video frame and the weight of the target video frame to extract the gait features in the target video frame.

[0136] In some alternative implementations, the feature extraction first unit includes:

[0137] The first subunit of the first feature extraction unit is used to determine the weights of each video frame using the following formula:

[0138]

[0139] Where, α t The weights L used to characterize the image in frame t. t L is used to characterize the illumination intensity of the t-th frame image. opt The preset light intensity is used to characterize the light sensitivity adjustment parameter, and σ is used to characterize the light sensitivity adjustment parameter.

[0140] In some alternative implementations, the feature extraction first unit includes:

[0141] The second subunit of the first feature extraction unit is used to extract gait features using the following formula:

[0142]

[0143] Among them, GEI adj (x,y) represents the grayscale value of the gait energy map at pixel coordinates (x,y), N represents the number of image frames used to calculate the gait energy map, and α t The weights C used to characterize the image in frame t are: t (x,y) is used to characterize the pixel value at pixel coordinates (x,y) of the image in frame t.

[0144] In some optional implementations, the data acquisition and feature extraction module 301 includes:

[0145] The confidence calculation unit is used to determine the weights of gait features and facial features based on the stability of gait features and facial features; to determine the confidence of gait features based on the product of matching degree and gait feature weights; and to determine the confidence of facial features based on the product of similarity and facial feature weights.

[0146] In some optional implementations, the confidence calculation unit includes:

[0147] The confidence calculation subunit is used to determine the gait feature weights using the following formula:

[0148]

[0149] Among them, w g Used to characterize gait feature weights Error variance estimates used to characterize facial features in video frames. Error variance estimates used to characterize gait features in video frames; facial feature weights are determined using the following formula:

[0150]

[0151] Among them, w f Used to characterize facial feature weights.

[0152] In some alternative implementations, the feature fusion module 302 includes:

[0153] The first unit of the feature fusion module is used to obtain the recognition result using the following formula:

[0154]

[0155] Where m(A) is used to characterize the confidence function value of the target belonging to category A in the recognition result, A is used to characterize the category of the target, B is used to characterize the gait recognition result, C is used to characterize the face recognition result, m1(B) is used to characterize the gait feature confidence, m2(C) is used to characterize the face feature confidence, and K is used to characterize the conflict coefficient.

[0156] In some alternative implementations, the control module 303 includes:

[0157] The first unit of the control module is used to obtain the behavioral intent using the following formula:

[0158]

[0159] Among them, V *(s) represents the target's behavioral intention in state s, s represents the target's state vector, a represents the target's behavior in state s, R(s,a) represents the reward value of the target performing behavior a in state s, γ represents the discount factor, s′ represents the next state the target transitions to after performing behavior a in state s, P(s′|s,a) represents the probability of transitioning to state s′ after performing behavior a in state s, and V * (s′) is used to characterize the target’s behavioral intention in state s′.

[0160] In some alternative implementations, the control module 303 includes:

[0161] The second control module is used to control the smart door to unlock if the identification result indicates that the target is an authorized user and the behavioral intent indicates that the target's behavior is "coming home"; if the identification result indicates that the target is a visitor and the behavioral intent indicates that the target's behavior is "visiting", the smart door sends a message to the authorized user asking whether to unlock, and controls the smart door to unlock or remain locked based on the authorized user's feedback to the message; if the identification result indicates that the target is a courier and the behavioral intent indicates that the target's behavior is "delivering a package", the smart door controls the package delivery slot to open.

[0162] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0163] In this embodiment, the device for intelligent door control is presented in the form of a functional unit. Here, a unit refers to an application-specific integrated circuit (ASIC) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0164] This invention also provides a computer device having the above-described features. Figure 3 The device shown is for intelligent door control.

[0165] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 4As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a graphical user interface on an external input / output device (such as a display device coupled to the interface). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 4 Take a processor 10 as an example.

[0166] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0167] The aforementioned memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0168] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0169] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0170] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.

[0171] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include a display device, auxiliary lighting device (e.g., light-emitting diode), and haptic feedback device (e.g., vibration motor). The aforementioned display device includes, but is not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0172] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0173] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0174] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for controlling intelligent doors, characterized in that, The method includes: Acquire video information within the field of view of the smart door, and extract gait features and facial features from the video information; determine the confidence level of the gait features based on the matching degree between the gait features and the gait features in the gait feature database; determine the confidence level of the facial features based on the similarity between the facial features and the facial features in the facial feature database. By fusing the gait feature confidence scores and the facial feature confidence scores, the target recognition result in the video is obtained; The target's behavioral intent is identified, and the smart door is controlled based on the identification result and the behavioral intent.

2. The method according to claim 1, characterized in that, The extraction of gait features from the video information includes: The weight of each video frame is determined based on the illumination intensity of each video frame in the video information. The target video frame and its weights are fused to extract gait features from the target video frame.

3. The method according to claim 2, characterized in that, Determining the weight of each video frame based on the illumination intensity of each video frame in the video information includes: The weights of each video frame are determined using the following formula: Where, α t The weights L used to characterize the image in frame t. t L is used to characterize the illumination intensity of the t-th frame image. opt The preset light intensity is used to characterize the light sensitivity adjustment parameter, and σ is used to characterize the light sensitivity adjustment parameter.

4. The method according to claim 2 or 3, characterized in that, The step of fusing the target video frame and its weights to extract gait features from the target video frame includes: The gait features are extracted using the following formula: Among them, GEI adj (x, y) represents the grayscale value of the gait energy map at pixel coordinates (x, y), N represents the number of image frames used to calculate the gait energy map, and α t The weights C used to characterize the image in frame t are: t (x,y) is used to characterize the pixel value at pixel coordinates (x,y) of the image in frame t.

5. The method according to claim 1, characterized in that, The process of determining gait feature confidence based on the matching degree between the gait features and gait features in the gait feature database, and determining facial feature confidence based on the similarity between the facial features and facial features in the facial feature database, includes: Based on the stability of the gait features and facial features, the weights of the gait features and facial features are determined. The confidence level of the gait feature is determined based on the product of the matching degree and the gait feature weights. The confidence level of the facial feature is determined based on the product of the similarity and the facial feature weight.

6. The method according to claim 5, characterized in that, The determination of gait feature weights and facial feature weights based on the stability of the gait features and facial features includes: The gait feature weights are determined using the following formula: Among them, w g Used to characterize gait feature weights Error variance estimates used to characterize facial features in video frames. Error variance estimates used to characterize the gait features in the video frame; The facial feature weights are determined using the following formula: Among them, w f Used to characterize facial feature weights.

7. The method according to claim 1, characterized in that, The fusion of the gait feature confidence and the facial feature confidence to obtain the target recognition result in the video includes: The recognition result is obtained using the following formula: Wherein, m(A) is used to characterize the confidence function value of the target belonging to category A in the recognition result, A is used to characterize the category of the target, B is used to characterize the gait recognition result, C is used to characterize the facial recognition result, m1(B) is used to characterize the confidence of the gait feature, m2(C) is used to characterize the confidence of the facial feature, and K is used to characterize the conflict coefficient.

8. The method according to claim 1, characterized in that, The identification of the target's behavioral intent includes: The behavioral intent can be obtained using the following formula: Among them, V * (s) represents the target's behavioral intention in state s, s represents the target's state vector, a represents the target's behavior in state s, R(s,a) represents the reward value for the target to perform behavior a in state s, and γ represents the discount factor. ′ P(s) is used to characterize the next state that the target transitions to after performing behavior a in state s. ′ |s,a) is used to characterize the probability of transitioning to state s′ after performing behavior a in state s, and V*(s′) is used to characterize the behavioral intention of the target in state s′.

9. The method according to claim 1, characterized in that, The control of the smart door based on the recognition result and the behavioral intent includes: If the recognition result indicates that the target is an authorized user, and the behavioral intent indicates that the target's behavior is to go home, then control the smart door to unlock; If the identification result indicates that the target is a visitor, and the behavioral intent indicates that the target's behavior is a visit, the smart door sends a message to the authorized user asking whether to unlock, and controls the smart door to unlock or remain locked based on the authorized user's feedback on the message. If the identification result indicates that the target is a courier, and the behavioral intent indicates that the target's behavior is to deliver a package, the package delivery port is opened based on the smart door control.

10. A device for intelligent door control, characterized in that, The device includes: The data acquisition and feature extraction module is used to acquire video information, extract gait features and facial features from the video information; determine the confidence level of the gait features based on the matching degree between the gait features and the gait features in the gait feature library; and determine the confidence level of the facial features based on the similarity between the facial features and the facial features in the facial feature library. The feature fusion module is used to fuse the gait feature confidence and the facial feature confidence to obtain the target recognition result in the video; The control module is used to identify the target's behavioral intent and control the smart door based on the identification result and the behavioral intent.

11. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the method for intelligent door control as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method for intelligent door control as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Access control system with integration of face recognition and gait recognition

    CN109920111A

  • Intelligent access control security method, device and system and storage medium

    CN112017345A

  • Access control method, device, medium and equipment

    CN117765652A

  • Webpage multi-modal information extraction method and system based on DS evidence theory

    CN118193816A

  • Identity recognition method and equipment based on access control and medium

    CN118212720A