Visual doorbell based on face recognition

By integrating a local recognition engine and a cloud-based collaborative module, and combining liveness detection and multimodal authentication, the system solves the security and user experience problems of traditional video doorbells, and realizes a fast, secure, and privacy-protected smart doorbell system.

CN122116510APending Publication Date: 2026-05-29HUNAN HONGGUANG ELECTRONIC TECH DEV CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN HONGGUANG ELECTRONIC TECH DEV CO LTD
Filing Date
2025-10-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional video doorbells have shortcomings in security and intelligent interaction. Network latency leads to delayed response, they are vulnerable to spoofing attacks, lack a tiered handling mechanism, have a poor user experience, rigid data processing, insufficient power supply system independence, and limited device collaboration capabilities.

Method used

It adopts a dual architecture of local recognition engine and cloud collaboration module, integrates liveness detection function, supports dynamic response strategy, combines voiceprint recognition for multimodal authentication, realizes privacy-encrypted transmission, and links with smart home devices.

Benefits of technology

It achieves rapid identity recognition, resists spoofing attacks, improves security and response efficiency, adapts to different user needs, optimizes recognition accuracy, meets privacy protection requirements, and ensures collaborative operation of devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116510A_ABST
    Figure CN122116510A_ABST
Patent Text Reader

Abstract

The application relates to the field of visual technology, in particular to a visual doorbell based on face recognition, which comprises a doorbell body, the front of the doorbell body is provided with a camera and a doorbell button, the doorbell body is provided with a doorbell system, the doorbell system comprises a face collection module, the face collection module is arranged on an outdoor unit and is used for capturing a visitor's face image in real time, a local recognition engine, a built-in pre-stored face database, feature extraction and identity matching of the collected image, a dynamic response unit, a differentiated response strategy is triggered according to a recognition result, when the recognition result is that the visitor is an authorized user, the dynamic response unit automatically unlocks the access control and pushes personalized welcome information to an indoor terminal, when the recognition result is that the visitor is an unauthorized person, the dynamic response unit starts real-time alarm and records a visitor behavior video, a cloud collaborative module, data synchronization with the local recognition engine, remote update of the face database and the recognition algorithm are supported, the visual doorbell based on face recognition automatically enhances the call volume and the interface display brightness for the old users, and a voice assistant is activated to guide the children.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video doorbell technology, specifically a video doorbell based on facial recognition. Background Technology

[0002] Traditional systems have significant shortcomings in security and intelligent interaction. Existing access control solutions generally rely on cloud servers for facial recognition, and network latency can easily lead to delayed responses, potentially causing recognition failures, especially in areas with unstable network signals. Most systems have weak defenses against spoofing attacks; printed photos or pre-recorded videos can easily deceive the basic facial verification module, resulting in significant security vulnerabilities. Door lock control functions typically employ a simple binary strategy: unlocking upon successful recognition and denying access upon failure, lacking a tiered handling mechanism based on different risk levels.

[0003] In terms of user experience, mechanical interaction methods struggle to meet diverse needs. Elderly users often struggle to hear conversations due to insufficient volume, and children are easily confused by complex interfaces; however, current technologies rarely consider the usability of these special groups. System updates rely on firmware updates from manufacturers, preventing users from customizing welcome strategies or integrating smart home devices according to their family's specific needs. Adding new visitors is cumbersome, requiring manual photo uploads and waiting for cloud synchronization, resulting in low efficiency.

[0004] The rigidity of data processing further restricts the system's ability to evolve. Most solutions use static face databases, which cannot optimize recognition accuracy based on dynamic factors such as changes in ambient lighting and user age. The problem of increasing false recognition rates after long-term use lacks a self-correcting mechanism. Privacy protection designs are also inadequate; directly uploading raw face images to the cloud poses a risk of leakage and does not comply with increasingly stringent data security regulations.

[0005] There is also room for improvement in the independence of the power supply system and the ability of equipment to work together. Components such as cameras, door locks, and indoor terminals in traditional video doorbells often require separate wiring, and there is a lack of real-time diagnosis and emergency measures in the event of power failure. Although some solutions propose two-wire integrated power supply technology, they do not solve the problem of intelligent decision-making and collaboration between devices based on facial recognition, such as how to maintain core security functions in the event of power anomalies.

[0006] To address the above issues, there is an urgent need to build a video doorbell system that integrates local rapid response, liveness attack defense, dynamic policy adaptation, and privacy-encrypted transmission, in order to balance security strength and user-friendly experience.

[0007] Therefore, we propose a video doorbell based on facial recognition. Summary of the Invention

[0008] One of the technical problems this application aims to solve is the urgent need to build a video doorbell system that integrates local rapid response, liveness attack defense, dynamic policy adaptation, and privacy-encrypted transmission, in order to balance security strength and user-friendly experience.

[0009] To address the aforementioned technical issues, this application provides a video doorbell based on face recognition, comprising a doorbell body, a camera and a doorbell button on the front of the doorbell body, and a doorbell system including a face capture module deployed in an outdoor unit for real-time capture of visitor facial images. The local recognition engine has a built-in pre-stored face database to extract features and match identities from captured images. The dynamic response unit triggers a differentiated response strategy based on the recognition results, including: When an authorized user is identified, the access control system will automatically unlock and a personalized welcome message will be pushed to the indoor terminal. When an unauthorized person is identified, a real-time alarm is triggered and video of the visitor's behavior is recorded; The cloud-based collaboration module synchronizes data with the local recognition engine, supporting remote updates to the face database and recognition algorithms.

[0010] In some embodiments, the local recognition engine integrates a liveness detection function, which analyzes facial micro-expressions and three-dimensional contour features to reject non-liveness attacks such as photos and videos.

[0011] In some embodiments, the dynamic response unit supports custom strategies, including: automatically increasing the indoor unit's call volume and brightening the display interface when an elderly user is detected; and triggering the voice assistant to provide interactive guidance when a child is detected.

[0012] In some embodiments, the cloud collaboration module has incremental learning capabilities and automatically optimizes the recognition confidence threshold of the local face database based on the identity of new visitors manually confirmed by the user.

[0013] In some embodiments, a multimodal fusion unit is also included, which combines voiceprint recognition with facial feature association verification, and enables secondary voiceprint authentication when the confidence level of facial recognition is lower than a preset threshold.

[0014] In some embodiments, voiceprint authentication uses a microphone array to directionally pick up visitor voice, extracts voiceprint features, and matches them with the bound user's voice template.

[0015] In some embodiments, the real-time alarm includes a tiered mechanism: when a person is identified as a high-risk individual, the location and video clips are immediately sent to a preset security contact; when a stranger is identified as a visitor, a notification is pushed to the user's mobile device and automatic recording and storage are initiated.

[0016] In some embodiments, the cloud collaboration module supports a privacy protection mode, where facial feature encryption is performed locally before uploading de-identified data, and the cloud only returns the recognition result label.

[0017] In some embodiments, the personalized welcome message includes at least one of user name recognition broadcast, indoor temperature and humidity prompts, and schedule reminders.

[0018] In some embodiments, the system is linked to a smart home gateway, and when a family member is detected returning home, the system automatically turns on the preset modes of the entryway lighting and air conditioning.

[0019] This invention has at least the following beneficial effects: 1. Leveraging a dual architecture of a local recognition engine and a cloud-based collaborative module, the system enables rapid processing and remote updates of facial information, overcoming the limitation of traditional doorbells that cannot effectively identify individuals offline. The local recognition engine integrates liveness detection, effectively resisting photo or video spoofing attacks by analyzing facial micro-expressions and 3D contour features, significantly improving the security of the access control system. The dynamic response unit triggers differentiated strategies based on different recognition results. For example, it automatically unlocks the door for authorized users and pushes personalized welcome messages containing the user's name and indoor environment information, while activating a tiered alarm mechanism for unauthorized personnel. When high-risk individuals are identified, their location and video clips are immediately sent to preset contacts; for ordinary strangers, notifications and recording are triggered, thus balancing response efficiency and security accuracy.

[0020] 2. For elderly users, the system automatically increases call volume and interface brightness; when interacting with children, it activates the voice assistant for guidance, improving the user experience for special groups. The cloud-based collaboration module supports an incremental learning mechanism, dynamically adjusting the recognition confidence threshold of the local facial database based on manually confirmed new visitor information, enabling continuous system optimization. When facial recognition confidence is insufficient, the system uses voiceprint authentication as a supplementary verification method. Voice is collected directionally through a microphone array and matched with a bound template to form multimodal fusion authentication, reducing the false recognition rate. All data uploaded to the cloud undergoes local facial feature encryption; only anonymized feature values ​​are transmitted, and the cloud provides recognition result labels without retaining the original data, fundamentally mitigating privacy risks.

[0021] 3. When a family member is detected returning home, the system automatically activates the entryway lighting and preset air conditioning modes, achieving a seamless homecoming experience. The entire system reduces reliance on the cloud and shortens response latency through localized data processing, while leveraging cloud-based algorithm iterations to enhance recognition accuracy, forming a dynamically optimized technological closed loop. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the video doorbell of the present invention; Figure 2 This is a schematic diagram of the doorbell system of the present invention; Figure 3 This is a schematic diagram of the operation process of the doorbell system of the present invention.

[0023] In the diagram, 100 is the doorbell body; 101 is the camera; and 102 is the doorbell button. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Example 1, see Figure 1 The present invention provides a technical solution: a video doorbell based on face recognition, including a doorbell body 100, a camera 101 and a doorbell button 102 arranged on the front of the doorbell body 100, and a doorbell system including a face acquisition module deployed in an outdoor unit for real-time capture of visitor facial images. The local recognition engine has a built-in pre-stored face database to extract features and match identities from captured images. The dynamic response unit triggers differentiated response strategies based on the recognition results, including: when the user is identified as an authorized user, automatically unlocking the access control and pushing a personalized welcome message to the indoor terminal; when the user is identified as an unauthorized person, activating a real-time alarm and recording video of the visitor's behavior. The cloud-based collaboration module synchronizes data with the local recognition engine, supporting remote updates to the face database and recognition algorithms.

[0026] The local recognition engine integrates liveness detection, which analyzes facial micro-expressions and 3D contour features to reject non-liveness attacks such as photos and videos.

[0027] The dynamic response unit supports custom strategies, including: automatically increasing the indoor unit's call volume and brightening the display when an elderly user is detected; and triggering the voice assistant to provide interactive guidance when a child is detected.

[0028] Specifically, the camera 101 on the front of the doorbell body 100 acts as a physical sensing terminal responsible for acquiring raw images. Its direct connection with the built-in face recognition system avoids the delay caused by image transfer in traditional solutions. The face acquisition module continuously monitors the activity range in front of the door, and immediately initiates high-definition capture once a facial outline is detected, ensuring that effective facial data can still be obtained in complex lighting conditions.

[0029] The core design is based on the local recognition engine's built-in face database, which operates independently on the device. This localized processing completely decouples the authentication process from network dependencies, enabling matching even when offline. Liveness detection is integrated into the local engine, effectively distinguishing between real faces and attacks on planar media through parallel computation of facial micro-expression changes and 3D depth information. This design directly addresses the vulnerability of traditional doorbells to spoofing, blocking unauthorized entry at its source.

[0030] The dynamic response unit, acting as the decision-making hub, embodies scenario-based thinking. For identified users, it goes beyond simply opening the door; instead, it triggers differentiated actions based on user attributes. Upon recognition of elderly users, it simultaneously increases call volume and interface brightness, addressing interaction barriers for those with hearing or vision impairments. For children, it activates voice assistant interaction guidance, preventing misjudgments caused by improper operation by young users. This special group adaptation strategy significantly improves efficiency in home settings by reducing manual intervention.

[0031] The cloud-based collaboration module employs a two-way incremental synchronization mechanism, allowing users to remotely add new visitor information via mobile devices while also silently deploying cloud-optimized algorithm models locally. This design retains the low-latency advantages of local processing while leveraging the powerful computational evolution capabilities of the cloud. When facial recognition confidence fluctuates, the system automatically adjusts threshold parameters to maintain recognition accuracy, addressing the issue of recognition rate decay caused by long-term use.

[0032] The doorbell button 102 retains the traditional triggering method as a backup channel, allowing manual activation of the system when the face capture module fails. Encrypted data bus transmission is used between modules; facial feature values ​​are anonymized before entering the cloud channel, while the original image remains in local storage. This data flow design complies with privacy regulations and meets the dual requirements of real-time response and continuous evolution in home security scenarios.

[0033] Example 2: The cloud-based collaboration module has incremental learning capabilities and automatically optimizes the recognition confidence threshold of the local face database based on the identity of new visitors manually confirmed by the user.

[0034] It also includes a multimodal fusion unit that combines voiceprint recognition with facial feature association verification. When the confidence level of facial recognition is lower than a preset threshold, secondary voiceprint authentication is enabled.

[0035] Voiceprint authentication uses a microphone array to directionally pick up visitor voices, extracts voiceprint features, and matches them against the bound user's voice template. Real-time alerts include a tiered mechanism: When identified as a high-risk individual, the system immediately sends the location and video clips to the preset security contact; when identified as an unknown visitor, a push notification is sent to the user's mobile device and automatic recording and storage are initiated.

[0036] Users can manually input facial features or associate information with a legal and compliant personnel database through authorized terminals to identify high-risk individuals. The system will only execute corresponding alarm policies for user-defined high-risk individuals to ensure the legality and security of the function.

[0037] Specifically, the cloud-based collaborative module of the video doorbell system employs an incremental learning mechanism to address the issue of database stagnation. After a user manually confirms a new visitor's identity, the system automatically extracts common facial feature parameters and dynamically adjusts the recognition confidence threshold range of the local database. This design allows the doorbell to adapt to changes in family members' appearances, such as growing a beard or changing glasses, avoiding the tedious operation of frequent manual calibration. The database optimization process is executed silently in the background, maintaining the real-time performance of local recognition while significantly reducing the false rejection rate through flexible threshold adjustment.

[0038] The multimodal fusion unit provides an auxiliary verification channel for critical scenarios. When rainy weather or side lighting causes fluctuations in facial matching confidence, voiceprint authentication is automatically activated as a secondary verification mechanism. A directional microphone array uses beamforming technology to filter environmental noise and accurately extract the frequency resonant features of the visitor's voice. The association between the voiceprint template and facial information ensures the consistency of the two factors; for example, a registered user's unique husky voice is included in the recognition elements. This composite verification mode effectively compensates for the physical limitations of single biometric identification, especially significantly improving the recognition success rate when visitors are wearing masks or looking down.

[0039] The tiered alarm system achieves precise security through risk prediction. Ordinary strangers trigger basic defense strategies, continuously recording video to preserve evidence for future tracing. High-risk individuals are identified, initiating a multi-threaded emergency response, simultaneously transmitting the visitor's geographic coordinates to a pre-set contact person with encrypted data. Location data is combined with information from the doorbell's built-in BeiDou module, and video clips are prioritized for segments exhibiting abnormal behavior. This threat-level-based triage mechanism avoids overreaction while ensuring rapid suppression of truly dangerous targets.

[0040] Privacy protection is ensured throughout the entire data processing workflow. All feature data uploaded to the cloud is encrypted using national cryptographic algorithms. After comparison, the cloud immediately destroys the feature file and only returns the result label. The local database employs domain-based storage technology, storing voiceprint templates and facial features separately and performing binary obfuscation on both. This design satisfies cross-modal verification requirements while eliminating the possibility of unauthorized restoration of original biometric information.

[0041] Contextualized decision-making enhances the system's fault tolerance. Incremental learning maintains long-term stability in recognition accuracy, two-factor authentication overcomes verification challenges in critical states, and tiered response achieves optimal allocation of security resources. Ultimately, this achieves the core goal of reliable operation in complex home environments.

[0042] Example 3: The cloud-based collaboration module supports a privacy protection mode. Facial features are encrypted locally before being uploaded as de-identified data; the cloud only returns the recognition result label. Personalized welcome messages include at least one of the following: user name recognition announcement, indoor temperature and humidity alerts, and schedule reminders. The system is linked to a smart home gateway; when a family member returns home, the system automatically turns on the entryway lighting and air conditioning preset modes.

[0043] Specifically, the privacy protection design of the video doorbell system follows the principle of data minimization. Facial features are vectorized and encrypted locally, generating irreversible de-identified feature codes before being transmitted to the cloud. The cloud server only performs feature code comparison tasks, returning a binary label of "authorized / unauthorized," without ever touching the original facial image. This technical approach fundamentally blocks the risk of facial data leakage, meets the compliance requirements of EU GDPR and other regulations on biometric processing, and reduces cloud storage pressure.

[0044] Personalized announcements enhance the interactive experience through multi-source data fusion. After identifying the user, the system automatically links to a personalized configuration library stored on the local chip and calls a pre-recorded name pronunciation snippet. Real-time temperature and humidity data collected by environmental sensors is converted into natural language descriptions via edge computing and dynamically concatenated with the user's schedule reminders to form complete sentences. This design transforms mechanical security equipment into a home information hub, simultaneously providing environmental status and task reminders when family members return home, reducing subsequent operational steps.

[0045] The home automation mechanism relies on a trusted execution environment between devices. After the doorbell identifies a family member, it sends a trigger command to the smart home gateway via an encrypted communication protocol. The gateway verifies the command signature and then activates a preset scene. The entryway lighting uses a gradual brightening mode to avoid glare, while the air conditioner selects either energy-saving or comfort mode based on historical usage data. This entire process avoids the cumbersome steps required by traditional solutions that necessitate manual operation of a mobile app, achieving a truly seamless experience.

[0046] The system's overall operation embodies a balance between edge intelligence and privacy computing. The core facial recognition algorithm is deployed within a secure chip on the doorbell itself, ensuring that critical data does not leave the device. The cloud only undertakes auxiliary incremental learning and model optimization tasks, processing only encrypted, non-sensitive data. User control over the smart home system remains with the local gateway, maintaining basic connectivity even during internet outages. This architecture ensures the security of biometric data while creating a natural human-centered interactive experience through efficient collaboration between devices.

[0047] Example 4: The operation of a face recognition-based video doorbell goes through the following four stages: The initial triggering phase begins when a visitor approaches the doorbell's monitoring range, automatically activating the outdoor camera's face capture program. Infrared sensors and dynamic detection algorithms filter out invalid targets (such as moving objects or animals), and once a facial outline is confirmed, high-definition image acquisition begins. Simultaneously, the traditional doorbell button remains available as a backup physical trigger.

[0048] Next comes the biometric verification stage, where raw image data is directly connected to the local recognition engine for two levels of security verification. First, liveness detection is performed, analyzing dynamic changes in eye reflections and differences in nose bridge height to intercept non-liveness attacks. After passing liveness verification, the system extracts a 128-dimensional facial feature vector, performs similarity matching with a local encrypted database, and generates a confidence score.

[0049] Next comes the decision-making and triage phase. If the confidence level is greater than or equal to a preset threshold (e.g., 98%), the authorized user response process is initiated. The system queries the user's attribute tags (e.g., elderly / child), and the dynamic response unit performs associated operations: it calls the pre-stored voice module to play a welcome message containing the user's name, and simultaneously sends ambient temperature and humidity information to the indoor terminal. At the same time, it sends an encrypted command to the smart home gateway via the home LAN to activate the entryway lights to gradually brighten and the air conditioner to a preset mode.

[0050] If the confidence level is less than the security threshold but greater than or equal to the critical value (e.g., 80-98%), multimodal authentication is activated. A microphone array collects the visitor's voice commands, and a voiceprint verification unit matches the feature frequency point set. The two-factor authentication result covers the confidence level fluctuations of facial recognition. If verification is successful, the same authorization process as for high-confidence authentication is triggered.

[0051] If the confidence level is below the threshold and voiceprint verification is not enabled or has failed, the unauthorized handling process begins. The system first encrypts and compresses the facial features, then uploads the anonymized data to the cloud for expanded database matching. If the cloud-returned result is associated with a high-risk tag (such as a fugitive), an emergency protocol is immediately activated: the location coordinates are obtained via the BeiDou module, a 15-second video clip of the behavior is extracted, and encrypted and pushed to the public security platform and the security contact's mobile phone. If it is only an ordinary stranger, a notification is pushed to the user's mobile device and local loop recording is initiated.

[0052] Next comes the data collaboration phase, which automatically starts cloud collaboration every morning at midnight. An incremental learning mechanism uploads the feature vector distribution characteristics of newly confirmed users for the day and receives optimized recognition model parameters. In privacy protection mode, all transmitted data is irreversibly encrypted with feature codes, and temporary data is destroyed immediately after the model training is completed in the cloud. Users can adjust response strategy parameters (such as volume gain values ​​for the elderly) at any time via mobile devices, and configuration commands are synchronized to the local device after being digitally signed.

[0053] Finally, there's the anomaly handling mechanism. In the event of a power outage, the backup battery sustains the core identification module for 90 minutes and sends a fault alarm to the user via the gateway. During network outages, the local database maintains basic identification functions, newly added visitor data is temporarily stored in an encrypted cache, and resumes transmission after network recovery.

[0054] The entire process utilizes a modular design to ensure efficiency along the critical path. Liveness verification and face matching are completed within 500ms, while voiceprint authentication, as an auxiliary channel, adds no more than 300ms of latency. Data processing adheres to the principle of usability without visibility throughout, ensuring that raw biometric information remains completely within its domain.

[0055] The main workflow of a video doorbell is as follows: The process begins with a visitor triggering the outdoor unit. When someone approaches the doorbell's range, the outdoor facial recognition module immediately activates a high-definition camera to capture the visitor's facial image. The captured image data is directly transmitted to the local recognition engine built into the doorbell device, which uses a pre-stored facial database for real-time feature extraction and matching. During the authentication phase, the system first performs liveness detection, analyzing subtle eye movements, facial muscle movements, and differences in three-dimensional depth contours to intercept simulated attacks using photos or videos.

[0056] If the liveness verification passes, the system proceeds to the identity matching process. The local recognition engine compares the visitor's facial features with the authorized user database to generate a matching confidence score. For identified users whose confidence score reaches a set threshold, the dynamic response unit triggers a preset strategy: when a family member is identified, the system links with the smart home gateway to automatically turn on the entryway lights and switch the air conditioner to a preset mode; for elderly users, the system triggers voice channel gain and interface brightness enhancement functions, while the indoor terminal broadcasts a personalized welcome message, incorporating the user's name, indoor temperature and humidity, and to-do reminders. If the identified user is a child, the voice assistant automatically intervenes to guide the interaction process.

[0057] For visitors whose matches fail, the system activates a multi-level response mechanism. For ordinary strangers, real-time video recording and storage are triggered, and a notification is pushed to the linked mobile device. Users can remotely view the video clips to decide whether to authorize temporary access. When facial features successfully match a high-risk database in the cloud (e.g., for fugitives), the system immediately activates an emergency protocol: encrypting and sending the visitor's location coordinates and a 15-second video of their behavior to pre-set security contacts and the public security system. Simultaneously, the outdoor unit emits an alarm sound to deter suspicious targets.

[0058] When the facial recognition confidence level reaches a critical range, the system seamlessly switches to multimodal authentication. A microphone array directionally collects the visitor's voice commands, and the voiceprint verification unit extracts frequency features and compares them with the voice template of the bound user. The two-factor authentication result is updated to the local database in real time. If confirmed as a new authorized visitor, the incremental learning mechanism automatically adjusts the person's recognition threshold parameters.

[0059] Privacy protection logic is implemented throughout the data processing process. All facial features are encrypted locally before being uploaded to the cloud. The cloud-based collaborative module only provides the de-identified recognition result labels and cannot access the original image. User-defined settings are implemented through indoor terminals or mobile applications, such as customizing interaction strategies for different groups or adjusting door lock opening and closing sensitivity. When a local algorithm update is required, the cloud pushes the optimized model to the device, automatically completing silent deployment at night. This closed-loop system achieves a dynamic balance between security accuracy and user experience through continuous algorithm iteration and scenario-based decision-making. It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0060] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

Claims

1. A video doorbell based on face recognition, comprising a doorbell body (100), characterized in that: The doorbell body (100) is equipped with a camera (101) and a doorbell button (102) on the front. The doorbell body (100) is equipped with a doorbell system, including: a face acquisition module, which is deployed in the outdoor unit and is used to capture visitor facial images in real time. The local recognition engine has a built-in pre-stored face database to extract features and match identities from captured images. The dynamic response unit triggers a differentiated response strategy based on the recognition results, including: When an authorized user is identified, the access control system will automatically unlock and a personalized welcome message will be pushed to the indoor terminal. When an unauthorized person is identified, a real-time alarm is triggered and video of the visitor's behavior is recorded; The cloud-based collaboration module synchronizes data with the local recognition engine, supporting remote updates to the face database and recognition algorithms.

2. A video doorbell based on face recognition according to claim 1, characterized in that: The local recognition engine integrates a liveness detection function, which analyzes facial micro-expressions and three-dimensional contour features to reject non-liveness attacks such as photos and videos.

3. A video doorbell based on face recognition according to claim 1, characterized in that: The dynamic response unit supports custom strategies, including: When an elderly user is detected, the indoor unit's call volume is automatically increased and the display screen is brightened. When a child is detected, the voice assistant is triggered to provide interactive guidance.

4. A video doorbell based on face recognition according to claim 3, characterized in that: The cloud-based collaborative module has incremental learning capabilities, automatically optimizing the recognition confidence threshold of the local face database based on the identity of new visitors manually confirmed by the user.

5. A video doorbell based on face recognition according to claim 1, characterized in that: It also includes a multimodal fusion unit that combines voiceprint recognition with facial feature association verification. When the confidence level of facial recognition is lower than a preset threshold, secondary voiceprint authentication is enabled.

6. A video doorbell based on face recognition according to claim 5, characterized in that: The voiceprint authentication uses a microphone array to pick up the visitor's voice in a directional manner, extracts voiceprint features, and matches them with the bound user's voice template.

7. A video doorbell based on face recognition according to claim 1, characterized in that: The real-time alarm includes a tiered mechanism: When a person is identified as a high-risk individual, their location and video clips are immediately sent to a pre-set security contact. When a visitor is identified as an unknown visitor, a push notification is sent to the user's mobile device and automatic video recording and storage are initiated.

8. A video doorbell based on face recognition according to claim 1, characterized in that: The cloud-based collaboration module supports a privacy protection mode, where facial features are encrypted locally before being uploaded to de-identify the data, and the cloud only returns the recognition result labels.

9. A video doorbell based on face recognition according to claim 1, characterized in that: The personalized welcome message includes at least one of the following: user name recognition broadcast, indoor temperature and humidity prompts, and schedule reminders.

10. A video doorbell based on face recognition according to claim 1, characterized in that: The system is linked to a smart home gateway, and automatically turns on the entryway lighting and air conditioning preset modes when a family member is detected returning home.