Doorbell hanging up method, device, equipment and computer-readable storage medium

By identifying the video frame image of the smart doorbell, judging the status of the door body and the human body, and automatically launching or ending the call, the problem of insensible and automated calls in the prior art is solved, and the user experience is improved.

CN115701096BActive Publication Date: 2025-08-26CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110805876.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-16
Publication Date
2025-08-26
Estimated Expiration
2041-07-16

AI Technical Summary

Technical Problem

The existing smart doorbell continues to call when the visitor leaves or the door is accidentally touched, resulting in the call being not intelligent and automated enough, which interferes with the life of the household head.

Method used

By acquiring video data from the image acquisition device, identifying multiple video frame images, determining the status of the target object (door body and human body), and automatically initiate or ending the door opening call according to preset conditions.

Benefits of technology

It realizes an intelligent and automated call process, avoids unnecessary call interference and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115701096B_ABST
    Figure CN115701096B_ABST
Patent Text Reader

Abstract

The present application discloses a doorbell hanging up method, device, equipment and computer-readable storage medium, the method comprising: in response to a pressing operation on the doorbell, acquiring video data currently acquired by an image acquisition device, and acquiring multiple video frame images based on the video data; when it is determined based on multiple video frame images within a preset time length that a call condition is met, initiating a door opening call to a target terminal; performing image processing on the multiple video frame images to determine the state of a target object, the target object including at least one of a door body and a human body; when it is determined based on the state of the target object that a call ending condition is met, ending the door opening call, and being able to automatically determine whether to initiate a door opening call before the call, and automatically determine whether to end the door opening call during the door opening call, thereby realizing intelligent calling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communications, and relates to, but is not limited to, a doorbell hanging up method, device, equipment, and computer-readable storage medium. Background Art

[0002] As IoT technology matures, sales of smart devices continue to rise. According to surveys, the smart doorbell market has a compound annual growth rate of 69%. In addition to providing video surveillance and motion detection, smart doorbells are increasingly integrating video calling features. This feature is primarily targeted at elderly people living alone, children, and those on business trips. It provides real-time video intercom communication without opening the door, allowing users to predict the identity of visitors and avoid opening the door to strangers.

[0003] In the related art, there is a problem that if the visitor leaves or the door is opened during the smart doorbell call, the call will still be maintained. Furthermore, if the doorbell is touched by mistake, the call will still be initiated to the homeowner, thereby disturbing the homeowner's life. Therefore, it is urgent to solve the problem that the call process is fixed and single, and the call is not smart and automated enough. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a doorbell hanging up method, apparatus, device, and computer-readable storage medium.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides a doorbell hanging up method, comprising:

[0007] In response to a pressing operation on the doorbell, acquiring video data currently captured by an image capture device, and acquiring a plurality of video frame images based on the video data;

[0008] When the call condition is determined to be met based on multiple video frame images within a preset time, a door opening call is initiated to the target terminal;

[0009] performing image processing on the plurality of video frame images to determine a state of a target object, the target object comprising at least one of a door and a human body;

[0010] When it is determined based on the state of the target object that a call termination condition is met, the door opening call is terminated.

[0011] The present invention provides a doorbell hanging device, which includes:

[0012] a response module, configured to obtain video data currently captured by the image capture device in response to a pressing operation on the doorbell, and to obtain a plurality of video frame images based on the video data;

[0013] A calling module is used to initiate a door opening call to a target terminal when a calling condition is determined to be met based on multiple video frame images within a preset time length;

[0014] a processing module, configured to perform image processing on the plurality of video frame images to determine a state of a target object, wherein the target object includes at least one of a door and a human body;

[0015] The ending module is used to end the door opening call when it is determined that the call ending condition is met based on the state of the target object.

[0016] An embodiment of the present application provides a doorbell hanging device, the doorbell hanging device comprising:

[0017] processor; and

[0018] a memory for storing a computer program executable on the processor;

[0019] Wherein, when the computer program is executed by the processor, the above-mentioned doorbell hanging up method is implemented.

[0020] An embodiment of the present application provides a computer-readable storage medium, wherein the computer storage medium stores computer-executable instructions, and the computer-executable instructions are configured to execute the above-mentioned doorbell hanging up method.

[0021] The embodiments of the present application provide a doorbell hanging up method, device, equipment and computer-readable storage medium. In response to a pressing operation on the doorbell, the video data currently collected by the image acquisition device is obtained, and the video data is intercepted and processed to obtain multiple video frame images; then, multiple video frame images within a preset time length are intercepted, and when the multiple video frame images within the preset time length meet the call conditions, a door opening call is initiated to the target terminal. In this way, it is possible to automatically determine whether a call needs to be initiated to the target terminal based on the video frame images, thereby improving intelligence and automation; then, image processing is performed on the multiple video frame images, and the state of the target object is determined based on the image processing results, wherein the target object includes at least one of a door body and a human body; finally, when the target object meets the call end condition, the door opening call will be ended. In this way, during the call process, the state of the target object can be timely judged, and when the target state meets the call end condition, the door opening call can be automatically ended, thereby avoiding disturbing the homeowner and realizing intelligent calling. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In the drawings, which are not necessarily drawn to scale, like reference numerals may describe similar components throughout the different views.The drawings illustrate generally, by way of example and not limitation, various embodiments discussed herein.

[0023] Figure 1 A schematic diagram of an implementation flow of the doorbell hanging up method provided in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of an interactive implementation flow of the doorbell hanging up method provided in an embodiment of the present application;

[0025] Figure 3 A schematic diagram of another interactive implementation flow of the doorbell hanging up method provided in an embodiment of the present application;

[0026] Figure 4 A schematic diagram of the structure of the audio and video call architecture provided in an embodiment of the present application;

[0027] Figure 5 A schematic diagram of an implementation flow of a method for automatically ending a door opening call provided in an embodiment of the present application;

[0028] Figure 6 A schematic diagram of an implementation flow of the method for automatically hanging up a door call provided in an embodiment of the present application;

[0029] Figure 7 A schematic diagram of an implementation flow of the method for automatically hanging up a call provided in an embodiment of the present application;

[0030] Figure 8 A schematic diagram of another structure of the doorbell hanging device provided in an embodiment of the present application;

[0031] Figure 9 A schematic diagram of the structure of the doorbell hang-up device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0033] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0034] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0036] To address the problems in the related art, embodiments of the present application provide a doorbell hanging up method. The method provided in embodiments of the present application can be implemented via a computer program. When executed, the computer program completes each step of the doorbell hanging up method provided in embodiments of the present application. In some embodiments, the computer program can be executed by a processor in a doorbell hanging up device. Figure 1 This is a schematic diagram of an implementation flow of the doorbell hang-up method provided in an embodiment of the present application. The method can be applied to a server, which can establish a communication connection between the doorbell and the target terminal. The server can also make a judgment based on the received information and control whether to initiate or end the door opening call based on the judgment result. Figure 1 As shown, the method includes:

[0037] Step S101 : In response to a pressing operation on a doorbell, video data currently captured by an image capture device is acquired, and a plurality of video frame images are acquired based on the video data.

[0038] Here, the doorbell is provided with a pressable module, and the pressing operation on the doorbell can be a pressing operation on the pressable module of the doorbell; the image acquisition device can be a camera, a video camera, etc., and the image acquisition device can be integrated with the doorbell or a monitoring device independent of the doorbell. The video data can be video data of the corridor outside the door, aisle, etc.

[0039] In an embodiment of the present application, after obtaining the currently collected video data, the server may capture video frame images from the video data according to a preset period. For example, if the preset period is 1 second, one image is captured every 1 second, thereby obtaining multiple video frame images. Alternatively, video frame images may be captured from the video image non-periodically. This embodiment of the present application does not limit the implementation method for obtaining multiple video frame images based on video data.

[0040] Step S102: When it is determined based on a plurality of video frame images within a preset time period that a call condition is met, a door opening call is initiated to a target terminal.

[0041] Here, the preset duration can be 1 second, 2 seconds, etc., and the preset duration can be a default value or a custom setting value. The video frame image within the preset duration includes at least one, and the number of video frame images within the preset duration is related to the preset duration and the video capture rule. For example, when the preset duration is 2 seconds and video data is captured at a 1-second periodicity, 2 video frame images can be captured within 2 seconds. In this case, there are 2 video frame images within the preset duration.

[0042] In this embodiment of the present application, the starting point of the preset duration is consistent with the starting point of the video data. That is, the multiple video frames within the preset duration are the video frames with earlier time points among the multiple video frames acquired in step S101. Here, the target terminal can be an indoor LCD display, a mobile phone, a smart wearable device, etc., which can identify the multiple video frames within the preset duration. When at least one video frame is identified as containing a human image, the call condition is considered to be met, and a door opening call is initiated to the target terminal.

[0043] In some embodiments, if it is recognized that no human image is contained in any video frame image, it is considered that the call condition is not met, and a door opening call is not initiated to the target terminal. When the doorbell is accidentally touched, a door opening call will not be initiated to the homeowner, thereby improving the intelligence of the call.

[0044] Step S103: performing image processing on the multiple video frame images to determine the state of the target object.

[0045] Here, the target object includes at least one of a door body and a human body, and the image processing may include: image feature extraction, image recognition, image detection, etc.

[0046] Among them, when the target object is a door body, each image feature of each video frame image can be extracted, and then the various difference information between each image feature and the reference feature can be determined, wherein the reference feature is a feature screen extracted from the reference image representing the closed door body; if there is difference information greater than the difference threshold, it indicates that the door body is not in a closed state, and the state of the door body is determined to be an open state; and when each difference information is less than or equal to the difference threshold, the state of the door body is determined to be a closed state.

[0047] When the target object is a human body, image recognition can be performed on each video frame image. When it is identified that none of the video frame images contain a human body image, it is determined that the human body is in a leaving state. When it is identified that at least one video frame image contains a human body image, the video frame image containing the human body image is determined as the target video frame image. The distance information of the human body from the image acquisition device in the target video frame image is determined. Then, the direction of movement of the human body is further judged in combination with the time information of the target video frame image and the determined distance information. When it is determined that the human body is moving away from the image acquisition device, it is considered that the human body is in a leaving state. When it is determined that the distance between the human body and the image acquisition device remains almost unchanged or is moving toward the image acquisition device, it is considered that the human body is in a present state.

[0048] Step S104: When it is determined based on the state of the target object that a call termination condition is met, the door opening call is terminated.

[0049] In an embodiment of the present application, when the target object is a door, the call termination condition can be considered to be satisfied when the door is in an open state; when the target object is a human body, the call termination condition can also be considered to be satisfied when the human body is in a away state. Then, when the call termination condition is satisfied, either the door is already in an open state, indicating that the purpose of the door opening call has been achieved and there is no need to continue initiating the door opening call; or the human body is in a away state, indicating that the human body no longer needs to call the door and there is no need to continue initiating the door opening call. Based on this, an end instruction can be sent to the target terminal to cause the target terminal to stop the door opening call; when the target objects are a door and a human body, when the door is in an open state or the human body is in a away state, the call termination condition is considered to be satisfied, thereby terminating the door opening call to the target terminal.

[0050] In some embodiments, when the target object is a door, if the door is in a closed state, it can be considered that the call end condition is not met; when the target object is a human body, if the human body is in an existing state, it can also be considered that the call end condition is not met. Then, if the call end condition is not met, either the door is still in a closed state, indicating that the purpose of the door opening call has not been achieved and the door opening call still needs to be initiated; or the human body is still in an existing state, indicating that the human body still has a door opening call demand and the door opening call still needs to be initiated. In this case, it can also be determined based on the state of the target object that the call end condition is not met, and the door opening call is continued to be initiated to the target terminal even if the call end condition is not met; when the target objects are a door and a human body, if the door is in a closed state and the human body is in an existing state, it is considered that the call end condition is not met, and the door opening call is continued to be initiated to the target terminal.

[0051] An embodiment of the present application provides a doorbell hanging up method, which responds to a pressing operation on the doorbell, obtains video data currently collected by the image acquisition device, and intercepts and processes the video data to obtain multiple video frame images; then, intercepts multiple video frame images within a preset time length, and when the multiple video frame images within the preset time length meet the call conditions, initiates a door opening call to the target terminal. In this way, it is possible to automatically determine whether a call needs to be initiated to the target terminal based on the video frame images, thereby improving intelligence and automation; then, image processing is performed on the multiple video frame images, and the state of the target object is determined based on the image processing results, wherein the target object includes at least one of a door body and a human body; finally, when the target object meets the call end condition, the door opening call will be ended. In this way, during the call process, the state of the target object can be timely judged, and the door opening call can be automatically ended when the target state meets the call end condition, thereby avoiding disturbing the homeowner and realizing intelligent calling.

[0052] Based on the above embodiment, the embodiment of the present application further provides an interactive method for doorbell hang-up, which can be applied to the doorbell, server and target terminal, and can automatically end the call during the process of initiating the door opening call, such as Figure 2 As shown, the method includes:

[0053] In step S201, the doorbell sends the received pressing operation to the server.

[0054] Here, the doorbell and the server can establish a communication connection via wired or wireless communication, based on which the doorbell can send the received pressing operation to the server.

[0055] In step S202 , the server obtains the video data currently captured by the image capture device in response to the pressing operation on the doorbell, and obtains a plurality of video frame images based on the video data.

[0056] Here, the server can establish information interaction with the image acquisition device based on the received pressing operation and obtain the video data currently acquired by the image acquisition module, that is, the video data is the video data when the doorbell receives the pressing operation.

[0057] In step S203 , the server performs image recognition on each video frame image within a preset time length to obtain each second recognition result.

[0058] Here, each second recognition result can be used to indicate whether each video frame image within a preset time length contains a human body image.

[0059] In an embodiment of the present application, images can be recognized through methods such as neural networks, wavelet moments, fractal features, etc. to identify whether each video frame image contains a human body image.

[0060] In step S204 , the server determines whether each second recognition result indicates that at least one video frame image in each video frame image within a preset time length contains a human body image.

[0061] Here, if at least one video frame image contains a human image, it indicates that there is a visitor outside the door, that is, the visitor is waiting outside the door to open the door, and the process goes to step S206; if each video frame image does not contain a human image, it indicates that there is no visitor outside the door, and the doorbell pressing operation is a wrong operation, and the process goes to step S205.

[0062] Step S205: The server does not initiate a door opening call to the target terminal.

[0063] At this time, there is no visitor outside the door, and pressing the doorbell is an erroneous operation. In order not to disturb the homeowner, the server does not initiate a door opening call to the target terminal.

[0064] Step S206: The server determines that the call condition is met and initiates a door opening call to the target terminal.

[0065] At this time, there is a visitor outside the door, that is, the visitor is waiting outside the door to open the door, and the event needs to be notified to the homeowner through a door opening call. Then, the server initiates a door opening call to the target terminal.

[0066] Step S207: The target terminal initiates a door opening call.

[0067] Here, the target terminal can initiate a door opening call through ringing, vibration, etc., so that the homeowner knows the time when the visitor is outside the door and makes preliminary preparations for opening the door and making a call.

[0068] Step S208: The server obtains reference features.

[0069] Here, the reference features are features extracted from a reference image representing the door closed state. Before step S202, the server obtains a reference image of the door closed state in advance through an image acquisition device and performs feature extraction on the reference image to obtain the reference features.

[0070] In step S209 , the server extracts image features of each video frame image and determines each difference information between each image feature and a reference feature.

[0071] Here, feature extraction methods such as oriented gradient histogram and scale-invariant feature transform can be used to extract features of each video frame image, so as to obtain the image features of each video frame image; then, the absolute value of the difference between each image feature and the reference feature, the square of the difference or the Euclidean distance, etc. can be obtained respectively, and the absolute value of the difference, the square of the difference or the Euclidean distance can be determined as each difference information.

[0072] In step S210 , the server determines whether there is at least one difference information among the difference information that is greater than a difference threshold.

[0073] Here, the difference threshold can be a default value or a custom value. For example, the difference threshold can be 10, 15, etc.; if it is determined that the difference information is greater than the difference threshold, it indicates that the door body is not in a closed state, and the process goes to step S211, that is, it is determined that the door body is in an open state; if it is determined that the difference information is less than or equal to the difference threshold, it indicates that the door body is in a closed state, and the process goes to step S213 to continue to judge the state of the human body.

[0074] In step S211, the server determines that the door is in an open state.

[0075] At this time, the difference information is greater than the difference threshold, indicating that the door body is not in a closed state, and thus, the door body is in an open state.

[0076] In step S212, the server determines that the call termination condition is met and terminates the door opening call.

[0077] Here, when the door body is in the open state, the purpose of the door opening call has been achieved. Therefore, it is determined that the call end condition is met at this time, and the door opening call is ended.

[0078] In step S213 , the server performs image recognition on each video frame image to obtain each first recognition result.

[0079] Here, each first recognition result is used to indicate whether each video frame image contains a human body image. Image recognition can be performed on each video frame image using methods such as neural networks, wavelet moments, and fractal features to obtain each first recognition result for each video frame image.

[0080] In step S214 , the server determines whether each first recognition result indicates that each video frame image does not contain a human body image.

[0081] Here, if it is determined that each video frame image in each first recognition result does not contain a human image, it indicates that there is no visitor outside the door at this time, and the process goes to step S218; if it is determined that at least one video frame image in the first recognition result contains a human image, it indicates that there is still a visitor outside the door at this time, and the state of the human body is further determined, and the process goes to step S215.

[0082] In step S215 , the server determines the video frame image containing the human body image as the target video frame image.

[0083] Here, if there is a video frame image containing a human body image, then the video frame image containing the human body image is determined as a target video frame image, as the video frame image that needs to be further analyzed next.

[0084] In step S216 , the server determines the direction of movement of the human body based on the time information and distance information of each target video frame image.

[0085] Here, the information of each target video frame image includes the time information of the video frame image and the distance information of the human body from the image acquisition device. The server can obtain the time information and distance information from each target video frame image by reading the instruction. Furthermore, if the distance between the human body and the image acquisition device increases with time, the direction of human movement is determined to be moving away from the image acquisition device; if the distance between the human body and the image acquisition device remains essentially unchanged or increases with time, the direction of human movement is determined to be moving toward the image acquisition device.

[0086] In step S217, the server determines whether the human body is moving in a direction away from the image acquisition device.

[0087] Here, when the direction of movement of the human body is moving away from the image acquisition device, step S218 is entered, which indicates that the state of the human body is leaving; if the human body does not move or the direction of movement is moving towards the image acquisition device, step S206 is returned to, that is, the door opening call is continued to be initiated to the target terminal.

[0088] In step S218, the server determines that the human body is in the away state.

[0089] At this time, it is determined that there is no human image outside the door or the human body is leaving, that is, there is no visitor outside the door or the visitor is leaving, then, the state of the human body is determined to be the leaving state, indicating that the visitor has given up the visit, then, it can return to step S212, that is, determine that the call end condition is met at this time, and end the door opening call.

[0090] Through the above steps S201 to S218, the doorbell sends the received pressing operation to the server. In response to the pressing operation, the server obtains the video data currently collected by the image acquisition device and intercepts the video data into multiple video frame images. Next, the server performs image recognition on the video frame images within a preset time period. If the recognition result indicates that at least one video frame image contains a human image, the server determines that the call condition is met and initiates a door opening call to the target terminal. Then, during the call process, the server also obtains a reference feature indicating that the door is in a closed state, extracts the image features of each video frame image, and determines the difference information between each image feature and the reference feature. If the difference information is greater than the difference threshold, it is determined that the door is in an open state, and the call end condition is determined to be met, and the door opening call is terminated. If the difference information is less than or equal to the difference threshold, it is determined that the door is in a closed state, and further identifies whether the video frame image contains a human image. If the video frame image does not contain a human image or the human body is away from the image acquisition device, it is determined that the human body is in a leaving state, and the call end condition is determined to be met, and the door opening call is terminated. If the video frame image contains a human image, the door opening call is continued to be initiated to the target terminal. Therefore, before initiating the door opening call, it is possible to intelligently determine whether to initiate the door opening call based on whether the human body is in the away state, and not initiate the door opening call to the target terminal when the human body is in the away state; in addition, during the call process, it is also possible to continue to intelligently determine whether to continue to initiate the door opening call based on whether the door body is in the open state and whether the human body is in the away state, and the door opening call can be ended when the door body is in the open state or the human body is in the away state, thereby avoiding disturbing the homeowner and realizing intelligent calling.

[0091] Based on the above embodiment, the embodiment of the present application provides another doorbell hang-up interaction method, which can be applied to the doorbell, the server and the target terminal, and can automatically end the call during the call, such as Figure 3 As shown, the method includes:

[0092] Step S301: The target terminal sends a call permission instruction to the server.

[0093] Here, the call permission instruction may be a pressing instruction for “answer” or a voice “answer” instruction; then, the target terminal sends the call permission instruction to the server through the established communication connection.

[0094] Step S302: The server establishes a call connection between the doorbell and the target terminal.

[0095] Here, the server can establish a call connection between the doorbell and the target terminal by enabling the communication pin between the doorbell and the target terminal, so that the doorbell and the target terminal can conduct an audio and video call based on the call connection.

[0096] Step S303: The doorbell collects target video data through an image acquisition device during the call.

[0097] Here, the doorbell is provided with an image acquisition device. During a call, the doorbell can control the image acquisition device to acquire video data of the actual situation outside the door. Here, the video of the actual situation outside the door is recorded as target video data.

[0098] Step S304: The doorbell sends the target video data to the server.

[0099] Here, based on the existing communication connection between the doorbell and the server, the doorbell sends the collected target video data to the server.

[0100] In step S305, the server determines whether the door body is in an open state based on the target video data until a preset number of determinations is reached, and obtains each determination result.

[0101] Here, in order to ensure the accuracy of the judgment result, the target video data will be judged multiple times, that is, a preset number of judgments will be performed, wherein the preset number of judgments can be 5, 6, 7, etc.

[0102] In the embodiment of the present application, taking a judgment process as an example, the server has obtained the target video data through the above steps. Then, the server can intercept the video frame image corresponding to the target video data, extract the features of the target image, and then obtain the target difference information between the target image feature and the reference feature. When the target difference information is less than or equal to the difference threshold, the door is judged to be in the closed state; when the difference information is greater than the difference threshold, the door is judged to be in the open state. In this way, a judgment is completed and a judgment result is obtained. According to a similar judgment method, the judgment process is performed a preset number of times to obtain each judgment result.

[0103] In some embodiments, a door closer may be provided on the door. Still taking a single determination process as an example, the door closer can be in two states: closed and open. The server can directly read the state of the door closer. When the door closer is read as closed, the server determines that the door is in the closed state; when the door closer is read as open, the server determines that the door is in the open state. This completes a single determination, resulting in a single determination result. This determination process is repeated a predetermined number of times using a similar determination method to obtain each determination result.

[0104] In step S306, the server determines the number of times the door is in the open state in each judgment result.

[0105] Here, taking the preset number of determinations as 5 as an example, if the 5 determination results are all that the door is in the open state, the number of determinations determined by the server is 5.

[0106] Step S307: The server determines whether the number of times reaches a threshold.

[0107] Here, the number of times threshold is less than or equal to the preset number of times. Assuming the number of times threshold is 5, continuing with the above example, the number of times is also 5, indicating that the number of times is equal to the number of times threshold, that is, the number of times reaches the number of times threshold, and then proceeds to step S308; assuming the number of times threshold is 5, the number of times determined by the server is 3, indicating that the number of times is less than the number of times threshold, that is, the number of times does not reach the number of times threshold, and then proceeds to step S310.

[0108] Step S308: The server determines that the door is in an open state.

[0109] Here, when the number of times reaches the number threshold, it indicates that the door is in an open state.

[0110] Step S309: The server determines that the call termination condition is met and terminates the call connection.

[0111] Here, the purpose of the call connection is to open the door. When it is determined that the door is in an open state, it indicates that the purpose of the call connection has been achieved. Then, the call connection can be automatically ended, thereby improving the intelligence and automation level of the call.

[0112] Step S310: The server determines that the call termination condition is not met and maintains the call connection.

[0113] Here, when it is determined that the door is in a closed state, it indicates that the purpose of the call connection has not been achieved, so it is determined that the call termination condition is not met at this time, and the call connection is still maintained.

[0114] In some embodiments, the call connection can also be ended by:

[0115] Based on the call permission instruction sent by the target terminal, the server establishes a call connection between the doorbell and the target terminal. The server then obtains audio information from the call. The doorbell and the target terminal may be equipped with a sound collection device that can collect the audio of the call between the visitor and the homeowner and transmit the collected audio to the server. Finally, the server determines whether the audio information contains a preset target closing phrase and, if so, terminates the call. The preset target closing phrase may be phrases such as "goodbye," "see you later," or "come in," and generally appears at the end of the call. The server can use pattern matching to match the audio information with the target closing phrase to determine whether the audio information contains the target closing phrase. If so, the server terminates the call if it does. If not, the server continues the call, thereby intelligently terminating the call.

[0116] Through steps S301 to S310, the target terminal sends a call permission instruction to the server, and the server establishes a call connection between the doorbell and the target terminal based on the call permission instruction; then, during the call, the doorbell collects target video data through the image acquisition device and sends the target video data to the server; then, the server judges whether the door body is in an open state based on the received target video data, wherein the judgment is performed a preset number of times to obtain each judgment result; the server continues to determine the number of times that the door body is in an open state in each judgment result; finally, when the server judges that the number reaches the number threshold, it determines that the door body is in an open state and ends the call connection, so that the state of the door body can be accurately identified through multiple judgment results, and when it is determined that the door body is in an open state, the call connection is automatically ended, thereby improving the intelligence and automation level of the call.

[0117] Based on the above embodiment, the embodiment of the present application further provides a method for hanging up a smart doorbell, which is applied to a smart doorbell system and can realize audio and video calls. The audio and video call architecture 400 is as follows: Figure 4As shown, it includes three parts: communication terminal 401, communication platform 402 and artificial intelligence (AI) capability platform 403. Among them, communication terminal 401 can be a smart doorbell 4011 or a speaker, mobile phone 4012, etc. By linking the smart doorbell with terminals such as speakers and mobile phones, users can view the door situation in real time and have real-time conversations with visitors remotely. The communication platform 402 adopts a signaling and media separation strategy. The signaling module 4021 mainly handles call logic such as calling, talking, and hanging up, and the media module 4022 mainly processes audio and video stream data. The AI ​​capability platform mainly processes the video stream during calls or conversations and returns the results. Here, human detection 4031 can detect whether there is a person in the video stream through a deep learning algorithm, and open / close door detection 4032 can detect whether the door is opened through an image processing algorithm. The image processing algorithm can compare video images over a period of time. If the image difference exceeds a threshold, it is considered that the door is opened, that is, the door has been moved during the period.

[0118] In some embodiments, the communication platform and the AI ​​capability platform may be integrated into the same server or into different servers.

[0119] In the actual use of the smart doorbell, the hang-up method provided in the embodiment of the present application can realize the following three hang-up processes: the first is that when the smart doorbell initiates a door opening call, it can automatically end the door opening call; the second is that when the communication platform initiates a door opening call to the called terminal, it can automatically end the door opening call; the third is that during the call between the caller and the called party, the call is automatically hung up.

[0120] The first one is that when the smart doorbell initiates a door opening call, it can automatically end the door opening call. In the embodiment of the present application, the first hang-up process is as follows: Figure 5 As shown:

[0121] Step S501, start.

[0122] Step S502: The visitor presses the smart doorbell button.

[0123] Here, the smart doorbell is able to detect the action of pressing itself.

[0124] Step S503: The audio and video call process is activated, and the smart doorbell initiates the call.

[0125] After the user presses the smart doorbell, the smart doorbell can activate its own audio and video call process based on the pressing operation, prepare for subsequent audio and video calls, and the smart doorbell also initiates a door opening call to the communication platform.

[0126] Step S504: The communication platform receives a call request.

[0127] Here, after receiving the door opening call request, the communication platform does not send it to the device first, but continues to execute the following step S505.

[0128] In step S505, the communication platform captures the video images within a preset time period and sends them to the artificial intelligence capability platform.

[0129] Here, the preset duration can be 1 second, 2 seconds, etc. The communication platform can capture the video image within 1 second and send it to the AI ​​capability platform.

[0130] Step S506: The artificial intelligence capability platform returns the detection result.

[0131] Here, the AI ​​capability platform uses intelligent algorithms to detect whether there is someone at the door and returns the result to the communication platform.

[0132] In step S507, the communication platform determines whether there is anyone within the monitoring range based on the detection result.

[0133] The communication platform determines whether there is someone at the door based on the results returned by the AI ​​capability platform. If there is no one at the door, the call is stopped and the process goes to step S509. Otherwise, if there is someone at the door, the communication platform initiates a door opening call to the target terminal and the process goes to step S508.

[0134] Step S508: initiate a door opening call.

[0135] Step S509: Stop the door opening call.

[0136] Step S510, end.

[0137] Through the above steps S501 to S510, when a visitor presses the smart doorbell, the smart doorbell initiates a door-opening call to the communication platform. After receiving the door-opening call request, the communication platform does not send it to the target device first. Instead, it captures a video image within a preset time period and sends it to the AI ​​capability platform. Next, the AI ​​capability platform uses an intelligent algorithm to detect whether there is someone at the door and returns the detection result to the communication platform. The communication platform then determines whether there is someone at the door based on the returned detection result. If there is no one at the door, the call stops. If there is someone at the door, the communication platform initiates a door-opening call to the target terminal. Through human body detection technology, it analyzes whether there is someone at the door. If there is no one at the door, the smart doorbell stops initiating the call, thus avoiding the problem of initiating a door-opening call and disturbing the homeowner if someone accidentally presses the smart doorbell.

[0138] The second type is that the communication platform initiates a door opening call to the called terminal, and the door opening call automatically hangs up the process. In the embodiment of the present application, the second hang-up process is as follows: Figure 6 As shown:

[0139] Step S601, start.

[0140] Step S602: The communication platform initiates a door opening call to the target terminal.

[0141] Here, the target terminal can be a mobile phone or device bound to the user.

[0142] Step S603: Determine whether the user answers the call.

[0143] If the user has answered the door opening call, the process proceeds to step S604; if the user has not answered the door opening call, the process proceeds to step S605.

[0144] Step S604: Enter the call process.

[0145] Step S605: During the door opening call process, the communication platform captures an image at a set time interval and sends it to the artificial intelligence capability platform.

[0146] Here, the set duration can be 1 second, 2 seconds, etc. Taking 1 second as an example, the communication platform captures a picture every 1 second and sends it to the AI ​​capability platform, so that the AI ​​capability platform can detect the image.

[0147] Step S606: The artificial intelligence capability platform returns the detection results.

[0148] Here, the AI ​​capability platform uses intelligent algorithms to detect the status of the door, which can be open or closed, and whether there is someone at the door, and returns the results to the communication platform.

[0149] Step S607: The communication platform determines whether the door is opened.

[0150] If it is determined that the door is open, go to step S608; if it is determined that the door is not open, go to step S609.

[0151] Step S608: Stop the door opening call.

[0152] In step S609, the communication platform determines whether there is anyone within the monitoring range.

[0153] If there is someone within the monitoring range, the process returns to step S602 and continues to initiate a door opening call to the target terminal; if there is no one within the monitoring range, the process proceeds to step S610.

[0154] Step S610: Whether it is determined for a preset number of consecutive times that there is no one within the monitoring range.

[0155] Here, the preset number of times can be 5 times, 6 times, etc. Taking the preset number of times as an example, if it is judged that there is no one within the monitoring range for 5 consecutive times, return to step S608, that is, stop the door opening call; if there is at least 1 time that it is judged that there is someone within the monitoring range, return to step S602 and continue to initiate the door opening call to the target terminal.

[0156] Step S611, end.

[0157] Through steps S601 to S611, the communication platform initiates a door-opening call to the target terminal, then determines whether the user answers the call on the target terminal. If the user answers the call, the call proceeds. If the user does not answer the call, the call continues, capturing an image at set intervals during the call to determine whether the door is open and whether there is someone at the door. If the door is open, the call is terminated. If the door is closed, the platform repeatedly determines whether there is no one within the monitoring range. If the multiple determinations indicate that no one is within the monitoring range, the call is terminated. Using human detection technology, the platform analyzes whether there is someone at the door. If no one is at the door, the smart doorbell terminates the door-opening call. This means that if a visitor leaves the door due to the homeowner's lack of response during the call, the call can be stopped without disturbing the homeowner. Furthermore, using door open / close detection technology, the platform analyzes whether the door is open. If the door is open, the call is terminated. If the homeowner opens the door without answering the call, the call is automatically terminated.

[0158] The third type is a process in which the call is automatically hung up during the call between the calling party and the called party. In the embodiment of the present application, the third hang-up process is as follows: Figure 7 As shown:

[0159] Step S701, start.

[0160] Step S702: The homeowner is talking to the visitor.

[0161] Step S703: During the process, the communication platform captures an image at a set time interval and sends it to the artificial intelligence capability platform.

[0162] Here, the set time length can be 1 second, 2 seconds, etc. Taking 1 second as an example, the communication platform captures a picture every 1 second and sends it to the AI ​​capability platform, so that the AI ​​capability platform can detect the image.

[0163] Step S704: The artificial intelligence capability platform returns the detection results.

[0164] Here, the AI ​​capability platform uses intelligent algorithms to detect the status of the door, which can be open or closed, and whether there is someone at the door, and returns the results to the communication platform.

[0165] Step S705: The communication platform determines whether the door is opened.

[0166] If it is determined that the door is open, go to step S706; if it is determined that the door is not open, return to step S702.

[0167] Step S706: Determine whether the door is opened for a preset number of consecutive times.

[0168] Here, in order to avoid errors caused by this judgment result, a preset number of judgments is performed, wherein the preset number of times can be 5 times, 6 times, etc. Taking the preset number of times as an example, if the door is judged to be open for 5 consecutive times, it means that the purpose of the call has been achieved, the visitor can enter the room or building, and the process goes to step S707; if there is at least 1 judgment that the door is not opened and is still in a closed state, the process returns to step S702.

[0169] Step S707: Stop the call.

[0170] Step S708, end.

[0171] Through the above steps S701 to S708, during the conversation between the homeowner and the visitor, the communication platform captures an image at a set time interval and sends it to the AI ​​capability platform; then, the AI ​​capability platform detects the image and returns the detection result to the communication platform; then, the communication platform determines whether the door is open based on the detection result. If the door is not open, the conversation between the homeowner and the visitor continues; if the door is open, it continues to determine whether the door is open multiple times, and stops the conversation when multiple judgments indicate that the door is open. Through the open / closed door detection technology, it analyzes whether the door is open. If the door is open, the call is stopped. During the call, the homeowner goes directly to the door to open it. At this time, there is no need to continue the call, and the call can be automatically stopped, thereby improving the automation level of the call and improving the user experience.

[0172] Based on the foregoing embodiments, an embodiment of the present application provides a doorbell hang-up device, wherein the modules included in the device and the units included in each module can be implemented by a processor in a computer device; of course, they can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0173] The embodiment of the present application further provides a doorbell hanging device, Figure 8 This is a schematic diagram of the structure of the doorbell hanging device provided in the embodiment of the present application, as shown in FIG. Figure 8 As shown, the doorbell hanging device 800 includes:

[0174] A response module 801 is configured to obtain video data currently captured by an image capture device in response to a pressing operation on the doorbell, and to obtain a plurality of video frame images based on the video data;

[0175] A calling module 802 is configured to initiate a door opening call to a target terminal when a calling condition is determined to be met based on multiple video frame images within a preset time period;

[0176] A processing module 803 is configured to perform image processing on the plurality of video frame images to determine a state of a target object, wherein the target object includes at least one of a door and a human body;

[0177] The ending module 804 is configured to end the door opening call when it is determined based on the state of the target object that a call ending condition is met.

[0178] In some embodiments, when the target object includes a door, the processing module 803 includes:

[0179] An acquisition submodule, configured to acquire reference features, wherein the reference features are features extracted from a reference image representing a closed state of the door;

[0180] An extraction submodule, configured to extract image features of each video frame image and determine each difference information between each image feature and the reference feature;

[0181] A first determining submodule is configured to determine that the door is in an open state when at least one difference information among the various difference information is greater than a difference threshold;

[0182] A second determining submodule is configured to determine that the door is in a closed state when all the difference information is less than or equal to the difference threshold;

[0183] Correspondingly, the doorbell hanging device 800 further includes:

[0184] The first determining module is configured to determine whether the door is in an open state, wherein when the door is in an open state, it is determined that a call termination condition is satisfied.

[0185] In some embodiments, when the target object includes a human body, the processing module 803 further includes:

[0186] an identification submodule, configured to perform image recognition on each video frame image to obtain each first recognition result, wherein each first recognition result is used to indicate whether each video frame image contains a human body image;

[0187] A third determining submodule is configured to determine that the state of the human body is a leaving state when each of the first recognition results indicates that each of the video frame images does not contain a human body image;

[0188] Correspondingly, the doorbell hanging device 800 further includes:

[0189] The second determining module is configured to determine whether the state of the human body is an away state, wherein when it is determined that the state of the human body is an away state, it is determined that a call end condition is satisfied.

[0190] In some embodiments, the processing module 803 further includes:

[0191] A fourth determination submodule is configured to determine the video frame image containing the human body image as the target video frame image when each of the first recognition results indicates that at least one of the video frame images contains a human body image;

[0192] a fifth determining submodule, configured to determine distance information between a human body and the image acquisition device in each target video frame image;

[0193] a sixth determining submodule, configured to determine a moving direction of a human body based on respective time information and respective distance information of respective target video frame images;

[0194] The seventh determining submodule is configured to determine that the human body is in a leaving state when it is determined that the moving direction of the human body is moving away from the image acquisition device.

[0195] In some embodiments, the doorbell hanging device 800 further includes:

[0196] a recognition module, configured to perform image recognition on each video frame image within a preset time length to obtain each second recognition result, wherein each second recognition result is used to indicate whether each video frame image within the preset time length contains a human body image;

[0197] The third determining module is configured to determine that a call condition is satisfied when each second recognition result indicates that at least one video frame image in each video frame image within a preset time period contains a human body image.

[0198] In some embodiments, the ending module 804 is further configured to end the call connection when a call ending condition is determined to be satisfied based on the door state; the doorbell hanging up device 800 further includes:

[0199] A first establishing module, configured to receive a call permission instruction sent by the target terminal and establish a call connection between the doorbell and the target terminal;

[0200] The fourth determination module is used to obtain the target video data collected by the image acquisition device during the call, and determine the state of the door body based on the target video data.

[0201] In some embodiments, the fourth determining module includes:

[0202] a judgment submodule, configured to judge whether the door body is in an open state based on the target video data until a preset number of judgments is reached, and obtain each judgment result;

[0203] an eighth determining submodule, configured to determine the number of times the door is in the open state in each judgment result;

[0204] The ninth determination submodule is configured to determine that the door body is in an open state when the number of times reaches a number threshold, wherein the number threshold is less than or equal to the preset judgment number of times.

[0205] In some embodiments, the doorbell hanging device 800 further includes:

[0206] A second establishing module is configured to receive a call permission instruction sent by the target terminal and establish a call connection between the doorbell and the target terminal;

[0207] The acquisition module is used to obtain the audio information during the call;

[0208] The fifth determining module is configured to determine whether the audio information includes a preset target ending word, wherein when the audio information includes the target ending word, the call connection is terminated.

[0209] It should be noted that the description of the doorbell hanging device in the embodiment of the present application is similar to the description of the above-mentioned method embodiment, and has similar beneficial effects as the method embodiment, so it will not be repeated here. For technical details not disclosed in the embodiment of the present device, please refer to the description of the method embodiment of the present application for understanding.

[0210] It should be noted that in the embodiment of the present application, if the above-mentioned doorbell hanging up method is implemented in the form of a software function module and is sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0211] Accordingly, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the doorbell hanging up method provided in the above embodiment are implemented.

[0212] The embodiment of the present application provides a doorbell hanging device, Figure 9 This is a schematic diagram of the structure of the doorbell hang-up device provided in the embodiment of the present application, as shown in FIG. Figure 9 As shown, the doorbell hang-up device 900 includes: a processor 901, at least one communication bus 902, a user interface 903, at least one external communication interface 904, and a memory 905. The communication bus 902 is configured to enable communication between these components. The user interface 903 may include a display screen, and the external communication interface 904 may include a standard wired interface and a wireless interface. The processor 901 is configured to execute a program for the doorbell hang-up method stored in the memory to implement the steps of the doorbell hang-up method provided in the above embodiment.

[0213] The description of the above doorbell hang-up device and storage medium embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the doorbell hang-up device and storage medium embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0214] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0215] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0216] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0217] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the embodiment of the present application.

[0218] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0219] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes: mobile storage devices, ROM, disks or optical disks, and other media that can store program codes.

[0220] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an AC to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks or optical disks.

[0221] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A doorbell hanging up method, characterized in that: The method comprises: In response to a pressing operation on the doorbell, acquiring video data currently captured by an image capture device, and acquiring a plurality of video frame images based on the video data; When the call condition is determined to be met based on multiple video frame images within a preset time, a door opening call is initiated to the target terminal; performing image processing on the plurality of video frame images to determine a state of a target object, the target object comprising at least one of a door and a human body; When it is determined based on the state of the target object that a call termination condition is met, terminating the door opening call; receiving a call permission instruction sent by the target terminal, and establishing a call connection between the doorbell and the target terminal; Acquire target video data captured by the image acquisition device during the call, and determine whether the door is in an open state based on the target video data until a preset number of determinations is reached, thereby obtaining each determination result; Determining the number of times the door is in the open state in each judgment result; When the number of times reaches a number threshold, determining that the state of the door body is an open state, wherein the number threshold is less than or equal to the preset judgment number; When the door is in the open state, the call connection is terminated.

2. The method according to claim 1, wherein When the target object includes a door, performing image processing on the multiple video frame images to determine the state of the target object includes: Acquire a reference feature, wherein the reference feature is a feature extracted from a reference image representing a closed state of the door; Extracting image features of each video frame image, and determining each difference information between each image feature and the reference feature; When at least one difference information among the various difference information is greater than a difference threshold, determining that the door body is in an open state; When each piece of difference information is less than or equal to the difference threshold, determining that the door is in a closed state; Correspondingly, the method further includes: Determine whether the state of the door body is an open state, wherein when it is determined that the state of the door body is an open state, it is determined that a call end condition is met.

3. The method according to claim 1, wherein When the target object includes a human body, performing image processing on the multiple video frame images to determine the state of the target object includes: Performing image recognition on each video frame image to obtain each first recognition result, wherein each first recognition result is used to indicate whether each video frame image contains a human body image; When each first recognition result indicates that each video frame image does not contain a human body image, determining that the state of the human body is a leaving state; Correspondingly, the method further includes: It is determined whether the state of the human body is an away state, wherein when it is determined that the state of the human body is the away state, it is determined that a call end condition is met.

4. The method according to claim 3, wherein The performing image processing on the plurality of video frame images to determine the state of the target object further includes: When each first recognition result indicates that at least one of the video frame images contains a human body image, determining the video frame image containing the human body image as a target video frame image; Determining each distance information between the human body and the image acquisition device in each target video frame image; Determine the direction of movement of the human body based on the time information and distance information of each target video frame image; When it is determined that the moving direction of the human body is moving away from the image acquisition device, it is determined that the human body is in the leaving state.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Performing image recognition on each video frame image within a preset time length to obtain each second recognition result, wherein each second recognition result is used to indicate whether each video frame image within the preset time length contains a human body image; When each second recognition result indicates that at least one video frame image in each video frame image within the preset time length includes a human body image, it is determined that the call condition is met.

6. The method according to claim 1, wherein The method further comprises: receiving a call permission instruction sent by the target terminal, and establishing a call connection between the doorbell and the target terminal; Get the audio information during the call; Determine whether the audio information includes a preset target ending word, wherein when the audio information includes the target ending word, end the call connection.

7. A doorbell hanging device, characterized in that: The doorbell hanging device comprises: a response module, configured to obtain video data currently captured by the image capture device in response to a pressing operation on the doorbell, and to obtain a plurality of video frame images based on the video data; A calling module is used to initiate a door opening call to a target terminal when a calling condition is determined to be met based on multiple video frame images within a preset time length; a processing module, configured to perform image processing on the plurality of video frame images to determine a state of a target object, wherein the target object includes at least one of a door and a human body; An end module is used to end the door opening call when it is determined that the call end condition is met based on the state of the target object; receive a call permission instruction sent by the target terminal, and establish a call connection between the doorbell and the target terminal; obtain the target video data collected by the image acquisition device during the call, and judge whether the door body state is open based on the target video data until a preset number of judgments is reached to obtain each judgment result; determine the number of times the door body is open in each judgment result; when the number reaches a number threshold, determine that the state of the door body is open, wherein the number threshold is less than or equal to the preset number of judgments; when the state of the door body is open, end the call connection.

8. A doorbell hanging device, characterized in that: The doorbell hang-up device includes: processor; and a memory for storing a computer program executable on the processor; Wherein, when the computer program is executed by a processor, the doorbell hanging up method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are configured to execute the doorbell hanging up method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method, device and system for realizing access control visual intercom service, and intelligent robot

    CN111405225A

  • Door bell system having human face detection function

    CN201233630Y

  • Intercom device

    JP2016063362A