Alarm method and device and electronic equipment
By collecting key videos of alarm trigger events monitored by the entry door lock equipment, using the description model to generate and rewrite the event description information, determine the target alarm level for accurate alarm, solving the problem of low alarm accuracy of the entry door lock equipment, and achieving higher alarm accuracy and user safety.
Patent Information
- Application Number
- CN202510397028.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The alarm accuracy of existing entrance door lock equipment is low and false alarms often occur.
Key videos are collected after monitoring the alarm trigger event, the first description model is used to generate event description information, and combined with detection information and prompt words, input it to the second description model for rewritten, and obtain target description information to distinguish the relationship between suspicious persons, legal persons and suspicious objects, and determine the target alarm level for accurate alarm.
It improves the alarm accuracy of the entry door lock equipment, ensures that the alarm information is more accurate, reduces false alarms, and improves user safety.
Smart Images

Figure CN120260193A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of safety detection technology, and in particular to an alarm method, device and electronic equipment. Background Art
[0002] Entrance door locks refer to smart door locks installed at the entrances of buildings such as residences and offices. They are mainly used to control and manage the unlocking methods of doors, and improve the security, convenience and intelligence of access control. With the continuous development of entrance door locks, the attention paid to user safety is also increasing.
[0003] At present, in practical applications, the alarm accuracy of the door lock device is low, and the alarm error often occurs. Based on this, how to improve the alarm accuracy of the door lock device is a technical problem that needs to be solved urgently. Summary of the invention
[0004] In view of this, the present application provides an alarm method, device and electronic device to improve the accuracy of the alarm.
[0005] The present application provides an alarm method, which includes:
[0006] Detecting the key video collected by the door lock device after detecting the alarm triggering event to obtain detection information; the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information;
[0007] Input the key video into the first description model to obtain event description information of the key video; the event description information is used to briefly describe the content in the key video; the event description information at least does not distinguish between suspicious person information, legal person information, target alarm level, alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons and suspicious objects;
[0008] Determining, based on the detection information, a prompt word that matches the detection information and is to be input into the second description model;
[0009] Determine the target alarm level corresponding to the key video from the set alarm levels;
[0010] The event description information, prompt words, target alarm level and detection information are input into the second description model to rewrite the event description information to obtain target description information; compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description between the legitimate person and the suspicious person and / or the suspicious object, and the target description information is used to accurately generate an alarm description;
[0011] Based on the target description information, an alarm is issued.
[0012] An embodiment of the present application further provides an alarm device, which includes:
[0013] A detection module, configured to detect the key video collected by the household door lock device after detecting an alarm trigger event, and obtain detection information; the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information;
[0014] An input module, configured to input the key video into a first description model to obtain event description information of the key video; the event description information is used to briefly describe the content in the key video; the event description information at least does not distinguish between suspicious person information, legal person information, target alarm level, alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons, and suspicious objects;
[0015] A first determination module, configured to determine a prompt word to be input into a second description model that matches the detection information based on the detection information;
[0016] A second determination module, configured to determine the target alarm level corresponding to the key video from the set alarm levels;
[0017] A rewriting module, configured to input the event description information, the prompt word, the target alarm level, and the detection information into the second description model to rewrite the event description information and obtain target description information; compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description between legal persons and suspicious persons and / or suspicious objects, and the target description information is used for accurate alarm;
[0018] An alarm module, configured to perform an alarm based on the target description information.
[0019] An embodiment of the present application further provides an electronic device, which includes:
[0020] A processor; and
[0021] A computer-readable storage medium, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the steps of the above method.
[0022] An embodiment of the present application further provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor executes the steps in the above method.
[0023] As can be seen from the above technical solutions, in the embodiments of the present application, by monitoring the alarm trigger event, then actively triggering the alarm detection based on the alarm trigger event, and then combining the suspicious detection of the key video after the alarm trigger event is monitored, multiple verifications can effectively improve the alarm accuracy.
[0024] Further, in the embodiments of the present application, the alarm detection is actively triggered by monitoring the alarm trigger event first. When performing the alarm detection on the key video after the alarm trigger event is monitored, with the help of the user scenario and the prompt words and adding prior information such as the detection information of the key video and the target alarm level corresponding to the key video, the alarm event description of the key video is corrected and rewritten to obtain more accurate alarm information, thereby further improving the alarm accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings herein are incorporated into the specification and form a part of this application, showing embodiments consistent with this application and used together with the specification to explain the principles of this application.
[0026] Figure 1 It is a schematic flowchart of the method provided by the embodiments of the present application.
[0027] Figure 2 It is a schematic diagram of the actions provided by the embodiments of the present application.
[0028] Figure 3 It is a schematic diagram of a video frame provided by the embodiments of the present application.
[0029] Figure 4 It is a schematic diagram of the device structure provided by the embodiments of the present application.
[0030] Figure 5 It is a schematic diagram of the electronic device structure provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0032] See Figure 1 , Figure 1 It is a flowchart of the method provided by the embodiments of the present application. In this embodiment, as an example, this method can be applied to a cloud device communicatively connected to an entrance door lock device, or can also be applied to the entrance door lock device, and is not specifically limited here. In this embodiment, the application scenario of this method is not specifically limited either. For example, as an example, this method can be applied to anti-abduction alarm detection in an entrance scenario.
[0033] In the embodiments of the present application, this method is described by taking its application to the above cloud device as an example. As Figure 1 shown, the process may include the following steps:
[0034] Step 101: Detect the key video collected by the entrance door lock device after detecting an alarm trigger event, and obtain detection information; the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information.
[0035] In this embodiment, as an example, the specific implementation of the above entrance door lock device detecting an alarm trigger event may be: a camera is equipped in the entrance door lock device, and the entrance door lock device detects the video frames collected by the camera. If it is detected that the video frame includes a set expression and / or gesture, it is determined that an alarm trigger event has been detected. As for how to specifically detect that the video frame includes a set expression and / or gesture, examples will be described below and will not be elaborated here for the time being.
[0036] In this embodiment, as an example, after the entrance door lock device detects an alarm trigger event, it will trigger the collection of key video and transmit the collected key video to the cloud device; then, after the cloud device receives the key video transmitted by the entrance door lock device, it can use the trained model to detect the key video to obtain the corresponding detection information, and the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information.
[0037] Optionally, the above suspicious person information may include, for example: the category of the suspicious person (such as the category is a suspicious person), the location information of the suspicious person (such as it can be the bounding box information of the suspicious person in one of the image frames of the key video), and the number of suspicious persons, etc. Among them, a suspicious person may refer to a strange person who may pose a danger to legal persons.
[0038] Optionally, the above suspicious object information may include, for example: the category of the suspicious object (such as the category is a certain dangerous item such as a knife), the location information of the suspicious object (such as it can be the bounding box information of the suspicious object in one of the image frames of the key video), and the number of suspicious objects, etc. Among them, a suspicious object may refer to an item that may pose a danger to legal persons.
[0039] Optionally, the above legal person information may include, for example: the category of the legal person (such as the category is a legal person), the location information of the legal person (such as it can be the bounding box information of the legal person in one of the image frames of the key video), and the number of legal persons, etc. Among them, a legal person can be understood as a non-suspicious person, such as a user and the user's family members, etc.
[0040] In this embodiment, as an example, the above-mentioned key video may refer to a video composed of image frames obtained by the household door lock device within a specified time period. Herein, the specified time period may refer to a period of time starting from the acquisition start time of the image frame that monitors the alarm trigger event, and this embodiment does not specifically limit it.
[0041] As for how to specifically detect the key video collected by the household door lock device after detecting the alarm trigger event to obtain the detection information in this step, it will be described by way of example below and will not be elaborated here for the time being.
[0042] Step 102: Input the above-mentioned key video into the first description model to obtain the event description information of the key video; the event description information is used to briefly describe the content in the key video; the event description information at least does not distinguish between suspicious person information, legal person information, target alarm level, alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons, and suspicious objects.
[0043] In this embodiment, the first description model may refer to a pre-trained large model for generating corresponding description information based on the input. In this embodiment, by using the above-mentioned key video as the input parameter and inputting it into the first description model, the event description information of the key video can be obtained. For example, the event description information may be: "There is a woman blinking in the video, there is a knife near her mouth, and there is also a serious-looking man behind her."
[0044] Optionally, this embodiment does not specifically limit the specific form of the first description model. For example, it may be a multimodal large language model, etc.
[0045] Step 103: Determine a prompt word that matches the detection information and is to be input into the second description model based on the above-mentioned detection information.
[0046] In this embodiment, the second description model may refer to a pre-trained large model for generating corresponding description information based on the input. Optionally, the above-mentioned first description model and the second description model may be the same model or different models, and this embodiment does not specifically limit it.
[0047] In this embodiment, in order to facilitate determining the prompt word that matches the detection information and is to be input into the second description model, the above-mentioned obtained detection information may be summarized and sorted according to a preset detection information template to obtain well-organized detection information.
[0048] Optionally, this embodiment does not specifically limit the above-mentioned preset detection information template, and it can be flexibly set based on actual application requirements; for example, it can be a template such as the following example:
[0049] "time:yyyy-MM-dd HH:mm:ss;
[0050] weapons: ["detection_box": {xxx, xxx, xxx, xxx}, "class": xx], [...];
[0051] person: ["detection_box": {xxx, xxx, xxx, xxx}, "id": xxx], [...]; ".
[0052] Based on the above description, as an example, the above-mentioned prompt word determined based on the detection information and matching the detection information to be input into the second description model can be, for example, when specifically implemented: filling the content of a preset prompt word template based on the detection information to obtain a prompt word that matches the detection information; wherein, the prompt word may include, but is not limited to: the content interpretation information of the above-mentioned well-organized detection information.
[0053] Optionally, in this embodiment, no specific limitation is imposed on the above-mentioned preset prompt word template, and it can be flexibly set based on actual application requirements; for example, it can be a template such as the following example:
[0054] "prompt word = There is now an original event description information "xxx (used to indicate the event description information of the above-mentioned key video)", and prior information "xxx (used to indicate the above-mentioned detection information)";
[0055] Among them, "weapons" in the prior information "xxx" represents the detected suspicious items, "detection_box" represents the location information (such as the bounding box coordinates), "class" represents the category of the suspicious items (such as 1 represents a knife, 2 represents..., no specific limitation is made here), "person" represents the detected person, and "id" represents the category of the person (such as 1 represents a suspicious person, 2 represents a legitimate person, etc.);
[0056] Please rewrite the original event description information "xxx" according to the input original event description information "xxx", prior information "xxx", and the target alarm level, in accordance with the preset alarm information template that matches the target alarm level, to obtain a more accurate alarm information.
[0057] Step 104, determine the target alarm level corresponding to the key video from the set alarm levels.
[0058] In this embodiment, as an example, N alarm levels can be preset based on actual application requirements for alarm purposes, where N is greater than 1; among them, different alarm levels have different alarm triggers; the alarm trigger corresponding to any alarm level includes an action sequence arranged in order. As for how to specifically determine the alarm trigger corresponding to the alarm level, it will be described by way of example below and will not be elaborated here for the time being.
[0059] Based on the above description, as an example, in this step, the target alarm level corresponding to the key video is determined from the preset alarm levels. In specific implementation, for example, it can be: first, input the above key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to legal personnel. Here, any alarm trigger detection result includes the action sequence of legal personnel arranged in the order of time stamps; then, compare the alarm trigger detection result with the alarm trigger corresponding to the preset alarm level to obtain the target alarm level corresponding to the key video, where the alarm trigger detection result and the alarm trigger corresponding to the target alarm level meet the set matching requirements.
[0060] In this embodiment, as an example, any of the above action sequences at least includes: a sequence of facial expression actions and / or gesture actions. Here, the specific content of the facial expression actions and gesture actions in this embodiment is not specifically limited and can be flexibly set based on actual application requirements; for example, the facial expression actions can include blinking and frowning, and the gesture actions can include making numbers, etc.
[0061] As for how to specifically input the above key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to legal personnel in this step, and how to specifically compare the alarm trigger detection result with the alarm trigger corresponding to the preset alarm level to obtain the target alarm level corresponding to the key video, it will be described by way of example below and will not be elaborated here for the time being.
[0062] Step 105: Input the event description information, prompt words, target alarm level, and detection information into the second description model to rewrite the event description information and obtain the target description information; compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant descriptions between legal personnel and suspicious personnel and / or suspicious objects. The target description information is used for accurate alarm.
[0063] In this embodiment, in this step, the event description information, the prompt, the target alarm level, and the detection information are input into the second description model, so that the second description model rewrites the event description information based on the rewriting requirements indicated by the prompt to obtain the target description information. Specifically, for example, the second description model rewrites the event description information according to the input event description information, detection information, and target alarm level, in accordance with the preset alarm information template matching the target alarm level, so as to obtain a more accurate alarm information, that is, the target description information.
[0064] Optionally, this embodiment may preset multiple different alarm levels; different alarm levels may correspond to different alarm information templates. This embodiment does not specifically limit the alarm information templates corresponding to each alarm level, and can be flexibly set based on actual application requirements.
[0065] For example, taking the alarm information template corresponding to one of the alarm levels, such as alarm level 1, as an example for illustration, the alarm information template can be the following example template:
[0066] "Alarm: Potential dangerous situation detected! Which person in the picture holds what suspicious item, the relative positions of the people, which alarm level the alarm is triggered, which may pose a threat to their personal safety, please take immediate measures!"
[0067] Based on the above description, exemplarily, assume that the event description information obtained by inputting the key video into the first description model is: "There is a woman blinking in the video, there is a knife near her mouth, and there is a serious-looking man behind her". It can be seen that this event description information only briefly describes the content in the key video, does not distinguish the identities of the woman and the man in the key video, that is, does not distinguish between suspicious person information and legal person information, nor does it distinguish the target alarm level, the alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons, and suspicious items.
[0068] Based on this, the event description information, prompt words, target alarm level, and detection information are input into the second description model to rewrite the event description information of the key video. The obtained target description information can be: "Alarm: Potential dangerous situation detected! A suspicious person in the video is holding a knife and appears behind Legal Person 1. Legal Person 1 triggers a Level 1 alarm, which may pose a threat to their personal safety. Please take immediate measures!". It can be seen that the target description information is more accurate than the above event description information. It not only distinguishes between suspicious person information and legal person information but also distinguishes the target alarm level (such as Level 1 alarm), the alarm information corresponding to the target alarm level (such as, Alarm: Potential dangerous situation detected! Legal Person 1 triggers a Level 1 alarm, which may pose a threat to their personal safety. Please take immediate measures!), and the relevant description between the legal person and the suspicious person and / or suspicious object (such as a suspicious person in the video is holding a knife and appears behind Legal Person 1), achieving a more accurate description of the alarm event and providing a more accurate basis for alarm reporting, thereby improving the accuracy of alarm reporting.
[0069] Step 106: Alarm based on the target description information.
[0070] In this embodiment, as an example, different alarm levels correspond to different alarm categories; among them, the higher the alarm level, the higher the urgency of the alarm category.
[0071] In this embodiment, the above alarm category can be used to indicate the push method of the target description information. As an example, the alarm category corresponding to the alarm level at least includes: sending the target description information to a specified application; and / or, sending the target description information to a specified user; and / or, sending the target description information to a specified alarm system. Among them, the urgency of sending the target description information to a specified alarm system is greater than that of sending the target description information to a specified application, and the urgency of sending the target description information to a specified user is greater than that of sending the target description information to a specified application.
[0072] Based on the above description, as an example, in this step, alarm is performed based on the target description information. In specific implementation, for example, it can be: pushing the target description information based on the alarm category corresponding to the target alarm level to achieve alarm.
[0073] Thus, the Figure 1 shown process is completed.
[0074] Through Figure 1 It can be seen from the shown process that in the embodiment of the present application, by monitoring the alarm trigger event, then actively triggering alarm detection based on the alarm trigger event, and then combining suspicious detection of the key video after monitoring the alarm trigger event, multiple validations can effectively improve the accuracy of alarm reporting.
[0075] Furthermore, in the embodiments of the present application, the alarm detection is actively triggered by monitoring the alarm trigger event. When performing alarm detection on the key video after detecting the alarm trigger event, with the help of the user scenario and the prompt words and adding prior information such as the detection information of the key video and the target alarm level corresponding to the key video, the alarm event description of the key video is corrected and rewritten to obtain alarm information with higher accuracy, thereby further improving the alarm accuracy.
[0076] The following describes how to detect that the video frame includes the set expression and / or gesture in step 101 above:
[0077] In this embodiment, as an example, for detecting that the video frame includes the set expression, in a specific implementation, it can be: first, perform face detection on the video frame through a locally trained face detection model to obtain the face detection result corresponding to the video frame; if the face detection result indicates that there is a face in the video frame, then based on the face key point information indicated by the face detection result, extract the image area associated with the set expression from the video frame. For example, if the set expression is an open - mouth expression, the image area associated with the set expression can be the mouth area; then, input the obtained image area associated with the set expression into the locally trained feature extraction model to obtain the expression feature vector corresponding to the video frame; then, perform similarity matching between the expression feature vector corresponding to the video frame and the reference expression feature vector corresponding to the set expression to obtain the expression similarity matching result. If the expression similarity matching result is greater than the set similarity threshold, it is determined that the video frame includes the set expression.
[0078] As for the monitoring process of the set gesture is similar to the monitoring process of the set expression above; for example, as an example, first, perform detection on the video frame through a locally trained hand detection model to obtain the hand detection result corresponding to the video frame; if the hand detection result indicates that there is a hand in the video frame, then input the image area corresponding to the hand bounding box indicated by the hand detection result into the locally trained feature extraction model to obtain the gesture feature vector corresponding to the video frame; then, perform similarity matching between the gesture feature vector corresponding to the video frame and the reference gesture feature vector corresponding to the set gesture to obtain the gesture similarity matching result. If the gesture similarity matching result is greater than the set similarity threshold, it is determined that the video frame includes the set gesture.
[0079] In this embodiment, the above-mentioned face detection model can be, for example, the Retina Face model, etc.; the above-mentioned hand detection model can be, for example, the YOLO model, etc.; the above-mentioned feature extraction model can be, for example, a graph convolutional neural network model, etc. This embodiment does not specifically limit this, and it can be flexibly set based on actual application requirements.
[0080] In this embodiment, as an example, the reference expression feature vector corresponding to the set expression and the reference gesture feature vector corresponding to the set gesture can be determined in advance based on the pictures uploaded by the user. Specifically, the user can upload pictures containing the set expression and pictures containing the set gesture to the cloud device in advance through a specified application, such as Figure 2 the picture containing an open-mouth expression and the picture containing a gesture of holding up the number 2 as shown.
[0081] After that, the cloud device processes the pictures uploaded by the user containing the set expression through the locally trained face detection model and feature extraction model to obtain the reference expression feature vector corresponding to the set expression; and processes the pictures uploaded by the user containing the set gesture through the locally trained hand detection model and feature extraction model to obtain the reference gesture feature vector corresponding to the set gesture; and transmits the obtained reference expression feature vector corresponding to the set expression and the reference gesture feature vector corresponding to the set gesture to the household door lock device for subsequent use. As for how to specifically determine the reference expression feature vector and the reference gesture feature vector, it is similar to the determination method of the expression feature vector and the gesture feature vector corresponding to the above video frame, and will not be elaborated here.
[0082] The following describes how to detect the key video collected by the household door lock device after detecting an alarm trigger event in step 101 to obtain detection information:
[0083] In this embodiment, there are many ways to detect the key video collected by the household door lock device after detecting an alarm trigger event to obtain detection information. For example, as an example, for each video frame in the key video, the locally trained suspicious object detection model is used to detect suspicious objects in the video frame to obtain the suspicious object detection result corresponding to the video frame; where if there is a suspicious object in the video frame, the suspicious object detection result corresponding to the video frame at least includes information such as the category of the suspicious object and the bounding box of the suspicious object.
[0084] Further, perform face detection on the video frame through a locally trained face detection model to obtain a face detection result corresponding to the video frame. Among them, if there is a face in the video frame, the face detection result corresponding to the video frame at least includes the face bounding box of each face in the video frame. At this time, the image region corresponding to the face bounding box in the face detection result can be input into the locally trained feature extraction model to obtain the face feature vector of each face in the video frame. Then, for each face in the video frame, perform a similarity match between the face feature vector of the face and the locally stored reference face feature vectors. If the face similarity match result is greater than the set similarity threshold, determine that the category of the face is a legitimate person; otherwise, determine that the category of the face is a suspicious person, and also use the category of each face in the video frame as the face detection result corresponding to the video frame.
[0085] After that, based on the suspicious object detection result and the face detection result corresponding to each video frame, obtain the detection information corresponding to the video frame. And based on the detection information corresponding to each video frame, select a video frame with the largest number of faces and suspicious objects from each video frame, and use the detection information corresponding to the selected video frame as the detection information corresponding to the key video. The detection information at least includes: suspicious person information and / or suspicious object information, and legitimate person information. For example, Figure 3 the video frame shown can be the selected video frame with the largest number of faces and suspicious objects, which includes 1 suspicious person, 1 legitimate person, and 1 suspicious object.
[0086] In this embodiment, the above-mentioned suspicious object detection model can be a target detection model such as the YOLO model, and there is no specific limitation here. As for the face detection model and the feature extraction model, reference can be made to the above relevant descriptions and will not be elaborated here.
[0087] In this embodiment, as an example, the above-mentioned locally stored reference face feature vectors can be determined in advance based on the face pictures uploaded by the user. Specifically, the user can upload at least one face picture to the cloud device through a specified application program in advance. Among them, the faces in each face picture uploaded by the user are different. For example, they can be the face pictures of the user and his family and friends, etc. There is no specific limitation in this embodiment.
[0088] After that, the cloud device performs face detection on the face pictures uploaded by the user through a locally trained face detection model to obtain the face bounding box corresponding to the face in the face picture, and inputs the image region corresponding to the obtained face bounding box into the locally trained feature extraction model to obtain the face feature vector corresponding to the face in the face picture as the reference face feature vector and store it locally.
[0089] This completes the further description of the above step 101.
[0090] Next, the alarm triggers for each alarm level will be described first:
[0091] In this embodiment, the alarm triggers corresponding to different alarm levels include an action sequence arranged in order. Based on this, as an embodiment, any alarm level corresponds to at least one alarm trigger, and any alarm trigger includes an action sequence arranged in order. The actions corresponding to the action sequences included in different alarm triggers are different. That is to say, any alarm level corresponds to at least one action.
[0092] Exemplarily, it is assumed that 3 alarm levels are set, namely, level 1 alarm, level 2 alarm, and level 3 alarm; among them, the level 1 alarm corresponds to two alarm triggers, which are respectively the alarm trigger for the blinking action (the action sequence it includes is, for example: "blink left, blink right, blink left, blink right"), and the alarm trigger for showing numbers (the action sequence it includes is, for example: "show number 1, show number 2, show number 3").
[0093] The level 2 alarm corresponds to two alarm triggers, which are respectively the alarm trigger for the left - right rotation action of the eyeballs (the action sequence it includes is, for example: "eyeball rotates to the left, eyeball rotates to the right, eyeball rotates to the left, eyeball rotates to the right"), and the alarm trigger for showing numbers (the action sequence it includes is, for example: "show number 4, show number 5, show number 6").
[0094] The level 3 alarm corresponds to two alarm triggers, which are respectively the alarm trigger for the frowning action (the action sequence it includes is, for example: "frown, unfrown, frown"), and the alarm trigger for showing numbers (the action sequence it includes is, for example: "show number 7, show number 8, show number 9").
[0095] Based on the above description, the following will describe how to specifically input the above - mentioned key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to a legitimate person in the above step 104:
[0096] As an embodiment, inputting the above - mentioned key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to a legitimate person may be specifically implemented as follows: for each video frame of the key video, if the face detection result of the video frame indicates that there is a legitimate person in the video frame, then for each action (such as an expression action or a gesture action) corresponding to each alarm level, determine whether the video frame contains the action through the detection model for detecting the action.
[0097] Afterwards, if there is at least one video frame containing the action in the key video, an action sequence corresponding to the action is formed based on the action detection results of the action in each video frame of the key video, so as to be used as the alarm trigger detection result corresponding to the action; wherein, the action detection results in the action sequence corresponding to the action are arranged in the order of time stamps.
[0098] In this embodiment, as an example, to determine whether the video frame contains the action through the detection model for detecting the action, in specific implementation, for example: if the action is an expression action, obtain the face key point information corresponding to the legal person from the face detection result of the video frame, so as to intercept the image region associated with the action from the video frame based on the face key point information. For example, assuming the action is a frowning action, the image region associated with the action is the eyebrow region; and input the intercepted image region into the detection model for detecting the action to obtain the action detection result corresponding to the action.
[0099] If the action is a gesture action, obtain the hand bounding box of the legal person from the face detection result of the video frame, and input the image region corresponding to the hand bounding box in the video frame into the detection model for detecting the action to obtain the action detection result corresponding to the action.
[0100] Afterwards, if the action detection result corresponding to the action indicates the existence of the action, it is determined that the video frame contains the action; otherwise, it is determined that the video frame does not contain the action.
[0101] In this embodiment, as an example, the detection model for detecting any action can be, for example, a classification model for identifying and classifying the action, and this embodiment does not specifically limit this.
[0102] In this embodiment, as an example, if the action is a gesture action, the action may include at least one gesture action. For example, the action may include three gesture actions: making the number 1 gesture, making the number 2 gesture, and making the number 3 gesture; based on this, it should be noted that if the action detection result corresponding to the action indicates the existence of one of the gesture actions corresponding to the action, it is determined that the action exists.
[0103] The following describes how to specifically compare the alarm trigger detection result with the alarm trigger corresponding to the set alarm level in step 104 above to obtain the target alarm level corresponding to the key video:
[0104] In this embodiment, as an example, the above-mentioned comparison of the alarm trigger detection result with the alarm trigger corresponding to the set alarm level is performed to obtain the target alarm level corresponding to the key video. In specific implementation, it may include, for example: traversing each set alarm level, and taking the traversed alarm level as the current alarm level; then, if the alarm trigger detection results corresponding to each action of the legitimate personnel include the alarm trigger detection results corresponding to each action of the current alarm level, check whether the alarm trigger of each action corresponding to the current alarm level and the alarm trigger detection result corresponding to each action meet the set matching requirements. If so, determine the current alarm level as the target alarm level; otherwise, continue to traverse each set alarm level.
[0105] In this embodiment, as an example, for each action corresponding to the current alarm level, check whether the alarm trigger of this action corresponding to the current alarm level and the alarm trigger detection result corresponding to this action meet the set matching requirements. In specific implementation, it may be, for example: traverse the alarm trigger detection results corresponding to this action in the order of time stamps. If an action sequence that matches the alarm trigger of this action corresponding to the current alarm level is found from the alarm trigger detection results corresponding to this action, determine that the alarm trigger of this action corresponding to the current alarm level and the alarm trigger detection result corresponding to this action meet the set matching requirements; otherwise, determine that the alarm trigger of this action corresponding to the current alarm level and the alarm trigger detection result corresponding to this action do not meet the set matching requirements.
[0106] Exemplarily, assume that the key video includes 10 key frames, and this action is frowning. The alarm trigger of this action corresponding to the current alarm level is the action sequence of "frowning, not frowning, frowning".
[0107] Based on this, assume that the alarm trigger detection result corresponding to this action is the action sequence of "not frowning, not frowning, frowning, frowning, not frowning, frowning, frowning, not frowning, not frowning, not frowning"; then traverse the alarm trigger detection results corresponding to this action in the order of time stamps, and the action sequence of "frowning, not frowning, frowning" can be found; therefore, it can be determined that the alarm trigger of this action and the alarm trigger detection result corresponding to this action meet the set matching requirements.
[0108] And assume that the alarm trigger detection result corresponding to this action is the action sequence of "not frowning, not frowning, frowning, frowning, not frowning, not frowning, not frowning, not frowning, not frowning, not frowning"; then traverse the alarm trigger detection results corresponding to this action in the order of time stamps, and the action sequence of "frowning, not frowning, frowning" cannot be found; therefore, it can be determined that the alarm trigger of this action and the alarm trigger detection result corresponding to this action do not meet the set matching requirements.
[0109] In another exemplary case, assume that the key video includes 10 key frames, and the action is to compare with the numbers 1, 2, and 3. The alarm trigger for this action corresponding to the current alarm level is the action sequence of "compare with the number 1, compare with the number 2, compare with the number 3".
[0110] Based on this, assume that the alarm trigger detection result corresponding to this action is the action sequence of "others (such as 0), others, compare with the number 1, compare with the number 1, others, compare with the number 2, compare with the number 2, others, compare with the number 3, compare with the number 3". Then, traversing the alarm trigger detection result corresponding to this action in chronological order of timestamps, the action sequence of "compare with the number 1, compare with the number 2, compare with the number 3" can be found. Therefore, it can be determined that the alarm trigger of this action and the alarm trigger detection result corresponding to this action meet the set matching requirements.
[0111] However, assume that the alarm trigger detection result corresponding to this action is the action sequence of "others, others, compare with the number 3, compare with the number 3, others, compare with the number 2, compare with the number 2, others, others, compare with the number 1". Then, traversing the alarm trigger detection result corresponding to this action in chronological order of timestamps, the action sequence of "compare with the number 1, compare with the number 2, compare with the number 3" cannot be found. Therefore, it can be determined that the alarm trigger of this action and the alarm trigger detection result corresponding to this action do not meet the set matching requirements.
[0112] Thus, the further description of the above step 104 is completed.
[0113] In this embodiment, in specific applications, there may be a situation where suspicious persons and suspicious objects are blocked by legitimate persons. In this case, the detection information obtained in the above step 101 does not include suspicious person information and suspicious object information. However, to avoid missing an alarm, the process still continues to determine the target alarm level corresponding to the key video from the set alarm levels, and input the event description information, prompt words, target alarm level, and detection information into the second description model to rewrite the event description information to obtain the target description information, and perform an alarm based on the target description information. Among them, compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description of the legitimate person.
[0114] The following is a further description of the above step 106:
[0115] First, describe the alarm categories corresponding to the alarm levels:
[0116] As described above, the alarm categories corresponding to the alarm levels at least include: sending the target description information to a specified application program; and / or, sending the target description information to a specified user; and / or, sending the target description information to a specified alarm system.
[0117] In this embodiment, the specified application refers to, for example, an application related to alarm that is bound to the household door lock device. As an embodiment, when specifically implementing the above-mentioned sending of the target description information to the specified application, it can be, for example: sending the target description information to the specified members and / or specified groups in the specified application to achieve alarm. This embodiment does not specifically limit the above-mentioned specified members and specified groups, and can be flexibly set based on actual application requirements.
[0118] In this embodiment, as an embodiment, when specifically implementing the above-mentioned sending of the target description information to the specified user, it can be, for example: sending the target description information to the specified user by means of text message and / or email to achieve alarm. This embodiment also does not specifically limit the above-mentioned specified user, and can be flexibly set based on actual application requirements.
[0119] In this embodiment, as an embodiment, when specifically implementing the above-mentioned sending of the target description information to the specified alarm system, it can be, for example: a specified alarm system can be jointly established with relevant departments in advance; based on this, the target description information can be sent to the specified alarm system to achieve alarm.
[0120] Thus, the description of the alarm categories corresponding to the alarm levels is completed. The following further describes the alarm based on the target description information:
[0121] In this embodiment, based on the above description, it can be known that the detection information obtained in step 101 is the detection information corresponding to a video frame in the key video. Based on this, there are many ways to specifically implement the above-mentioned alarm based on the target description information; for example, as an embodiment, based on the alarm category corresponding to the target alarm level, the video frame corresponding to the detection information and the target description information are pushed to achieve alarm. Another example is that, as another embodiment, the video frame corresponding to the detection information is used as the cover of the key video, and based on the alarm category corresponding to the target alarm level, the key video and the target description information are pushed to achieve alarm.
[0122] To facilitate understanding of the specific implementation process of the above-mentioned alarm method, the following is described by way of specific examples.
[0123] In this embodiment, first, user A uploads, in advance, pictures of expressions and gestures for triggering alarm detection to the cloud device through the specified application. For example, it can be pictures Figure 2 such as the picture containing the open-mouth expression and the picture containing the gesture of making the number 2 as shown.
[0124] As an example, user A can also upload at least one face image to the cloud device in advance through a specified application. As described above, the faces in each face image are different. For example, they can be the face images of user A and his family members and friends, etc.
[0125] In addition, user A can also set at least one alarm level and the corresponding alarm triggers and alarm categories in the cloud device in advance through a specified application for subsequent determination of the target alarm level and pushing of alarm information, etc. As for the specific setting content of the alarm level, alarm triggers, and alarm categories, reference can be made to the above relevant descriptions and will not be elaborated here.
[0126] The cloud device processes the pictures uploaded by the user for triggering alarm detection through the locally trained face detection model and feature extraction model to obtain the corresponding reference expression feature vector and reference gesture feature vector, and transmits the reference expression feature vector and reference gesture feature vector to the entrance door lock device for subsequent detection of alarm trigger events.
[0127] The cloud device processes the face images uploaded by the user through the locally trained face detection model and feature extraction model to obtain the face feature vector corresponding to the face in the face image as the reference face feature vector and stores it locally for subsequent detection of suspicious persons and legitimate persons. As for how to specifically obtain the face feature vector corresponding to the face in the face image, reference can be made to the above relevant descriptions and will not be elaborated here.
[0128] After that, if an event such as user A being held hostage occurs behind user A's door. User A will make the expressions and / or gesture actions for triggering alarm detection as described above, such as opening the mouth and / or making the finger gesture of number 2, and will also make the alarm trigger action corresponding to a certain alarm level, such as the alarm trigger action corresponding to the third-level alarm described above: frowning twice, and continuously making the finger gestures of numbers 7, 8, and 9.
[0129] The camera installed on user A's door will collect video frames and transmit them to the entrance door lock device.
[0130] When the entrance door lock device receives the video frames transmitted by the camera, it will detect whether the video frames include the above-mentioned expressions and / or gesture actions for triggering alarm detection based on the reference expression feature vector and / or reference gesture feature vector stored locally. If so, it determines that an alarm trigger event has been detected; if not, it continues to detect. As for how to specifically detect alarm trigger events based on the reference expression feature vector and / or reference gesture feature vector stored locally, reference can be made to the above relevant descriptions and will not be elaborated here.
[0131] After the door lock device for home entry detects an alarm trigger event, it will take the key video collected after detecting the alarm trigger event, such as the 20 - second video collected after detecting the alarm trigger event, as the key video and transmit it to the cloud device for further alarm detection.
[0132] After the cloud device receives the key video uploaded by the door lock device for home entry, it will use the above - stored reference face feature vectors locally and the trained models (such as the above - mentioned face detection model, suspicious object detection model, feature extraction model) to detect the key video transmitted by the door lock device for home entry to obtain detection information. Here, the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information.
[0133] After the cloud device receives the key video, it will also determine the target alarm level corresponding to the above - mentioned key video from the set alarm levels.
[0134] After the cloud device receives the key video, it will also input the above - mentioned key video into the first description model to obtain the event description information of the key video. Among them, the event description information of the key video only briefly describes the content in the key video. For example, the event description information is: "There is a woman blinking in the video, there is a knife near the mouth, and there is also a serious - looking man behind."
[0135] After obtaining the event description information of the key video, the cloud device will input the event description information, prompt words, target alarm level, and detection information into the second description model to rewrite the event description information to obtain the target description information. For example, the target description information is: "Warning: Potential dangerous situation detected! A suspicious person in the picture is holding a knife and appears behind legal person 1. Legal person 1 triggers a level - one alarm, which may pose a threat to their personal safety. Please take immediate measures!" It can be seen that the rewritten target description information is more accurate than the event description information of the above - mentioned key video. It distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description between legal persons and suspicious persons and / or suspicious objects for accurate alarm.
[0136] After obtaining the target description information, the cloud device will perform information push on the target description information based on the alarm category of the target alarm level to achieve alarm. For example, assume that the target alarm level is a level - three alarm, and the alarm category corresponding to the level - three alarm is to send the target description information to a specified alarm system. Based on this, the cloud device can send the target description information to the specified alarm system to achieve alarm.
[0137] So far, the description of the method provided by the embodiment of the present application is completed. Next, the device provided by the embodiment of the present application will be described:
[0138] SeeFigure 4 , Figure 4 is a schematic structural diagram of an alarm device provided by an embodiment of the present application. As Figure 4 shown, the device 400 includes a detection module 401, an input module 402, a first determination module 403, a second determination module 404, a rewriting module 405, and an alarm module 406;
[0139] The detection module 401 is configured to detect the key video collected by the household door lock device after detecting an alarm trigger event, and obtain detection information; the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information;
[0140] The input module 402 is configured to input the key video into a first description model to obtain event description information of the key video; the event description information is used to briefly describe the content in the key video; the event description information at least does not distinguish between suspicious person information, legal person information, target alarm level, alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons, and suspicious objects;
[0141] The first determination module 403 is configured to determine a prompt word to be input into the second description model that matches the detection information based on the detection information;
[0142] The second determination module 404 is configured to determine the target alarm level corresponding to the key video from the preset alarm levels;
[0143] The rewriting module 405 is configured to input the event description information, the prompt word, the target alarm level, and the detection information into the second description model to rewrite the event description information and obtain target description information; compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description between legal persons and suspicious persons and / or suspicious objects, and the target description information is used for accurate alarm;
[0144] The alarm module 406 is configured to perform an alarm based on the target description information.
[0145] As an embodiment, the second determination module 404 is specifically configured to: input the key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to a legal person; compare the alarm trigger detection result with the alarm trigger corresponding to the preset alarm level to obtain the target alarm level corresponding to the key video; the alarm trigger detection result and the alarm trigger corresponding to the target alarm level meet the set matching requirements.
[0146] As an embodiment, different alarm levels correspond to different alarm categories; the higher the alarm level, the higher the urgency of the alarm category.
[0147] As an embodiment, the alarm categories corresponding to the target alarm level at least include: sending target description information to a specified application; and / or, sending target description information to a specified user; and / or, sending target description information to a specified alarm system.
[0148] As an embodiment, any alarm trigger detection result includes an action sequence of legitimate personnel arranged in chronological order of timestamps; the alarm trigger corresponding to any alarm level includes an action sequence arranged in order; any action sequence at least includes: a sequence of facial expression actions and / or gesture actions.
[0149] As an embodiment, the household door lock device detecting an alarm trigger event means that: when the household door lock device detects that the same video frame includes a set facial expression and / or gesture, it is determined that an alarm trigger event is detected.
[0150] Thus, the completion Figure 4 of the structural description of the shown device.
[0151] For the implementation processes of the functions and roles of each module in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0152] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0153] Please refer to Figure 5 , which is a schematic hardware structure diagram of an electronic device provided by an exemplary embodiment of this application. The electronic device may include a processor 501, a communication interface 502, a computer-readable storage medium 503, and a communication bus 504. The processor 501, the communication interface 502, and the computer-readable storage medium 503 complete communication with each other through the communication bus 504. Among them, computer program instructions are stored on the computer-readable storage medium 503; the processor 501 can execute the steps of the method described in the above embodiment by executing the computer program instructions stored on the computer-readable storage medium 503. According to the actual functions of the electronic device, the electronic device may further include other hardware, which will not be elaborated here.
[0154] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium, on which a number of computer program instructions are stored. When the computer program instructions are executed by a processor, the method disclosed in the above examples of the present application can be implemented.
[0155] Exemplarily, the above computer-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the computer-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof. The processor and the memory can be supplemented by or incorporated into dedicated logic circuits.
[0156] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. An alarm method, characterized in that, The method includes: Detecting the key video collected by the entrance door lock device after detecting an alarm trigger event to obtain detection information; the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information; Inputting the key video into a first description model to obtain event description information of the key video; the event description information is used to briefly describe the content in the key video; the event description information at least does not distinguish between suspicious person information, legal person information, the target alarm level, the alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons, and suspicious objects; Determining a prompt word to be input into a second description model that matches the detection information based on the detection information; Determining the target alarm level corresponding to the key video from the preset alarm levels; Inputting the event description information, the prompt word, the target alarm level, and the detection information into the second description model to rewrite the event description information and obtain target description information; compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description between legal persons and suspicious persons and / or suspicious objects, and the target description information is used for accurate alarm; based on the target description information, an alarm is made.
2. The method according to claim 1, characterized in that The determining the target alarm level corresponding to the key video from the preset alarm levels includes: Inputting the key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to a legal person; Comparing the alarm trigger detection result with the alarm trigger corresponding to the preset alarm level to obtain the target alarm level corresponding to the key video; the alarm trigger detection result and the alarm trigger corresponding to the target alarm level meet the set matching requirements.
3. The method according to claim 1, wherein Different alarm levels correspond to different alarm categories; The higher the alarm level, the higher the urgency of the alarm category.
4. The method according to claim 1 or 3, wherein The alarm category corresponding to the target alarm level at least includes: Sending the target description information to a specified application; and / or, sending the target description information to a specified user; and / or, sending the target description information to a specified alarm system.
5. The method according to claim 2, wherein Any alarm trigger detection result includes an action sequence of the legal person arranged in chronological order of timestamps; The alarm trigger corresponding to any alarm level includes an action sequence arranged in order; Any action sequence at least includes: a sequence of facial expression actions and / or gesture actions.
6. The method according to claim 1, wherein The entrance door lock device detecting an alarm trigger event means that the entrance door lock device determines that an alarm trigger event is detected when it detects a preset facial expression and / or gesture in the same video frame.
7. An alarm device, characterized in that, The device includes: A detection module, configured to detect the key video collected by the household door lock device after detecting an alarm trigger event, and obtain detection information; the detection information at least includes: suspicious person information and / or suspicious object information, and legal person information; An input module, configured to input the key video into a first description model to obtain event description information of the key video; the event description information is used to briefly describe the content in the key video; the event description information at least does not distinguish between suspicious person information, legal person information, a target alarm level, the alarm information corresponding to the target alarm level, and the relationship between suspicious persons, legal persons, and suspicious objects; A first determination module, configured to determine, based on the detection information, a prompt word to be input into a second description model that matches the detection information; A second determination module, configured to determine, from the preset alarm levels, the target alarm level corresponding to the key video; A rewriting module, configured to input the event description information, the prompt word, the target alarm level, and the detection information into the second description model to rewrite the event description information and obtain target description information; compared with the event description information of the key video, the target description information distinguishes the target alarm level, the alarm information corresponding to the target alarm level, and the relevant description between legal persons and suspicious persons and / or suspicious objects, and the target description information is used for accurate alarm; An alarm module, configured to perform an alarm based on the target description information.
8. The device according to claim 7, wherein the second determination module is specifically configured to: input the key video into at least one detection model to obtain at least one alarm trigger detection result corresponding to a legal person; compare the alarm trigger detection result with the alarm trigger corresponding to the preset alarm level to obtain the target alarm level corresponding to the key video; the alarm trigger detection result and the alarm trigger corresponding to the target alarm level meet the set matching requirements; and / or different alarm levels correspond to different alarm categories; the higher the alarm level, the higher the urgency of the alarm category; and / or the alarm category corresponding to the target alarm level at least includes: sending the target description information to a specified application; and / or, sending the target description information to a specified user; and / or, sending the target description information to a specified alarm system; and / or any alarm trigger detection result includes an action sequence of the legal person arranged in chronological order of timestamps; the alarm trigger corresponding to any alarm level includes an action sequence arranged in order; any action sequence at least includes: a sequence of facial expression actions and / or gesture actions; and / or the household door lock device detecting an alarm trigger event means that: the household door lock device determines that an alarm trigger event is detected when it detects a preset facial expression and / or gesture in the same video frame.
9. An electronic device, characterized in that, The electronic device includes: a processor; and A computer-readable storage medium stores computer program instructions, and when the computer program instructions are run by the processor, the processor executes the steps in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are run by the processor, the processor executes the steps in any one of claims 1 to 6.
Citation Information
Cited By
Information analysis method and device, electronic equipment and storage medium
CN121117154A