Alarm information processing method, device and system, and alarm information auxiliary processing method, device and system
By extracting multimodal features from camera alarm videos and matching decisions determined by user-labeled visual samples, personalized alarm push decisions are generated, solving the problem of user preferences being unable to dynamically adapt in existing technologies and achieving accurate alarm processing that meets user expectations.
Patent Information
- Application Number
- CN202511027354.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-12
AI Technical Summary
Existing camera alarm recognition technology is unable to provide expected alarm content based on user personalized preferences, resulting in high false alarm rates, rough alarm levels, low user participation, and lack of personalized adaptation capabilities.
By receiving the alarm video reported by the camera device, multimodal feature extraction is performed to generate a multimodal feature vector, and based on the decision matching objects determined by the client-annotated visual samples, a push decision is generated to realize user-personalized alarm information processing.
The accuracy of alarm information and the degree of compliance with user expectations are improved, the false alarm rate is reduced, and the precision of alarm levels is improved.
Smart Images

Figure CN120640083A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of security technology, and in particular to a method, device and system for processing and assisting in processing alarm information. Background Art
[0002] Currently, cameras installed in homes, offices and other places generally have risk identification and alarm capabilities such as motion detection and human recognition. They can identify security risks in the surrounding environment and inform users in the first time, becoming a good helper for users to promptly discover and deal with related security risks.
[0003] However, due to the complex environment in which the camera is located and the diverse user scenarios, the traditional alarm recognition algorithm built into the camera often fails to understand the user's personalized preferences, resulting in a large number of invalid alarms or alarm levels that are too coarse. This makes it difficult to provide alarm content that meets user expectations and requires improvement. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, and system for processing and assisting in processing alarm information to solve the problem that related technologies cannot provide alarm content that meets user expectations.
[0005] To solve the above technical problems, the embodiments of the present application are implemented as follows:
[0006] In a first aspect, a method for processing alarm information is provided, which is applied to a server, and the method includes:
[0007] Receiving first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video;
[0008] Performing multimodal feature extraction on the first alarm video to generate a first multimodal feature vector;
[0009] Based on the first multimodal feature vector and the decision matching object, a push decision for the first alarm information is generated, wherein the decision matching object is determined based on the annotated visual sample uploaded by the client, and the annotated visual sample is obtained by the user annotating the visual sample on the client.
[0010] In a second aspect, a method for assisting in processing alarm information is provided, which is applied to a client, and the method includes:
[0011] In response to the labeling operation on the visual sample, labeling the visual sample to obtain a labeled visual sample containing labeling data;
[0012] Uploading the labeled visual sample to the server;
[0013] The visual annotation sample is used by the server to update or determine a decision matching object, and the decision matching object is used to determine a push decision of the first alarm information from the first camera device.
[0014] In a third aspect, an alarm information processing device is provided, which is applied to a server, and the device includes:
[0015] An information receiving module, configured to receive first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video;
[0016] a first feature extraction module, configured to perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector;
[0017] A decision generation module is used to generate a push decision for the first alarm information based on the first multimodal feature vector and a decision matching object, wherein the decision matching object is determined based on the annotated visual sample uploaded by the client, and the annotated visual sample is obtained by the user annotating the visual sample on the client.
[0018] In a fourth aspect, an alarm information processing device is provided, which is applied to a client, and the device includes:
[0019] a labeling module, configured to label the visual sample in response to a labeling operation on the visual sample, to obtain a labeled visual sample containing labeling data;
[0020] An uploading module, used for uploading the labeled visual sample to a server;
[0021] The visual annotation sample is used by the server to update or determine a decision matching object, and the decision matching object is used to determine a push decision of the first alarm information from the first camera device.
[0022] In a fifth aspect, an alarm system is provided, the system comprising a first camera device, a client, and a server;
[0023] The client is configured to, in response to a labeling operation on a target visual sample, label the target visual sample, obtain a labeled visual sample containing the labeling data, and upload the labeled visual sample to the server;
[0024] The server is used to receive first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video; perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector; and generate a push decision for the first alarm information based on the first multimodal feature vector and a decision matching object, wherein the decision matching object is determined based on the labeled visual sample.
[0025] According to a sixth aspect, an electronic device is provided, including:
[0026] processor;
[0027] a memory for storing instructions executable by the processor;
[0028] The processor is configured to execute the instructions to implement the method as described in the first aspect or the second aspect.
[0029] In a seventh aspect, a computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in the first aspect or the second aspect.
[0030] In an eighth aspect, a computer program product comprising instructions is provided, wherein when a computer runs the instructions of the computer program product, the computer executes the method described in the first aspect or the second aspect.
[0031] In an embodiment of the present application, after receiving the first alarm information including the first alarm video reported by the first camera device, the server can perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector; and based on the first multimodal feature vector and the decision matching object, generate a push decision for the first alarm information. Since the decision matching object is determined based on the user-annotated visual sample uploaded by the client, it can better reflect the alarm information characteristics expected by the user, so the push decision generated based on the decision matching object is more in line with the user's personalized alarm needs. Processing the first alarm information according to the push decision can provide more accurate alarm content that is more in line with user expectations, thereby effectively reducing the false alarm rate and improving the precision of the alarm level. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0033] Figure 1 This is a schematic diagram of the architecture of an alarm system provided in an embodiment of the present application.
[0034] Figure 2 This is a flowchart of an alarm information processing method provided in an embodiment of the present application.
[0035] Figure 3This is a flowchart of an alarm information processing method provided in an embodiment of the present application.
[0036] Figure 4 This is a flowchart of an alarm information processing method provided in an embodiment of the present application.
[0037] Figure 5 This is a flowchart of a method for assisting in alarm processing provided in an embodiment of the present application.
[0038] Figure 6 This is a schematic diagram of the interactive flow of an alarm information processing method provided in an embodiment of the present application.
[0039] Figure 7 It is a structural diagram of an electronic device according to an embodiment of the present application.
[0040] Figure 8 It is a structural diagram of an alarm information processing device provided in an embodiment of the present application.
[0041] Figure 9 It is a structural diagram of an alarm information processing device provided in an embodiment of the present application.
[0042] Figure 10 It is a structural diagram of an alarm information processing device provided in an embodiment of the present application.
[0043] Figure 11 This is a structural diagram of a device for assisting in processing alarm information provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.
[0045] The terms "first," "second," etc., in this application and the claims are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that the numerals used in this manner are interchangeable where appropriate so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, the term "and / or" in this application and the claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0046] In current intelligent security systems, the push of alarm information mainly relies on the following technical solutions, but these solutions have significant shortcomings in meeting user personalization and accuracy requirements:
[0047] 1) Identification scheme based on preset fixed rules
[0048] This solution uses fixed recognition algorithms built into smart cameras (such as human recognition based on target detection, boundary intrusion detection based on region calibration, and motion detection based on pixel changes) to identify and trigger specific event alarms. This solution has the following problems:
[0049] Rigid rules and poor adaptability: Preset fixed rules cannot dynamically adapt to the specific needs of different user scenarios;
[0050] High false alarm rate: Lack of ability to distinguish normal activities outside of preset fixed rules (such as family members walking around, pets moving around), resulting in a large number of unexpected alarms (false alarms);
[0051] Low user participation: Alarm filtering logic is hard-coded in smart cameras or servers, and users cannot participate in defining or adjusting the rules according to their actual needs.
[0052] 2) Manual configuration based on list and time period
[0053] In this solution, users can manually configure list rules on the client (e.g., a "whitelist" based on facial recognition that allows specific objects not to trigger alarms, or a "blacklist" that blocks alarms for specific devices) and alarm time periods (alarms are triggered only during the set time period). This solution has the following problems:
[0054] Complicated configuration and lack of intelligence: Rules rely heavily on manual user setup and maintenance, lacking the ability to intelligently learn and adapt based on user preferences or scenario changes.
[0055] Lack of semantic understanding: The system cannot understand the specific behavioral semantics of the alarm event (for example, it cannot distinguish whether a pet wandering at the door is normal behavior or a potential threat), resulting in rough alarm judgment;
[0056] Insufficient scenario flexibility: For the same object or event at different times, locations, and situations, the alarm processing method (such as whether to push and the alarm level) cannot be flexibly and dynamically adjusted according to the context.
[0057] 3) Filtering solution based on a fixed unified model deployed on the server
[0058] This solution processes reported alarm information by deploying a general behavior analysis model on the server. For example, it can merge similar alarms that occur frequently within a short period of time to reduce push frequency; or it can use the model to predict the "value score" of alarms and not push alarms that fall below a preset threshold. This solution has the following problems:
[0059] Weak personalization capabilities: The model focuses on generalization capabilities for common scenarios and is unable to effectively learn and capture the individual preferences and concerns of different users;
[0060] Difficulty learning user preferences: The lack of an effective mechanism for continuous learning and responding to user feedback makes it difficult for push notifications to accurately match users' true expectations.
[0061] Insufficient controllability: Users lack effective intervention and control channels over the decision-making process of filtering rules and models, and model behavior is a "black box" to users.
[0062] Clearly, existing alert solutions suffer from rigid rules and models, low user engagement, weak personalized adaptation capabilities, high false alarm rates, and coarse alert granularity. The core problem lies in the lack of a technical solution that can effectively integrate personalized user needs, dynamically learn user preferences, and grant users flexible control over alert filtering logic. This results in alerts that ultimately fail to accurately match user expectations, leading to a poor user experience.
[0063] To address the above-mentioned issues, the present application proposes an alarm information processing method, a method for assisting in the processing of alarm information, an alarm system, a device, and a computer-readable storage medium. The method can be executed by an electronic device or software installed in an electronic device. The electronic device includes, but is not limited to, any of the following smart devices: a smartphone, a personal computer (PC), a laptop, a tablet, an e-reader, an internet-connected television, a wearable device, and the like.
[0064] like Figure 1 As shown, an alarm system provided by an embodiment of the present application may include: a client 11, a server 12 and a camera device 13.
[0065] In some embodiments, the client 11 may be an application (APP) installed on a user's terminal device, and the camera 13 may be a smart camera installed at a security site. For example, the client 11 may be an APP installed on a user's mobile phone, and the camera 13 may be a smart camera installed in the user's home. Both the client 11 and the camera 13 may communicate with the server 12.
[0066] In some embodiments, the number of the clients 11 may be one or more; the number of the camera devices 13 may also be one or more, and the first camera device described below may be one of them.
[0067] The following describes an alarm information processing method provided in an embodiment of the present application.
[0068] like Figure 2 As shown, an alarm information processing method proposed in the embodiment of the present application can be applied to Figure 1 As shown in the server 12, the method may include:
[0069] Step 201: Receive first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video.
[0070] In some embodiments, the first alarm information may be generated and uploaded to the server after the first camera device identifies an alarm event using a locally built-in recognition algorithm. The first alarm information may include a video clip corresponding to the time period during which the alarm event occurred, namely, the first alarm video. The recognition algorithm built into the first camera device may include, but is not limited to, at least one of human detection, area intrusion detection, and sound source detection algorithms.
[0071] Step 202: Perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector.
[0072] It can be understood that video data usually contains features of multiple modalities such as vision, semantics, and action and behavior. Therefore, multimodal feature extraction is performed on the first alarm video to generate a first multimodal feature vector that can represent the first multimodal feature.
[0073] In some embodiments, the first multimodal feature vector may include at least two of the following:
[0074] Visual feature vector;
[0075] Action and behavior feature vectors;
[0076] Semantic embedding feature vector.
[0077] In some embodiments, since the first warning video newly generated by the first camera device has no user annotation and therefore no semantic label, step 202 may specifically include:
[0078] Extracting key frames from the first alarm video;
[0079] Extracting a feature vector of the target visual object from the key frame based on the target detection model to obtain a visual feature vector of the first alarm video: V_obj_new, such as a visual feature vector formed by the facial features and appearance contour features of the circled “family”;
[0080] Based on the key frame and the video frames before and after the key frame, the motion trajectory of the target visual object is identified, and the action and behavior feature vector of the first alarm video is generated: V_act_new, such as action trajectories such as "slow movement" and "staying in place";
[0081] A binary feature quantity (V_obj_new, V_act_new) consisting of the visual feature vector and the action and behavior feature vector of the first alarm video is used as the first multimodal feature vector.
[0082] In some embodiments, step 202 may specifically include:
[0083] Extracting a feature vector of the target visual object from the key frame based on the target detection model to obtain a visual feature vector of the first alarm video: V_obj_new, such as a visual feature vector formed by the facial features and appearance contour features of the circled “family”;
[0084] Based on the key frame and the video frames before and after the key frame, the motion trajectory of the target visual object is identified, and the action and behavior feature vector of the first alarm video is generated: V_act_new, such as action trajectories such as "slow movement" and "staying in place";
[0085] Performing semantic content inference on video images (e.g., key frames) or key objects (e.g., visual objects in key frames) in the first warning video using a multimodal model to generate a semantic embedding feature vector V_txt_predicted for the first warning video; wherein the multimodal model may be contrastive language-image pre-training (CLIP);
[0086] The triplet feature quantity (V_obj_new, V_act_new, V_txt_predicted) consisting of the visual feature vector, the action and behavior feature vector, and the semantic embedding feature vector of the first alarm video is used as the first multimodal feature vector.
[0087] Step 203: Generate a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object, wherein the decision matching object is determined based on the annotated visual sample uploaded by the client, and the annotated visual sample is obtained by the user annotating the visual sample on the client.
[0088] The decision matching object may include but is not limited to at least one of a decision matching rule library and a decision matching model, and the decision matching rule library contains at least one decision matching rule constructed according to at least one annotated visual sample.
[0089] In some embodiments, the decision object includes a decision matching rule base. Before step 203, Figure 3 As shown, the alarm information processing method proposed in this application may also include:
[0090] Step 204: Receive the annotated visual sample uploaded by the client.
[0091] The annotated visual sample is obtained by the user annotating the client. The annotated visual sample may include but is not limited to at least one of the following:
[0092] 1) The server pushes a second alarm video in the historical alarm information to the client.
[0093] In some embodiments, the second alarm video in the historical alarm information pushed by the server to the client may be a video included in the false alarm information. After the user receives and views the historical alarm information uploaded by the first camera device on the client, if the alarm information is determined to be a false alarm, the client can provide a "marking mode" for the user to mark it.
[0094] 2) The client pulls a third alarm video from the server in the alarm information suppressed from being pushed by the server.
[0095] 3) The local video of the client.
[0096] 4) The local image of the client.
[0097] In some embodiments, the user may also select a video or picture in the gallery in the client as a visual sample and then annotate it.
[0098] In some embodiments, if the labeled visual sample is the second warning video or the third warning video, its labeling data may include at least one of the following:
[0099] 1) Decision-making support tags
[0100] Auxiliary decision labels can be video-level positive or negative labels. Positive labels are used to label positive samples, which are alarm videos that meet user expectations, while negative labels are used to label negative samples, which are typically false alarm videos. For example, a user can directly mark the second or third alarm video as an invalid event (negative label), indicating that no alarm will be triggered for this scenario (such as "routine patrol," "acquaintance returning home," or "child falling"). Alternatively, a user can directly mark the second or third alarm video as a key event (positive label), indicating that the alarm level should be increased for this scenario.
[0101] 2) Video-level semantic tags
[0102] The user can select one or more pre-set semantic tags to label the second or third warning video. The user can also enter a semantic tag to label the second or third warning video to indicate the user's subjective warning intention.
[0103] 3) Context summary of the annotated video frame, such as the frame index and alarm ID of the annotated video frame.
[0104] 4) Visually annotated objects in the annotated video frames
[0105] During playback of the annotated video (such as the second or third alarm video), the user can select key frames and circle specific visual objects or visual areas as visual annotation objects. These visual annotation objects can include at least one of the following:
[0106] People (e.g., family members, visitors);
[0107] Animals (such as pets);
[0108] Equipment or items (such as sweeping robots, express boxes, etc.).
[0109] 5) The coordinates of the visual annotation area in the annotated video frame and the corresponding timestamp
[0110] The visually annotated region in the annotated video frame may be a bounding box. The coordinates of the visually annotated region in the annotated video frame may be pixel coordinates of the bounding box. The timestamp corresponding to the annotated video frame may be the timestamp when the annotated video frame was captured.
[0111] 6) Semantic labels of visually annotated objects in the annotated video frames
[0112] Users can select one or more pre-set semantic tags to label the visually annotated objects in the annotated video frame. Users can also enter their own semantic tags to label the visually annotated objects in the annotated video frame to indicate their subjective alert intention. For example, users can assign one of the following semantic tags to the visually annotated objects in the annotated video frame:
[0113] a. Reverse: "Permanent staff", "Family", "Pet activities", "Normal cleaning", etc.
[0114] b. Reverse: Custom entries such as "decoration construction" and "ignorable events".
[0115] c. Positive: “The child enters the kitchen” etc.
[0116] d. Positive: "Customized title" is similar to an instruction, requiring the addition of personalized titles based on semantic understanding to this type of adaptation or all videos.
[0117] 7) Marking method information, where the marking method may include but is not limited to at least one of manual input and voice input.
[0118] In some embodiments, if the annotated visual sample is a local video of the client, its annotated data may also include at least one of the above seven items of data.
[0119] In some embodiments, if the annotated visual sample is a local image on the client, its annotated data may include at least one of the following:
[0120] 1) Decision-making support labels;
[0121] 2) Image-level semantic labels;
[0122] 3) visually annotated objects in the image;
[0123] 4) Coordinates of the visual annotation area in the image;
[0124] 5) semantic labels of visually annotated objects in the image;
[0125] 6) Marking method information.
[0126] The annotated visual samples obtained by users on the client, including the annotated data, can be considered visual samples that contain "user-personalized intent expressions." Such visual samples offer enhanced interpretability, personalization, and content adaptability. Using these visual samples to construct decision matching objects allows for accurate screening of alert information that meets user expectations.
[0127] Step 205: Perform multimodal feature extraction on the labeled visual sample to generate a second multimodal feature vector.
[0128] After receiving the annotated visual samples uploaded by the client, the server can start the multimodal feature parsing process to convert the visual annotated samples into machine-recognizable multimodal feature vectors.
[0129] In some embodiments, the second multimodal feature vector may include at least two of the following:
[0130] Visual feature vector;
[0131] Action and behavior feature vectors;
[0132] Semantic labels are embedded into feature vectors.
[0133] In some embodiments, step 202 may specifically include:
[0134] 1) Extracting feature vectors and semantic vectors of the labeled visual objects from the labeled video frames of the labeled visual samples based on the target detection model, and generating a visual feature vector.
[0135] In some embodiments, You Only Look Once version 8 (YOLOv8) or Faster Region-based Convolutional Neural Network (Faster R-CNN) can be used as a target detection model to extract the category (such as people / pets, etc.) and appearance visual features (such as color, texture, etc.) of the annotated visual object from the image frame selection area in the annotated video frame to obtain a feature vector of the annotated visual object; a backbone model such as Residual Network (ResNet) or Shifted Window Transformer (Swin Transformer) can be used to extract the semantic vector of the annotated visual object; then the feature vector and semantic vector of the annotated visual object are fused to obtain a visual feature vector of the annotated visual object: V_obj.
[0136] 2) Based on the annotated video frame in the annotated visual sample and the video frames before and after the annotated video frame, identify the motion trajectory of the annotated visual object and generate action and behavior feature vectors.
[0137] In some embodiments, a sliding window can be constructed and slid from the video frame before the annotated video frame to the annotated video frame and the video frame after the annotated video frame. Then, during the sliding of the sliding window, a 3D convolution model (such as I3D or SlowFast) is used to identify the motion trajectory of the annotated visual object and output the action behavior encoding vector: V_act.
[0138] 3) Based on the semantic labels of the annotated visual samples, generate a semantic label embedding feature vector: V_txt.
[0139] In some embodiments, a Chinese language model (such as BERT or Ernie) or a graph-text model (such as CLIP) may be used to embed the semantic tags input by the user and output a semantic tag embedding feature vector:
[0140] In some embodiments, a triplet feature quantity (V_obj, V_act, V_txt) consisting of the visual feature vector, the action and behavior feature vector, and the semantic label embedding feature vector of the annotated visual sample can be used as the second multimodal feature vector.
[0141] Step 206: Generate a decision matching rule based on the second multimodal feature vector and store it in the decision matching rule library.
[0142] In some embodiments, step 206 may specifically include: generating a decision matching rule based on the auxiliary decision label of the annotated visual sample, the second multimodal feature vector, and the matching threshold, and storing it in the decision matching rule library. The matching threshold is used to determine whether the first multimodal feature vector of the first alarm video contained in the first alarm information matches the multimodal feature vector in the decision matching rule. Generally, if the degree of matching between the first multimodal feature vector and the second multimodal feature vector in a decision matching rule is higher than the matching threshold, it means that the first alarm video in the first alarm information matches the decision matching rule.
[0143] For example, based on the triplet feature value (V_obj, V_act, V_txt) of the annotated visual sample generated in step 205 and the preset matching threshold, the server can generate the following decision matching rule:
[0144]
[0145] Then, the server may store the above decision matching rules and the corresponding auxiliary decision labels (forward labels / reverse labels) of the annotated visual samples in a decision matching rule library.
[0146] In some embodiments, different users and / or clients are provided with their own decision matching rule libraries. Therefore, the server can store the above-mentioned decision matching rules and the corresponding auxiliary decision labels (forward labels / reverse labels) of the annotated visual samples in the decision matching rule library of the corresponding user and / or client, which is equivalent to building a personalized decision matching rule library for different users.
[0147] In some embodiments, before step 203, the alarm information processing method proposed in this application may further include:
[0148] Clustering the decision matching rules in the decision matching rule library;
[0149] The decision matching rules belonging to the same cluster center in the decision matching rule library are merged, and a matching threshold is set for the merged decision matching rules.
[0150] The purpose of this embodiment is to merge similar decision matching rules in a decision matching rule base to reduce the number of decision matching rules in the decision matching rule base, thereby improving decision matching efficiency.
[0151] As an example, for decision matching rules with the same auxiliary decision label in the decision matching rule library, clustering can be performed based on the similarity between the second multimodal vectors in different decision matching rules. For multiple decision matching rules in the decision matching rule library that belong to the same cluster center, the second multimodal feature vectors of the multiple decision matching rules can be fused according to the corresponding modality (e.g., summed and averaged) to generate a new second multimodal feature vector, and the matching threshold for the new second multimodal feature vector can be reset to obtain a clustered decision matching rule. The reset matching threshold can be the average of the matching thresholds in the multiple clustered decision matching rules.
[0152] After the matching decision database is constructed, step 203 may include:
[0153] sequentially determining a comprehensive similarity between the first multimodal feature vector and a second multimodal feature vector in each decision matching rule in the decision matching rule library to obtain a target decision matching rule, wherein the comprehensive similarity between the second multimodal feature vector in the target decision matching rule and the first multimodal feature vector is higher than a corresponding matching threshold;
[0154] After obtaining the target decision matching rule, a push decision for the first alarm information is generated based on the auxiliary decision tag of the target decision matching rule.
[0155] In some embodiments, if the first multimodal feature vector is the above-mentioned triple feature quantity (V_obj_new, V_act_new, V_txt_predicted), and the second multimodal feature vector in any decision matching rule in the decision matching rule library is the above-mentioned triple feature quantity (V_obj, V_act, V_txt), then the step of sequentially determining the comprehensive similarity between the first multimodal feature vector and the multimodal feature vectors in each decision matching rule in the decision matching rule library may include:
[0156] For each decision matching rule in the decision matching rule library, determining the similarity between the visual feature vector in the first multimodal feature vector and the visual feature vector in the decision matching rule to obtain a first similarity; determining the similarity between the action and behavior feature vector in the first multimodal feature vector and the action and behavior feature vector in the decision matching rule to obtain a second similarity; determining the similarity between the semantic label embedding feature vector in the first multimodal feature vector and the semantic label embedding feature vector in the decision matching rule to obtain a third similarity;
[0157] Determine a comprehensive similarity between the first multimodal feature vector and the second multimodal feature vector in the decision matching rule based on the first similarity, the second similarity, and the third similarity.
[0158] In some embodiments, a weighted average of the first similarity, the second similarity, and the third similarity may be used as the comprehensive similarity between the first multimodal feature vector and the second multimodal feature vector in the decision matching rule. The weights may be set in advance based on the importance of vector features in different dimensions.
[0159] In some embodiments, the first similarity, the second similarity, and the third similarity may all be cosine similarities.
[0160] In some embodiments, when the auxiliary decision tag of the target decision matching rule is a forward tag, the push decision may include: pushing the first alarm information or pushing the first alarm information after raising the alarm level; or, when the auxiliary decision tag of the target decision matching rule is a reverse tag, the push decision may include: not pushing the first alarm information or pushing the first alarm information after lowering the alarm level.
[0161] For example, the following similarity can be calculated for the first alarm video and the i-th decision matching rule:
[0162] 1) Similarity of visual feature vectors
[0163] For example, calculate the cosine similarity between V_obj_new and V_obj_i: cos(V_obj_new, V_obj_i). Assuming that V_obj_new can represent the facial feature vector of a person in the first alarm video, then the similarity between V_obj_new and the facial feature vector V_obj_i of the "family" in the visual sample that the user has labeled can be compared.
[0164] 2) Similarity between action and behavior feature vectors
[0165] For example, the cosine similarity between V_act_new and V_act_i can be calculated: cos(V_act_new, V_act_i), where the behaviors and actions represented by V_act_new and V_act_i may include: whether it is "repeatedly entering the kitchen" or "repeatedly walking around", etc.
[0166] 3) Similarity of semantic label embedding feature vectors
[0167] If the current image contains a face and the cosine similarity of the semantic embedding feature vector is: cos(V_txt_new,V_txt_i)≥0.9, it can be judged that the "acquaintance" intention has been expressed.
[0168] Afterward, if the weighted average (comprehensive similarity) of cos(V_obj_new, V_obj_i), cos(V_act_new, V_act_i), and cos(V_txt_new, V_txt_i) exceeds the matching threshold of the i-th decision matching rule (e.g., comprehensive similarity ≥ 0.85), the i-th decision matching rule is deemed the target decision matching rule, and the first alarm video meets the "personalized intent expression item" annotated by the user in the annotated visual sample corresponding to the i-th decision matching rule. Therefore, a push decision for the first alarm information can be generated based on the auxiliary decision label corresponding to the i-th decision matching rule. Specifically, if the auxiliary decision label corresponding to the i-th decision matching rule is a positive label, the generated push decision for the first alarm information can be: push the first alarm information or push the first alarm information after raising the alarm level. If the auxiliary decision label corresponding to the i-th decision matching rule is a negative label, the generated push decision for the first alarm information can be: not push the first alarm information or push the first alarm information after lowering the alarm level. Whether to increase or decrease the alarm level can be determined based on the semantic label in the annotated visual sample corresponding to the i-th decision matching rule.
[0169] In some embodiments, the alarm levels may be pre-classified according to the severity of the alarm event. For example, the alarm levels may be pre-classified into the following three levels according to the severity of the alarm event:
[0170] High risk: such as elderly people falling;
[0171] Medium risk: such as children holding dangerous objects;
[0172] Low risk: normal movement of family members, changes in optical fiber, etc.
[0173] When pushing alarm information of different alarm levels to the client, they can be distinguished by sound, color, etc.
[0174] For example, suppose a user circles the face of an elderly family member in a video (a labeled visual sample) and labels it with "family" + "normal activity" + "reverse label." The server extracts the facial feature vector V_obj_oldmother and the action and behavior feature vector V_act_slowly, embeds this feature vector with the "family" label into the feature vector V_txt_family, and forms a multimodal feature vector triple. A decision matching rule is generated and stored in the decision matching rule library. When the same person appears again in a similar background in a subsequent alarm video generated and uploaded by the camera, the server compares the new alarm video's triplet feature vector {V_obj_new, V_act_new, V_txt_new} with the triplet feature vector (V_obj_child, V_act_holding dangerous object, V_txt_family) in the decision matching rule and finds that the combined similarity exceeds the corresponding matching threshold. The server then generates a push decision to either not push the first alarm information or push the first alarm information after lowering the alarm level.
[0175] As another example, suppose a user circles the face of a child in a video (a labeled visual sample) and labels it with "family" + "dangerous activity" + "positive label." The system extracts the facial feature vector V_obj_child and the action and behavior feature vector V_act_holding dangerous objects, and combines this with the semantic embedding feature vector V_txt_family for the "family" label to form a triplet of multimodal feature vectors. A decision matching rule is then generated and stored in the decision matching rule library. If the same person appears again in a similar background in a subsequent alert video generated and uploaded by the camera, the server compares the triplet feature vector {V_obj_new, V_act_new, V_txt_new} of the new alert video with the triplet feature vector (V_obj_child, V_act_holding dangerous objects, V_txt_family) in the decision matching rule and finds that the combined similarity between the triplet feature vector {V_obj_child, V_act_holding dangerous objects, V_txt_family} in the decision matching rule exceeds the corresponding matching threshold, then a push decision is generated: "Raise the alert level and push."
[0176] In some embodiments, the decision object includes a decision matching model. Before step 203, Figure 4 As shown, the alarm information processing method proposed in this application may also include:
[0177] Step 204: Receive the annotated visual sample uploaded by the client.
[0178] Step 205: Perform multimodal feature extraction on the labeled visual sample to generate a second multimodal feature vector.
[0179] It should be noted that the specific implementation of step 204 and step 205 is the same as Figure 3 The embodiment shown is consistent, please refer to the above Figure 3 The description of the illustrated embodiment will not be repeated here.
[0180] Step 207: training the decision matching model according to the second multimodal feature vector.
[0181] The decision matching model can be a neural network. During training, the input of the decision matching model can be the second multimodal feature vector of the labeled visual sample, and the output can be the push decision of the labeled visual sample; during application, the input of the decision matching model can be the alarm video in the alarm information uploaded by the camera device, and the output can be the push decision of the alarm information.
[0182] At this time, the above-mentioned step 203 may include: inputting the first multimodal feature vector into the decision matching model, and using the decision matching model to generate a push decision for the first alarm information.
[0183] The push decision may be to push the first alarm information or to push the first alarm information after increasing the alarm level, or the push decision may be not to push the first alarm information or to push the first alarm information after decreasing the alarm level.
[0184] In some embodiments, after step 203, the alarm information processing method proposed in this application may further include: processing the first alarm information according to the push decision.
[0185] Specifically, processing the first alarm information according to the push decision may include:
[0186] In a case where the push decision is to push the first alarm information, pushing the first alarm information to the client;
[0187] In a case where the push decision is to push the first alarm information after raising the alarm level, pushing the first alarm information after raising the alarm level of the first alarm information;
[0188] When the push decision is not to push the first alarm information, the push of the first alarm information is suppressed and the first alarm information is stored.
[0189] In some embodiments, an alarm information processing method proposed in the present application may also include: iteratively updating the decision matching object based on at least one of the newly uploaded annotated visual samples by the client, the suppressed alarm information restored by the client, and the false alarm information recorded by the client.
[0190] For example, under specific settings, the server can use few-shot learning methods (such as Siamese Network and ProtoNet) to optimize the decision matching rules in the decision matching rule library based on the user's labeling behavior, recovery behavior, and long-term false alarm records, thereby improving the generalization ability of the decision matching rules, such as making the decision matching rules adaptable to similar scenarios or similar objects; or, use few-shot learning methods to iteratively update the decision matching model to improve the generalization ability of the decision matching model.
[0191] In some embodiments, an alarm information processing method proposed in the present application may also include: synchronizing the decision matching object to other clients, that is, sharing the decision matching object across devices to achieve shared alarm preferences for the whole family, thereby avoiding the trouble of repeatedly constructing decision matching objects for different clients in the same security scenario.
[0192] In an alarm information processing method provided in an embodiment of the present application, after receiving the first alarm information including the first alarm video reported by the first camera device, the server can perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector; and based on the first multimodal feature vector and the decision matching object, generate a push decision for the first alarm information. Since the decision matching object is determined based on the user-annotated visual sample uploaded by the client, it can better reflect the alarm information characteristics expected by the user, so the push decision generated based on the decision matching object is more in line with the user's personalized alarm needs. Processing the first alarm information based on the push decision can provide more accurate and user-expected alarm content, thereby effectively reducing the false alarm rate and improving the precision of the alarm level, solving the problems commonly found in related technologies such as rigid rules / models, low user participation, weak personalized adaptation capabilities, high false alarm rates, and coarse alarm granularity.
[0193] From the above description, it can be seen that the alarm information processing method proposed in the embodiment of the present application is a technical solution that can effectively integrate user personalized needs, dynamically learn user preferences to generate decision matching objects, and give users flexible control over the alarm filtering logic, so that the alarm information finally pushed can accurately match the user's expectations, thereby improving the user experience.
[0194] like Figure 5 As shown, the embodiment of the present application also proposes a method for assisting in processing alarm information, which can be applied to Figure 1 In the client 11 shown, the method includes:
[0195] Step 501 : In response to a labeling operation on a visual sample, label the visual sample to obtain a labeled visual sample containing labeling data.
[0196] The annotated visual sample may include but is not limited to at least one of the following:
[0197] 1) The server pushes a second alarm video in the historical alarm information to the client.
[0198] In some embodiments, the second alarm video in the historical alarm information pushed by the server to the client may be a video included in the false alarm information. After the user receives and views the historical alarm information uploaded by the first camera device on the client, if the alarm information is determined to be a false alarm, the client can provide a "marking mode" for the user to mark it.
[0199] 2) The client pulls a third alarm video from the server in the alarm information suppressed from being pushed by the server.
[0200] 3) The local video of the client.
[0201] 4) The local image of the client.
[0202] In some embodiments, the user may also select a video or picture in the gallery in the client as a visual sample and then annotate it.
[0203] In some embodiments, the client can provide the user with an "alarm details page" and provide an annotation entry on the alarm details page. For example, a user can click on a received alarm message to enter the alarm details page. The alarm details page can provide an "I want to mark" button (such as in the upper right corner / bottom of the page). After the user clicks the "I want to mark" button, it switches to the annotation mode interface.
[0204] In some embodiments, the client may further provide a marking mode interface for the user, and provide, in the interface, functions including but not limited to at least one of the following for supporting user marking:
[0205] 1) The playback control supports frame-level forward and backward movement;
[0206] 2) Provide an image selection tool, allowing users to drag and select people / animals / objects;
[0207] 3) Provide a tag selection bar so users can quickly click on preset semantic tags such as "family" and "pets";
[0208] 4) Provide the function of inputting custom labels (such as "decoration worker", "delivery to regular customers", etc.);
[0209] 5) At the bottom of the page, click "Submit Annotated Visual Samples" to save the annotated data and upload it to the server.
[0210] In some embodiments, the annotated visual samples uploaded to the server may carry additional information, which may include user ID, host device ID of the client, annotation timestamp and other information, for the server to perform personalized archiving and storage of decision matching objects constructed based on the annotated visual samples.
[0211] For example, the annotation mode interface can provide a variety of interactive styles such as highlighted borders and multi-label selection to improve the convenience of annotation.
[0212] In a case where the annotated visual sample is the second alarm video or the third alarm video, the annotated data of the annotated visual sample includes at least one of the following:
[0213] 1) Decision-making support tags
[0214] Auxiliary decision labels can be video-level positive labels or negative labels. Positive labels are used to label positive samples, which are alarm videos that meet user expectations. Negative labels are used to label negative samples, which are alarm videos that are false alarms.
[0215] 2) Video-level semantic tags
[0216] The user can select one or more pre-set semantic tags to label the second or third warning video. The user can also enter a semantic tag to label the second or third warning video to indicate the user's subjective warning intention.
[0217] 3) Context summary of the annotated video frame, such as the frame index and alarm ID of the annotated video frame.
[0218] 4) Visually annotated objects in the annotated video frames
[0219] During the playback of the annotated video (such as the second alarm video or the third alarm video), the user can select key frames and circle specific visual objects or visual areas as visual annotation objects for annotation.
[0220] 5) The coordinates of the visual annotation area in the annotated video frame and the corresponding timestamp
[0221] The visually annotated region in the annotated video frame may be a bounding box. The coordinates of the visually annotated region in the annotated video frame may be pixel coordinates of the bounding box. The timestamp corresponding to the annotated video frame may be the timestamp when the annotated video frame was captured.
[0222] 6) Semantic labels of visually annotated objects in the annotated video frames
[0223] Users can select one or more pre-set semantic tags to label the visually annotated objects in the annotated video frame. Users can also enter their own semantic tags to label the visually annotated objects in the annotated video frame to indicate their subjective alert intention.
[0224] 7) Marking method information, where the marking method may include but is not limited to at least one of manual input and voice input.
[0225] In some embodiments, if the annotated visual sample is a local video of the client, its annotated data may also include at least one of the above seven items of data.
[0226] In some embodiments, if the annotated visual sample is a local image on the client, its annotated data may include at least one of the following:
[0227] 1) Decision-making support labels;
[0228] 2) Image-level semantic labels;
[0229] 3) visually annotated objects in the image;
[0230] 4) Coordinates of the visual annotation area in the image;
[0231] 5) semantic labels of visually annotated objects in the image;
[0232] 6) Marking method information.
[0233] The annotated visual samples obtained by users on the client, including the annotated data, can be considered visual samples that contain "user-personalized intent expressions." Such visual samples offer enhanced interpretability, personalization, and content adaptability. Using these visual samples to construct decision matching objects allows for accurate screening of alert information that meets user expectations.
[0234] Step 502: Upload the annotated visual sample to the server.
[0235] The visual annotation sample is used by the server to update or determine a decision matching object, and the decision matching object is used to determine a push decision of the first alarm information from the first camera device.
[0236] In some embodiments, the decision matching object includes a decision matching model, which is trained based on the labeled visual samples, and the decision matching model is used to generate a push decision for the first alarm information.
[0237] In some embodiments, the decision matching object includes a decision matching rule library, which contains at least one decision matching rule constructed based on at least one of the annotated visual samples, and the decision matching rule library is used to determine the push decision of the first alarm information.
[0238] Regarding how the server trains a decision matching model based on labeled visual samples and how it constructs a decision matching rule base based on labeled visual samples, please refer to the above description of an alarm information processing method applied to the server. The description will not be repeated here.
[0239] A method for assisting in processing alarm information proposed in an embodiment of the present application can, in response to a labeling operation on a visual sample, label the visual sample to obtain a labeled visual sample containing labeled data; and upload the labeled visual sample to a server, so that the server can update or determine a decision matching object that meets the user's personalized alarm needs. The push decision of the first alarm information from the first camera device generated based on such a decision matching object can give a push decision result that meets the user's expectations, thereby solving the problems commonly existing in related technologies such as rigid rules / models, low user participation, weak personalized adaptation capabilities, high false alarm rate, and coarse alarm granularity.
[0240] In some embodiments, the method for assisting in processing alarm information proposed in this application may further include:
[0241] The first alarm information is received, wherein the first alarm information is pushed by the server when the push decision is to push the first alarm information or push the first alarm information after raising the alarm level. It is easy to understand that this push is more in line with the user's expectations.
[0242] In some embodiments, the method for assisting in processing alarm information proposed in this application may further include:
[0243] Pulling the suppressed second alarm information from the server and displaying it;
[0244] In response to the restoration alarm operation for the second alarm information, restoration information for the second alarm information is sent to the server, wherein the restoration information is used by the server to update the decision matching object.
[0245] In some embodiments, the client may provide a filtering record interface, through which the following functions may be implemented:
[0246] 1) Display all alarm messages that have been "intelligently filtered" (suppressed from immediate push);
[0247] 2) The alarm entry displays the "reason for being filtered", for example, "highly matches decision matching rule X";
[0248] 3) A "restore to normal alarm" button is provided. In response to the user's click operation on restoring the alarm to normal, the alarm will be moved to the message center and a matching threshold adjustment prompt for the corresponding decision matching rule X will be triggered.
[0249] Through the above embodiments, the decision matching object can be iteratively updated according to the dynamic changes of the user's alarm needs, so that the decision matching object is more in line with the user's personalized alarm matching needs, and the accuracy of alarm information push is further improved.
[0250] In some embodiments, the client may further provide a decision matching object management interface, which may be located in the settings center of the client. The user may use the decision matching object management interface to implement the following functions:
[0251] 1) In response to the user's decision matching rule display operation, the decision matching rules created by the user are displayed (including creation time, key screenshots, tag information, etc.);
[0252] 2) In response to the user's decision matching rule modification operation, perform operations such as deletion, modification, and deactivation on the decision matching rule;
[0253] 3) In response to the user's decision matching rule server-side validation operation, the decision matching rule is synchronized to the server-side for validation.
[0254] In the method of assisting in processing alarm information proposed in the embodiment of this application, when users can interact with the server through the client and express their intention to filter / strengthen alarms, several interface changes and interaction designs are involved. These changes mainly revolve around "reviewing alarm content → active annotation → uploading annotated visual samples → deciding matching objects", which can significantly improve the user experience. For example:
[0255] 1) After receiving the alert information pushed by the server, the client can click to enter the alert details page; alternatively, the user can use the video / image in the client's local gallery as a sample to be labeled;
[0256] 2) During the client-side playback of the alarm video, users can determine whether the event is a false alarm or requires increased attention;
[0257] 3) If it is determined to be a "false alarm" or requires "enhanced attention", you can actively enter the "marking mode";
[0258] 4) After completing visual selection and semantic annotation in annotation mode, submit it to the server to generate personalized decision matching rules and display them on the client;
[0259] 5) Users can view, edit or cancel previously defined decision matching rules at any time in the settings center;
[0260] 6) All filtered alarm information can be viewed in the "Filter Record" and retroactively restored.
[0261] Figure 6 This is a schematic diagram of the interactive process of an alarm information processing method provided by an embodiment of the present application. Figure 6 As shown, an alarm information processing method provided in an embodiment of the present application may include:
[0262] In step 601 , the first camera device 13 detects an alarm event and generates second alarm information including a second alarm video after detecting the alarm event.
[0263] Step 602 : The first camera device 13 uploads the second alarm information to the server 12 .
[0264] Step 603 : The server 12 pushes the second alarm information to the client 11 .
[0265] In step 604 , the client 11 annotates the second alarm video in response to the user's annotation operation to obtain an annotated visual sample.
[0266] In some embodiments, the client 11 may annotate the second alarm video in response to a marking operation performed by the user when confirming that the second alarm video is a false alarm or requires special attention, to obtain an annotated visual sample.
[0267] In step 605 , the client 11 uploads the labeled visual sample to the server 12 .
[0268] In step 606 , the server 12 performs multimodal feature extraction on the labeled visual sample to generate a second multimodal feature vector.
[0269] Step 607: construct a decision matching object based on the second multimodal feature vector.
[0270] In some embodiments, the decision matching object is a decision matching rule library, and step 607 may specifically include: generating a decision matching rule based on the second multimodal feature vector, the auxiliary decision label of the annotated visual sample and a matching threshold and storing it in the decision matching rule library.
[0271] In some embodiments, the decision matching object is a decision matching model, and step 607 may specifically include: training the decision matching model based on the second multimodal feature vector and the auxiliary decision label of the annotated visual sample.
[0272] In step 608 , after detecting the alarm event, the first camera device 13 generates first alarm information including a first alarm video.
[0273] Step 609 : The first camera device 13 uploads the first alarm information to the server 12 .
[0274] In step 610 , the server 12 performs multimodal feature extraction on the first alarm video to obtain a first multimodal feature vector, and generates a push decision for the first alarm information based on the first multimodal feature vector and a decision matching object.
[0275] Step 611 : When the push decision is to push the first alarm information, push the first alarm information to the client 11 .
[0276] In some embodiments, when the push decision is not to push the first alarm information, the server may suppress the push of the first alarm information and save the first alarm information, and then the client 11 may execute step 612 .
[0277] In step 612, the client 11 performs backtracking of the filtered alarm information.
[0278] In some embodiments, the client may provide a filtering record interface, through which the following functions may be implemented:
[0279] 1) Display all alarm messages that have been "intelligently filtered" (suppressed from immediate push);
[0280] 2) The alarm entry displays the "reason for being filtered", for example, "highly matches decision matching rule X";
[0281] 3) A "restore to normal alarm" button is provided. In response to the user's click operation on restoring the alarm to normal, the alarm will be moved to the message center and a matching threshold adjustment prompt for the corresponding decision matching rule X will be triggered.
[0282] Through the above embodiments, the decision matching object can be iteratively updated according to the dynamic changes of the user's alarm needs, so that the decision matching object is more in line with the user's personalized alarm matching needs, and the accuracy of alarm information push is further improved.
[0283] As can be seen from the above description, the alarm information processing method proposed in the embodiment of the present application can achieve at least one of the following beneficial effects:
[0284] 1) Users can participate in the construction of alarm information screening rules (decision matching objects) through intuitive interaction on the client;
[0285] 2) Different from the fixed rules / static models built into the camera device, the embodiments of the present application can achieve active personalized matching;
[0286] 3) By performing multimodal content understanding (image + action + semantics) on alarm videos, matching accuracy can be enhanced, thereby improving the accuracy of alarm information push, and the pushed alarm information is more in line with user expectations;
[0287] 4) Suppressed or filtered alarm information can be traced back and restored, which improves the controllability and security of alarms.
[0288] An embodiment of the present application further provides an alarm system, which includes a first camera device, a client, and a server.
[0289] The client is configured to, in response to a labeling operation on a target visual sample, label the target visual sample, obtain a labeled visual sample containing the labeling data, and upload the labeled visual sample to the server;
[0290] The server is used to receive first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video; perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector; and generate a push decision for the first alarm information based on the first multimodal feature vector and a decision matching object, wherein the decision matching object is determined based on the labeled visual sample.
[0291] An alarm system provided in an embodiment of the present application can achieve the same technical effect as an alarm information processing method provided in an embodiment of the present application, and will not be described in detail.
[0292] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0293] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 7 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.
[0294] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0295] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0296] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming an alarm information processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0297] Receiving first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video;
[0298] Performing multimodal feature extraction on the first alarm video to generate a first multimodal feature vector;
[0299] Based on the first multimodal feature vector and the decision matching object, a push decision for the first alarm information is generated, wherein the decision matching object is determined based on the annotated visual sample uploaded by the client, and the annotated visual sample is obtained by the user annotating the visual sample on the client.
[0300] Alternatively, the processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a device for assisting in processing alarm information at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0301] In response to the labeling operation on the visual sample, labeling the visual sample to obtain a labeled visual sample containing labeling data;
[0302] Uploading the labeled visual sample to the server;
[0303] The visual annotation sample is used by the server to update or determine a decision matching object, and the decision matching object is used to determine a push decision of the first alarm information from the first camera device.
[0304] The above application Figure 7 The methods performed by the alarm information processing device / device assisting in processing alarm information disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0305] The electronic device may also perform Figure 2 、 Figure 3 or Figure 4 The method described above realizes the alarm information processing device in Figure 8 、 Figure 9 or Figure 10 Alternatively, the electronic device may also perform Figure 5 The method and the device for assisting in processing the alarm information are implemented in Figure 11 The functions in the illustrated embodiments will not be described in detail in the embodiments of the present application.
[0306] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0307] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by a portable electronic device including multiple target applications, can enable the portable electronic device to execute Figure 2 、 Figure 3 、 Figure 4 or Figure 5 The method of the embodiment shown.
[0308] The embodiment of the present application further proposes a computer program product comprising instructions, wherein when a computer runs the instructions of the computer program product, the computer executes the following Figure 2 、 Figure 3 、 Figure 4 or Figure 5 The method of the embodiment shown.
[0309] Figure 8 This is a structural diagram of an alarm information processing device 800 provided in an embodiment of the present application. Figure 8 In a software implementation, the alarm information processing device 800 may include: an information receiving module 801, a first feature extraction module 802 and a decision generation module 803.
[0310] The information receiving module 801 is configured to receive first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video.
[0311] In some embodiments, the first alarm information may be generated and uploaded to the server after the first camera device identifies an alarm event using a locally built-in recognition algorithm. The first alarm information may include a video clip corresponding to the time period during which the alarm event occurred, namely, the first alarm video. The recognition algorithm built into the first camera device may include, but is not limited to, at least one of human detection, area intrusion detection, and sound source detection algorithms.
[0312] The first feature extraction module 802 is configured to perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector.
[0313] In some embodiments, the first multimodal feature vector may include at least two of the following:
[0314] Visual feature vector;
[0315] Action and behavior feature vectors;
[0316] Semantic embedding feature vector.
[0317] In some embodiments, since the first warning video newly generated by the first camera device has no user annotation and therefore no semantic label, the first feature extraction module 802 may be specifically configured to:
[0318] Extracting key frames from the first alarm video;
[0319] Extracting a feature vector of a target visual object from the key frame based on a target detection model to obtain a visual feature vector of the first alarm video;
[0320] Based on the key frame and the video frames before and after the key frame, identifying the motion trajectory of the target visual object, and generating the action and behavior feature vector of the first warning video;
[0321] A binary feature quantity consisting of the visual feature vector and the action and behavior feature vector of the first alarm video is used as the first multimodal feature vector.
[0322] In some embodiments, the first feature extraction module 802 may be specifically configured to:
[0323] Extracting a feature vector of a target visual object from the key frame based on a target detection model to obtain a visual feature vector of the first alarm video;
[0324] Based on the key frame and the video frames before and after the key frame, identifying the motion trajectory of the target visual object, and generating the action and behavior feature vector of the first warning video;
[0325] Using a multimodal model to perform semantic content reasoning on video images (such as key frames) or key objects (such as visual objects in key frames) in the first alarm video to generate a semantic embedding feature vector of the first alarm video;
[0326] A triplet feature quantity consisting of the visual feature vector, the action and behavior feature vector, and the semantic embedding feature vector of the first alarm video is used as the first multimodal feature vector.
[0327] The decision generation module 803 is used to generate a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object, wherein the decision matching object is determined based on the annotated visual sample uploaded by the client, and the annotated visual sample is obtained by the user annotating the visual sample on the client.
[0328] The decision matching object may include but is not limited to at least one of a decision matching rule library and a decision matching model, and the decision matching rule library contains at least one decision matching rule constructed according to at least one annotated visual sample.
[0329] In some embodiments, the decision object includes a decision matching rule library, such as Figure 9 As shown, an alarm information processing device 800 proposed in this application may further include: a sample receiving module 804 , a second feature extraction module 805 and a rule establishment module 806 .
[0330] The sample receiving module 804 is configured to receive annotated visual samples uploaded by the client.
[0331] The annotated visual sample is obtained by the user annotating the client. The annotated visual sample may include but is not limited to at least one of the following:
[0332] 1) The server pushes a second alarm video in the historical alarm information to the client;
[0333] 2) a third alarm video in the alarm information that is pulled by the client from the server and that is suppressed from being pushed by the server;
[0334] 3) local video of the client;
[0335] 4) The local image of the client.
[0336] In some embodiments, if the labeled visual sample is the second warning video or the third warning video, its labeling data may include at least one of the following:
[0337] 1) Decision-making support labels;
[0338] 2) Video-level semantic tags;
[0339] 3) contextual summary of the annotated video frame;
[0340] 4) visually annotated objects in the annotated video frames;
[0341] 5) The coordinates of the visual annotation area in the annotated video frame and the corresponding timestamp;
[0342] 6) semantic labels of visually annotated objects in the annotated video frames;
[0343] 7) Marking method information.
[0344] In some embodiments, if the annotated visual sample is a local video of the client, its annotated data may also include at least one of the above seven items of data.
[0345] In some embodiments, if the annotated visual sample is a local image on the client, its annotated data may include at least one of the following:
[0346] 1) Decision-making support labels;
[0347] 2) Image-level semantic labels;
[0348] 3) visually annotated objects in the image;
[0349] 4) Coordinates of the visual annotation area in the image;
[0350] 5) semantic labels of visually annotated objects in the image;
[0351] 6) Marking method information.
[0352] The second feature extraction module 805 is configured to perform multimodal feature extraction on the annotated visual sample to generate a second multimodal feature vector.
[0353] After receiving the annotated visual samples uploaded by the client, the server can start the multimodal feature parsing process to convert the visual annotated samples into machine-recognizable multimodal feature vectors.
[0354] In some embodiments, the multimodal feature vector for annotating a visual sample may include at least two of the following:
[0355] Visual feature vector;
[0356] Action and behavior feature vectors;
[0357] Semantic labels are embedded into feature vectors.
[0358] The rule establishing module 806 is configured to generate a decision matching rule based on the second multimodal feature vector and store the rule in the decision matching rule library.
[0359] In some embodiments, the rule establishment module 806 can be specifically used to: generate a decision matching rule based on the auxiliary decision label of the annotated visual sample, the second multimodal feature vector and the matching threshold, and store it in the decision matching rule library. The matching threshold is used to determine whether the first multimodal feature vector of the first alarm video contained in the first alarm information matches the multimodal feature vector in the decision matching rule. Generally, if the degree of matching between the first multimodal feature vector and the second multimodal feature vector in a decision matching rule is higher than the matching threshold, it means that the first alarm video in the first alarm information matches the decision matching rule.
[0360] In some embodiments, different users and / or clients are provided with their own decision matching rule libraries. Therefore, the server can store the above-mentioned decision matching rules and the corresponding auxiliary decision labels (forward labels / reverse labels) of the annotated visual samples in the decision matching rule library of the corresponding user and / or client, which is equivalent to building a personalized decision matching rule library for different users.
[0361] In some embodiments, the alarm information processing device 800 proposed in this application may further include:
[0362] A rule clustering module, used for clustering the decision matching rules in the decision matching rule library;
[0363] The rule merging module is used to merge the decision matching rules belonging to the same cluster center in the decision matching rule library, and set a matching threshold for the merged decision matching rules.
[0364] After the matching decision database is constructed, the decision generation module 803 can be used to:
[0365] sequentially determining a comprehensive similarity between the first multimodal feature vector and a second multimodal feature vector in each decision matching rule in the decision matching rule library to obtain a target decision matching rule, wherein the comprehensive similarity between the second multimodal feature vector in the target decision matching rule and the first multimodal feature vector is higher than a corresponding matching threshold;
[0366] After obtaining the target decision matching rule, a push decision for the first alarm information is generated based on the auxiliary decision tag of the target decision matching rule.
[0367] In some embodiments, when the auxiliary decision tag of the target decision matching rule is a forward tag, the push decision may include: pushing the first alarm information or pushing the first alarm information after raising the alarm level; when the auxiliary decision tag of the target decision matching rule is a reverse tag, the push decision may include: not pushing the first alarm information or pushing the first alarm information after lowering the alarm level.
[0368] In some embodiments, the decision object includes a decision matching model, such as Figure 10 As shown, an alarm information processing device 800 proposed in this application may further include: a sample receiving module 804, a second feature extraction module 805 and a model training module 807.
[0369] The sample receiving module 804 is configured to receive annotated visual samples uploaded by the client.
[0370] The second feature extraction module 805 is configured to perform multimodal feature extraction on the annotated visual sample to generate a second multimodal feature vector.
[0371] The model training module 807 is used to train the decision matching model according to the second multimodal feature vector.
[0372] The decision matching model can be a neural network. During training, the input of the decision matching model can be the labeled visual sample, and the output can be the push decision of the labeled visual sample; during application, the input of the decision matching model can be the alarm video in the alarm information uploaded by the camera device, and the output can be the push decision of the alarm information.
[0373] At this time, the decision generation module 803 can be specifically configured to: input the first multimodal feature vector into the decision matching model, and use the decision matching model to generate a push decision for the first alarm information. The push decision can be to push the first alarm information or to push the first alarm information after raising the alarm level, or the push decision can be to not push the first alarm information or to push the first alarm information after lowering the alarm level.
[0374] In some embodiments, the alarm information processing device 800 proposed in the present application may further include: an information processing module, which processes the first alarm information according to the push decision.
[0375] Specifically, the information processing module may be configured to: push the first alarm information to the client if the push decision is to push the first alarm information; or push the first alarm information after raising the alarm level if the push decision is to push the first alarm information;
[0376] When the push decision is not to push the first alarm information, the push of the first alarm information is suppressed and the first alarm information is stored.
[0377] In some embodiments, an alarm information processing device 800 proposed in the present application may also include: a decision matching object update iteration module, which is used to iteratively update the decision matching object based on at least one of the newly uploaded annotated visual samples by the client, the suppressed alarm information restored by the client, and the false alarm information recorded by the client.
[0378] In some embodiments, the alarm information processing device 800 proposed in the present application may further include: a sharing module, which is used to synchronize the decision matching object to other clients, that is, to share the decision matching object across devices.
[0379] The alarm information processing device 800 provided in the embodiment of the present application can also execute Figure 2 、 Figure 3 or Figure 4 method and implement Figure 2 、 Figure 3 or Figure 4 The functions of the embodiments shown in the figure are the same and the technical effects are achieved, so the embodiments of the present application will not be described in detail here.
[0380] Figure 11 This is a schematic diagram of the structure of an apparatus 1100 for assisting in processing alarm information provided by an embodiment of the present application. Figure 11 In a software implementation, the alarm information processing device 1100 may include: a marking module 1101 and an uploading module 1102.
[0381] The labeling module 1101 is configured to label the visual sample in response to a labeling operation on the visual sample, and obtain a labeled visual sample containing labeling data.
[0382] The annotated visual sample may include but is not limited to at least one of the following:
[0383] 1) The server pushes a second alarm video in the historical alarm information to the client;
[0384] 2) a third alarm video in the alarm information that is pulled by the client from the server and that is suppressed from being pushed by the server;
[0385] 3) local video of the client;
[0386] 4) The local image of the client.
[0387] In some embodiments, if the annotated visual sample is the second alarm video or the third alarm video, its annotation data may include at least one of the following: an auxiliary decision label, a video-level semantic label, a context summary of the annotated video frame, the visually annotated object in the annotated video frame, the visually annotated area coordinates and corresponding timestamp in the annotated video frame, the semantic label of the visually annotated object in the annotated video frame, and annotation method information.
[0388] In some embodiments, if the annotated visual sample is a local video of the client, its annotated data may also include at least one of the above seven items of data.
[0389] In some embodiments, if the annotated visual sample is a local picture of the client, its annotation data may include at least one of the following: auxiliary decision labels, picture-level semantic labels, visual annotation objects in the picture, visual annotation area coordinates in the picture, semantic labels of visual annotation objects in the picture, and annotation method information.
[0390] The uploading module 1102 is configured to upload the annotated visual sample to a server.
[0391] The visual annotation sample is used by the server to update or determine a decision matching object, and the decision matching object is used to determine a push decision of the first alarm information from the first camera device.
[0392] In some embodiments, the decision matching object includes a decision matching model, which is trained based on the labeled visual samples, and the decision matching model is used to generate a push decision for the first alarm information.
[0393] In some embodiments, the decision matching object includes a decision matching rule library, which contains at least one decision matching rule constructed based on at least one of the annotated visual samples, and the decision matching rule library is used to determine the push decision of the first alarm information.
[0394] In some embodiments, the apparatus 1100 for assisting in processing alarm information proposed in this application may further include: an information receiving module configured to receive the first alarm information, wherein the first alarm information is pushed by the server when the push decision is to push the first alarm information or to push the first alarm information after raising the alarm level. It is understood that this push is more in line with user expectations.
[0395] In some embodiments, the device 1100 for assisting in processing alarm information proposed in this application may further include: a display module and a recovery module.
[0396] a display module, configured to pull the suppressed second alarm information from the server and display it;
[0397] A recovery module is used to send recovery information for the second alarm information to the server in response to a recovery alarm operation for the second alarm information, wherein the recovery information is used by the server to update the decision matching object.
[0398] The alarm information processing device 1100 provided in the embodiment of the present application can also execute Figure 5 method and implement Figure 5 The functions of the embodiments shown in the figure are the same and the technical effects are achieved, so the embodiments of the present application will not be described in detail here.
[0399] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0400] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0401] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0402] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0403] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.
Claims
1. A method for processing alarm information, characterized in that: Applied to the server, the method includes: Receiving first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video; Performing multimodal feature extraction on the first alarm video to generate a first multimodal feature vector; Based on the first multimodal feature vector and the decision matching object, a push decision for the first alarm information is generated, wherein the decision matching object is determined based on the annotated visual sample uploaded by the client, and the annotated visual sample is obtained by the user annotating the visual sample on the client.
2. The method according to claim 1, characterized in that The decision matching object includes a decision matching rule library, wherein the decision matching rule library includes at least one decision matching rule constructed based on at least one of the annotated visual samples. Before generating a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object, the method further includes: receiving the annotated visual sample uploaded by the client; performing multimodal feature extraction on the annotated visual sample to generate a second multimodal feature vector; Based on the second multimodal feature vector, a decision matching rule is generated and stored in the decision matching rule library.
3. The method according to claim 2, characterized in that The annotated data of the annotated visual sample includes an auxiliary decision label, wherein generating a decision matching rule based on the second multimodal feature vector includes: A decision matching rule is generated based on the auxiliary decision label of the annotated visual sample, the second multimodal feature vector, and a matching threshold.
4. The method according to claim 3, characterized in that The generating a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object includes: sequentially determining a comprehensive similarity between the first multimodal feature vector and a second multimodal feature vector in each decision matching rule in the decision matching rule library to obtain a target decision matching rule, wherein the comprehensive similarity between the second multimodal feature vector in the target decision matching rule and the first multimodal feature vector is higher than a corresponding matching threshold; After obtaining the target decision matching rule, a push decision for the first alarm information is generated based on the auxiliary decision tag of the target decision matching rule.
5. The method according to claim 4, characterized in that In the case where the auxiliary decision tag of the target decision matching rule is a positive tag, the push decision includes: pushing the first alarm information or pushing the first alarm information after raising the alarm level; or, In the case that the auxiliary decision tag of the target decision matching rule is a reverse tag, the push decision includes: not pushing the first alarm information or pushing the first alarm information after lowering the alarm level.
6. The method according to claim 3, characterized in that Before generating a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object, the method further includes: Clustering the decision matching rules in the decision matching rule library; The decision matching rules belonging to the same cluster center in the decision matching rule library are merged, and a matching threshold is set for the merged decision matching rules.
7. The method according to claim 1, characterized in that The decision matching object includes a decision matching model, and the decision matching model is trained based on the labeled visual sample. Before generating a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object, the method further includes: receiving the annotated visual sample uploaded by the client; performing multimodal feature extraction on the annotated visual sample to generate a second multimodal feature vector; The decision matching model is trained according to the second multimodal feature vector.
8. The method according to claim 7, characterized in that The generating a push decision for the first alarm information based on the first multimodal feature vector and the decision matching object includes: The first multimodal feature vector is input into the decision matching model, and the decision matching model is used to generate a push decision for the first alarm information.
9. The method according to claim 2 or 7, characterized in that The annotated visual sample includes any of the following: The server pushes a second alarm video in the historical alarm information to the client; a third alarm video in the alarm information pulled by the client from the server and suppressed from being pushed by the server; The local video of the client; The local image of the client.
10. The method according to claim 9, characterized in that In a case where the annotated visual sample is the second alarm video or the third alarm video, the annotated data of the annotated visual sample includes at least one of the following: Decision-making support labels; Video-level semantic labeling; Contextual summary of the annotated video frame; Visually annotated objects in the annotated video frames; The coordinates of the visual annotation area in the annotated video frame and the corresponding timestamp; Semantic labels of visually annotated objects in the annotated video frames; Marking method information.
11. The method according to claim 10, characterized in that The second multimodal feature vector includes: Visual feature vector; Action and behavior feature vectors; Semantic labels are embedded into feature vectors.
12. The method according to claim 11, characterized in that The performing multimodal feature extraction on the annotated visual sample to generate a second multimodal feature vector includes: Extracting feature vectors and semantic vectors of the annotated visual objects from the annotated video frames of the annotated visual samples based on the object detection model, and generating a visual feature vector; Based on the annotated video frame in the annotated visual sample and the video frames before and after the annotated video frame, identifying the motion trajectory of the annotated visual object and generating an action and behavior feature vector; Based on the semantic labels of the annotated visual samples, a semantic label embedding feature vector is generated.
13. The method according to any one of claims 1-8, 10-12, characterized in that Also includes: Process the first alarm information according to the push decision.
14. The method according to any one of claims 1-8, 10-12, characterized in that Also includes: The decision matching object is iteratively updated based on at least one of the annotated visual sample newly uploaded by the client, the suppressed alarm information restored by the client, and the false alarm information recorded by the client.
15. A method for assisting in processing alarm information, characterized in that: Applied to a client, the method includes: In response to the labeling operation on the visual sample, labeling the visual sample to obtain a labeled visual sample containing labeling data; Uploading the labeled visual sample to the server; The visual annotation sample is used by the server to update or determine a decision matching object, and the decision matching object is used to determine a push decision of the first alarm information from the first camera device.
16. The method according to claim 15, characterized in that The decision matching object includes a decision matching model, the decision matching model is trained based on the annotated visual sample, and the decision matching model is used to generate a push decision for the first alarm information; and / or, The decision matching object includes a decision matching rule library, which contains at least one decision matching rule constructed according to at least one of the annotated visual samples. The decision matching rule library is used to determine the push decision of the first alarm information.
17. The method according to claim 15, characterized in that The annotated visual sample includes any of the following: The server pushes a second alarm video in the historical alarm information to the client; a third alarm video in the alarm information pulled by the client from the server and suppressed from being pushed by the server; The local video of the client; The local image of the client.
18. The method according to claim 17, characterized in that In a case where the annotated visual sample is the second alarm video or the third alarm video, the annotated data of the annotated visual sample includes at least one of the following: Decision-making support labels; Video-level semantic labeling; Contextual summary of the annotated video frame; Visually annotated objects in the annotated video frames; The coordinates of the visual annotation area in the annotated video frame and the corresponding timestamp; Semantic labels of visually annotated objects in the annotated video frames; Marking method information.
19. The method according to claim 15, characterized in that Also includes: The first alarm information is received, wherein the first alarm information is pushed by the server when the push decision is to push the first alarm information or to push the first alarm information after increasing the alarm level.
20. The method according to any one of claims 15 to 19, characterized in that: Also includes: Pulling the suppressed second alarm information from the server and displaying it; In response to the restoration alarm operation for the second alarm information, restoration information for the second alarm information is sent to the server, wherein the restoration information is used by the server to update the decision matching object.
21. An alarm system, characterized in that: The system includes a first camera device, a client and a server; The client is configured to, in response to a labeling operation on a target visual sample, label the target visual sample, obtain a labeled visual sample containing the labeling data, and upload the labeled visual sample to the server; The server is used to receive first alarm information reported by a first camera device, wherein the first alarm information includes a first alarm video; perform multimodal feature extraction on the first alarm video to generate a first multimodal feature vector; and generate a push decision for the first alarm information based on the first multimodal feature vector and a decision matching object, wherein the decision matching object is determined based on the labeled visual sample.
22. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 20.
23. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 20.
24. A computer program product comprising instructions, characterized in that When a computer runs the instructions of the computer program product, the computer performs the method according to any one of claims 1 to 20.
Citation Information
Patent Citations
Monitoring method, device and equipment
CN118736478A
Multi-modal data processing method and device, equipment and storage medium
CN119293518A
Monitoring and early warning method and device based on multi-modal large model, equipment and medium
CN119946226A
Multimodal heterogeneous feature fusion-based compact video event description method
WO2023050295A1
Video feature extraction method and apparatus, video generation method and apparatus, and medium and device
WO2025092911A1
Cited By
Alarm method, electronic equipment and program product
CN121686366A