Video monitoring and warning method and device, computer equipment and storage medium
By fusing video and text rule features into a multimodal large model, the problem of insufficient intelligence in property security monitoring systems has been solved, enabling timely identification and risk classification of abnormal events, and improving the level of intelligence in property security.
Patent Information
- Application Number
- CN202510924477.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-28
AI Technical Summary
Existing property security monitoring systems lack sufficient intelligence, have slow response times and high error rates, and are unable to effectively identify complex behavioral patterns and conduct dynamic risk assessments.
By collecting structured data related to property security, collecting and preprocessing image data to form video features and text rule features, and embedding adapters in the visual big model, the multimodal big model is used for weighted fusion to determine abnormal events and conduct dynamic risk assessment.
It enables the identification and risk classification of abnormal events in property scenarios, improves the intelligence level and risk response capability of property security monitoring, reduces the reliance on a large amount of labeled data, and enhances the ability to identify complex behaviors.
Smart Images

Figure CN120853102A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of large-scale artificial intelligence models and computer vision technology, and in particular to a video monitoring and alarm method, apparatus, computer equipment, and storage medium. Background Art
[0002] Currently, property security monitoring systems generally face the challenge of insufficient intelligence, a problem particularly prominent in specific scenarios such as non-motorized vehicle management and personnel access control. Traditional monitoring systems have relatively limited functionality, only capable of passively recording video and lacking the ability to identify abnormal behavior in real time. For example, in a real-world scenario, there was a case where an electric scooter was pushed out of Gate 5 by a stranger; there was a significant time lag between the incident and the resident noticing and reporting it. Moreover, relying on manual review of surveillance footage is highly susceptible to missed detections.
[0003] While existing computer vision algorithms (such as object detection algorithms) can identify vehicles and people in videos, they lack the ability to understand business rules. This prevents these algorithms from effectively linking textual rules like "non-motorized vehicles are prohibited in car lanes" with actual cart-pushing behavior in videos, hindering their ability to achieve greater security benefits.
[0004] Furthermore, widely used general-purpose video big data models (such as VideoLLaMA) have many limitations due to a lack of specific optimization for property scenarios. For example, they struggle to identify complex behavioral patterns such as "tailgating" and "pretending to be a resident." More critically, current property security monitoring systems lack multimodal information fusion capabilities. They cannot organically combine visual data (such as unlocked vehicles), spatiotemporal context information (such as special time periods and weather conditions like nighttime and rainy days), and domain knowledge (such as access control permissions), thus failing to achieve dynamic risk assessment. This lack of capability leads to frequent security risks such as "unauthorized personnel infiltrating," posing a significant challenge to property security management. Summary of the Invention
[0005] This invention provides a video monitoring and alarm method, device, computer equipment, and storage medium, aiming to solve the problems of insufficient intelligence, slow response, and high false judgment rate in existing property security monitoring systems.
[0006] In a first aspect, embodiments of the present invention provide a video monitoring and alarm method based on a multimodal large model, including:
[0007] Collect structured data related to property security, including multiple text rules;
[0008] Image data from property video surveillance is collected, and the image data is preprocessed to obtain video features; the structured data is preprocessed to obtain text rule features.
[0009] An adapter is embedded in the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Anomaly events are determined based on the fused features to obtain the determination result.
[0010] If an abnormal event is found in the determination result, the abnormal event is classified into risk levels through the dynamic risk assessment module to obtain the classification result, and an alarm is triggered based on the classification result.
[0011] Secondly, embodiments of the present invention provide a video monitoring and alarm device based on a multimodal large model, comprising:
[0012] A collection unit is used to collect structured data related to property security, including business rules in text form;
[0013] The preprocessing unit is used to collect image data from property video surveillance, preprocess the image data to obtain video features, and preprocess the structured data to obtain text rule features.
[0014] The fusion unit is used to embed an adapter into the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Abnormal events are judged based on the fused features to obtain the judgment result.
[0015] The assessment and alarm unit is used to classify the risk of the abnormal event through the dynamic risk assessment module if there is an abnormal event in the judgment result, obtain the classification result, and trigger an alarm according to the classification result.
[0016] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video monitoring and alarm method based on a multimodal large model as described in the first aspect.
[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the video monitoring and alarm method based on a multimodal large model as described in the first aspect.
[0018] This invention provides a video monitoring and alarm method, device, computer equipment, and storage medium. It collects structured data (including multiple text rules) related to property security and acquires and preprocesses property video surveillance image data to obtain video features. Simultaneously, it preprocesses the structured data to obtain text rule features. Then, an adapter is embedded in the main structure of a visual large model to form a multimodal large model. Transfer learning is used to weightedly fuse the video features and text rule features to obtain fused features, which are then used to determine abnormal events. If an abnormal event exists, a dynamic risk assessment module performs risk classification and triggers an alarm. This achieves deep fusion of text rules and video image data, effectively identifying abnormal events in property scenarios, classifying risks of abnormal events, and providing timely alarms, thereby improving the intelligence level and risk response capability of property security monitoring. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating a video monitoring and alarm method based on a multimodal large model, provided in an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of a sub-process of a video monitoring and alarm method based on a multimodal large model provided in an embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of another sub-process of a video monitoring and alarm method based on a multimodal large model provided in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of another sub-process of a video monitoring and alarm method based on a multimodal large model provided in an embodiment of the present invention;
[0024] Figure 5 This is a schematic diagram of the structure of a video monitoring and alarm device based on a multimodal large model provided in an embodiment of the present invention;
[0025] Figure 6 A schematic diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0028] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] Please see Figure 1 , Figure 1 This is a flowchart illustrating a video monitoring and alarm method based on a multimodal large model, provided by an embodiment of the present invention. The method includes steps S100 to S400:
[0031] S100. Collect structured data related to property security, wherein the structured data includes multiple text rules;
[0032] In this embodiment, firstly, various business rules, operational procedures, and security requirements in property security management are reviewed to identify the types of abnormal behaviors that need to be monitored, such as non-motorized vehicles entering car lanes, unauthorized personnel tailgating, and unauthorized vehicle parking. Then, these rules are converted into structured text data, such as explicit text rules like "Non-motorized vehicles are prohibited from entering car lanes" and "Unregistered personnel are prohibited from entering," and organized into a structured dataset. Simultaneously, other structured information related to these rules is collected, such as access control permission lists, vehicle registration information, and personnel whitelists, ensuring that the structured data covers all aspects of property security management and provides foundational data support for subsequent feature extraction and model training.
[0033] S200. Collect image data from the property video surveillance, and preprocess the image data to obtain video features; preprocess the structured data to obtain text rule features;
[0034] In this embodiment, high-definition cameras are deployed in key areas of the property to collect video surveillance image data in real time, ensuring coverage of all scenes requiring monitoring, such as community entrances and exits, parking lots, and corridors. The collected image data undergoes preprocessing, including noise reduction, contrast enhancement, and resolution adjustment, to improve image quality. Subsequently, computer vision technologies, such as object detection and behavior recognition algorithms, are used to extract key features from the preprocessed image data, such as vehicle type, personnel movement trajectories, and object positions, forming video features. Simultaneously, the structured text rule data collected in step S100 is preprocessed using natural language processing techniques, such as word embedding and text vectorization, to convert the text rules into numerical text rule features for subsequent fusion analysis with video features.
[0035] In one embodiment, if Figure 2 As shown, step S200 includes steps S210 to S212:
[0036] S210. Preprocess the image data using a SlowFast network structure to obtain video features F. v ;
[0037] S211. Encode the text rules using the BERT language encoding model, and convert each text rule into a text rule vector according to the following formula:
[0038]
[0039] in, d t 3D real vector space;
[0040] S212. Use average pooling to aggregate the text rule vectors to form text rule features F. t .
[0041] In this embodiment, in step S210, a SlowFast network structure is used to preprocess the acquired property video surveillance image data. The SlowFast network consists of a slow path (Slow branch) and a fast path (Fast branch). The slow path extracts semantic features of the video at a lower frame rate, focusing on scene content (such as person identity and background environment), and is used to capture spatial semantic information in the image. The fast path extracts behavioral details at a higher frame rate, capturing fast actions (such as pushing a cart, running, following, etc.), and is used to process temporal motion information. Through dual-path feature extraction, static scene features and dynamic behavioral features in the video are effectively obtained, and the final video feature F is output. v It comprehensively represents the video content.
[0042] In step S211, the text rules in the structured data are encoded using the BERT language encoding model. The BERT model, through pre-training, possesses rich language representation capabilities, enabling it to map each text rule to a high-dimensional semantic space. Specifically, according to the formula above, the text rules are converted into a dimension d... t A real vector, where This represents the semantic representation of text rules after BERT encoding, ensuring that the semantic information of the text rules is accurately extracted and quantified.
[0043] Furthermore, the text rule vectors generated in step S211 are aggregated using the average pooling method. By calculating the average value of all text rule vectors across each dimension, multiple independent text rule vectors are compressed into a comprehensive text rule feature vector. This method, while preserving the semantic information of each rule, reduces the feature dimensionality, forming a unified and compact text rule feature representation, which facilitates subsequent fusion analysis with video features.
[0044] S300. An adapter is embedded in the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Anomaly events are determined based on the fused features to obtain a determination result.
[0045] In this embodiment, firstly, a pre-trained large-scale visual model (such as ResNet, VisionTransformer, etc.) is selected, and an adapter module is embedded in its backbone network structure. The adapter adopts a lightweight neural network structure, such as a bottleneck structure or a gating mechanism, to adjust the fusion ratio of different modal features.
[0046] Next, the video features and text rule features obtained in step S200 are input into the modified multimodal large model through transfer learning. During model training, a weighted fusion strategy is used to dynamically combine video features (visual information) and text rule features (semantic rules) through learnable weight parameters, enabling the model to adaptively balance the importance of the two modalities.
[0047] Furthermore, the fused features enter the model's classification layer, and after processing by a fully connected layer and activation function, the abnormal event determination result is output. The determination result is expressed in probability form. When the probability of a certain type of abnormal event exceeds a preset threshold, it is determined that the abnormal event has occurred, such as a non-motorized vehicle entering a car lane or an outsider tailgating into the vehicle.
[0048] In summary, through this weighted fusion and transfer learning approach, the model can fully utilize multimodal information to improve its ability to identify complex and abnormal behaviors in property scenarios, while reducing its reliance on large amounts of labeled data, improving training efficiency and model generalization ability, and achieving deep collaborative reasoning between business rules and visual data.
[0049] In one embodiment, step S300 includes step S310:
[0050] S310. The video feature F is calculated according to the following formula. v With the text rule feature F t The input is fed into the multimodal large model for weighted fusion to obtain fused features:
[0051] F = α·F v +β·F t
[0052] Where α and β are the fusion weights of visual features and text rule features, respectively, and α+β=1.
[0053] In this embodiment, during model training, the weight parameters α and β are automatically adjusted using the backpropagation algorithm, enabling the model to dynamically balance the contributions of video features and text rule features based on the characteristics of the input data. For example, when identifying the rule "No non-motorized vehicles allowed in car lanes," if a non-motorized vehicle is detected entering the car lane in the video, the model will assign a higher weight to the text rule features to strengthen the rule constraint; while in conventional monitoring scenarios, video features may be more relied upon for behavior recognition. Ultimately, the fused feature F combines visual information and semantic rules, providing a more comprehensive basis for subsequent anomaly event determination.
[0054] In one embodiment, if Figure 3 As shown, step S300 includes steps S320 to S321:
[0055] S320. Input the fused features into the BERT language encoding model and vectorize them according to the following formula to obtain the observation semantic vector:
[0056] f obs =F=Encoder(event_text)
[0057] Among them, f obs ∈R d R d Let represent a d-dimensional real vector space, Encoder represent the encoder, event represent video events, and text represent semantic text;
[0058] S321. Input the text rules into the language encoder according to the following formula to vectorize them, and obtain a set of rule vectors:
[0059]
[0060] Each rule vector The semantic representation of the i-th text rule, text i) This represents the i-th semantic text.
[0061] In this embodiment, in step S320, the fused features obtained in step S300 are input into the BERT language encoding model for further vectorization processing. The encoder of the BERT model maps the fused features to a higher-level semantic space, generating observation semantic vectors.
[0062] In step S321, the original structured text rules are individually input into the language encoder (which can be the same BERT encoder as in step S320 or another dedicated encoder) for vectorization processing to generate a set of rule vectors. Each text rule is encoded and converted into a corresponding rule vector. All rule vectors together constitute a set of rule vectors, which are used for subsequent comparative analysis or rule matching with observed semantic vectors to enhance the model's understanding and execution capabilities of security rules.
[0063] It should be noted that when performing the first abnormal event judgment, the rule base has not yet been built in the multimodal large model. A set of rule vectors, that is, the rule base, needs to be built through step S321.
[0064] In one embodiment, step S300 further includes steps S330 to S331:
[0065] S330. Perform similarity matching between the observed semantic vector and the rule vector set according to the following formula:
[0066]
[0067] Among them, S i Indicates similarity; The observation semantic vector f obs With regular vectors The dot product of vectors, ||·|| denotes the Euclidean norm of the vector;
[0068] S331, When the similarity S i When the threshold θ is exceeded, the abnormal event is determined to match the i-th text rule.
[0069] In this embodiment, in step S330, the observation semantic vector generated in step S320 is matched with each rule vector in the rule vector set obtained in step S321 for similarity. Cosine similarity is used as the matching metric, and the calculation formula is shown above.
[0070] In step S331, a similarity threshold θ is set (typically ranging from [0,1]). When the similarity between the observed semantic vector and a rule vector exceeds this threshold, it is determined that the currently detected abnormal event matches the i-th text rule. For example, if the similarity between the observed semantic vector and the rule vector "No non-motorized vehicles allowed in car lane" exceeds the threshold, it is determined that the non-motorized vehicle entering the car lane in the current video violates this rule, thereby triggering the corresponding abnormal alarm process. In this way, the visual detection results can be accurately correlated with the preset security rules, improving the accuracy and interpretability of abnormal event determination.
[0071] In one embodiment, step S300 further includes steps S340 to S341:
[0072] S340. The judgment result is obtained by using the logical rule matching algorithm according to the following formula:
[0073]
[0074] Among them, w i Weights for the importance of text rules;
[0075] S341. When the determination result R exceeds the set threshold τ, the abnormal event is determined to be a high-risk event.
[0076] In this embodiment, a logical rule matching algorithm is used to comprehensively determine the similarity matching results obtained in step S330. Specifically, for cross-dimensional rule combinations (such as "stranger" + "restricted area" + "nighttime"), a knowledge graph containing entity nodes (people, regions, time) and rule edges is constructed, and the logical rule matching algorithm is used to complete relationship path mining and composite rule identification. This process can be modeled using a risk-weighted aggregation function, where R represents the determination result or a preliminary risk index.
[0077] A threshold τ is set (usually adjusted based on actual security needs). When the judgment result R exceeds this threshold τ, the current abnormal event is judged as a high-risk event. For example, if multiple high-weight rules (such as "unauthorized personnel entering a restricted area" or "abnormal detection of hazardous materials") are triggered simultaneously, the judgment result will significantly exceed the threshold, immediately marking it as a high-risk event and triggering an emergency alarm process. Through weight allocation and threshold judgment, the system can distinguish the severity of abnormal events and achieve risk-level response.
[0078] S400. If there is an abnormal event in the determination result, the abnormal event is classified into risk levels through the dynamic risk assessment module to obtain the classification result, and an alarm is triggered according to the classification result.
[0079] In this embodiment, the dynamic risk assessment module performs real-time risk classification of abnormal events based on multi-dimensional features. The core process includes: extracting key features from abnormal events, such as event type (intrusion, illegal parking, etc.), occurrence time (peak hours / night), occurrence area (core area / peripheral area), and historical frequency (number of times similar events have occurred).
[0080] Specifically, the preliminary risk index R obtained in step S340 reflects its deviation in the rule space; in order to further improve the context adaptability of the multimodal large model, dynamic information such as time, environment, and historical behavior are introduced to correct the preliminary risk index R, which can also be understood as risk classification to form a classification result.
[0081] Furthermore, different response strategies are adopted based on the risk level classification: For example, high-risk events are alerted by immediately triggering audible and visual alarms, sending SMS messages or notifications via mobile app to security personnel, and automatically locking cameras in the relevant area. The specific response process requires security personnel to confirm and handle the situation within 2 minutes and record the results. Medium-risk events are alerted by sending system notifications to security personnel and displaying event details on the monitoring screen. The specific response process requires security personnel to handle the situation within 10 minutes, with the system providing periodic reminders. Low-risk events are alerted by recording the event in a logbook, with only a summary notification in the daily report. The specific response process does not require immediate action but can be included in long-term analysis.
[0082] In one embodiment, if Figure 4 As shown, step S400 includes step S410:
[0083] S410. Calculate the grading results according to the following formula:
[0084] R′=γ·C+δ·T+∈·E
[0085] Where c represents the behavioral credibility of the currently detected abnormal event, T is the time weight, E is the environmental risk coefficient, γ, v and ∈ are pre-set adjustment weights, and γ+δ+∈=1.
[0086] In this embodiment, the dynamic risk assessment module comprehensively considers contextual factors such as time (e.g., nighttime), environment (e.g., rainy days), and frequency of historical events to dynamically adjust the monitoring sensitivity.
[0087] Furthermore, a multi-factor comprehensive assessment is used to classify abnormal events. Specifically, the behavioral credibility C, time weight T, and environmental risk coefficient E of the currently detected abnormal event are weighted and summed. Behavioral credibility C reflects the reliability of the abnormal event's behavior itself; time weight T considers the impact of the event's occurrence time on the risk level; and environmental risk coefficient E reflects the risk inherent in the event's environment. γ, δ, and ∈ are pre-set adjustment weights used to adjust the proportion of each factor in the final classification result, satisfying γ + δ + ∈ = 1. This ensures a reasonable allocation of weights for each factor, allowing the final classification result R' to comprehensively and reasonably reflect the overall risk level of the abnormal event, thus providing a basis for subsequent appropriate measures based on the classification results.
[0088] like Figure 5 As shown, this embodiment of the invention also provides a video monitoring and alarm device 500 based on a multimodal large model, including a collection unit 501, a preprocessing unit 502, a fusion unit 503, and an evaluation and alarm unit 504.
[0089] Collection unit 501 is used to collect structured data related to property security, the structured data including business rules in text form;
[0090] The preprocessing unit 502 is used to collect image data from property video surveillance, preprocess the image data to obtain video features, and preprocess the structured data to obtain text rule features.
[0091] The fusion unit 503 is used to embed an adapter in the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Abnormal events are judged based on the fused features to obtain a judgment result.
[0092] The assessment and alarm unit 504 is used to classify the abnormal event by the dynamic risk assessment module if there is an abnormal event in the judgment result, obtain the classification result, and trigger an alarm according to the classification result.
[0093] The video monitoring and alarm device based on a multimodal large model provided in this invention collects structured data (including multiple text rules) related to property security and collects and preprocesses property video surveillance image data to obtain video features. At the same time, it preprocesses the structured data to obtain text rule features. Then, an adapter is embedded in the main structure of the visual large model to form a multimodal model. Transfer learning is used to weightedly fuse the video features and text rule features to obtain fused features, which are used to determine abnormal events. If an abnormal event exists, a dynamic risk assessment module is used to classify the risk and trigger an alarm. This achieves deep fusion of text rules and video image data, effectively identifies abnormal events in property scenarios, classifies the risk of abnormal events, and provides timely alarms, thereby improving the intelligence level and risk response capability of property security monitoring.
[0094] This invention also provides a video monitoring and alarm device based on a multimodal large model, which can be implemented as a computer program. This computer program can be used in various ways, such as... Figure 6 It runs on the computer device shown.
[0095] Please see Figure 6 , Figure 6 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device 600 is a server, which can be a standalone server or a server cluster composed of multiple servers.
[0096] See Figure 6 The computer device 600 includes a processor 602, a memory, and a network interface 605 connected via a system bus 601. The memory may include a non-volatile storage medium 603 and internal memory 604.
[0097] The non-volatile storage medium 603 can store an operating system 6031 and a computer program 6032. When the computer program 6032 is executed, it enables the processor 602 to execute a video monitoring and alarm method based on a multimodal large model.
[0098] The processor 602 provides computing and control capabilities to support the operation of the entire computer device 600.
[0099] The internal memory 604 provides an environment for the operation of the computer program 6032 in the non-volatile storage medium 603. When the computer program 6032 is executed by the processor 602, the processor 602 can execute a video monitoring and alarm method based on a multimodal large model.
[0100] This network interface 605 is used for network communication, such as providing data transmission. Those skilled in the art will understand that... Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device 600 to which the present invention is applied. The specific computer device 600 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0101] The processor 602 is used to run a computer program 6032 stored in the memory to perform the following functions:
[0102] Collect structured data related to property security, including multiple text rules;
[0103] Image data from property video surveillance is collected, and the image data is preprocessed to obtain video features; the structured data is preprocessed to obtain text rule features.
[0104] An adapter is embedded in the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Anomaly events are determined based on the fused features to obtain the determination result.
[0105] If an abnormal event is found in the determination result, the abnormal event is classified into risk levels through the dynamic risk assessment module to obtain the classification result, and an alarm is triggered based on the classification result.
[0106] Those skilled in the art will understand that Figure 6 The embodiments of the computer device shown do not constitute a limitation on the specific configuration of the computer device. In other embodiments, the computer device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may include only memory and a processor. In such embodiments, the structure and function of the memory and processor are different from those shown. Figure 1 The embodiments shown are consistent and will not be described again here.
[0107] It should be understood that, in this embodiment of the invention, the processor 602 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0108] In another embodiment of the invention, a computer-readable storage medium is provided. This computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, performs the following steps:
[0109] Collect structured data related to property security, including multiple text rules;
[0110] Image data from property video surveillance is collected, and the image data is preprocessed to obtain video features; the structured data is preprocessed to obtain text rule features.
[0111] An adapter is embedded in the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Anomaly events are determined based on the fused features to obtain the determination result.
[0112] If an abnormal event is found in the determination result, the abnormal event is classified into risk levels through the dynamic risk assessment module to obtain the classification result, and an alarm is triggered based on the classification result.
[0113] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0114] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.
[0115] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.
[0116] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0117] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.
[0118] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A video monitoring and alarm method based on a multimodal large model, characterized in that, Includes the following steps: Collect structured data related to property security, including multiple text rules; Collect image data from property video surveillance and preprocess the image data to obtain video features; The structured data is preprocessed to obtain text rule features; An adapter is embedded in the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Anomaly events are determined based on the fused features to obtain the determination result. If an abnormal event is found in the determination result, the abnormal event is classified into risk levels through the dynamic risk assessment module to obtain the classification result, and an alarm is triggered based on the classification result.
2. The video monitoring and alarm method based on a multimodal large model according to claim 1, characterized in that, The process involves collecting image data from property video surveillance and preprocessing the image data to obtain video features. The structured data is preprocessed to obtain text rule features, including: The image data is preprocessed using a SlowFast network structure to obtain video features F. v ; The text rules are encoded using the BERT language encoding model, and each text rule is converted into a text rule vector according to the following formula: in, Indicates d t 3D real vector space; The text rule vectors are aggregated using average pooling to form the text rule features F. t .
3. The video monitoring and alarm method based on a multimodal large model according to claim 2, characterized in that, The step of inputting the video features and the text rule features into the multimodal large model for weighted fusion through transfer learning includes: The video feature F is calculated according to the following formula. v With the text rule feature F t The input is fed into the multimodal large model for weighted fusion to obtain fused features: F=α·F v +β·F t Where α and β are the fusion weights of visual features and text rule features, respectively, and α+β=1.
4. The video monitoring and alarm method based on a multimodal large model according to claim 1, characterized in that, The step of determining abnormal events based on the fusion features to obtain a determination result includes: The fused features are input into the BERT language encoding model and vectorized according to the following formula to obtain the observation semantic vector: f obs =F=Encoder(event\_text) Among them, f obs ∈R d R d Let represent a d-dimensional real vector space, Encoder represent the encoder, event represent video events, and text represent semantic text; The text rules are input into the language encoder and vectorized according to the following formula to obtain a set of rule vectors: Each rule vector The semantic representation of the i-th text rule, text (i) This represents the i-th semantic text.
5. The video monitoring and alarm method based on a multimodal large model according to claim 4, characterized in that, The step of determining abnormal events based on the fusion features and obtaining the determination result further includes: The observed semantic vector and the set of rule vectors are matched for similarity according to the following formula: Among them, S i Indicates similarity; The observation semantic vector f represents obs With regular vectors The dot product of vectors, ||·|| denotes the Euclidean norm of the vector; When the similarity S i When the threshold θ is exceeded, the abnormal event is determined to match the i-th text rule.
6. The video monitoring and alarm method based on a multimodal large model according to claim 5, characterized in that, The step of determining abnormal events based on the fusion features and obtaining the determination result further includes: The judgment result is obtained by using a logical rule matching algorithm according to the following formula: Among them, w i Assign importance weights to text rules; When the determination result R exceeds the set threshold τ, the abnormal event is determined to be a high-risk event.
7. The video monitoring and alarm method based on a multimodal large model according to claim 1, characterized in that, The step of classifying the abnormal events through a dynamic risk assessment module to obtain classification results includes: The grading results are calculated using the following formula: R′=γ·C+δ·T+∈·E Where C represents the behavioral credibility of the currently detected abnormal event, T is the time weight, E is the environmental risk coefficient, γ, δ and ∈ are pre-set adjustment weights, and γ+δ+∈=1.
8. A video monitoring and alarm method, apparatus, computer equipment, and storage medium, characterized in that, include: A collection unit is used to collect structured data related to property security, including business rules in text form; The preprocessing unit is used to acquire image data from property video surveillance and preprocess the image data to obtain video features; The structured data is preprocessed to obtain text rule features; The fusion unit is used to embed an adapter into the backbone structure of the visual large model to obtain a multimodal large model. The video features and the text rule features are input into the multimodal large model through transfer learning for weighted fusion to obtain fused features. Abnormal events are judged based on the fused features to obtain the judgment result. The assessment and alarm unit is used to classify the risk of the abnormal event through the dynamic risk assessment module if there is an abnormal event in the judgment result, obtain the classification result, and trigger an alarm according to the classification result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video monitoring and alarm method based on a multimodal large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the video monitoring and alarm method based on a multimodal large model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Sensitive information discovery and automatic classification and grading method based on multi-modal fusion
CN116049397A
Video anomaly detection system and method, computer equipment and storage medium
CN118485954A
Dangerous behavior identification and early warning method based on multi-modal analysis
CN119360278A
Video processing method, electronic device and storage medium
US20220027634A1