Visual language model system under VLM security and protection monitoring scene

By using a visual language model system, combined with multimodal information and model optimization technology, the system addresses the shortcomings of security monitoring systems in intelligent recognition of complex scenes and high-order semantic events, achieving efficient and accurate security monitoring capabilities suitable for embedded devices.

CN121963034APending Publication Date: 2026-05-01HANGZHOU HUMPBACK WHALE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HUMPBACK WHALE TECHNOLOGY CO LTD
Filing Date
2026-01-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing security monitoring systems have limited intelligence in understanding complex scenarios and recognizing high-order semantic events, relying on manual operation and judgment, and are difficult to automatically retrieve target events or perform semantic-level analysis.

Method used

A visual language model system is adopted, which receives image frames and depth maps through a multimodal fusion module, identifies high-order semantic events through an abnormal event recognition module, and generates standardized event data through a structured event output module. The model is optimized by using chain-thinking CoT training data and teacher-student distillation technology to adapt to the security monitoring needs of fixed scenarios.

Benefits of technology

It achieves high-accuracy and low-latency intelligent understanding and alarm capabilities for security scenarios, improves the model's accuracy in recognizing key semantic events, enhances spatial reasoning capabilities, supports structured and traceable event output, and is compatible with embedded device deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963034A_ABST
    Figure CN121963034A_ABST
Patent Text Reader

Abstract

The invention discloses a visual language model system under a VLM security and protection monitoring scene, and relates to the technical field of visual language models, and the system comprises a multi-modal fusion module which is used for receiving and processing an image frame, a depth map and scene information, an abnormal event recognition module which is used for analyzing a monitoring image according to a visual language model and recognizing a high-order semantic event, and a VLM security and protection monitoring module. The structured event output module is used for generating standardized event data including event types, timestamps, position coordinates and confidence fields, the visual language model optimization method oriented to the fixed scene is provided for the application bottleneck of a current visual language model in a security and protection monitoring scene, multi-source perceptual information is fused, and the security and protection monitoring efficiency is improved. According to the method and the system, the high-accuracy and low-delay security scene intelligent understanding and warning capability is realized by means of targeted training means such as scene graph description, depth map and time sequence features and the like, such as chained thinking CoT data design, model distillation and cutting and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Visual Language Model System for VLM Security Monitoring Scenarios Technical Field

[0001] This invention relates to the field of visual language model technology, specifically a visual language model system for VLM security monitoring scenarios. Background Technology

[0002] In current security monitoring systems, users generally face the problems of information overload and high costs of manual monitoring. With the development of urbanization and digitalization, a large number of video surveillance devices have been deployed in various scenarios such as commerce, industry, and residence. These devices often continuously collect high-definition video data and generate massive amounts of video content, which brings great challenges to the subsequent manual retrieval, screening, and backtracking work.

[0003] Traditional security monitoring systems mainly rely on manual inspections or rule-based visual detection algorithms (such as target detection, motion detection, and area intrusion detection), but their level of intelligence is limited, especially in understanding complex scenarios and recognizing high-order semantic events (such as "customers and employees arguing" or "abnormal items being discarded").

[0004] To improve efficiency, users typically use multi-screen playback and playback speed adjustment functions:

[0005] Multi-screen playback: Simultaneously displays the images from multiple monitoring positions, allowing users to observe the situation on site from different perspectives;

[0006] Speed ​​up playback: Increase video playback speed and quickly locate target events in long-term monitoring records.

[0007] Nevertheless, these methods still rely on manual operation and judgment, making it difficult to automate the retrieval of target events or perform semantic-level analysis.

[0008] In recent years, multimodal large models, especially visual language models (VLMs), have demonstrated strong generalization and understanding capabilities in tasks such as image-text alignment, image question answering, and video understanding. These models are trained on massive image-text pairs and have strong image semantic understanding and text generation capabilities, enabling them to answer high-order semantic questions such as "What happened in the picture?", "Are people fighting?", and "Are there any suspicious items?".

[0009] Introducing VLM technology into the field of security monitoring can achieve the following capabilities:

[0010] Natural language search for surveillance footage: such as quickly filtering videos by "finding footage of strangers breaking in";

[0011] Video Q&A and summary: such as "What happened in this video?", "Did anyone fall?", etc.;

[0012] Behavior and scene recognition: Identifying complex events or multi-agent interactions from images;

[0013] Generate structured alert information: Output and push abnormal events found in the video as structured text. Summary of the Invention

[0014] This invention provides a visual language model system for VLM security monitoring scenarios, which can effectively solve the problems mentioned in the background art.

[0015] To achieve the above objectives, the present invention provides the following technical solution: a visual language model system for VLM security monitoring scenarios, comprising:

[0016] The multimodal fusion module is used to receive and process image frames, depth maps, and scene information;

[0017] The abnormal event recognition module is used to analyze the monitoring screen and identify high-order semantic events based on the visual language model;

[0018] The structured event output module is used to generate standardized event data, including event type, timestamp, location coordinates, and confidence field.

[0019] According to the above technical solution, the structured event output module converts natural language responses into structured data in JSON format, supporting integration with alarm systems, video indexing platforms, and quality inspection platforms.

[0020] According to the above technical solution, the abnormal event recognition module includes VLM-based functions for personnel behavior recognition, abnormal item detection, and intrusion behavior judgment.

[0021] According to the above technical solution, the system method includes the following steps:

[0022] Acquire environmental information for a fixed scene, including camera layout, area division, and depth map information;

[0023] For each monitoring scene, generate multimodal input features and perform inference using a visual language model;

[0024] The reasoning results are converted into structured event data, which facilitates alarms and system integration.

[0025] According to the above technical solution, the structured event data includes event type, timestamp, spatial location and confidence value, supporting real-time alarm and event tracing.

[0026] According to the above technical solution, the multimodal input features include image frames, depth maps, and scene annotations.

[0027] According to the above technical solution, a visual language model system is applied to a security monitoring event recognition method in a fixed scenario. The method optimizes the visual language model (VLM) in the security monitoring scenario by fusing multi-source perception information and targeted training methods. The method includes the following steps:

[0028] Collect and model fixed environmental information in security monitoring scenarios, including camera location information, depth map, scene area division and camera view;

[0029] Multimodal input features are generated based on this fixed scene information;

[0030] We designed chain-based CoT training data and conducted multiple rounds of reasoning training to improve the semantic understanding and reasoning ability of VLM.

[0031] Teacher and student distillation techniques are used to transfer the reasoning capabilities of large general-purpose models to lightweight models, and model distillation optimizes computational resource consumption.

[0032] It outputs structured event data and supports integration with real-time alarm and post-processing systems.

[0033] According to the above technical solution, the chain-thinking CoT training data includes image and natural language question-and-answer pairs, and each question-and-answer pair contains multiple rounds of reasoning steps to enhance the model's logical reasoning ability.

[0034] According to the above technical solution, the teacher-student distillation technology uses a large VLM model (such as BLIP-2, LLaVA, etc.) as the teacher model to transfer the knowledge of the domain data to a lightweight student model, thereby realizing real-time deployment of security equipment.

[0035] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention addresses the current bottlenecks in the application of visual language models in security monitoring scenarios by proposing a visual language model optimization method for fixed scenarios. By integrating multi-source perception information, such as scene graph description, depth map, and temporal features, with targeted training methods, such as chain-thinking CoT data design, model distillation, and pruning, it achieves high-accuracy and low-latency intelligent understanding and alarm capabilities for security scenarios.

[0036] Specific objectives include:

[0037] Improve the model's accuracy in identifying key semantic events in security (such as fighting, falling, theft, unauthorized intrusion, etc.);

[0038] Enhance the model's understanding of scene structure and spatial reasoning ability;

[0039] Reduce model size and inference latency to adapt to embedded or edge devices;

[0040] It supports structured, traceable event output, improving system interpretability and integration capabilities.

[0041] The core innovation lies in:

[0042] Fixed scene adaptation modeling: By integrating scene map, depth map information, camera view and other information, input features with spatial structure perception capabilities are constructed to improve the model's spatial reasoning ability under actual security monitoring layout;

[0043] Chain-of-Thought (CoT) fine-tuning: For high-order semantic events in security (such as people falling, items being abandoned, and people occluding their actions), we design multi-turn inference training samples to improve the model’s ability to answer logic and semantic understanding.

[0044] Teacher-student model distillation technology: Using a large-scale multimodal model as the teacher network, knowledge distillation is performed by combining domain data to compress reasoning capabilities into a lightweight model and enable edge deployment;

[0045] Structured output interface design: Convert natural language output into structured event data, including event type, time, spatial location, confidence level, etc., to facilitate integration of the alarm system with the backend logic.

[0046] The key idea of ​​this invention is to improve the practicality and applicability of VLM in security scenarios by using the technical approach of "scene customization + spatial enhancement + mind chain fine-tuning + model distillation". Attached Figure Description

[0047] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0048] In the attached diagram:

[0049] Figure 1 is a schematic diagram of the method steps of the present invention. Detailed Implementation

[0050] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0051] Example: As shown in Figure 1, the present invention provides a technical solution, a visual language model system for VLM security monitoring scenarios, the system including the following operation steps:

[0052] Step 1: Fixed Scene Modeling and Multimodal Data Preparation

[0053] Acquire surveillance video frame images and establish a "fixed scene configuration" for each monitoring point, including:

[0054] Camera location and orientation information;

[0055] Scene area division (e.g., entrance, cashier area, exit);

[0056] Depth map / spatial layout map (generated by depth camera or monocular estimation);

[0057] Image generation includes multimodal inputs such as image frames, depth maps, and scene annotations.

[0058] Step Two: Constructing Chain Thinking (CoT) Training Data:

[0059] Design multi-round inference samples, example:

[0060] Q: Why did this person suddenly fall down?

[0061] A: There was water on his feet before he slipped, which may have caused the fall.

[0062] Q: Has anyone failed to wear their name tag as required?

[0063] A: The employee on the right side against the wall did not have a clearly visible employee badge on his chest.

[0064] Construct a triple of image + question + thought chain answer for fine-tuning the general VLM.

[0065] Step 3: Fine-tuning the visual language model and knowledge distillation:

[0066] Domain-specific fine-tuning is performed using basic models such as BLIP-2 and LLaVA;

[0067] Train lightweight models (such as Mini-VLM, Tiny-VLM) as student models;

[0068] By using distillation techniques to transfer domain knowledge to small models, inference latency and computational resource consumption are significantly reduced.

[0069] Step 4: Security Scenario Reasoning and Execution:

[0070] During the online operation phase, video frames and scene structure information are input into the model;

[0071] You can input natural language questions (such as: "Has anyone fallen in area X?") or directly trigger timed analysis;

[0072] The model outputs a natural language response plus optional structured event results.

[0073] Step 5: Structured Event Output and Integration

[0074] The model output is parsed to generate standard structured data:

[0075] To improve the system's adaptability and inference accuracy under different security operation scenarios, the following improvements were made to the original process:

[0076] Video frames are processed in time windows, and incremental inference is used to reduce the amount of computation each time.

[0077] Asynchronous processing of multiple camera streams enables parallel computing.

[0078] High-precision analysis is performed on key frames first, while lightweight models are used for fast scanning of non-key frames.

[0079] Spatial enhancement and fixed scene modeling optimization;

[0080] Introduce depth map / spatial layout auxiliary information to map 3D spatial coordinates to the model input.

[0081] Multi-camera fusion formula:

[0082]

[0083] in:

[0084] This represents the fused multi-view feature vector;

[0085] For the first Feature vectors extracted from each camera (viewpoint);

[0086] The fusion weights for corresponding cameras are usually adaptively determined based on dynamic evaluation indicators such as viewpoint quality, occlusion degree, confidence or semantic consistency.

[0087] The weights satisfy the normalization constraint:

[0088]

[0089] Multi-turn inference (CoT) optimization:

[0090] Problem: Single-round reasoning is not accurate enough for identifying complex events (such as falls or disputes).

[0091] improve:

[0092] Constructing multi-round question-answer chains enhances logical reasoning ability.

[0093] Multi-round confidence accumulation formula:

[0094]

[0095] in:

[0096] Indicates the first Round of reasoning (or the first round of reasoning) The event confidence level output by each independent detection module;

[0097] This represents the cumulative confidence level after fusing the results of N rounds of inference;

[0098] This formula is based on a probabilistic model that assumes at least one positive detection, and it assumes that each round of inference is independent. It can be regarded as the first The probability that the wheel does not detect the target event, therefore This represents the probability that no event is detected in all rounds, and its complement is the cumulative confidence level of the event that is eventually detected.

[0099] Lightweight model distillation and edge deployment:

[0100] Problem: General-purpose large-scale models require a large amount of computation and are difficult to deploy in real time.

[0101] improve:

[0102] Teacher-student distillation:

[0103]

[0104] in:

[0105] This represents the total training loss of the student model;

[0106] The supervision loss (such as cross-entropy loss, mean squared error, etc.) for the student model on the target task;

[0107] This is the knowledge distillation loss, used to measure the consistency between the student model output and the teacher model output (such as KL divergence, mean squared error, etc.).

[0108] It is an adjustable hyperparameter used to balance the relative importance of task learning and knowledge transfer.

[0109] Structured event output optimization:

[0110] Problem: Natural language output is difficult to integrate into alarm systems.

[0111] improve:

[0112] Add coordinate, time, and confidence fields to the output events.

[0113] Event scoring formula:

[0114]

[0115] in:

[0116] This represents a comprehensive score indicating the overall reliability of the target event.

[0117] This represents the model's raw prediction confidence level for the event (such as classification or detection probability);

[0118] A score is given for the consistency between the event and its contextual information (such as the logic of events before and after, semantic coherence, etc.).

[0119] This refers to the degree of matching between the event and the prior knowledge of the current scenario (such as scenario category, spatiotemporal rationality, domain knowledge, etc.).

[0120] The weights are non-negative and satisfy the normalization constraints:

[0121]

[0122] Improvement strategies for different operating scenarios

[0123]

[0124] To address the issues of low accuracy, high latency, and unsuitability for real-world scenarios in existing technologies, this paper introduces fixed-scene modeling, multimodal input enhancement, chained inference sample training, and model distillation optimization, resulting in the following beneficial effects:

[0125] Significantly improved accuracy: Through customized CoT training and scenario modeling, the model is made closer to the actual business semantics, the recognition accuracy is improved by more than 10%, and the false alarm rate is significantly reduced;

[0126] Enhanced spatial understanding: With the help of auxiliary information such as depth maps and region maps, the model can better determine 3D semantic events such as occlusion, falling direction, and boundary crossing behavior;

[0127] Supports lightweight deployment: The distilled and compressed student model can run on edge boxes and embedded devices, with strong adaptability and a response latency of less than 1 second;

[0128] Structured and traceable output results: A unified structured event format is used to facilitate integration with alarm systems, BI analysis systems, and other systems;

[0129] Easy to expand and reuse: It can be quickly adapted through fixed scenario definition and automatic sample generation mechanism.

[0130] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A visual language model system for VLM security monitoring scenarios, characterized in that: include: The multimodal fusion module is used to receive and process image frames, depth maps, and scene information; The abnormal event recognition module is used to analyze the monitoring screen and identify high-order semantic events based on the visual language model; The structured event output module generates standardized event data, including event type, timestamp, location coordinates, and confidence level fields. This system method includes the following steps: acquiring environmental information of a fixed scene, including camera layout, area division, and depth map information; generating multimodal input features for each monitoring scene and performing inference using a visual language model; converting the inference results into structured event data for easy alarm and system integration. The visual language model system is applied to a fixed-scene security monitoring event recognition method. This method optimizes the visual language model (VLM) in security monitoring scenes by fusing multi-source perception information with targeted training methods. The method includes the following steps: collecting and modeling fixed environmental information in the security monitoring scene, including camera location information, depth map, scene area division, and camera viewpoint; generating multimodal input features based on this fixed scene information; designing chain-thinking CoT training data and performing multi-round inference training to improve the semantic understanding and inference capabilities of the VLM; using teacher and student distillation techniques to transfer the inference capabilities of a large-scale general model to a lightweight model and optimizing computational resource usage through model distillation; and outputting structured event data to support real-time alarm and post-processing system integration.

2. The visual language model system for VLM security monitoring scenarios according to claim 1, characterized in that, The structured event output module converts natural language responses into structured data in JSON format, supporting integration with alarm systems, video indexing platforms, and quality inspection platforms.

3. The visual language model system for VLM security monitoring scenarios according to claim 1, characterized in that, The abnormal event recognition module includes VLM-based functions for personnel behavior recognition, abnormal item detection, and intrusion behavior judgment.

4. The visual language model system for VLM security monitoring scenarios according to claim 1, characterized in that, The structured event data includes event type, timestamp, spatial location, and confidence value, supporting real-time alarms and event tracing.

5. The visual language model system for VLM security monitoring scenarios according to claim 1, characterized in that, The multimodal input features include image frames, depth maps, and scene annotations.

6. The visual language model system for VLM security monitoring scenarios according to claim 5, characterized in that, The Chain Thinking CoT training data includes image and natural language question-and-answer pairs, and each question-and-answer pair contains multiple rounds of reasoning steps.

7. The visual language model system for VLM security monitoring scenarios according to claim 5, characterized in that, The teacher-student distillation technique uses a large VLM model as the teacher model to transfer knowledge from domain data to a lightweight student model.