Large-scene monitoring video abnormal event early warning method based on multi-modal large model
By combining multimodal large models with deep fusion and analysis of video, audio, and sensor data, the problems of delay and low accuracy in abnormal event early warning in large-scale scene monitoring have been solved, achieving efficient and intelligent abnormal event detection and early warning.
Patent Information
- Application Number
- CN202511460997.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-09
AI Technical Summary
Existing technologies for large-scale scene monitoring suffer from problems such as single-modal data processing, insufficient real-time performance, limited types of abnormal event detection, and insufficient depth of multimodal data fusion, resulting in delayed abnormal event warnings, low accuracy, and limited coverage.
By employing a multimodal large model that combines video, audio, and sensor data, and using the Transformer architecture for deep fusion and analysis, we can achieve synchronous processing and feature extraction of multi-source data. We can also use deep learning models to detect various abnormal events in real time and trigger early warnings.
It improves the accuracy and robustness of abnormal event identification, enables real-time processing of high-resolution multi-channel video streams, enhances monitoring coverage and data quality, meets diverse security monitoring needs, and reduces false alarm and missed alarm rates.
Smart Images

Figure CN121305463A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of abnormal event early warning technology and multimodal understanding, and provides a method for early warning of abnormal events in large-scene surveillance videos based on a multimodal large model. Background Technology
[0002] With the development of the social economy and the acceleration of urbanization, the demand for security management in large public places such as airports and train stations is increasing. Traditional monitoring systems mainly rely on single video data for real-time monitoring, but when faced with complex and ever-changing scenarios and massive amounts of video data, they often suffer from problems such as low information processing efficiency and insufficient accuracy in event detection.
[0003] In recent years, the development of artificial intelligence technologies, especially deep learning and multimodal learning, has provided new solutions for video surveillance and analysis. Multimodal large models can integrate and process data from different sources, such as video, audio, and sensor information, achieving a comprehensive understanding and accurate analysis of complex scenes through cross-modal information fusion. These models have already achieved significant results in areas such as image recognition, speech processing, and natural language understanding, and are gradually being applied to security monitoring systems.
[0004] Several key advancements have been made in the field of multimodal understanding. First, the scale of models and the amount of training data have continued to expand. For example, the Flamingo model has reached 80 bytes of parameters, the VALOR model jointly models four modalities: image, video, text, and audio, while ImageBind connects six modalities—text, audio, vision, infrared, and inertial measurement unit (IMU) signals—through contrastive learning centered on vision. Furthermore, the types of tasks supported have become more diverse. The OFA model jointly models various understanding and generation tasks, including visual question answering, image generation, image captioning, text tasks, and object detection. Simultaneously, fine-grained perception capabilities for open worlds have been enhanced; for example, GLIPv2 can achieve object detection with an open vocabulary. Finally, the integration of large language models and the introduction of instruction learning have become development trends. For instance, LLavA trains large models to better follow human instructions by constructing instruction fine-tuning datasets.
[0005] Existing event warning methods largely rely on single-modal data analysis, such as detecting abnormal behavior solely based on video images or combining simple sensor data for auxiliary judgment. However, these methods often face challenges such as data processing latency and low warning accuracy when dealing with large-scale, complex real-time surveillance videos. Furthermore, video summarization technology is not yet fully utilized in existing systems, causing monitoring personnel to spend a significant amount of time on video playback and event retrieval, reducing work efficiency.
[0006] Therefore, there is an urgent need for a method that can comprehensively utilize the advantages of multimodal large models to efficiently generate summaries and accurately warn of events from surveillance videos in large scenarios, so as to improve the intelligence level and response capability of security monitoring systems.
[0007] Prior art, publication number: CN119027882A
[0008] Invention Title: A Fusion Method for Dynamic Target Tracking in Large-Scene Monitoring
[0009] A method for tracking and fusing dynamic targets in large-scale scene monitoring is disclosed. This invention relates to the field of image processing technology and includes a panoramic camera for acquiring global monitoring images and multiple detail cameras for capturing detailed monitoring images. The fusing method includes identifying and tracking a dynamic target in the global monitoring image; determining the location of the dynamic target in the detail monitoring image and the global target image; tracking the dynamic target in the detail monitoring image to obtain a detail target image; calculating the mapping matrix between the global target image and the detail target image, and transforming the detail target image to obtain a transformed detail target image; and fusing the transformed detail target image and the global monitoring image in real time to obtain a fused global monitoring image. This invention achieves accurate tracking and high-definition display of dynamic targets in large-scale monitoring scenes through the collaborative work of a panoramic camera and multiple detail cameras.
[0010] The technical problem this proposal aims to solve
[0011] Existing technical solutions mainly involve the collaborative operation of panoramic cameras and multiple detail cameras to achieve accurate tracking and high-definition display of dynamic targets in large-scale monitoring scenes. Its core steps include:
[0012] (1) Identify and track dynamic targets in the global monitoring image.
[0013] (2) Determine the location of the dynamic target in the detailed monitoring image and continue to track the target in the detailed image.
[0014] (3) Calculate the mapping matrix between the global target image and the detail target image, and transform the detail target image.
[0015] (4) The transformed detailed target image is fused with the global monitoring image in real time to generate the fused global monitoring image.
[0016] While existing methods have certain advantages in tracking and displaying dynamic targets in high definition, they have significant shortcomings in the following aspects, which are also the technical problems that this application aims to solve:
[0017] (1) Single-modal data processing: Existing methods mainly rely on video image data for tracking and fusion of dynamic targets, lacking comprehensive utilization of other modal data (such as audio and sensor data). This limits the system's ability to identify abnormal events in complex scenarios.
[0018] (2) Insufficient real-time performance and processing efficiency: Existing methods have performance bottlenecks in real-time decoding, stitching, and synchronous processing of multi-camera video streams, making it difficult to achieve low-latency real-time processing under high-resolution multi-channel video streams. In large-scale scenarios, the amount of video data is enormous, and existing methods cannot meet the needs of real-time early warning, resulting in delays in the detection and early warning of abnormal events, affecting the system's immediate response capability.
[0019] (3) The abnormal event detection function is limited. Existing methods mainly focus on tracking and displaying dynamic targets, and lack specialized detection and early warning mechanisms for various abnormal events such as falls, crowding, and fights. It is impossible to fully cover and identify various types of abnormal events, which limits the application scope and practicality of the system and makes it difficult to meet the diverse needs of actual security monitoring.
[0020] (4) Insufficient depth of data fusion and analysis: Existing methods perform simple image fusion using mapping matrices, lacking large-scale data fusion and intelligent analysis based on deep learning, making it difficult to achieve high-precision event recognition. The insufficient depth and intelligence of data fusion result in low accuracy and reliability of abnormal event recognition in complex scenarios, leading to high false alarm and false negative rates. Summary of the Invention
[0021] This invention solves the problems of delayed, low-accuracy, and limited coverage of abnormal event warnings in large-scene surveillance videos caused by single-modal data processing, insufficient real-time performance, limited types of abnormal event detection, and insufficient depth of multimodal data fusion.
[0022] To achieve the above objectives, the present invention employs the following technical solution:
[0023] A method for early warning of abnormal events in large-scene surveillance videos based on a multimodal large model includes the following steps:
[0024] Step 1: Perform preprocessing such as stitching on multiple high-resolution surveillance videos from the same large scene to obtain unified panoramic video data;
[0025] Step 2: Acquire and preprocess audio data and sensor data to obtain synchronized multimodal data;
[0026] Step 3: Use a traditional small model to detect the surveillance video stream and obtain the task keyframes. Input the keyframes and preprocessed multimodal data into a large multimodal model for data fusion and feature extraction.
[0027] Step 4: Utilize the fused features to detect abnormal events, such as falls, crowding, and fights;
[0028] Step 5: When an abnormal event is detected, trigger the early warning mechanism and send early warning information to relevant management personnel;
[0029] Step 6: Record and store detailed information about the abnormal event for subsequent analysis and tracking.
[0030] Step 1 above specifically includes the following steps:
[0031] Step 1.1: Multi-channel video stream acquisition. Through network or wired connection, the video streams acquired by multiple high-resolution cameras deployed in the same large scene are transmitted to the central processing unit to obtain multiple high-resolution surveillance video streams from different angles and positions.
[0032] Step 1.2: Video stitching and synchronization. High-precision timestamps are added to each video stream to ensure temporal synchronization. Video stitching algorithms (such as feature matching and image fusion) are used to seamlessly stitch multiple video streams into a unified panoramic video stream, resulting in temporally and spatially synchronized high-resolution panoramic video data.
[0033] Step 1.2.1. Perform time synchronization processing on N video streams under the same monitoring scenario, for the first... Add high-precision timestamps to video streams ,in As the base timestamp, For the first Video stream transmission delay compensation amount, For the first Video frame rate;
[0034] Step 1.2.2. Establish spatial mapping relationships between adjacent video streams using a feature matching algorithm, through feature point pairs. Calculate the homography matrix ,in The coordinates of feature points in the source image. The coordinates of the feature points in the target image;
[0035] Step 1.2.3. Generate panoramic video frames based on dynamic weight fusion algorithm ,in The coordinate transformation function obtained in step (2) is... Based on the distance from the pixel to the stitching boundary Calculated fusion weights;
[0036] Step 1.2.4. Output a spatiotemporally synchronized panoramic video stream through the timestamp remapping module. ,in This represents the maximum alignment value of the timestamps from multiple video streams.
[0037] Step 1.3: Video preprocessing. Efficient video decoding and compression algorithms (H.265) are used to reduce data volume and improve transmission efficiency. Image processing techniques such as noise reduction, contrast adjustment, and color correction are applied to improve video quality. The video frame rate is adjusted according to actual needs to balance processing efficiency and video smoothness.
[0038] Step 1.3.1. Decode and compress the panoramic video stream using the H.265 / HEVC standard. Dynamically adjust the quantization parameter (QP) using the rate-distortion optimization (RDO) algorithm to balance compression ratio and video quality. Apply entropy coding (CABAC) to generate a compressed video stream, significantly reducing data transmission bandwidth.
[0039] Step 1.3.2. Combine the nonlocal mean (NLM) algorithm with wavelet thresholding denoising technology to process low-frequency noise and high-frequency noise respectively. Improve the clarity and detail retention of video frames by calculating pixel block similarity weights and hard thresholding of high-frequency components.
[0040] Step 1.3.3. Based on the Limiting Contrast Adaptive Histogram Equalization (CLAHE) algorithm, the video frame is divided into local sub-regions and the histogram distribution is limited. Global contrast equalization is achieved through bilinear interpolation to avoid excessive noise enhancement.
[0041] Step 1.3.4. White balance correction is performed using the Gray World Assumption, adjusting the RGB channel gain to equalize the global color mean, and optimizing the color difference distribution in conjunction with the CIE Lab color space to ensure the authenticity and consistency of color reproduction.
[0042] Step 1.3.5. Based on the actual scenario requirements and hardware processing capabilities, dynamically adjust the video frame rate, and balance processing efficiency and video smoothness through frame sampling or interpolation techniques to ensure optimal real-time performance and resource utilization.
[0043] Step 2 above specifically includes the following steps:
[0044] Step 2.1: Audio data acquisition. Real-time acquisition of environmental audio data to obtain a high-quality real-time environmental audio stream.
[0045] Step 2.2: Collect various environmental parameters in real time to obtain real-time environmental sensor data.
[0046] Step 2.3: Perform audio signal enhancement and noise reduction processing, extract key audio features (spectrum, pitch, etc.), clean and normalize the sensor data, remove noise and outliers, and finally obtain optimized audio data and sensor data to ensure data quality and consistency.
[0047] Step 2.4: Multimodal data synchronization. Through high-precision timestamps and data buffering mechanisms, the synchronization of different modal data in time is ensured, resulting in a multimodal dataset that is fully synchronized in time and space, suitable for subsequent fusion and analysis.
[0048] Step 2.4.1. Generate a unified reference timestamp. Configure Network Time Protocol (NTP) for all data sources (video, audio, sensors) to synchronize the system clocks of each device to microsecond-level accuracy. Using the central processing unit clock as the reference, append a globally unified high-precision timestamp to each data stream. The timestamp format is as follows: ,in As the base time, This is the clock offset compensation amount. This is a clock drift state correction value.
[0049] Step 2.4.2. Dynamic compensation for transmission delay: Real-time monitoring of the transmission delay of each data stream (including network delay and device processing delay), and calculation of the average delay using a sliding window algorithm: Where W and W' are the window sizes, This represents the k-th delayed sample value of the i-th data stream. The timestamp is dynamically corrected. .
[0050] Step 2.4.3. Establish a circular buffer with timestamp indexes for each type of data (video frames, audio clips, sensor readings), and store data packets in ascending time order. Use a greedy matching algorithm to extract data groups from the buffer that meet the following conditions: ,in and For the preset synchronization tolerance threshold (e.g.) ≤10 milliseconds (≤10 milliseconds). This is a synchronization tolerance threshold where the difference between the timestamps of a video frame and the corresponding audio data packet must not exceed 10 milliseconds. In other words, when the greedy matching algorithm extracts "data packets" from the circular buffer of each modality, it will only proceed if the difference between the timestamp of a video frame and the timestamp of the corresponding audio data packet is less than or equal to 10 milliseconds. Take here Only after 10ms will a record be considered "synchronous" and grouped into the same group for further processing.
[0051] Step 2.4.4. Missing Data Interpolation and Error Tolerance: If data for a certain modality is missing within the time window, temporary data is generated using a Kalman filter or linear interpolation.
[0052] Video frame missing: Based on optical flow method, predict the motion trajectory between adjacent frames to generate virtual frames.
[0053] Missing audio / sensor data: Use the mean or trend of the preceding and following data points to fit and complete the data.
[0054] An alarm is triggered for consecutive missing events exceeding the threshold (e.g., more than 3 frames), and an anomaly log is recorded.
[0055] Step 2.4.5. Synchronization Verification and Dynamic Adjustment: Verify synchronization accuracy through cross-modal correlation analysis (such as the temporal correlation between video actions and audio voiceprints). If a synchronization deviation is detected (e.g., |t actual alignment − t theoretical alignment| > ϵ), dynamically adjust the buffer window size or correct the timestamp compensation parameters, iteratively optimizing until the synchronization requirements are met.
[0056] Step 3 above specifically includes the following steps:
[0057] Step 3.1: Use traditional small models (such as YOLOv8) based on deep learning for object detection and multi-object tracking to process the spliced surveillance video stream frame by frame, detect important features in each frame of video, and identify key frames.
[0058] Step 3.1.1. Video Stream Input and Frame Normalization: The stitched panoramic video stream is divided into frames, uniformly adjusted to a fixed resolution (e.g., 640×640), and normalized. , where μ and σ are the mean and standard deviation of the training dataset.
[0059] Step 3.1.2. Object detection model inference: Load the pre-trained YOLOv8 model, perform real-time object detection for each frame, and output detection boxes. Category tags and confidence level The detection results were filtered using non-maximum suppression (NMS) to remove overlapping boxes (IoU threshold set to 0.5), retaining high-confidence targets. ≥0.6).
[0060] Step 3.1.3. Multi-target tracking initialization and data association: Assign a unique ID to each detected target and perform tracking based on the DeepSORT algorithm.
[0061] Motion prediction: Predict the target's state vector in the next frame using a Kalman filter.
[0062]
[0063] Appearance feature extraction: Extracting the appearance descriptor of the target using a ReID network.
[0064] Data association: Combining Mahalanobis distance (motion similarity) and cosine similarity (appearance similarity), the cost matrix is calculated:
[0065]
[0066] Where λ=0.7, the Hungarian algorithm is used to achieve optimal matching, and the cost matrix is... It is only used when the Hungarian algorithm completes the one-to-one matching of the detection box and the predicted trajectory; in subsequent steps such as keyframe determination, metadata storage and multimodal fusion, the system only relies on the matching result (target ID and trajectory information) and no longer directly calls the matrix.
[0067] Step 3.1.4: Keyframe Determination and Dynamic Threshold Adjustment, defining keyframe determination rules:
[0068] Event-triggered: Detects a preset anomaly category (such as a fall, weapon possession) or a certain confidence level. .
[0069] Density-triggered type: Sudden increase in the number of targets in a single frame (e.g.) Or the area density exceeds a threshold (e.g., >5 people / m²). Trajectory anomaly: Sudden change in target movement speed (e.g., ... (or the trajectory deviates from the normal path (Hausdorff distance > 1.5 m).
[0070] Dynamically adjust thresholds: Update judgment parameters in real time based on scene complexity (such as changes in lighting and occlusion ratio) to avoid misjudgments.
[0071] Step 3.1.5: Keyframe Tagging and Metadata Storage: Attach timestamps (tktk), target detection results, and tracking IDs to keyframes to generate structured data. The keyframe data is compressed into JSON format and stored in a cache queue for real-time access by the multimodal large model.
[0072] Step 3.2: Simultaneously input video keyframes, audio data, and sensor data into the multimodal large model. The multimodal large model fuses the multimodal data and extracts more comprehensive and richer scene features by utilizing the complementarity of image features, audio features, and sensor data.
[0073] Step 3.2.1: Multimodal data time alignment, based on the synchronized dataset generated in Step 2.4 Extract the video keyframe Vt, audio clip At, and sensor data St, all with strictly aligned timestamps t. Perform a final check on timestamp discrepancies to ensure they meet the synchronization tolerance threshold.
[0074]
[0075] Step 3.2.2: Cross-modal attention fusion. Features from each modality are input into the multimodal Transformer layer, and information interaction is achieved through a cross-modal attention mechanism: Query-key-value projection: Q, K, and V matrices are generated for fv, fa, and fs respectively. Attention weight calculation:
[0076] ,
[0077] Intermodal feature enhancement: Captures cross-modal associations such as vision-audio and vision-sensor through multi-head attention, and outputs fused features.
[0078] Step 3.2.3. Unified feature vector generation, which integrates the features. Dimensionality reduction and standardization are performed on the input fully connected layer:
[0079] ,
[0080] in Output a unified multimodal feature vector.
[0081] Step 3.3: Integrate the features extracted from visual, audio, and sensor data to form a unified multimodal feature vector, providing comprehensive input for subsequent anomaly detection tasks.
[0082] Step 3.3.1. Feature Dimension Alignment: Align the visual, audio, and sensor fusion features output from Step 3.2. Dimension matching is performed by mapping each modality feature to the same dimensional space (e.g., 512 dimensions) through a fully connected layer to ensure compatibility for subsequent integration.
[0083] Step 3.3.2. Weighted Feature Fusion: Dynamically assign weights based on modal importance. Weight Calculation: Based on the current scene context (e.g., lighting conditions, noise level), generate visual, audio, and sensor weight coefficients using a lightweight network. ,satisfy .
[0084] Feature weighting: Weighted summation of multimodal features according to their weights.
[0085]
[0086] Step 3.3.3. Feature Dimensionality Reduction and Standardization: Use fully connected layers to reduce the dimensionality of the weighted features to the target dimension (e.g., 256 dimensions) to reduce redundant information. Perform Layer Normalization on the dimensionality-reduced features to eliminate dimensional differences between modalities and improve the stability of model training.
[0087] Step 3.3.4. Temporal Context Embedding: Concatenate the current frame features with the historical frame features (through a sliding window buffer) and input them into a Temporal Convolutional Network (TCN) to capture short-term dependencies and enhance the temporal coherence of the features.
[0088] Step 3.3.5: Unified Feature Vector Generation. The temporally enhanced features are input into the residual connection module to further optimize the feature representation capability and output the final unified multimodal feature vector. Add metadata such as timestamps and modality source identifiers to the feature vectors to ensure traceability to the original data.
[0089] Step 3.3.6: Feature Storage and Interface Calls. The unified feature vector is stored in a high-speed cache queue, and a lightweight compression algorithm (such as Zstandard) is used to reduce memory usage. A standardized API interface is provided to support the anomaly detection model in Step 4 in real-time reading of feature data, ensuring low latency in end-to-end processing.
[0090] Step 4 above specifically includes the following steps:
[0091] Step 4.1: Input the fused features into the abnormal event detection model, perform inference calculations, and obtain preliminary detection results for various abnormal events, including event type and location.
[0092] Step 4.1.1: Model Loading and Initialization. Load the pre-trained multimodal anomaly detection model (such as a Transformer-based spatiotemporal network), and load the model weights and configuration file. Initialize the input interface to receive the unified multimodal feature vector from Step 3.3. And allocate computing resources (such as GPU memory).
[0093] Step 4.1.2: Feature input and preprocessing. Dynamically standardize the input features to eliminate the bias in feature distribution. Divide the feature sequence into time windows (e.g., 5-second windows) and fill or truncate it to a fixed length to adapt to the model input dimension.
[0094] Step 4.1.3: Spatiotemporal correlation reasoning. A Temporal Convolutional Network (TCN) is used to extract short-term dependencies in feature sequences, capturing dynamic patterns in event evolution. An attention mechanism is combined to enhance key regions (such as densely populated areas), generating spatially sensitive feature representations. Fully connected layers and a Softmax function are used to output the probability distribution for each type of abnormal event (falls, crowding, fights, etc.).
[0095]
[0096] Step 4.1.4: Preliminary result generation and filtering, retaining probability For events with a confidence level ≥ 0.7, low-confidence detection results are filtered out. Non-maximum suppression (NMS) is applied to multiple bounding boxes for the same event (e.g., IoU > 0.4), retaining the most significant result.
[0097] Step 4.2: Classify and confirm the inference results according to the predefined abnormal event categories (such as falls, crowding, fighting, etc.), and clarify the detected abnormal event categories and their specific locations.
[0098] Step 4.2.1: Load predefined event categories. Load the predefined abnormal event category configuration file (such as fall, crowding, fighting, etc.), which contains the semantic description, feature template and confidence threshold of each event category.
[0099] Step 4.2.2: Joint verification of multimodal evidence. Combining the multimodal features generated in Step 3.3, cross-modal verification is performed on the preliminary detection results of Step 4.1:
[0100] Visual-audio association: For example, when a "fall" event is detected, verify whether there is a collision sound or a cry for help in the audio data.
[0101] Vision-sensor correlation: For example, when a "crowding" event is detected, verify whether the area temperature or people flow sensor data is abnormally high.
[0102] Step 4.2.3: Classification model inference. Input the unified feature vector into the pre-trained multi-classification model (such as a hierarchical classification network based on attention mechanism) and output the event category probability distribution.
[0103] Step 4.3: Filter the detection results through a rule engine or secondary verification mechanism to eliminate false alarms and irrelevant events, thereby improving the accuracy of the detection.
[0104] Step 4.3.1: Rule engine condition filtering. Read scene-related filtering conditions (such as time range, regional restrictions, and event duration thresholds) from the predefined rule library. Match the classification results from Step 4.2 against each rule in the rule library and remove events that violate logical conditions (such as high-density alarms during non-working hours).
[0105] Step 4.3.2: Lightweight secondary validation model inference. Load the pre-trained lightweight validation model (such as random forest or small neural network), and input the multimodal feature vector and preliminary detection results. Output the corrected confidence score. If the corrected score is lower than the threshold... It will then be marked as a false alarm.
[0106] Step 4.3.3: Multimodal evidence consistency check. If the event relies on visual detection (e.g., "fighting"), verify whether there are fighting sounds or abnormal voiceprints in the audio data. Events with conflicting modal evidence (e.g., only visual detection but no audio support) are downweighted or filtered.
[0107] Step 5 above specifically includes the following steps:
[0108] Step 5.1: Based on the detected abnormal event category and location, generate corresponding early warning information, including detailed information such as event type, occurrence time, and occurrence location, forming structured early warning information.
[0109] Step 5.2: Send the early warning information to relevant management personnel in real time so that they can take prompt countermeasures.
[0110] Step 5.3: Record and store the detailed data of the early warning information and related abnormal events in the database to facilitate subsequent analysis and tracking, establish a complete abnormal event file, and support post-event review and security management.
[0111] Step 6 above specifically includes the following steps:
[0112] Step 6.1: Organize relevant data of abnormal events, including video clips, audio clips, sensor data, early warning information, etc., to form a complete event dataset.
[0113] Step 6.2: Store the event data in the local storage system, and use efficient storage technology to ensure fast data retrieval and access.
[0114] Step 6.3: Employ data backup and encryption technologies to ensure data security and reliability, prevent data loss and leakage, and guarantee the long-term preservation and secure use of event data.
[0115] Because the present invention employs the above-mentioned technical means, it has the following beneficial effects:
[0116] 1. Integrating multimodal data processing to improve the accuracy and robustness of anomaly event recognition: This method not only relies on video data but also integrates audio data and data from various environmental sensors, performing deep fusion and analysis through a multimodal large model based on the Transformer architecture. This comprehensive utilization of multi-source information significantly improves the system's ability to recognize complex anomaly events (such as fights, falls, etc.). Existing technical solutions mainly rely on single-modal video data for dynamic target tracking and fusion, lacking comprehensive utilization of other modal data, resulting in low accuracy in anomaly event recognition in complex scenes and susceptibility to environmental interference.
[0117] 2. Real-time processing of high-resolution multi-channel video streams to ensure monitoring coverage and data quality in large-scale scenarios: This method achieves real-time stitching, compression, and preprocessing of multiple 4K high-resolution monitoring video streams through edge computing devices, specifically corresponding to the core process in step 1 (video preprocessing): First, high-performance hardware (such as NVIDIA Jetson AGX Orin) is deployed at edge nodes to add nanosecond-level timestamps to multiple video streams using the PTP protocol and dynamically compensate for transmission delays (step 1.2.1), ensuring that the time synchronization error is ≤10ms; Second, a seamless 4K panoramic video stream (resolution ≥7680×4320) is generated based on deep learning feature matching (SuperPoint+SuperGlue) and dynamic weight fusion algorithm (steps 1.2.2-1.2.3); Subsequently, hardware-accelerated H.265 encoding (step 1.3.1) is used in conjunction with rate-distortion optimization (RDO) and CABAC entropy. Encoding compresses the bitrate to ≤20Mbps@4K, while improving video quality through non-local mean denoising (step 1.3.2) and CLAHE contrast enhancement (step 1.3.3). Finally, based on an edge-cloud collaborative architecture, lightweight object detection (YOLOv8-Tiny) and adaptive frame rate adjustment are performed at the edge (step 1.3.5), while the cloud focuses on multimodal depth analysis. End-to-end latency is ≤200ms, significantly solving the performance bottleneck of existing technologies in real-time processing of high-resolution multi-channel video, achieving high-efficiency, low-latency large-scene monitoring. This not only ensures broad monitoring coverage and image quality but also reduces data transmission volume through edge processing, improving overall processing efficiency. Existing technologies suffer from performance bottlenecks in real-time processing of high-resolution and multi-channel video streams, making it difficult to process multiple high-resolution videos simultaneously. This results in incomplete monitoring coverage or high data processing latency in large scenes, affecting the effectiveness of real-time early warning. A breakthrough was achieved through the key technical solutions in Step 1 (video preprocessing): On edge computing devices (such as NVIDIA Jetson AGX Orin), nanosecond-level timestamps were added to multiple 4K video streams based on the PTP protocol (Step 1.2.1), combined with the dynamic latency compensation formula. To correct network jitter and control time synchronization error within 10ms, eliminating multipath coverage blind spots; and to quickly calculate the homography matrix using SuperPoint feature extraction and SuperGlue matching algorithm (step 1.2.2). And based on dynamic weight fusion The system generates seamless 4K panoramic video (resolution ≥7680×4320) to ensure full coverage of large scenes. At the same time, it adopts hardware-accelerated H.265 encoding (step 1.3.1), and compresses the bit rate to ≤20Mbps@4K through rate-distortion optimization (RDO) and CABAC entropy encoding. Combined with edge-side lightweight object detection (YOLOv8-Tiny) and dynamic frame rate adjustment (step 1.3.5), the system achieves a total latency of ≤200ms, which significantly reduces data transmission and computing load. Finally, the system solves the problem of real-time processing of multiple high-resolution videos through an edge-cloud collaborative architecture (edge processing of raw streams and cloud-based deep analysis), thereby improving monitoring coverage and early warning timeliness.
[0118] 3. Enhancing Data Fusion Depth and Intelligent Analysis Capabilities through Transformer-Based Multimodal Large Models: Employing a Transformer-based multimodal large model leverages its self-attention mechanism for deep fusion and feature extraction of multi-source data, achieving more efficient and intelligent anomaly event identification. Existing technologies mostly employ traditional image processing and simple multimodal fusion methods, lacking the application of deep learning, especially the Transformer architecture, in multimodal data fusion. This results in insufficient depth and intelligence in data fusion, limiting the accuracy and reliability of anomaly event identification.
[0119] 4. Diverse anomaly detection and early warning mechanisms, covering a wider range of application scenarios: This method uses deep learning models to detect various anomalies in real time, such as falls, crowding, and fights, meeting the safety monitoring needs of different scenarios. Existing technologies mainly focus on the tracking and display of dynamic targets, lacking dedicated detection and early warning mechanisms for multiple anomalies, failing to fully cover diverse safety monitoring needs, and offering only single early warning methods with insufficient timeliness. Attached Figure Description
[0120] Figure 1 This is a simplified flowchart of the present invention. Detailed Implementation
[0121] The embodiments of the present invention will be described in detail below. Although the present invention will be described and illustrated in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, any modifications or equivalent substitutions made to the present invention should be covered within the scope of the claims of the present invention.
[0122] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art will understand that the present invention can be practiced without these specific details.
[0123] To address the shortcomings of existing technical solutions, this invention proposes a method for early warning of abnormal events in large-scene surveillance videos based on a multimodal large model. The method effectively solves the aforementioned technical problems through the following innovative technical solutions:
[0124] (1) Integrated multimodal data processing: By introducing a large multimodal model, video, audio, and sensor data are integrated to achieve deep fusion and comprehensive analysis of multi-source data. This improves the system's ability and accuracy in identifying complex abnormal events, enabling comprehensive monitoring and analysis of different types of abnormal events and enhancing the system's intelligence level.
[0125] (2) High-efficiency real-time processing and optimization: edge computing devices are used for preliminary data processing and optimization (such as video compression, image enhancement, and frame rate adjustment). Combined with high-performance GPUs and hardware acceleration technology, low-latency real-time processing of high-resolution multi-channel video streams is achieved.
[0126] (3) Diversified abnormal event detection and early warning mechanism: using multimodal fusion data, the multimodal large model can detect various abnormal events such as falls, crowding, and fighting in real time and automatically trigger early warning.
[0127] (4) Deep data fusion and intelligent analysis: Deep data fusion and intelligent analysis are performed using a Transformer-based multimodal large model to achieve high-precision anomaly event identification in complex scenarios. This improves the depth and intelligence of data fusion, reduces false alarms and false negatives, and enhances the reliability and stability of the system.
[0128] The core innovation of this method lies in achieving deep fusion and intelligent analysis of video, audio, and sensor data through a multimodal large model based on the Transformer architecture, significantly improving the accuracy and robustness of anomaly event identification. Its key technologies and corresponding implementation schemes are as follows:
[0129] 1. Synchronous acquisition and preprocessing of multimodal data (corresponding to step 2): Spatiotemporal synchronization of multimodal data (video, audio, sensor) is achieved through high-precision timestamps (step 2.4.1) and dynamic transmission delay compensation (step 2.4.2); audio signal enhancement (step 2.3) and sensor data cleaning (step 2.3) are used to improve data quality and lay the foundation for subsequent multimodal fusion.
[0130] 2. Deep Fusion of Multimodal Large Models Based on Transformer (corresponding to steps 3.2-3.3): Cross-modal attention fusion (step 3.2.2) utilizes the self-attention mechanism of Transformer to achieve interactive enhancement of visual, audio and sensor features; Dynamic weighted feature fusion (step 3.3.2) adaptively allocates modal weights according to scene context (such as illumination, noise) to optimize feature representation ability; Temporal context embedding (step 3.3.4): captures short-term dependencies through temporal convolutional networks (TCN) to enhance temporal coherence.
[0131] 3. Multimodal feature-driven anomaly detection (corresponding to step 4): Combine cross-modal verification mechanisms (step 4.2.2), such as verifying visually detected "fall" events through audio voiceprints, or verifying "crowding" events through sensor data; based on spatiotemporal correlation reasoning of multimodal unified feature vectors (step 4.1.3), significantly improve the detection accuracy of anomalies (such as fighting, crowding) in complex scenarios.
[0132] 4. The essential difference from traditional technologies: Existing technologies rely on single-modality (such as video only) or shallow multimodal fusion, while this method achieves deep interaction and semantic alignment of multi-source data through the Transformer architecture (steps 3.2-3.3), which solves the problems of insufficient data fusion and low recognition accuracy in existing technologies.
[0133] Example 1
[0134] A method for early warning of abnormal events in large-scene surveillance videos based on a multimodal large model includes the following steps:
[0135] Step 1: Perform preprocessing such as stitching on multiple high-resolution surveillance videos from the same large scene to obtain unified panoramic video data;
[0136] Step 2: Acquire and preprocess audio data and sensor data to obtain synchronized multimodal data;
[0137] Step 3: Use a traditional small model to detect the surveillance video stream and obtain the task keyframes. Input the keyframes and preprocessed multimodal data into a large multimodal model for data fusion and feature extraction.
[0138] Step 4: Utilize the fused features to detect abnormal events, such as falls, crowding, and fights;
[0139] Step 5: When an abnormal event is detected, trigger the early warning mechanism and send early warning information to relevant management personnel;
[0140] Step 6: Record and store detailed information about the abnormal event for subsequent analysis and tracking.
[0141] Step 1 above specifically includes the following steps:
[0142] Step 1.1: Multi-channel video stream acquisition. Through network or wired connection, the video streams acquired by multiple high-resolution cameras deployed in the same large scene are transmitted to the central processing unit to obtain multiple high-resolution surveillance video streams from different angles and positions.
[0143] Step 1.2: Video stitching and synchronization. High-precision timestamps are added to each video stream to ensure temporal synchronization. Video stitching algorithms (such as feature matching and image fusion) are used to seamlessly stitch multiple video streams into a unified panoramic video stream, resulting in temporally and spatially synchronized high-resolution panoramic video data.
[0144] Step 1.2.1. Perform time synchronization processing on N video streams under the same monitoring scenario, for the first... Add high-precision timestamps to video streams ,in As the base timestamp, For the first Video stream transmission delay compensation amount, For the first Video frame rate;
[0145] Step 1.2.2. Establish spatial mapping relationships between adjacent video streams using a feature matching algorithm, through feature point pairs. Calculate the homography matrix ,in The coordinates of feature points in the source image. The coordinates of the feature points in the target image;
[0146] Step 1.2.3. Generate panoramic video frames based on dynamic weight fusion algorithm ,in The coordinate transformation function obtained in step (2) is... Based on the distance from the pixel to the stitching boundary Calculated fusion weights;
[0147] Step 1.2.4. Output a spatiotemporally synchronized panoramic video stream through the timestamp remapping module. ,in This represents the maximum alignment value of the timestamps from multiple video streams.
[0148] Step 1.3: Video preprocessing. Efficient video decoding and compression algorithms (H.265) are used to reduce data volume and improve transmission efficiency. Image processing techniques such as noise reduction, contrast adjustment, and color correction are applied to improve video quality. The video frame rate is adjusted according to actual needs to balance processing efficiency and video smoothness.
[0149] Step 1.3.1. Decode and compress the panoramic video stream using the H.265 / HEVC standard. Dynamically adjust the quantization parameter (QP) using the rate-distortion optimization (RDO) algorithm to balance compression ratio and video quality. Apply entropy coding (CABAC) to generate a compressed video stream, significantly reducing data transmission bandwidth.
[0150] Step 1.3.2. Combine the nonlocal mean (NLM) algorithm with wavelet thresholding denoising technology to process low-frequency noise and high-frequency noise respectively. Improve the clarity and detail retention of video frames by calculating pixel block similarity weights and hard thresholding of high-frequency components.
[0151] Step 1.3.3. Based on the Limiting Contrast Adaptive Histogram Equalization (CLAHE) algorithm, the video frame is divided into local sub-regions and the histogram distribution is limited. Global contrast equalization is achieved through bilinear interpolation to avoid excessive noise enhancement.
[0152] Step 1.3.4. White balance correction is performed using the Gray World Assumption, adjusting the RGB channel gain to equalize the global color mean, and optimizing the color difference distribution in conjunction with the CIE Lab color space to ensure the authenticity and consistency of color reproduction.
[0153] Step 1.3.5. Based on the actual scenario requirements and hardware processing capabilities, dynamically adjust the video frame rate, and balance processing efficiency and video smoothness through frame sampling or interpolation techniques to ensure optimal real-time performance and resource utilization.
[0154] Step 2 above specifically includes the following steps:
[0155] Step 2.1: Audio data acquisition. Real-time acquisition of environmental audio data to obtain a high-quality real-time environmental audio stream.
[0156] Step 2.2: Collect various environmental parameters in real time to obtain real-time environmental sensor data.
[0157] Step 2.3: Perform audio signal enhancement and noise reduction processing, extract key audio features (spectrum, pitch, etc.), clean and normalize the sensor data, remove noise and outliers, and finally obtain optimized audio data and sensor data to ensure data quality and consistency.
[0158] Step 2.4: Multimodal data synchronization. Through high-precision timestamps and data buffering mechanisms, the synchronization of different modal data in time is ensured, resulting in a multimodal dataset that is fully synchronized in time and space, suitable for subsequent fusion and analysis.
[0159] Step 2.4.1. Generate a unified reference timestamp. Configure Network Time Protocol (NTP) for all data sources (video, audio, sensors) to synchronize the system clocks of each device to microsecond-level accuracy. Using the central processing unit clock as the reference, append a globally unified high-precision timestamp to each data stream. The timestamp format is as follows: ,in As the base time, This is the clock offset compensation amount. This is a clock drift state correction value.
[0160] Step 2.4.2. Dynamic compensation for transmission delay: Real-time monitoring of the transmission delay of each data stream (including network delay and device processing delay), and calculation of the average delay using a sliding window algorithm: Where W and W' are the window sizes, This represents the k-th delayed sample value of the i-th data stream. The timestamp is dynamically corrected. .
[0161] Step 2.4.3. Establish a circular buffer with timestamp indexes for each type of data (video frames, audio clips, sensor readings), and store data packets in ascending time order. Use a greedy matching algorithm to extract data groups from the buffer that meet the following conditions: ,in and For the preset synchronization tolerance threshold (e.g.) ≤10 milliseconds ≤10 milliseconds).
[0162] Step 2.4.4. Missing Data Interpolation and Error Tolerance: If data for a certain modality is missing within the time window, temporary data is generated using a Kalman filter or linear interpolation.
[0163] Video frame missing: Based on optical flow method, predict the motion trajectory between adjacent frames to generate virtual frames.
[0164] Missing audio / sensor data: Use the mean or trend of the preceding and following data points to fit and complete the data.
[0165] An alarm is triggered for consecutive missing events exceeding the threshold (e.g., more than 3 frames), and an anomaly log is recorded.
[0166] Step 2.4.5. Synchronization Verification and Dynamic Adjustment: Verify synchronization accuracy through cross-modal correlation analysis (such as the temporal correlation between video actions and audio voiceprints). If a synchronization deviation is detected (e.g., |t actual alignment − t theoretical alignment| > ϵ), dynamically adjust the buffer window size or correct the timestamp compensation parameters, iteratively optimizing until the synchronization requirements are met.
[0167] Step 3 above specifically includes the following steps:
[0168] Step 3.1: Use traditional small models (such as YOLOv8) based on deep learning for object detection and multi-object tracking to process the spliced surveillance video stream frame by frame, detect important features in each frame of video, and identify key frames.
[0169] Step 3.1.1. Video Stream Input and Frame Normalization: The stitched panoramic video stream is divided into frames, uniformly adjusted to a fixed resolution (e.g., 640×640), and normalized. , where μ and σ are the mean and standard deviation of the training dataset.
[0170] Step 3.1.2. Object detection model inference: Load the pre-trained YOLOv8 model, perform real-time object detection for each frame, and output detection boxes. Category tags and confidence level The detection results were filtered using non-maximum suppression (NMS) to remove overlapping boxes (IoU threshold set to 0.5), retaining high-confidence targets. ≥0.6).
[0171] Step 3.1.3. Multi-target tracking initialization and data association: Assign a unique ID to each detected target and perform tracking based on the DeepSORT algorithm.
[0172] Motion prediction: Predict the target's state vector in the next frame using a Kalman filter.
[0173]
[0174] Appearance feature extraction: Extracting the appearance descriptor of the target using a ReID network.
[0175] Data association: Combining Mahalanobis distance (motion similarity) and cosine similarity (appearance similarity), the cost matrix is calculated:
[0176]
[0177] Where λ=0.7, the Hungarian algorithm is used to complete the optimal matching.
[0178] Step 3.1.4: Keyframe Determination and Dynamic Threshold Adjustment, defining keyframe determination rules:
[0179] Event-triggered: Detects a preset anomaly category (such as a fall, weapon possession) or a certain confidence level. .
[0180] Density-triggered type: Sudden increase in the number of targets in a single frame (e.g.) Or the area density exceeds a threshold (e.g., >5 people / m²). Trajectory anomaly: Sudden change in target movement speed (e.g., ... (or the trajectory deviates from the normal path (Hausdorff distance > 1.5 m).
[0181] Dynamically adjust thresholds: Update judgment parameters in real time based on scene complexity (such as changes in lighting and occlusion ratio) to avoid misjudgments.
[0182] Step 3.1.5: Keyframe Tagging and Metadata Storage: Attach timestamps (tktk), target detection results, and tracking IDs to keyframes to generate structured data. The keyframe data is compressed into JSON format and stored in a cache queue for real-time access by the multimodal large model.
[0183] Step 3.2: Simultaneously input video keyframes, audio data, and sensor data into the multimodal large model. The multimodal large model fuses the multimodal data and extracts more comprehensive and richer scene features by utilizing the complementarity of image features, audio features, and sensor data.
[0184] Step 3.2.1: Multimodal data time alignment, based on the synchronized dataset generated in Step 2.4 Extract the video keyframe Vt, audio clip At, and sensor data St, all with strictly aligned timestamps t. Perform a final check on timestamp discrepancies to ensure they meet the synchronization tolerance threshold.
[0185]
[0186] Step 3.2.2: Cross-modal attention fusion. Features from each modality are input into the multimodal Transformer layer, and information interaction is achieved through a cross-modal attention mechanism: Query-key-value projection: Q, K, and V matrices are generated for fv, fa, and fs respectively. Attention weight calculation:
[0187] ,
[0188] Intermodal feature enhancement: Captures cross-modal associations such as vision-audio and vision-sensor through multi-head attention, and outputs fused features.
[0189] Step 3.2.3. Unified feature vector generation, which integrates the features. Dimensionality reduction and standardization are performed on the input fully connected layer:
[0190] ,
[0191] in Output a unified multimodal feature vector.
[0192] Step 3.3: Integrate the features extracted from visual, audio, and sensor data to form a unified multimodal feature vector, providing comprehensive input for subsequent anomaly detection tasks.
[0193] Step 3.3.1. Feature Dimension Alignment: Align the visual, audio, and sensor fusion features output from Step 3.2. Dimension matching is performed by mapping each modality feature to the same dimensional space (e.g., 512 dimensions) through a fully connected layer to ensure compatibility for subsequent integration.
[0194] Step 3.3.2. Weighted Feature Fusion: Dynamically assign weights based on modal importance. Weight Calculation: Based on the current scene context (e.g., lighting conditions, noise level), generate visual, audio, and sensor weight coefficients using a lightweight network. ,satisfy .
[0195] Feature weighting: Weighted summation of multimodal features according to their weights.
[0196]
[0197] Step 3.3.3. Feature Dimensionality Reduction and Standardization: Use fully connected layers to reduce the dimensionality of the weighted features to the target dimension (e.g., 256 dimensions) to reduce redundant information. Perform Layer Normalization on the dimensionality-reduced features to eliminate dimensional differences between modalities and improve the stability of model training.
[0198] Step 3.3.4. Temporal Context Embedding: Concatenate the current frame features with the historical frame features (through a sliding window buffer) and input them into a Temporal Convolutional Network (TCN) to capture short-term dependencies and enhance the temporal coherence of the features.
[0199] Step 3.3.5: Unified Feature Vector Generation. The temporally enhanced features are input into the residual connection module to further optimize the feature representation capability and output the final unified multimodal feature vector. Add metadata such as timestamps and modality source identifiers to the feature vectors to ensure traceability to the original data.
[0200] Step 3.3.6: Feature Storage and Interface Calls. The unified feature vector is stored in a high-speed cache queue, and a lightweight compression algorithm (such as Zstandard) is used to reduce memory usage. A standardized API interface is provided to support the anomaly detection model in Step 4 in real-time reading of feature data, ensuring low latency in end-to-end processing.
[0201] Step 4 above specifically includes the following steps:
[0202] Step 4.1: Input the fused features into the abnormal event detection model, perform inference calculations, and obtain preliminary detection results for various abnormal events, including event type and location.
[0203] Step 4.1.1: Model Loading and Initialization. Load the pre-trained multimodal anomaly detection model (such as a Transformer-based spatiotemporal network), and load the model weights and configuration file. Initialize the input interface to receive the unified multimodal feature vector from Step 3.3. And allocate computing resources (such as GPU memory).
[0204] Step 4.1.2: Feature input and preprocessing. Dynamically standardize the input features to eliminate the bias in feature distribution. Divide the feature sequence into time windows (e.g., 5-second windows) and fill or truncate it to a fixed length to adapt to the model input dimension.
[0205] Step 4.1.3: Spatiotemporal correlation reasoning. A Temporal Convolutional Network (TCN) is used to extract short-term dependencies in feature sequences, capturing dynamic patterns in event evolution. An attention mechanism is combined to enhance key regions (such as densely populated areas), generating spatially sensitive feature representations. Fully connected layers and a Softmax function are used to output the probability distribution for each type of abnormal event (falls, crowding, fights, etc.).
[0206]
[0207] Step 4.1.4: Preliminary result generation and filtering, retaining probability For events with a confidence level ≥ 0.7, low-confidence detection results are filtered out. Non-maximum suppression (NMS) is applied to multiple bounding boxes for the same event (e.g., IoU > 0.4), retaining the most significant result.
[0208] Step 4.2: Classify and confirm the inference results according to the predefined abnormal event categories (such as falls, crowding, fighting, etc.), and clarify the detected abnormal event categories and their specific locations.
[0209] Step 4.2.1: Load predefined event categories. Load the predefined abnormal event category configuration file (such as fall, crowding, fighting, etc.), which contains the semantic description, feature template and confidence threshold of each event category.
[0210] Step 4.2.2: Joint verification of multimodal evidence. Combining the multimodal features generated in Step 3.3, cross-modal verification is performed on the preliminary detection results of Step 4.1:
[0211] Visual-audio association: For example, when a "fall" event is detected, verify whether there is a collision sound or a cry for help in the audio data.
[0212] Vision-sensor correlation: For example, when a "crowding" event is detected, verify whether the area temperature or people flow sensor data is abnormally high.
[0213] Step 4.2.3: Classification model inference. Input the unified feature vector into the pre-trained multi-classification model (such as a hierarchical classification network based on attention mechanism) and output the event category probability distribution.
[0214] Step 4.3: Filter the detection results through a rule engine or secondary verification mechanism to eliminate false alarms and irrelevant events, thereby improving the accuracy of the detection.
[0215] Step 4.3.1: Rule engine condition filtering. Read scene-related filtering conditions (such as time range, regional restrictions, and event duration thresholds) from the predefined rule library. Match the classification results from Step 4.2 against each rule in the rule library and remove events that violate logical conditions (such as high-density alarms during non-working hours).
[0216] Step 4.3.2: Lightweight secondary validation model inference. Load the pre-trained lightweight validation model (such as random forest or small neural network), and input the multimodal feature vector and preliminary detection results. Output the corrected confidence score. If the corrected score is lower than the threshold... It will then be marked as a false alarm.
[0217] Step 4.3.3: Multimodal evidence consistency check. If the event relies on visual detection (e.g., "fighting"), verify whether there are fighting sounds or abnormal voiceprints in the audio data. Events with conflicting modal evidence (e.g., only visual detection but no audio support) are downweighted or filtered.
[0218] Step 5 above specifically includes the following steps:
[0219] Step 5.1: Based on the detected abnormal event category and location, generate corresponding early warning information, including detailed information such as event type, occurrence time, and occurrence location, forming structured early warning information.
[0220] Step 5.2: Send the early warning information to relevant management personnel in real time so that they can take prompt countermeasures.
[0221] Step 5.3: Record and store the detailed data of the early warning information and related abnormal events in the database to facilitate subsequent analysis and tracking, establish a complete abnormal event file, and support post-event review and security management.
[0222] Step 6 above specifically includes the following steps:
[0223] Step 6.1: Organize relevant data of abnormal events, including video clips, audio clips, sensor data, early warning information, etc., to form a complete event dataset.
[0224] Step 6.2: Store the event data in the local storage system, and use efficient storage technology to ensure fast data retrieval and access.
[0225] Step 6.3: Employ data backup and encryption technologies to ensure data security and reliability, prevent data loss and leakage, and guarantee the long-term preservation and secure use of event data.
Claims
1. A method for early warning of abnormal events in large-scene surveillance video based on a multimodal large model, characterized in that, Includes the following steps: Step 1: Perform preprocessing such as stitching on multiple high-resolution surveillance videos from the same large scene to obtain unified panoramic video data; Step 2: Acquire and preprocess audio data and sensor data to obtain synchronized multimodal data; Step 3: Use a traditional small model to detect the surveillance video stream and obtain the task keyframes. Input the keyframes and preprocessed multimodal data into a large multimodal model for data fusion and feature extraction. Step 4: Utilize the fused features to detect abnormal events, such as falls, crowding, and fights; Step 5: When an abnormal event is detected, trigger the early warning mechanism and send early warning information to relevant management personnel; Step 6: Record and store detailed information about the abnormal event for subsequent analysis and tracking.
2. The method according to claim 1, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Multi-channel video stream acquisition. Through network or wired connection, the video streams acquired by multiple high-resolution cameras deployed in the same large scene are transmitted to the central processing unit to obtain multiple high-resolution surveillance video streams from different angles and positions. Step 1.2: Video stitching and synchronization. Add a high-precision timestamp to each video stream to ensure that the videos are synchronized in time. Use a video stitching algorithm to seamlessly stitch multiple video streams into a unified panoramic video stream to obtain a panoramic high-resolution video data that is synchronized in time and space. Step 1.3: Video preprocessing. Use video decoding and compression algorithms to reduce data volume and improve transmission efficiency. Apply image processing techniques such as noise reduction, contrast adjustment, and color correction to improve video quality. Adjust the video frame rate according to actual needs to balance processing efficiency and video smoothness.
3. The method according to claim 2, characterized in that, Step 1.2 specifically includes the following steps: Step 1.2.
1. Perform time synchronization processing on N video streams under the same monitoring scenario, for the first... Add high-precision timestamps to video streams ,in As the base timestamp, For the first Video stream transmission delay compensation amount, For the first Video frame rate; Step 1.2.
2. Establish spatial mapping relationships between adjacent video streams using a feature matching algorithm, through feature point pairs. Calculate the homography matrix ,in The coordinates of feature points in the source image. The coordinates of the feature points in the target image; Step 1.2.
3. Generate panoramic video frames based on dynamic weight fusion algorithm ,in The coordinate transformation function obtained in step (2) is... Based on the distance from the pixel to the stitching boundary Calculated fusion weights; Step 1.2.
4. Output a spatiotemporally synchronized panoramic video stream through the timestamp remapping module. ,in This represents the maximum alignment value of the timestamps from multiple video streams.
4. The method according to claim 2, characterized in that, Step 1.3 specifically includes the following steps: Step 1.3.
1. Decode and compress the panoramic video stream using the H.265 / HEVC standard, dynamically adjust the quantization parameters through a rate-distortion optimization algorithm to balance the compression ratio and video quality, and apply entropy coding to generate a compressed video stream, significantly reducing data transmission bandwidth; Step 1.3.
2. Combining the nonlocal mean algorithm and wavelet thresholding denoising technology, low-frequency noise and high-frequency noise are processed respectively. By calculating the pixel block similarity weight and hard thresholding of high-frequency components, the clarity and detail retention of video frames are improved. Step 1.3.
3. Based on the limited contrast adaptive histogram equalization algorithm, the video frame is divided into local sub-regions and the histogram distribution is limited. Global contrast equalization is achieved through bilinear interpolation to avoid excessive noise enhancement. Step 1.3.
4. Perform white balance correction using the grayscale world assumption, adjust the RGB channel gain to equalize the global color mean, and optimize the color difference distribution using the CIE Lab color space to ensure the authenticity and consistency of color reproduction; Step 1.3.
5. Based on the actual scenario requirements and hardware processing capabilities, dynamically adjust the video frame rate, and balance processing efficiency and video smoothness through frame sampling or interpolation techniques to ensure optimal real-time performance and resource utilization.
5. The method according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: Audio data acquisition, real-time acquisition of environmental audio data to obtain a high-quality real-time environmental audio stream; Step 2.2: Collect various environmental parameters in real time to obtain real-time environmental sensor data; Step 2.3: Perform audio signal enhancement and noise reduction processing, extract key audio features, clean and normalize sensor data to remove noise and outliers, and finally obtain optimized audio data and sensor data to ensure data quality and consistency. Step 2.4: Multimodal data synchronization. Through high-precision timestamps and data buffering mechanisms, the synchronization of different modal data in time is ensured, resulting in a multimodal dataset that is fully synchronized in time and space, suitable for subsequent fusion and analysis.
6. The method according to claim 5, characterized in that, Step 2.4 specifically includes the following steps: Step 2.4.
1. Generate a unified reference timestamp. Configure a network time protocol for all data sources, synchronize the system clocks of each device to microsecond-level precision, and attach a globally unified high-precision timestamp to each data stream based on the central processing unit clock. The timestamp format is as follows: , in As the base time, This is the clock offset compensation amount. This is a clock drift state correction value; Step 2.4.
2. Dynamic compensation for transmission delay: Real-time monitoring of the transmission delay of each data stream, and calculation of the average delay using a sliding window algorithm: Where W and W' are the window sizes, For the first The first path of data stream The timestamp is dynamically corrected based on the second-delayed sample value. ; Step 2.4.
3. Establish a circular buffer with timestamp indexes for each type of data, store data packets in ascending time order, and use a greedy matching algorithm to extract data groups from the buffer that meet the following conditions: , in and The preset synchronization tolerance threshold, Indicates the timestamp of a video frame. Represents the timestamp of an audio segment. The timestamp represents the sensor data; Step 2.4.
4. Missing Data Interpolation and Error Tolerance: If data for a certain modality is missing within the time window, temporary data is generated using a Kalman filter or linear interpolation. Video frame missing: Based on optical flow method, predict the motion trajectory between adjacent frames to generate virtual frames; Missing audio / sensor data: Complete the missing data by fitting the mean or trend of the preceding and following data points; For consecutive events exceeding the threshold, such as missing events exceeding 3 frames, an alarm is triggered and an anomaly log is recorded; Step 2.4.
5. Synchronization Verification and Dynamic Adjustment: Synchronization accuracy is verified through cross-modal correlation analysis. If synchronization deviation is detected, the buffer window size is dynamically adjusted or the timestamp compensation parameters are corrected. Iterative optimization is performed until synchronization requirements are met, generating the final synchronization dataset. .
7. The method according to claim 6, characterized in that, Step 3 specifically includes the following steps: Step 3.1: Use traditional small models such as deep learning-based object detection and multi-object tracking to process the spliced surveillance video stream frame by frame, detect important features in each frame of video, and identify key frames; Step 3.1.
1. Video Stream Input and Frame Normalization: The stitched panoramic video stream is divided into frames, uniformly adjusted to a fixed resolution of 640×640, and normalized. , in This represents the pixel value at coordinates (x, y) in the normalized image. This represents the pixel value of the original image at coordinates (x, y), where μ and σ are the mean and standard deviation of the training dataset; Step 3.1.
2. Object detection model inference: Load the pre-trained YOLOv8 model, perform real-time object detection for each frame, and output detection boxes. Category Tags Confidence level The detection results were filtered using non-maximum suppression to remove overlapping bounding boxes. The IoU threshold was set to 0.5, while the confidence level was preserved. The goal; Step 3.1.
3. Multi-target tracking initialization and data association: Assign a unique ID to each detected target and perform tracking based on the DeepSORT algorithm: Motion prediction: Predicting the target's state vector in the next frame using a Kalman filter. The width of the target bounding box and height , For the goal and Speed of motion in a certain direction; Appearance feature extraction: Extracting the appearance descriptor of the target using a ReID network. ; Data association: Calculate the cost matrix by combining Mahalanobis distance and cosine similarity. ; in The optimal matching is achieved using the Hungarian algorithm. The Mahalanobis distance represents the motion similarity between the predicted target state and the detection result. Represents cosine distance, which measures the similarity of the appearance features of targets; Step 3.1.4: Keyframe Determination and Dynamic Threshold Adjustment, defining keyframe determination rules: Event-triggered: Detects a preset anomaly category or confidence level. , Density-triggered type: Sudden increase in the number of targets in a single frame Or the area density exceeds the threshold of more than 5 people / m². This represents the instantaneous change in the number of targets detected between adjacent frames; Anomalous trajectory: Sudden change in the target's velocity Or the trajectory deviates from the normal path by more than 1.5 m from the Hausdorff distance; Dynamically adjust thresholds: Update judgment parameters in real time based on scene complexity to avoid misjudgments; Step 3.1.5: Keyframe Marking and Metadata Storage, Attaching Timestamps to Keyframes Target detection results and tracking IDs are used to generate structured data. ; Keyframe data is compressed into JSON format and stored in a cache queue for real-time access by large multimodal models. Indicates the timestamp of the keyframe. The coordinates of the object detection bounding box. For confidence level, Indicates the target category label. This represents a unique tracking identifier for the target, used to associate the same target across frames. Step 3.2: Simultaneously input video keyframes, audio data, and sensor data into the multimodal large model. The multimodal large model fuses the multimodal data and extracts more comprehensive and richer scene features by utilizing the complementarity of image features, audio features, and sensor data. Step 3.2.1: Multimodal data time alignment, based on the synchronized dataset generated in Step 2.4 Extract timestamp Strictly aligned video keyframes Audio clips and sensor data A final check is performed on the timestamp deviation to ensure that the synchronization tolerance threshold is met. : ; Step 3.2.2: Cross-modal attention fusion. Features from each modality are input into the multimodal Transformer layer, and information interaction is achieved through a cross-modal attention mechanism: query-key-value projection. Generate query matrices respectively , Key matrix, Value matrix, attention weight calculation: , This represents the visual features extracted from video keyframes. Indicates audio features, Features extracted from environmental sensor data The dimension of the key vector. It is a normalization function; Intermodal feature enhancement: Captures cross-modal correlations between vision and audio, and vision and sensors through multi-head attention, and outputs fused features. , , It is the dimension of the feature vector; Step 3.2.
3. Unified feature vector generation, which integrates the features. Dimensionality reduction and standardization are performed on the input fully connected layer: , in Output a unified multimodal feature vector. Representation layer normalization, Indicates a modified linear unit. Represents the bias vector; Step 3.3: Integrate the features extracted from visual, audio, and sensor data to form a unified multimodal feature vector, providing comprehensive input for subsequent anomaly detection tasks; Step 3.3.
1. Feature Dimension Alignment: Align the visual, audio, and sensor fusion features output from Step 3.
2. Dimension matching is performed by mapping the features of each modality to the same dimensional space through a fully connected layer; Step 3.3.
2. Weighted feature fusion, dynamically assigning weights based on modal importance: Weight calculation: Based on the current scene context, a lightweight network generates weight coefficients for vision, audio, and sensors. ,satisfy ; Feature weighting: Weighted summation of multimodal features according to their weights. , Step 3.3.
3. Feature dimensionality reduction and standardization: Use a fully connected layer to reduce the dimensionality of the weighted features to the target dimension, reduce redundant information, and perform layer normalization on the dimensionality-reduced features; Step 3.3.
4. Temporal Context Embedding: The features of the current frame and the features of the historical frames are concatenated through a sliding window buffer and input into a temporal convolutional network to capture short-term dependencies and enhance the temporal coherence of the features; Step 3.3.5: Unified Feature Vector Generation. The temporally enhanced features are input into the residual connection module to optimize feature representation capabilities and output the final unified multimodal feature vector. Add metadata such as timestamps and modality source identifiers to the feature vectors to ensure traceability to the original data; Step 3.3.6: Feature storage and interface call. The unified feature vector is stored in a high-speed cache queue. A lightweight compression algorithm is used to reduce memory usage. A standardized API interface is provided to support the anomaly detection model in Step 4 to read feature data in real time, ensuring low latency in end-to-end processing.
8. The method according to claim 1, characterized in that, Step 4 specifically includes the following steps: Step 4.1: Input the fused features into the anomaly detection model for inference calculation to obtain preliminary detection results for various anomalies, including event type and location; Step 4.1.1: Model loading and initialization, load the pre-trained multimodal anomaly detection model, load model weights and configuration files, initialize the input interface, and receive the unified multimodal feature vector from Step 3.
3. And allocate computing resources; Step 4.1.2: Feature Input and Preprocessing. Dynamically standardize the input features to eliminate feature distribution bias. Segment the feature sequence according to time windows, padding or truncating to a fixed length to adapt to the model input dimension, resulting in a standardized input feature vector. ; Step 4.1.3: Spatiotemporal correlation reasoning. Short-term dependencies in feature sequences are extracted using a temporal convolutional network to capture dynamic patterns in event evolution. Attention mechanisms are used to enhance key regions, generating spatially sensitive feature representations. Fully connected layers and a Softmax function are employed to output the probability distribution for each type of anomalous event. , Represents the classification weight matrix. This represents the classification bias vector. Indicates the first Class of exceptions, Step 4.1.4: Preliminary result generation and filtering, retaining probability For events, low-confidence detection results are filtered out, and non-maximum suppression is applied to multiple detection boxes for the same event to retain the most significant results; Step 4.2: Classify and confirm the inference results according to the predefined abnormal event categories, and clarify the detected abnormal event categories and their specific locations; Step 4.2.1: Load predefined event categories. Load the predefined abnormal event category configuration file, which contains the semantic description, feature template and confidence threshold of each event category; Step 4.2.2: Joint verification of multimodal evidence, combining the multimodal feature vectors generated in step 3.
3. Cross-modal validation of the preliminary detection results from step 4.1: Visual-audio correlation: When a "fall" event is detected, verify whether there are collision sounds or cries for help in the audio data; Vision-sensor correlation: When a "crowding" event is detected, verify whether the area temperature or people flow sensor data is abnormally high; Step 4.2.3: Classification model inference. Input the unified feature vector into the pre-trained multi-classification model and output the event category probability distribution. Step 4.3: Filter the detection results through a rule engine or secondary verification mechanism to eliminate false alarms and irrelevant events, thereby improving the accuracy of the detection; Step 4.3.1: Rule engine condition filtering. Read scene-related filtering conditions from the predefined rule library, match the classification results of step 4.2 with the rule library one by one, and remove events that violate the logical conditions. Step 4.3.2: Lightweight secondary validation model inference. Load the pre-trained lightweight validation model, input the multimodal feature vector and the preliminary detection results, and output the corrected confidence score. If the corrected score is lower than the threshold, i.e. Then mark it as a false alarm; Step 4.3.3: Multimodal evidence consistency check. If the event relies on visual detection, verify whether there are abnormal voiceprints in the audio data, and reduce the weight or filter events with conflicting modal evidence.
9. The method according to claim 1, characterized in that, Step 5 specifically includes the following steps: Step 5.1: Based on the detected abnormal event category and location, generate corresponding early warning information, including detailed information such as event type, occurrence time, and occurrence location, to form structured early warning information; Step 5.2: Send the early warning information to relevant management personnel in real time so that they can take prompt countermeasures; Step 5.3: Record and store the detailed data of the early warning information and related abnormal events in the database to facilitate subsequent analysis and tracking, establish a complete abnormal event file, and support post-event review and security management.
10. The method according to claim 1, characterized in that, Step 6 specifically includes the following steps: Step 6.1: Organize relevant data of abnormal events, including video clips, audio clips, sensor data, early warning information, etc., to form a complete event dataset; Step 6.2: Store the event data in the local storage system, and use efficient storage technology to ensure fast data retrieval and access; Step 6.3: Employ data backup and encryption technologies to ensure data security and reliability, prevent data loss and leakage, and guarantee the long-term preservation and secure use of event data.
Citation Information
Patent Citations
Dynamic target tracking fusion method for large-scale scene monitoring
CN119027882A
Cited By
Multi-scene voice alarm system based on video AI behavior analysis
CN121545265A
Adaptive encryption transmission method for multi-mode transmitter data
CN121547286A
Crowd density detection method fusing optical flow and texture features
CN121661600A
High-diamagnetic passbook printing anomaly detection system based on large model
CN121686493A
Method and system for fast searching and automatic networking of large AES network equipment
CN121727944A