Video monitoring system, video monitoring method and spherical camera
Through edge computing node clusters, ultra-high-definition video acquisition modules and 5G network slicing technology, combined with cloud distributed computing platform, the existing video surveillance system has been solved in real-time, data fusion accuracy and system collaboration efficiency, and efficient multimodal data processing and intelligent decision-making are achieved, suitable for smart cities and industrial security.
Patent Information
- Application Number
- CN202510785198.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-22
AI Technical Summary
The existing video surveillance systems have shortcomings in real-time, data fusion accuracy and system collaboration efficiency, and it is difficult to meet the needs of efficient and intelligent decision-making in complex scenarios.
Deploy an edge computing node cluster, combine ultra-high-definition video acquisition module, multi-type environmental sensor array and 5G network slicing technology, and realize efficient collaborative processing of video and sensor data and multi-modal fusion decisions through a cloud distributed computing platform.
It realizes millisecond event detection and high-reliability alarms in complex scenarios, improves the real-time, accuracy and reliability of the system, and is suitable for smart cities and industrial security fields.
Smart Images

Figure CN120358330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video surveillance, and specifically, to a video surveillance system, a video surveillance method, and a spherical camera. Background Art
[0002] Existing video surveillance systems usually consist of cameras, transmission networks, and back-end analysis platforms. Among them, spherical cameras are widely used in fields such as security and traffic management due to their 360° panoramic surveillance capabilities. Traditional video surveillance methods mainly rely on cloud-based centralized processing, that is, video data collected by cameras is transmitted through the network to the server for analysis, while spherical cameras usually use mechanical rotation or fisheye lenses to achieve wide-angle coverage. In terms of data processing, existing technologies mostly use fixed-threshold detection or simple multi-sensor data stitching, such as the simple superposition of environmental data such as temperature and sound with video, to enhance the perception ability of the surveillance scene. However, these methods still have obvious deficiencies in terms of real-time performance, data fusion accuracy, and system cooperation efficiency.
[0003] The main problems faced by current video surveillance systems include: (1) insufficient real-time performance. Due to the large amount of video data and complex processing processes, traditional cloud-based centralized computing is difficult to respond to emergencies in a timely manner; (2) the data fusion algorithm is rough. Existing methods usually only perform simple stitching or weighted averaging on multi-source data, and fail to fully explore the correlation between video and sensor data, resulting in low reliability of the analysis results; (3) poor spatio-temporal calibration accuracy. Existing systems have large errors in time synchronization and spatial coordinate alignment, making it difficult to accurately match multi-sensor data and affecting the decision-making accuracy of the overall surveillance system. These problems seriously restrict the application effect of video surveillance systems in complex scenarios (such as smart cities, industrial security, etc.), and there is an urgent need for an efficient, accurate, and intelligent decision-making solution. Summary of the Invention
[0004] The main object of the present invention is to provide a video surveillance system, a video surveillance method, and a spherical camera, aiming to solve the problems in terms of real-time performance, data fusion accuracy, and system cooperation efficiency in the existing technology.
[0005] The technical solution of the present invention is as follows: A video surveillance system, comprising: A cluster of edge computing nodes deployed in the surveillance area, including GPU and DSP heterogeneous computing units; A super-high-definition video acquisition module connected to the cluster of edge computing nodes, with an 8K image sensor and a data cache unit built-in; An array of multi-type environmental sensors, including temperature, sound, and vibration sensors, communicating with the cluster of edge computing nodes through a high-speed interface; A transmission channel constructed based on 5G network slicing technology, with a data transmission priority policy set. A cloud distributed computing platform, integrating a real-time stream processing engine and a multi-modal fusion decision-making module.
[0006] In a possible embodiment, the execution of the edge computing node cluster includes: Using a hardware-accelerated video encoder to perform H.265 compression on the original video stream, and synchronously marking the dynamic target region ROI during the compression process; Performing normalization processing on the sensor data to generate a standardized data frame containing timestamps.
[0007] In a possible embodiment, the cloud distributed computing platform includes: A spatio-temporal convolutional neural network module for extracting the spatio-temporal motion features of the target in the video stream; A Mel frequency cepstral coefficient analysis unit for analyzing the spectral features of the sound sensor; A fusion decision tree generator for dynamically adjusting the weights of the feature splitting nodes to construct a multi-modal joint inference model.
[0008] In a possible embodiment, the fusion decision tree generator performs: Calculating the mutual information amount between the video features and the sensor features, and dynamically selecting the splitting node weights, where the video features include the spatio-temporal motion features extracted by the spatio-temporal convolutional neural network; Sorting the features according to the mutual information amount, selecting the feature with the highest correlation as the root node, and generating a multi-modal joint inference model; During the construction of the decision tree, combining the dynamic changes of the video features and the sensor data, activating the joint decision-making path of the multi-modal features, and comprehensively determining the joint decision-making path based on the relevance between the abnormal features of the environmental parameters and the video detection target.
[0009] In a possible embodiment, it further includes: A satellite time synchronization module for adding sub-millisecond accuracy timestamps to the video frames and sensor data; A three-dimensional laser scanning modeling unit for establishing a spatial coordinate system of the monitoring area and mapping it to the sensor position.
[0010] A video monitoring method running on any of the above video monitoring systems, including: Parallelly executing through the edge computing node cluster: a. Performing ROI region detection and encoding compression based on hardware acceleration on the 8K video stream; b. Performing unit unification and outlier filtering on the multi-source sensor data; Transmit the processed data to the cloud through the transmission channel constructed by the 5G network slicing technology; Perform spatio-temporal alignment of multi-modal feature fusion analysis in the cloud to generate a joint decision result.
[0011] In a possible embodiment, the multi-modal feature fusion analysis includes: Extract the spatio-temporal convolutional features of the target in the video stream; Calculate the MFCC coefficients and high-frequency energy mutation points of the sound sensor data; Construct a fusion decision tree, dynamically adjust the splitting node weights based on the mutual information of video features and sensor features, and select the features with mutual information exceeding the preset threshold as the key splitting nodes.
[0012] In a possible embodiment, the construction process of the fusion decision tree includes: Calculate the mutual information of video dynamic features and sensor data, and the video dynamic features are extracted by a spatio-temporal convolutional neural network; Dynamically adjust the splitting nodes of the decision tree according to the mutual information threshold, and select the features with mutual information exceeding the preset threshold as the key splitting nodes; Construct a multi-modal joint inference model. When the sensor data detects abnormal features of environmental parameters, perform cross-verification with the video analysis results. If the multi-modal features meet the joint determination conditions, trigger the corresponding event alarm.
[0013] In a possible embodiment, it further includes: Align the video frames and sensor data according to the satellite timestamp, and start data resynchronization when the time deviation exceeds 1 ms; Based on the three-dimensional space coordinate mapping, exclude the abnormal sensor data with a distance from the target object exceeding 2 meters.
[0014] A spherical camera, including the video monitoring system described in any one of the above, and the video monitoring system runs the video monitoring method described in any one of the above.
[0015] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: The technical solution of the present invention realizes the local parallel processing of video and sensor data by deploying an edge computing node cluster (including GPU and DSP heterogeneous computing units), solving the problem of insufficient real-time performance caused by cloud centralized computing; realizes the high-precision synchronous acquisition of multi-modal data through the collaborative work of an ultra-high-definition video acquisition module and a multi-type environmental sensor array, breaking through the perception limitation of traditional single data sources; constructs a hierarchical transmission channel through 5G network slicing technology to achieve low-latency priority transmission of key data, ensuring the instant response ability to emergencies; realizes the in-depth correlation analysis of video spatio-temporal features and sensor features through a cloud distributed computing platform integrating a real-time stream processing engine and a multi-modal fusion decision-making module, solving the problem of insufficient fusion accuracy caused by simple data splicing; finally, through the systematic integration of the above technologies, it realizes millisecond-level event detection, cross-modal accurate decision-making and highly reliable alarm output in complex scenarios, comprehensively overcoming the technical bottlenecks of traditional monitoring systems in terms of real-time performance, fusion accuracy and collaborative efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0017] Figure 1 It is the block diagram of the video monitoring system in Embodiment 1; Figure 2 It is the flowchart of video data processing in Embodiment 1; Figure 3 It is the flowchart of sensor data processing in Embodiment 1; Figure 4 It is the flowchart of constructing a fusion decision tree in Embodiment 1; Figure 5 It is the flowchart of the video monitoring method in Embodiment 2.
[0018] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the object, technical solution and advantages of the present application more clear, the following will further describe the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Embodiment 1
[0020] As Figures 1 to 4 shown, this embodiment proposes a video monitoring system, including: A cluster of edge computing nodes deployed in the monitoring area, including GPU and DSP heterogeneous computing units.
[0021] This cluster is deployed at the monitoring front end. The GPU is responsible for parallel computing and processing of video data (such as object detection and feature extraction), while the DSP is dedicated to real-time signal processing of sensor data (such as filtering and spectrum analysis). The two types of processors work together through a high-speed bus. When the GPU processes an 8K video stream, it calls the DSP to preprocess the sensor data, forming a heterogeneous computing pipeline to ensure efficient parallel processing of video and sensor data.
[0022] Through hardware-level task division, the parallel computing ability of the GPU and the low-power consumption characteristics of the DSP complement each other, enabling the edge nodes to achieve high-speed data processing while maintaining low energy consumption, significantly improving real-time performance (low latency), and at the same time reducing the overall power consumption of the system, which is suitable for long-running monitoring scenarios.
[0023] An ultra-high-definition video acquisition module connected to the edge computing node cluster, with an 8K image sensor and a data cache unit built-in.
[0024] The image sensor acquires video data at ultra-high resolution, and the cache unit temporarily stores the original video stream. When the edge computing node cluster extracts key frames as needed, a high-speed interface ensures rapid data transmission to the computing unit for processing.
[0025] Ultra-high resolution significantly improves the accuracy of object detection. Combined with an intelligent cache mechanism, it can maintain local emergency analysis capabilities during network fluctuations, enhancing the stability and reliability of the system in complex environments.
[0026] An array of multi-type environmental sensors, including temperature, sound, and vibration sensors, communicates with the edge computing node cluster through a high-speed interface.
[0027] Each type of sensor acquires environmental data at different sampling rates, packs the data uniformly through a high-speed interface, and transmits it to the edge computing node cluster, where the DSP performs preprocessing and time-domain alignment.
[0028] Multi-modal data collaborative acquisition expands the dimension of environmental perception. Combined with video analysis, it significantly improves the detection accuracy of abnormal events (such as intrusion and fire), while reducing the false alarm rate, making the monitoring system more intelligent.
[0029] A transmission channel built based on 5G network slicing technology, with a data transmission priority policy set.
[0030] When the edge computing node cluster transmits data through the 5G network, 5G network slicing technology assigns different priority transmission channels to video and sensor data, ensuring high-rate and low-latency transmission of critical data.
[0031] The 5G network slicing technology effectively avoids the impact of network congestion on data transmission, ensures the real-time and stability of monitoring data, and is especially suitable for large-scale monitoring systems with high-density deployment.
[0032] Cloud distributed computing platform, integrating real-time stream processing engine and multi-modal fusion decision-making module.
[0033] After the cloud platform receives the data transmitted by the edge computing node cluster, the real-time stream processing engine analyzes the video and sensor data in parallel, and the multi-modal fusion module synthesizes various features and outputs structured events.
[0034] The distributed architecture of the cloud distributed computing platform supports high-concurrency processing, and the multi-modal fusion algorithm greatly improves the recognition accuracy of complex events, enabling the monitoring system to have more powerful intelligent analysis capabilities and meet the high-end application requirements such as smart cities and industrial security.
[0035] Each component in the video monitoring system works together through a high-precision synchronization mechanism to form a closed loop of "edge real-time preprocessing - network reliable transmission - cloud intelligent decision-making". While ensuring ultra-high-definition video monitoring, it realizes high-rate and low-latency data processing, significantly improving the real-time, accuracy and reliability of the system.
[0036] In this embodiment, the edge computing node cluster performs the following: Use a hardware-accelerated video encoder to perform H.265 compression on the original video stream, and synchronously mark the dynamic target region ROI during the compression process.
[0037] The edge computing node cluster uses a dedicated hardware encoder (such as the encoding unit built into the GPU) to perform real-time compression on the original video stream, adopting the H.265 / HEVC encoding standard, which greatly reduces the data volume while ensuring the image quality.
[0038] Dynamic target region (ROI) marking: ROI definition: The dynamic target region refers to the region where motion changes occur in the video frame (such as pedestrians, vehicles, etc.), which is detected in real time through background difference method or optical flow method.
[0039] Synchronization marking mechanism: During the encoding process, the encoder combines the output of the target detection algorithm (such as YOLO), assigns a higher bit rate (retaining more details) to the detected dynamic region, and compresses the non-ROI region with a lower bit rate.
[0040] Perform normalization processing on the sensor data to generate a standardized data frame containing timestamps.
[0041] Operating process of sensor data normalization processing: Unify the data format, and convert sensor data in different units such as temperature (°C), sound (dB), and vibration (g) into standardized values (such as in the range of 0-1).
[0042] Filter out outliers, and use the sliding window statistical method (such as the 3σ principle) to remove noise data to ensure the validity of the data.
[0043] Embed timestamps, and add high-precision timestamps (aligned with the video frame timestamps) to each sensor data for subsequent multi-modal data fusion.
[0044] The edge computing node cluster significantly reduces the video data volume through hardware-accelerated video encoding and dynamic ROI marking technology, while ensuring the detail integrity of the key target areas, and solves the problem of image quality loss caused by traditional global compression; the normalization processing of sensor data (including format unification, outlier filtering, and time synchronization) eliminates the heterogeneity of multi-source data and provides highly consistent input for subsequent multi-modal fusion analysis in the cloud. The collaborative implementation of the two enables the system to complete data optimization at the edge side, reduce the occupancy of transmission bandwidth, and improve the analysis efficiency in the cloud. Overall, it achieves a balance between high-precision monitoring and resource overhead, and is applicable to large-scale deployed smart city and industrial Internet of Things scenarios.
[0045] In this embodiment, the cloud distributed computing platform includes: A spatio-temporal convolutional neural network module for extracting the spatio-temporal motion features of targets in the video stream.
[0046] Execution steps of the spatio-temporal convolutional neural network (TCN) module: Input preprocessing: Receive the compressed video stream from the edge computing node cluster, and generate a continuous video frame sequence after decoding.
[0047] Spatio-temporal feature extraction: In the spatial dimension, use a 3D convolutional kernel to extract static features such as the shape and texture of the target on a single-frame image. In the time dimension, capture the motion trajectory and behavior patterns of the target (such as walking direction and speed change) through cross-frame temporal convolutional operations.
[0048] Feature fusion: Concatenate the spatial feature map and the temporal feature map in the channel dimension to form a unified spatio-temporal feature tensor.
[0049] Dimensionality reduction output: Generate a low-dimensional feature vector through a fully connected layer for use by subsequent decision-making modules.
[0050] The spatio-temporal convolutional neural network module can accurately identify the dynamic behavior features of targets (such as wandering and running) by jointly analyzing the spatio-temporal dimension information of the video. Compared with traditional 2D CNN algorithms, the accuracy of judging behavior intentions is improved.
[0051] Mel Frequency Cepstral Coefficient Analysis Unit, used to analyze the spectral characteristics of the sound sensor.
[0052] Execution steps of the Mel Frequency Cepstral Coefficient (MFCC) Analysis Unit: Pre-emphasis processing: Apply high-frequency enhancement filtering to the original waveform of the sound sensor to improve the signal-to-noise ratio in the speech frequency band.
[0053] Framing and windowing: Segment the continuous audio into short-time frames of 20 ms, and use a Hamming window to reduce spectral leakage.
[0054] Fourier transform: Perform FFT conversion on each frame of the signal to obtain the power spectrum.
[0055] Mel filter bank: Simulate the auditory characteristics of the human ear through 40 triangular band-pass filters to extract the Mel spectrum.
[0056] Cepstrum analysis: Take the logarithm of the Mel spectrum and then perform DCT transformation, retaining the first 13 coefficients as the feature vector.
[0057] The Mel Frequency Cepstral Coefficient Analysis Unit effectively distinguishes the categories of environmental sounds (such as the sound of broken glass and thunder) through the feature extraction method of bionic hearing. It can still achieve a high accuracy rate for abnormal sound recognition under the condition of background noise of 60 dB, and the false alarm rate is significantly reduced compared with the traditional FFT spectrum analysis method, providing a reliable audio basis for safety event early warning.
[0058] Fusion decision tree generator, dynamically adjusts the weights of feature splitting nodes to construct a multi-modal joint inference model.
[0059] Execution steps of dynamically adjusting the weights of feature splitting nodes: Mutual information calculation: Calculate the mutual information value between each candidate feature (such as the TCN feature of the video, the MFCC coefficient of the sound) and the target variable.
[0060] Weight assignment: Dynamically set the splitting weights according to the proportion of mutual information. For example, in fire detection, the weight of the temperature gradient feature is W = 0.6, while the weight of the smoke color feature is W = 0.4.
[0061] Adaptive splitting: When splitting nodes, preferentially select features with high weights to ensure that the decision path points to the most significant feature combination.
[0062] Execution steps of constructing the multi-modal joint inference model: Feature pool construction: Store the video features (such as "movement speed = 1.2 m / s") output by the spatio-temporal convolutional neural network module, the audio features (such as "sudden increase in energy at 8 kHz") output by the Mel Frequency Cepstral Coefficient Analysis Unit, and the sensor features (such as "temperature rise rate 0.5 °C / s") into a unified feature pool.
[0063] Decision tree training: Root node selection, select the feature with the largest mutual information (such as the temperature rise rate) as the basis for the first-layer split. Branch generation, set judgment conditions according to the feature type (such as temperature rise rate > 0.3℃ / s → left branch, otherwise right branch). Leaf node, associate with specific decision results (such as "fire probability = 85%").
[0064] Online update: Recalculate the feature weights with new data every 24 hours to maintain the adaptability of the model.
[0065] The fusion decision tree generator constructs an interpretable and fast-response multi-modal inference model by quantifying the statistical correlation between features. It can quickly identify compound events (such as "theft + glass breakage"), with an accuracy higher than that of traditional SVM methods, and the model update does not require full retraining, which can effectively reduce computing resources.
[0066] The cloud distributed computing platform realizes spatio-temporal modeling of video behavior through the spatio-temporal convolutional neural network module, completes the bio-inspired analysis of sound features by the mel-frequency cepstral coefficient analysis unit, and the fusion decision tree generator establishes a dynamic inference mechanism for multi-modal features. The three work together to form a complete intelligent chain of "perception - analysis - decision". It can effectively control the event detection delay in the scenario of multiple cameras running concurrently, improve the system robustness in complex environments (such as rainy and foggy weather), and reduce the operation and maintenance costs, meeting the triple requirements of real-time, accuracy, and economy for smart city-level or industrial Internet of Things-level monitoring systems.
[0067] In this embodiment, the fusion decision tree generator executes:[[]]END]] Calculate the mutual information of video features and sensor features, and dynamically select the split node weights. The video features include spatio-temporal motion features extracted by the spatio-temporal convolutional neural network; Sort the features according to the mutual information, select the feature with the highest correlation as the root node, and generate a multi-modal joint inference model; During the construction of the decision tree, combine the dynamic changes of video features and sensor data, and activate the joint decision-making path of multi-modal features. The joint decision-making path is comprehensively determined based on the correlation between the abnormal features of environmental parameters and the video detection targets.
[0068] Execution process of calculating the mutual information of video features and sensor features: Video data extracts spatio-temporal motion features (such as target motion speed, direction change, etc.) through the spatio-temporal convolutional neural network (TCN). Sensor data (such as temperature, vibration, etc.) generates a standardized feature vector (such as temperature change rate, vibration intensity, etc.) after preprocessing. Through the high-precision time stamp of the satellite time service module, ensure that the video frames and sensor data are strictly aligned on the time axis.
[0069] Probability Distribution Modeling: Statistically model the joint distribution of video features and sensor features in historical data (such as the probability of the simultaneous occurrence of "high temperature gradient" and "flame color in the video"). Calculate the probabilities of the independent distributions of each feature and analyze the statistical correlation between the two.
[0070] Mutual Information Calculation: By comparing the differences between the joint distribution and their respective independent distributions, quantify the association strength between video features and sensor features. For example, if a certain sensor feature (such as a sudden increase in temperature) and a video feature (such as smoke diffusion) frequently co-occur, the mutual information is relatively high.
[0071] The mutual information calculation can quantify the association strength between video and sensor features, avoid interference from irrelevant features in decision-making, effectively enhance the sensitivity of the model to key features, and reduce the impact of noise on the analysis results.
[0072] By quantifying the association strength between video and sensor data, screening out feature combinations with high correlation, excluding the interference of irrelevant or noisy data, significantly improving the accuracy of multi-modal analysis, while reducing the risk of misjudgment caused by redundant features, enabling the system to focus on key information and enhancing the ability to analyze complex scenarios.
[0073] Process of Dynamically Selecting and Generating a Multi-modal Joint Inference Model for Split Nodes: Feature Ranking and Root Node Selection: Rank all candidate features (videos, sounds, temperatures, etc.) from high to low according to the mutual information, and select the feature with the highest mutual information as the root node of the decision tree (for example, preferentially select "temperature gradient" rather than "ambient humidity" as the initial basis for fire detection).
[0074] Dynamic Weight Allocation: Dynamically adjust the node weights according to the change of the mutual information of features in the real-time data stream. For example, if the correlation between "vibration intensity" and "human movement in the video" has increased recently, increase the weight of this feature. In the branches of the decision tree, features with high weights participate in split judgments first, forming key decision-making paths.
[0075] Model Generation and Update: Construct a multi-level decision tree, with each node corresponding to a feature judgment condition (such as "temperature change rate > threshold"). Recalculate the feature mutual information with new data regularly (such as every 24 hours), and update the node weights and split order.
[0076] Dynamically selecting split nodes can optimize the decision-making path according to real-time data, avoid model rigidity caused by fixed rules, significantly enhance the adaptability of the system to complex scenarios, and at the same time reduce redundant computing overhead.
[0077] Process of Activating the Joint Decision-making Path of Multi-modal Features: Real-time data dynamic monitoring: Continuously receive video and sensor data transmitted from edge nodes, extract features in real time and calculate the mutual information. Monitor the abnormal features of environmental parameters (such as sudden temperature changes, abnormal vibration frequencies).
[0078] Joint decision path triggering: When the sensor detects abnormal parameters (such as a sudden rise in temperature), automatically activate the relevant video analysis module (such as smoke diffusion detection). Through decision tree traversal, comprehensively consider the video target features (such as smoke texture) and sensor data (such as temperature gradient) to determine whether the joint alarm condition is met.
[0079] Comprehensive judgment and output: If the multi-modal features meet the preset relevance (such as high temperature accompanied by smoke), trigger an event alarm; if only a single feature is abnormal (such as a temperature rise but no video anomaly), it is determined as a low-risk event or noise interference.
[0080] Through the cross-validation mechanism of video and sensor data (such as the correlation analysis between temperature anomalies and video target behaviors), effectively eliminate false alarms or missed alarms of a single sensor, significantly improve the reliability of event detection, and at the same time reduce the impact of environmental interference on the decision-making results, enabling the system to maintain high-precision alarms in complex scenarios (such as crowded areas, bad weather).
[0081] The fusion decision tree generator realizes the in-depth fusion analysis of video and sensor data through the mutual information-driven dynamic feature selection, adaptive weight adjustment and multi-modal joint decision-making mechanism, breaks through the performance bottleneck of traditional single-modal or static models, significantly improves the recognition accuracy and response speed of complex events (such as intrusion, equipment failure), and at the same time reduces system resource consumption and false alarm rate, providing efficient and reliable intelligent decision-making support for scenarios such as intelligent security and industrial monitoring.
[0082] In this embodiment, it further includes: A satellite timing module that adds sub-millisecond precision timestamps to video frames and sensor data.
[0083] Specifically, by receiving the atomic clock signals of multi-satellite systems (GPS / Beidou / Galileo), a nanosecond-level time synchronization network is established between edge nodes and the cloud using the Precision Time Protocol (PTP). When each video frame and sensor data packet is collected, a satellite-calibrated timestamp is injected by the local clock module, and the transmission delay is fine-tuned and compensated through the Network Time Protocol (NTP), ultimately achieving the control of the overall system time deviation within sub-milliseconds.
[0084] The satellite timing module eliminates the problem of time asynchronization caused by local clock drift in traditional systems, ensures that video data and sensor data are strictly aligned at the millisecond level of precision, significantly improves the accuracy of multi-modal data fusion, and at the same time reduces the risk of misjudgment caused by time deviation.
[0085] 3D laser scanning and modeling unit, which establishes a spatial coordinate system for the monitoring area and maps it to the sensor position.
[0086] Specifically, a rotary lidar is used to perform panoramic scanning on the monitoring area, and a centimeter-level accurate 3D digital twin model is generated through point cloud stitching algorithm. The installation positions of each monitoring device and sensor are registered with spatial coordinates through calibration target balls to establish a unified global coordinate system. When the system runs, the pixel coordinates in the video are mapped to the 3D model in real time through camera calibration parameters, and the sensor data is associated with specific positions according to the preset spatial relationship matrix.
[0087] The 3D modeling technology realizes the accurate mapping between the physical space and the digital space, solves the spatial relationship problems such as height and occlusion that cannot be handled by traditional 2D monitoring systems, significantly improves the accuracy of target positioning and environmental perception, and at the same time reduces the impact of monitoring blind spots caused by perspective limitations. Embodiment 2
[0088] As Figure 5 shown, this embodiment proposes a video monitoring method running on any of the above video monitoring systems, including: Parallel execution through the edge computing node cluster: a. Perform ROI region detection and encoding compression based on hardware acceleration on the 8K video stream; b. Implement unit unification and outlier filtering on multi-source sensor data.
[0089] Step a specifically includes: Use the NVENC encoder to perform downsampling (compressed to 4K) on the non-ROI region; Keep the 8K original resolution for the ROI region where the dynamic target is located and add behavior feature markers.
[0090] By parallelly executing video ROI detection and sensor data preprocessing on the edge side, the amount of raw data is significantly reduced, the network transmission load is reduced, and at the same time the key information is retained, enabling the system to still achieve efficient real-time processing in resource-constrained environments, improving the overall response speed and reducing the cloud computing pressure.
[0091] Transmit the processed data to the cloud through the transmission channel constructed by 5G network slicing technology.
[0092] Use 5G network slicing technology to allocate exclusive transmission channels for monitoring data, ensure the stable and low-latency transmission of high-priority data (such as dynamic target video streams), effectively solve the data packet loss problem caused by traditional network congestion, and ensure the real-time and reliability of key monitoring information.
[0093] Perform spatio-temporal alignment and multi-modal feature fusion analysis in the cloud to generate joint decision results.
[0094] Integrate video spatiotemporal features and sensor features in the cloud for joint analysis, breaking through the data limitations of a single modality. Through cross-modal information complementarity, the recognition accuracy of complex events (such as intrusions and fires) is significantly improved, while reducing the risk of misjudgment caused by environmental interference.
[0095] Based on the fusion of multimodal features, structured event descriptions are output to provide security personnel with actionable and accurate alarm information, avoiding the information overload problem caused by simple alarms in traditional systems and greatly improving the decision-making effectiveness of the monitoring system.
[0096] In this embodiment, the multimodal feature fusion analysis includes: The spatiotemporal convolutional features of the target in the video stream are extracted, and the spatiotemporal convolutional features are generated by a spatiotemporal convolutional neural network module.
[0097] Extracting spatiotemporal convolution features of video streams: After the video stream is input into the spatiotemporal convolutional neural network (TCN) module, the spatial features of a single frame (such as target shape and texture) and the temporal features across frames (such as motion trajectory and speed change) are extracted simultaneously through multiple layers of 3D convolution kernels. The convolution operation captures the static properties of the target in the spatial dimension, analyzes the dynamic evolution between consecutive frames in the temporal dimension, and finally outputs a feature vector that integrates spatiotemporal information. For example, for the walking behavior of a person in the monitoring screen, TCN can simultaneously identify the human body contour (spatial features) and the direction of movement (temporal features).
[0098] Extracting spatiotemporal convolutional features of video streams Through the spatiotemporal convolutional neural network (TCN), the spatial static features (such as target shape and texture) and temporal dynamic features (such as motion trajectory and speed change) of video data are simultaneously captured, breaking through the limitation of traditional 2D convolution that only analyzes a single frame of image, thereby significantly improving the recognition accuracy of complex behaviors (such as wandering and climbing), while enhancing the system's ability to parse continuous target actions, providing highly reliable video feature input for subsequent multimodal decision-making.
[0099] Analyze the spectral characteristics and dynamic change trends of sensor data, where the sound sensor data is processed by the Mel-frequency cepstral coefficient analysis unit.
[0100] Analysis of sensor spectrum characteristics and dynamic trend execution process: Sound sensor processing: After the audio data is pre-emphasized and filtered to enhance the high-frequency components, it is framed and windowed (e.g., 20ms / frame) and Fourier transformed. The MFCC coefficients are extracted through the Mel filter bank to generate a spectral feature vector that represents the sound characteristics.
[0101] Other sensor processing: Data such as temperature and vibration are used to calculate the change trend (such as temperature gradient, vibration frequency fluctuation) through the sliding window statistical method and are transformed into standardized features. For example, after the vibration sensor data undergoes fast Fourier transform (FFT), the energy distribution of the main frequency bands is extracted as a dynamic feature.
[0102] By analyzing the spectral features and dynamic trends of the sensor, Mel-frequency cepstral coefficients (MFCC) analysis is performed on the sound sensor data to extract spectral features that mimic the human ear's auditory characteristics. Combining the dynamic trends (such as change rate, fluctuation pattern) of sensor data such as temperature and vibration, a multi-dimensional environmental perception model is constructed, which can effectively distinguish real events (such as the sound of broken glass) from environmental noise (such as the sound of wind), significantly improving the sensitivity and specificity of abnormal signal detection, and providing sensor feature inputs with high consistency for multi-modal fusion.
[0103] Construct a fusion decision tree, dynamically adjust the splitting node weights based on the mutual information between video features and sensor features, and select features with mutual information exceeding a preset threshold as key splitting nodes.
[0104] The process of constructing a fusion decision tree and dynamically adjusting node weights: Mutual information calculation: Statistically analyze the joint distribution of video spatio-temporal features (such as target movement speed) and sensor features (such as temperature change rate) in historical data, and calculate the correlation strength between the two.
[0105] Node weight assignment: Sort the features according to the mutual information, and preferentially select features with high correlation (such as temperature mutation and video smoke diffusion) as the root node of the decision tree, and dynamically adjust its weight. For example, automatically increase the decision priority of temperature-related nodes during peak fire seasons.
[0106] Decision tree generation: Construct multi-level judgment conditions from high to low according to the weights, each branch corresponding to a feature threshold (such as "temperature change rate > set value"), and the leaf nodes are associated with specific event labels (such as "fire", "intrusion").
[0107] Dynamically allocate node weights based on the mutual information between video and sensor features, and preferentially select features with high correlation (such as temperature mutation and video smoke diffusion) to construct decision paths, breaking through the rigid defects of traditional fixed-rule models, enabling the system to adapt to changes in data distribution in different scenarios (such as day-night temperature difference, equipment sensitivity fluctuation), improving the determination accuracy of complex events (such as fire, intrusion) while ensuring the inference speed, and reducing resource waste caused by redundant calculations.
[0108] Multi-modal feature fusion analysis realizes the deep fusion of multi-source data from underlying features to high-level decisions through the full-link collaboration of video spatio-temporal feature extraction, sensor dynamic parsing, and decision tree dynamic construction, breaking through the technical bottleneck of single-modal or simple splicing in traditional monitoring systems, significantly improving the event detection accuracy and response speed in complex scenarios (such as nighttime intrusion, industrial equipment failure), while reducing the false alarm and missed alarm rates caused by environmental interference, and providing efficient, reliable, and scalable intelligent analysis capabilities for diversified scenarios such as intelligent security and industrial monitoring.
[0109] In this embodiment, the construction process of the fusion decision tree includes: Calculate the mutual information of the video dynamic features and the sensor data, and the video dynamic features are extracted by a spatio-temporal convolutional neural network.
[0110] The execution process of calculating the mutual information of the video dynamic features and the sensor data: Feature alignment: Add a unified timestamp to the video frames and sensor data through a satellite time synchronization module to ensure strict temporal synchronization between the two.
[0111] Joint distribution statistics: Statistically analyze the co-occurrence frequency of video features (such as target movement speed) and sensor features (such as temperature change rate) in historical data to construct a joint probability distribution model.
[0112] Mutual information calculation: Quantify the correlation strength between the video and sensor data by comparing the differences between the joint distribution and the independent distributions of each feature.
[0113] Filter out high-correlation combinations of video and sensor features through mutual information, eliminate irrelevant noise interference, improve the accuracy of multi-modal analysis, and at the same time reduce the misjudgment risk caused by redundant data, enabling the system to focus on key information and enhance decision reliability.
[0114] Dynamically adjust the splitting nodes of the decision tree according to the mutual information threshold, and select the features with mutual information exceeding the preset threshold as the key splitting nodes.
[0115] The execution process of dynamically adjusting the splitting nodes and selecting key features: Feature ranking: Rank the video and sensor features from high to low according to the mutual information, and preferentially select the features with the highest correlation (such as sudden temperature change and video smoke diffusion).
[0116] Threshold filtering: Set the mutual information threshold, and only retain the features exceeding the threshold as candidate splitting nodes (such as features with mutual information > 0.7).
[0117] Dynamic weight allocation: Dynamically adjust the node weights according to the statistical distribution changes of the features in the real-time data stream (such as increasing the weight of temperature-related nodes during high fire incidence periods).
[0118] The dynamic adjustment mechanism enables the decision tree to adapt to the data distribution in different scenarios (such as the diurnal temperature difference and the fluctuation of device sensitivity), preferentially select the most significant features for decision-making, improve the recognition efficiency of complex events (such as intrusion and device failure), and reduce the consumption of invalid computing resources at the same time.
[0119] Construct a multi-modal joint reasoning model. When the sensor data detects abnormal features of environmental parameters, cross-verify with the video analysis results. If the multi-modal features meet the joint determination conditions, trigger the corresponding event alarm.
[0120] The execution process of constructing the multi-modal joint reasoning model: Decision tree generation: Use the feature with the highest mutual information as the root node, and construct multi-level judgment conditions layer by layer (such as "temperature change rate > threshold"), and the leaf nodes are associated with specific event labels (such as "fire" and "intrusion").
[0121] Model update: Recalculate the mutual information with new data regularly (such as daily), update the node weights and splitting order to maintain the adaptability of the model.
[0122] Through the construction of a decision tree driven by mutual information, a multi-modal reasoning model with strong interpretability and fast response speed is formed, breaking through the limitations of traditional fixed rules, significantly improving the robustness of the system to dynamic environments (such as weather changes and equipment aging), and reducing the model update and maintenance costs at the same time.
[0123] The execution process of cross-modal cross-validation and event alarm triggering: Real-time verification: When the sensor detects an anomaly (such as a sudden increase in vibration intensity), the system automatically associates the video data with the same timestamp and checks for the existence of a matching target (such as a climbing behavior).
[0124] Joint determination: If the video features (such as human movement trajectories) and sensor features (such as abnormal vibrations and metallic sounds) meet the preset association conditions (such as position overlap and time synchronization) in the spatio-temporal dimension, it is determined as a valid event.
[0125] Alarm output: Generate structured alarm information (such as "Intrusion on the east side wall - Confidence level 95%"), and push it to the monitoring platform through the API, and synchronously mark the relevant video segments for manual review.
[0126] Through the spatio-temporal cross-validation of multi-modal data, false alarms of a single sensor (such as temperature false alarms caused by device failures) are effectively excluded, the accuracy of alarm events is significantly improved, and at the same time, actionable decision-making information (such as location, type, and confidence level) is output, greatly reducing the workload of manual review.
[0127] The fusion decision tree realizes the efficient deep fusion analysis of video and sensor data through the dynamic feature selection driven by mutual information, the adaptive weight adjustment, and the cross-modal cross-validation mechanism, breaks through the performance bottleneck of traditional single-modal or static models, significantly improves the detection accuracy and response speed of complex events (such as intrusion, fire, equipment failure), and at the same time reduces the false alarm rate and system resource consumption, providing high-reliability and scalable intelligent decision support capabilities for scenarios such as intelligent security and industrial monitoring.
[0128] In this embodiment, it further includes: Align the video frames and sensor data according to the satellite timestamp, and start data resynchronization when the time deviation exceeds 1 ms.
[0129] The process of aligning the video frames and sensor data according to the satellite timestamp: The system injects high-precision timestamps into the video frames and sensor data through the satellite timing module (GPS / Beidou), and the edge node and the cloud platform perform time synchronization using the Precision Time Protocol (PTP). When the timestamp deviation between the video frames and sensor data is detected to exceed the set threshold (such as 1 ms), the system automatically triggers the data resynchronization mechanism, and the data resynchronization mechanism includes: Timestamp correction: Based on the satellite timing signal, recalibrate the local clock to align the data time axis.
[0130] Data buffer alignment: Store the lagging data in the temporary buffer, and wait for the leading data to arrive and then reorder them according to the timestamp.
[0131] Through the high-precision time synchronization mechanism, ensure that the video and sensor data are strictly aligned at the millisecond level of accuracy, eliminate the misalignment problem of multi-modal data caused by timing deviation, significantly improve the accuracy and reliability of event analysis, and avoid misjudgment or missed judgment caused by time asynchronization.
[0132] Based on the three-dimensional space coordinate mapping, exclude the abnormal sensor data whose distance from the target object exceeds 2 meters.
[0133] The process of excluding irrelevant sensor data based on the three-dimensional space coordinate mapping: Three-dimensional space modeling: Deploy lidar to perform panoramic scanning of the monitoring area, generate a three-dimensional point cloud model with centimeter-level accuracy, and mark the spatial coordinates of each camera and sensor.
[0134] Spatial relationship mapping: Through the coordinate transformation matrix, map the pixel coordinates of the target in the video to the physical position in the three-dimensional model (such as the XYZ coordinates of someone in the coordinate system).
[0135] Distance threshold filtering: When the distance between the sensor data (such as abnormal temperature) and the target in the video exceeds a preset threshold (such as 2 meters), it is determined as irrelevant data and automatically excluded. For example, if a person is detected at point A in the video, and the location of the abnormal data of temperature sensor B is more than 2 meters away from point A, then this abnormal data of sensor B is ignored.
[0136] Through three-dimensional space coordinate mapping and distance threshold filtering, the video target and sensor data are accurately associated, eliminating the problem of false association caused by sensor installation position deviation or environmental interference, significantly improving the accuracy of multi-modal data fusion, reducing the false alarm rate, and enhancing the decision-making reliability of the system in complex space environments (such as multi-story buildings, occluded areas). Embodiment 3
[0137] This embodiment proposes a spherical camera, including any of the above video surveillance systems, and the video surveillance system runs any of the above video surveillance methods.
[0138] The spherical camera proposed in this embodiment should have the basic functions and structures of existing spherical cameras. The difference from the prior art is that the spherical camera in this embodiment is equipped with the video surveillance system in this application and runs the video surveillance method in this application.
[0139] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A video surveillance system, characterized in that, Including: An edge computing node cluster deployed in the monitoring area, including GPU and DSP heterogeneous computing units; A super high-definition video acquisition module connected to the edge computing node cluster, with an 8K image sensor and a data cache unit built-in; A multi-type environmental sensor array, including temperature, sound, and vibration sensors, communicating with the edge computing node cluster through a high-speed interface; A transmission channel built based on 5G network slicing technology, setting a data transmission priority policy; A cloud distributed computing platform, integrating a real-time stream processing engine and a multi-modal fusion decision-making module.
2. The video surveillance system according to claim 1, wherein The edge computing node cluster performs the following: Using a hardware-accelerated video encoder to perform H.265 compression on the original video stream, and synchronously marking the dynamic target area ROI during the compression process; Normalizing the sensor data to generate a standardized data frame containing timestamps.
3. The video surveillance system according to claim 1, characterized in that, The cloud distributed computing platform includes: A spatio-temporal convolutional neural network module for extracting the spatio-temporal motion features of targets in the video stream; A Mel-frequency cepstral coefficient analysis unit for analyzing the spectral features of the sound sensor; A fusion decision tree generator for dynamically adjusting the weights of feature splitting nodes to build a multi-modal joint inference model.
4. The video surveillance system according to claim 3, wherein, The fusion decision tree generator performs the following: Calculating the mutual information of video features and sensor features, and dynamically selecting the splitting node weights. The video features include spatio-temporal motion features extracted by the spatio-temporal convolutional neural network; Sorting the features according to the mutual information, selecting the feature with the highest correlation as the root node, and generating a multi-modal joint inference model; During the construction of the decision tree, combining the dynamic changes of video features and sensor data, activating the joint decision-making path of multi-modal features. The joint decision-making path is comprehensively determined based on the correlation between the abnormal features of environmental parameters and the video detection targets.
5. The video monitoring system according to claim 4, characterized in that Further including: A satellite time synchronization module for adding sub-millisecond accuracy timestamps to video frames and sensor data; A three-dimensional laser scanning modeling unit for establishing a spatial coordinate system of the monitoring area and mapping it to the sensor positions.
6. A video surveillance method running on any of the video surveillance systems of claims 1-5, characterized in that, Including: Parallelly executed by the edge computing node cluster: a. Performing ROI area detection and encoding compression on the 8K video stream based on hardware acceleration; b. Performing unit unification and outlier filtering on multi-source sensor data; Transmitting the processed data to the cloud through the transmission channel built by the 5G network slicing technology; Performing spatio-temporal alignment of multi-modal feature fusion analysis in the cloud to generate a joint decision result.
7. The video monitoring method according to claim 6, characterized in that The multi-modal feature fusion analysis includes: Extracting the spatio-temporal convolutional features of targets in the video stream; Calculating the MFCC coefficients and high-frequency energy mutation points of the sound sensor data; Constructing a fusion decision tree, dynamically adjusting the splitting node weights based on the mutual information of video features and sensor features, and selecting the features with mutual information exceeding the preset threshold as the key splitting nodes.
8. The video monitoring method according to claim 7, characterized in that, The construction process of the fusion decision tree includes: Calculating the mutual information of video dynamic features and sensor data. The video dynamic features are extracted by the spatio-temporal convolutional neural network; Dynamically adjusting the splitting nodes of the decision tree according to the mutual information threshold, and selecting the features with mutual information exceeding the preset threshold as the key splitting nodes; Build a multi-modal joint reasoning model. When abnormal features of environmental parameters are detected in sensor data, cross-verify with the video analysis results. If the multi-modal features meet the joint determination conditions, trigger corresponding event alarms.
9. The video surveillance method according to claim 6, characterized in that, Further include: Align video frames with sensor data according to satellite timestamps, and start data re-synchronization when the time deviation exceeds 1 ms; Based on three-dimensional space coordinate mapping, exclude abnormal sensor data with a distance from the target object exceeding 2 meters.
10. A spherical camera, characterized in that, Include the video surveillance system according to any one of claims 1-5, and the video surveillance system runs the video surveillance method according to any one of claims 6-9.
Citation Information
Patent Citations
Management and control system and method for hot-line work robot
CN112549020A
Informationized smart park management system
CN117221361A
Intelligent monitoring system and method
CN118433362A
Three-dimensional motion analysis system based on dynamic multi-sensor cross-modal fusion
CN118470785A
Self-adaptive video coding method based on unmanned aerial vehicle
CN118573867A
Cited By
Behavior recognition method, medium and electronic equipment
CN121305663A
Bridge cantilever assembly construction monitoring method, device and equipment and storage medium
CN121346901A
Video monitoring transmission time delay detection method and system based on image recognition
CN121442087A