Real-time video monitoring intelligent analysis system based on deep learning
By using multimodal data fusion and hierarchical processing architecture based on deep learning technology, the problems of low efficiency and high misjudgment rate in traditional video surveillance systems are solved, enabling real-time and accurate video analysis in complex scenarios and improving the intelligence level of the surveillance system.
Patent Information
- Application Number
- CN202511612307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-01-02
AI Technical Summary
Traditional video surveillance systems rely on manual analysis, which is inefficient, slow to respond, and has a high error rate. Furthermore, analysis based on a single video stream is insufficient to meet the real-time monitoring needs of complex scenarios. Existing systems often rely on visible light video data, which is susceptible to environmental interference and limits the comprehensive perception of complex scenarios.
A real-time video surveillance intelligent analysis system based on deep learning is adopted. The system acquires real-time monitoring video streams through the video acquisition unit, and extracts dynamic object masks, scene semantic maps, thermal infrared features and dynamic environmental variable information using the scene parsing unit and the acquisition unit. Combined with the semantic segmentation unit, multimodal data fusion is performed to generate accurate semantic segmentation results.
It enables accurate analysis of surveillance videos, improves the intelligence level of video surveillance systems, meets the real-time, accuracy and reliability requirements in complex scenarios, and supports edge computing and cloud collaborative working modes.
Smart Images

Figure CN121259699A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a real-time video monitoring intelligent analysis system based on deep learning. BACKGROUND
[0002] With the improvement of social security demand and the rapid development of information technology, video monitoring systems have been widely used in security management, traffic management, smart park and other fields. Traditional video monitoring systems mainly rely on manual analysis of video streams, which has low efficiency, slow response, high misjudgment rate and other problems, and is difficult to meet the real-time monitoring demand in complex scenarios. In recent years, the introduction of deep learning technology has significantly improved the intelligent level of video monitoring, realizing automatic analysis and understanding of video content.
[0003] However, most of the schemes are based on visible light video data for analysis, which is easily affected by environmental interference, and the existing systems mostly rely on single video stream for analysis, limiting the comprehensive perception of complex scenes, resulting in poor video analysis effect. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a real-time video monitoring intelligent analysis system based on deep learning, which can accurately analyze the monitoring video. The specific scheme is as follows:
[0005] A real-time video monitoring intelligent analysis system based on deep learning, comprising: a video acquisition unit configured to acquire a real-time monitoring video stream to be analyzed and extract a target video frame sequence from the real-time monitoring video stream; a scene analysis unit configured to process the target video frame sequence using a pre-trained scene analysis model to obtain a dynamic object mask and a scene semantic graph corresponding to the real-time monitoring video stream; an acquisition unit configured to acquire thermal infrared features and dynamic environmental variable information of a monitoring area corresponding to the real-time monitoring video stream; a semantic segmentation unit configured to input the dynamic object mask, the scene semantic graph, the thermal infrared features and the dynamic environmental variable information into a pre-trained semantic segmentation model to obtain a semantic segmentation result output by the semantic segmentation model.
[0006] The system described above, optionally, the scene analysis unit comprises: The scene analysis subunit is configured to input the target video frame sequence into a scene analysis model, perform convolution processing on the target video frame sequence by a backbone network in the scene analysis model, and obtain convolution features; output a basic feature map based on the convolution features by a multi-scale feature extraction network in the scene analysis model; output a dynamic object mask and a dynamic context feature based on the convolution features and the basic feature map by a time difference convolution network in the scene analysis model; and obtain a scene semantic map based on the basic feature map, the dynamic object mask, and the dynamic context feature by a semantic enhancement network in the scene analysis model.
[0007] Optionally, the system further includes a first model training unit. The first model training unit is configured to obtain a first training sample set and an initial scene analysis model, the first training sample set includes a plurality of first training samples and a sample label of each first training sample, the first training sample includes a historical target monitoring video frame sequence, and the initial scene analysis model is trained by using the first training sample set.
[0008] Optionally, the semantic segmentation unit includes: The semantic segmentation subunit is configured to input the dynamic object mask, the scene semantic map, the thermal infrared feature, and the dynamic environmental variable information into a pre-trained semantic segmentation model, perform spatial registration and cross-modal association on the thermal infrared feature and the scene semantic map by a feature alignment module in the semantic segmentation model, and obtain a fusion feature; output a spatio-temporal enhanced feature based on the fusion feature by a spatio-temporal convolution network in the semantic segmentation model; and output a semantic segmentation result based on the spatio-temporal enhanced feature, the dynamic environmental variable information, and the dynamic object mask by a dynamic gating network in the semantic segmentation model.
[0009] Optionally, the system further includes a second model training unit. The second model training unit is configured to obtain a second training sample set and an initial semantic segmentation model, the second training sample set includes a plurality of second training samples and a sample label of each second training sample, the second training sample includes a historical dynamic object mask, a historical scene semantic map, a historical thermal infrared feature, and historical dynamic environmental variable information, and the initial semantic segmentation model is trained by using the second training sample set.
[0010] Optionally, the obtaining unit includes: The receiving subunit is configured to receive a thermal imaging data stream of a monitoring area corresponding to the real-time monitoring video stream. A feature extraction subunit is configured to perform feature extraction on the thermal imaging data stream by using a convolution network to obtain thermal infrared features.
[0011] The system described above, optionally, the acquisition unit comprises: An acquisition subunit is configured to acquire illumination data of a monitoring area corresponding to the real-time monitoring video stream, and calculate a spatial gradient matrix based on the illumination data by using a Sobel operator. An identification subunit is configured to identify motion state data of a target tracking object in the target video frame sequence, and calculate an interaction energy matrix based on the motion state data and the dynamic object mask. A first execution subunit is configured to receive an audio stream collected by a microphone array of the monitoring area, and obtain a sound field fluctuation coefficient based on the audio stream. A second execution subunit is configured to form the dynamic environment variable information based on the spatial gradient matrix, the interaction energy matrix, and the sound field fluctuation coefficient.
[0012] The system described above, optionally, further comprises a transmission unit. The transmission unit is configured to encrypt the semantic segmentation result, and send the encrypted semantic segmentation result to a monitoring center server.
[0013] The system described above, optionally, further comprises an abnormality alarm unit. The abnormality alarm unit is configured to output alarm information when an abnormal confidence map in the semantic segmentation result exceeds a preset abnormal threshold.
[0014] The system described above, optionally, further comprises a multi-modal event generation unit. The multi-modal event generation unit is configured to match the semantic segmentation result with a preset knowledge graph to generate a structured event description.
[0015] Based on the above, the application provides a real-time video monitoring intelligent analysis system based on deep learning, which comprises a video acquisition unit, a scene analysis unit, an acquisition unit and a semantic segmentation unit. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 The application provides a real-time video monitoring intelligent analysis system based on deep learning. Figure 2 The application provides an acquisition unit. Figure 3 The application provides another acquisition unit. Figure 4 The application provides a real-time video monitoring intelligent analysis method based on deep learning. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] In this application, the terms "comprising" or "including" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0020] Referring to Figure 1 A structure schematic diagram of a real-time video monitoring intelligent analysis system based on deep learning provided for an embodiment of the application, the system comprises: A video acquisition unit 101 is configured to acquire a real-time monitoring video stream to be analyzed, and extract a target video frame sequence from the real-time monitoring video stream.
[0021] In this embodiment, the real-time monitoring video stream can come from a camera device deployed in a monitoring area, which can be a high-definition network camera, an infrared camera, or a multi-modal sensor array, etc. The target video frame sequence is a video segment designated by a user or a system for intelligent analysis. In this application, the target video frame sequence can be a continuous frame sequence extracted in chronological order, or a frame set with significant changes extracted based on a key frame detection algorithm. The extraction of the target video frame sequence can be based on a fixed time interval or dynamic adaptive sampling, such as adjusting the sampling rate based on motion detection, without specific limitation.
[0022] A scene analysis unit 102 is configured to process the target video frame sequence using a pre-trained scene analysis model to obtain a dynamic object mask and a scene semantic graph corresponding to the real-time monitoring video stream.
[0023] The scene analysis model can be a deep learning-based computer vision model, such as a convolutional neural network, a recurrent neural network, or a graph convolution network, etc. The dynamic object mask is a binary or probabilistic matrix used to identify the location and contour of the moving target in the video frame. The scene semantic graph is a multi-channel feature map used to represent the semantic class distribution of the video frame. In this application, the scene analysis model can be trained by supervised learning, and the training data includes a video dataset labeled with object boundaries and semantic labels. Optionally, the semantic class distribution can be a scene element such as road, building, vegetation, etc.
[0024] An acquisition unit 103 is configured to acquire thermal infrared features and dynamic environmental variable information of the monitoring area corresponding to the real-time monitoring video stream.
[0025] The thermal infrared feature can be represented by feature extraction network processing temperature distribution data collected by an infrared sensor or a thermal imaging device, and is used to reflect the thermodynamic characteristics of the monitoring area. The dynamic environment variable information includes but is not limited to illumination intensity variation, audio spectrum feature, motion vector field and other environmental parameters. These parameters can be collected in real time by corresponding environmental sensors and feature extraction is performed through signal processing algorithms. In this application, the thermal infrared feature and the dynamic environment variable information can be fused with the video data in multiple modalities to enhance the adaptability of the system under different environmental conditions.
[0026] The semantic segmentation unit 104 is configured to input the dynamic object mask, the scene semantic map, the thermal infrared feature and the dynamic environment variable information into a pre-trained semantic segmentation model to obtain a semantic segmentation result output by the semantic segmentation model.
[0027] The semantic segmentation model can be a deep learning network based on multi-modal feature fusion, such as an encoder-decoder architecture with an attention mechanism or a spatio-temporal Transformer model. The semantic segmentation result is a pixel-level classification output, which is used to accurately identify the semantic category to which each pixel in the video frame belongs, and can also include confidence information for abnormal area detection. In this application, the semantic segmentation model can be trained in an end-to-end manner, and the optimization objectives include segmentation accuracy and inference efficiency. The output of the semantic segmentation unit 104 can be used for real-time monitoring and early warning, situation analysis and decision support, etc.
[0028] In this embodiment, the semantic segmentation result can include a semantic label map, an abnormal confidence map and a timestamp, etc. The semantic label map includes a two-dimensional matrix (H x W) corresponding to the spatial resolution of the input video frame. Each pixel in the matrix has a corresponding classification label, which is used to indicate the semantic category, such as road, pedestrian, vehicle, building, vegetation, etc. The abnormal confidence map includes a two-dimensional matrix with the same resolution as the semantic label map, and the value of each pixel is a continuous confidence score (e.g. between 0 and 1), which is used to represent the probability of an abnormal event occurring at that location, such as the probability of reverse driving, crowd gathering, fire, etc.
[0029] The system provided in this application can accurately perform semantic-level understanding and analysis of the monitoring video through multi-modal data fusion and deep learning technology, and improve the intelligent level of the video monitoring system. At the same time, the system adopts a hierarchical processing architecture, supports edge computing and cloud collaborative working mode, and meets the application requirements of real-time, accuracy and reliability. Through effective fusion of multi-source information and collaborative processing of deep learning models, the system can achieve accurate video semantic analysis in complex scenarios.
[0030] In an embodiment provided by the present application, based on the above scheme, the scene analysis unit 102 comprises, optionally: The scene analysis subunit is configured to input the target video frame sequence into a scene analysis model, perform convolution processing on the target video frame sequence by a backbone network in the scene analysis model, obtain convolution features, output a basic feature map based on the convolution features by a multi-scale feature extraction network in the scene analysis model, output a dynamic object mask and a dynamic context feature based on the convolution features and the basic feature map by a time difference convolution network in the scene analysis model, and obtain a scene semantic map based on the basic feature map, the dynamic object mask and the dynamic context feature by a semantic enhancement network in the scene analysis model.
[0031] The backbone network can be a deep convolutional neural network for feature extraction, such as ResNet, VGG or MobileNet, etc. The backbone network extracts convolution features with rich semantic information by performing multi-layer convolution operations on the input target video frame sequence. The convolution features can include low-level edge texture features and high-level semantic features, providing a basic feature representation for subsequent processing.
[0032] Optionally, the multi-scale feature extraction network can be a feature pyramid network containing a multi-branch structure, or a spatial pyramid pooling module using a dilated convolution. The multi-scale feature extraction network obtains a basic feature map containing multi-scale information by processing the convolution features at different scales or different receptive fields. The basic feature map can contain both detailed information and global context information, which helps to improve the processing effect of different size targets.
[0033] In the present embodiment, the time difference convolution network uses three-dimensional convolution and optical flow calculation and other time series modeling techniques to accurately identify moving targets in the scene by analyzing the feature change law between consecutive video frames. The time difference convolution network can be a time series modeling network based on three-dimensional convolution or optical flow calculation. The time difference convolution network identifies moving targets in the scene by analyzing the feature difference between consecutive video frames and generates a dynamic object mask. Specifically, the dynamic object mask is a binary mask or a probability map, which is used to identify the moving area in the video frame. The dynamic context feature is a feature representation containing motion patterns and time series relationships, which is used to describe the behavior characteristics of the moving target.
[0034] In the present embodiment, the semantic enhancement network can be a network structure based on attention mechanism or feature fusion. The semantic enhancement network enhances the understanding ability of the scene semantics by fusing the basic feature map, the dynamic object mask and the dynamic context feature. Optionally, the semantic enhancement network can use channel attention mechanism to highlight important features, or use spatial attention mechanism to focus on key areas, and finally output a scene semantic map with rich semantic information.
[0035] In the present application, through the cooperation of the above-mentioned multiple networks, the spatio-temporal features in the video sequence can be effectively extracted, the moving target can be accurately identified, and a high-quality semantic understanding result can be generated, thereby providing reliable input features for subsequent semantic segmentation tasks.
[0036] In an embodiment provided in the present application, based on the above-mentioned scheme, optionally, the method further comprises a first model training unit. The first model training unit is configured to obtain a first training sample set and an initial scene analysis model; the first training sample set comprises a plurality of first training samples and a sample label of each first training sample; the first training sample comprises a historical target monitoring video frame sequence; and the initial scene analysis model is trained by using the first training sample set.
[0037] The first training sample set is composed of a plurality of video data collected by a historical monitoring system, and the video data is processed by frame extraction to form continuous video frame sequences, and each video frame sequence corresponds to a first training sample.
[0038] Optionally, the sample label is generated by pixel-level labeling of a dynamic target region in the video frame by using a labeling tool, and the labeling content comprises target bounding box coordinates and semantic category information.
[0039] In the present embodiment, the initial scene analysis model adopts an encoder-decoder architecture, wherein the encoder part is initialized by using pre-trained ResNet-50 network weights on an ImageNet dataset, and the decoder part adopts a feature up-sampling module based on transposed convolution.
[0040] Optionally, in the training process, a stochastic gradient descent algorithm is used to iteratively optimize the model parameters, for example, the initial learning rate can be set to 0.001, and the learning rate is decayed to 0.5 times of the original value every 10 training periods. Data enhancement processing is implemented on the input video frame sequence during training, including random horizontal flipping, brightness adjustment, and scale transformation operations, to improve the generalization ability of the model. The loss function adopts a weighted combination of cross-entropy loss and Dice loss, for example, the weight coefficient of the cross-entropy loss is set to 0.7, and the weight coefficient of the Dice loss is set to 0.3.
[0041] In the present embodiment, the number of samples in each training batch is set to 16, and a total of 100 training periods of iterative training are performed. Optionally, during the training process, after completing one training period, the model performance is evaluated on the validation set, and when the average intersection over union index of the model on the validation set does not appear for 5 consecutive training periods, the training process is terminated in advance. The model parameters that finally show the best performance on the validation set are reserved as the trained scene analysis model, which is used for subsequent video analysis tasks.
[0042] In an embodiment provided by the present application, based on the above scheme, optionally, the semantic segmentation unit 102 comprises: The semantic segmentation subunit is configured to input the dynamic object mask, the scene semantic map, the thermal infrared feature, and the dynamic environment variable information into a pre-trained semantic segmentation model, so that the thermal infrared feature and the scene semantic map are spatially registered and cross-modality correlated by a feature alignment module in the semantic segmentation model to obtain a fusion feature; the fusion feature is used as input by a spatio-temporal convolution network in the semantic segmentation model to output a spatio-temporal enhanced feature; and the spatio-temporal enhanced feature, the dynamic environment variable information, and the dynamic object mask are used as input by a dynamic gating network in the semantic segmentation model to output a semantic segmentation result.
[0043] The dynamic object mask is a binary matrix generated by the scene analysis model and used to identify a moving target region in the video frame; the scene semantic map is a multi-channel feature map generated by the scene analysis model and used to represent a semantic class distribution at each position in the video frame; the thermal infrared feature is a temperature distribution feature obtained by collecting the temperature distribution by the infrared sensor and processing the temperature distribution by a feature extraction network; and the dynamic environment variable information is multi-dimensional environment parameters including an illumination gradient and a sound field fluctuation coefficient, which are collected by an environment sensor.
[0044] In the embodiment, the semantic segmentation model is a pre-trained multi-modal fusion neural network, and the semantic segmentation model comprises the feature alignment module, the spatio-temporal convolution network, and the dynamic gating network. The feature alignment module is configured to perform spatial registration processing on the thermal infrared feature and the scene semantic map, and the spatial registration processing is implemented by calculating affine transformation parameters between the two feature maps to achieve spatial alignment of the feature maps; and the feature alignment module is also configured to perform cross-modality correlation processing, and the cross-modality correlation processing is implemented by calculating a cross-correlation matrix between feature channels to establish a feature correlation between modalities, so as to finally obtain the fusion feature. The spatio-temporal convolution network is configured to perform convolution operation on the fusion feature in the spatial and temporal dimensions, and the convolution operation is implemented by using a three-dimensional convolution kernel to extract features in the spatial and temporal dimensions at the same time, so as to output the spatio-temporal enhanced feature.
[0045] Optionally, the dynamic gating network is configured to receive the spatio-temporal enhanced feature, the dynamic environment variable information, and the dynamic object mask, and the dynamic gating network calculates weight coefficients of the input features by using a gating mechanism, concatenates and transforms the weighted features, and finally outputs the semantic segmentation result. The semantic segmentation result is a pixel-level classification probability map, which is used to represent a probability distribution of each pixel belonging to each semantic class in the video frame.
[0046] In an embodiment provided by the present application, based on the above scheme, optionally, the second model training unit is further included. The second model training unit is configured to obtain a second training sample set and an initial semantic segmentation model, the second training sample set comprises a plurality of second training samples and a sample label of each second training sample, and each second training sample comprises a historical dynamic object mask, a historical scene semantic graph, a historical thermal infrared feature and historical dynamic environmental variable information; and the initial semantic segmentation model is trained by using the second training sample set.
[0047] In the embodiment, the historical dynamic object mask is a binary motion region identifier generated by processing a historical monitoring video by using a scene analysis model.
[0048] Optionally, the historical scene semantic graph is a feature representation containing a semantic category distribution extracted by using the scene analysis model, and the historical thermal infrared feature is a temperature distribution feature obtained by collecting, by using an infrared sensing device, and processing by using a feature extraction network.
[0049] Optionally, the historical dynamic environmental variable information is multi-dimensional environmental parameters containing an illumination gradient matrix and a sound field fluctuation coefficient collected by using an environmental sensing device. The initial semantic segmentation model adopts an encoder-decoder architecture, wherein the encoder part contains a feature alignment module and a spatio-temporal convolution network, and the decoder part contains a dynamic gating network and an up-sampling module.
[0050] In the embodiment, in the training process, a stochastic gradient descent algorithm is used to optimize the model parameters, the initial learning rate is set to 0.001, and the learning rate is decayed to 0.8 times of the original value every 20 training periods. During training, data enhancement processing is performed on the input second training sample, including random morphological transformation on the historical dynamic object mask, channel random discarding on the historical scene semantic graph, Gaussian noise addition on the historical thermal infrared feature, and the like. The loss function adopts a linear combination of a weighted cross-entropy loss and a boundary consistency loss, wherein the weight coefficient of the weighted cross-entropy loss is set to 0.6, and the weight coefficient of the boundary consistency loss is set to 0.4. The number of samples in each training batch is set to 8, and a total of 150 training periods of iterative training are performed.
[0051] Optionally, in the training process, after completing one training period, the model performance is evaluated on an independent validation set, and when the average intersection over union index of the model on the validation set does not appear significant improvement for 8 consecutive training periods, the early termination mechanism is started. The model parameters finally selected to obtain the best performance on the validation set are taken as the trained semantic segmentation model.
[0052] In an embodiment provided in the application, based on the above scheme, as shown in Figure 2 , the structure diagram of the obtaining unit 103 includes: The receiving sub-unit 1031 is configured to receive thermal imaging data stream of a monitoring area corresponding to a real-time monitoring video stream. The feature extraction subunit 1032 is configured to perform feature extraction on the thermal imaging data stream by using a convolutional network to obtain thermal infrared features.
[0053] The thermal imaging data stream is transmitted in the form of a sequence of continuous frames. Optionally, the spatial resolution of each frame of thermal imaging data is 640x480 pixels, and the sampling frequency is 30 frames / s.
[0054] In this embodiment, the feature extraction subunit can use a lightweight convolutional neural network architecture to perform feature extraction on the thermal imaging data stream. The convolutional neural network includes 5 convolutional layers and 2 pooling layers. Each convolutional layer uses a 3x3 convolutional kernel for feature mapping, and the pooling layer uses a 2x2 max-pooling operation for feature dimension reduction.
[0055] Optionally, during the feature extraction process, the input thermal imaging data stream is first normalized to convert the pixel values to the [0, 1] interval, and then the features are extracted layer by layer through the convolutional neural network to output 256-dimensional thermal infrared features. The thermal infrared features are represented in the form of a feature map, and the spatial size of the feature map is 40x30 pixels. Each pixel position corresponds to a 256-dimensional feature vector, which is used to describe the thermal distribution characteristics and the spatiotemporal variation pattern of the position.
[0056] In this embodiment, the convolutional neural network uses a ReLU activation function for nonlinear transformation, and a global average pooling operation is used in the last layer to convert the feature map to a fixed-dimensional feature representation.
[0057] In an embodiment provided in the present application, based on the above-mentioned scheme, the structure diagram of the acquisition unit 103 is shown in FIG. 10, which includes: Figure 3 The acquisition subunit 1033 is configured to acquire illumination data of a monitoring area corresponding to a real-time monitoring video stream, calculate the spatial gradient matrix based on the Sobel operator on the illumination data; The recognition subunit 1034 is configured to recognize the motion state data of the target tracking object in the target video frame sequence, and calculate the interaction energy matrix based on the motion state data and the dynamic object mask; The first execution subunit 1035 is configured to receive an audio stream collected by a microphone array of the monitoring area, and obtain the sound field fluctuation coefficient based on the audio stream; The second execution subunit 1036 is configured to form dynamic environment variable information based on the spatial gradient matrix, the interaction energy matrix, and the sound field fluctuation coefficient.
[0058] In this embodiment, the illumination data represents the illumination intensity values of each position in the monitoring area in the form of a two-dimensional matrix.
[0059] The Sobel operator includes a horizontal convolution kernel Gx and a vertical convolution kernel Gy, and the gradient components of the illumination data in the horizontal and vertical directions are calculated through convolution operation, and finally the gradient components in the two directions are synthesized to obtain the spatial gradient matrix.
[0060] In the embodiment, the motion state data includes position coordinates, motion speed and motion direction information of the target.
[0061] In the embodiment, the identification subunit can generate an interaction energy matrix according to the motion state data and the dynamic object mask by calculating the relative distance and the relative speed between the targets, using an interaction energy calculation formula, and the interaction energy matrix is used to represent the interaction strength between the targets.
[0062] Optionally, the first execution subunit receives audio stream data collected by the microphone array, performs frame processing and windowing operation on the audio stream data, then converts the time domain signal into frequency domain representation through fast Fourier transform, and finally calculates the fluctuation of the sound field energy in a specific frequency band range to obtain the sound field fluctuation coefficient.
[0063] In the embodiment, the second execution subunit 1036 can perform data fusion on the spatial gradient matrix, the interaction energy matrix and the sound field fluctuation coefficient to form dynamic environment variable information, and the dynamic environment variable information is represented in the form of a multi-dimensional feature vector and is used to describe the comprehensive state characteristics of the monitoring environment.
[0064] In an embodiment provided in the application, based on the above-mentioned scheme, optionally, the method further comprises a transmission unit. The transmission unit is configured to encrypt the semantic segmentation result and send the encrypted semantic segmentation result to the monitoring center server.
[0065] The transmission unit includes a data encryption module and a data transmission module.
[0066] Optionally, the data encryption module encrypts the semantic segmentation result by using an asymmetric encryption algorithm, the asymmetric encryption algorithm uses an RSA encryption scheme, the key length is 2048 bits, and in the encryption process, the semantic segmentation result is first subjected to data serialization processing, the pixel-level classification probability map is converted into a binary data stream, and then the binary data stream is subjected to encryption operation by using the public key of the receiver to generate an encrypted data packet.
[0067] In the embodiment, the data transmission module transmits the encrypted data packet to the monitoring center server through a secure communication protocol, the secure communication protocol adopts a TLS 1.2 protocol, performs bidirectional identity authentication when establishing a transmission link, adds a timestamp and a digital signature to the data packet in the transmission process, and ensures the integrity and non-repudiation of data transmission. In specific implementation, the transmission unit also includes a cache management module, which temporarily stores the encrypted data packet when transmission fails, and automatically reinitiates transmission after network recovery, while recording the transmission status log. After receiving the encrypted data packet, the monitoring center server decrypts it using the corresponding private key to restore the original semantic segmentation result data, which includes pixel-level classification information and abnormal confidence information for subsequent analysis and decision-making. Through the implementation of the above transmission unit, the security and reliability of the semantic segmentation result in the transmission process can be ensured, preventing data leakage and tampering, while ensuring the timeliness and integrity of the monitoring data.
[0068] In an embodiment provided in the present application, based on the above scheme, optionally, further comprising: an abnormal alarm unit; The abnormal alarm unit is configured to output alarm information when the abnormal confidence map in the semantic segmentation result exceeds a preset abnormal threshold.
[0069] The abnormal confidence map is a two-dimensional probability matrix output by the semantic segmentation model, used to represent the probability distribution of abnormal events occurring at each spatial position in the monitoring scene. The abnormal confidence map is generated through an abnormal detection branch in the semantic segmentation model, and its numerical range is between 0 and 1. The preset abnormal threshold is a critical value for triggering an alarm, which is set by the system in advance. The threshold can be dynamically adjusted according to the safety requirements of different application scenarios, and the value is 0.85. When the value of any pixel point in the abnormal confidence map exceeds the preset abnormal threshold, the abnormal alarm unit starts the alarm process, which includes generating alarm information, recording alarm logs, and triggering alarm devices. The alarm information includes abnormal type, abnormal position coordinates, abnormal occurrence time, and abnormal confidence value. The alarm information is provided to the monitoring personnel through various ways such as audible and visual alarms, SMS notifications, and system pop-up windows.
[0070] In an embodiment provided in the present application, based on the above scheme, optionally, further comprising a multi-modal event generation unit; The multi-modal event generation unit is configured to match the semantic segmentation result with a preset knowledge graph to generate a structured event description.
[0071] The preset knowledge graph is a domain knowledge base containing scene objects, event types, and spatio-temporal relationships. The knowledge graph is represented using a resource description framework, including three types of elements: entities, attributes, and relationships.
[0072] In the embodiment, the multi-modal event generation unit extracts key information in the semantic segmentation result through a semantic parsing algorithm, including a detected target class, a target spatial distribution, and an abnormal region position, matches the key information with entities in a knowledge graph, and finds an optimal matching path through a graph matching algorithm. After the matching is completed, the multi-modal event generation unit generates a structured event description according to the matching result, the event description is organized in a JSON format, and includes an event type, a participating object, a time of occurrence, a spatial position, and a confidence, and the like. Through the above processing procedure, the multi-modal event generation unit can convert a low-level pixel-level segmentation result into a high-level semantic event description, and provide a structured data basis for subsequent event analysis and decision support.
[0073] The embodiment of the application provides a real-time video monitoring intelligent analysis system based on deep learning, which is applied to an electronic device, a method flowchart of the method is as shown in Figure 4 The embodiment of the application provides a real-time video monitoring intelligent analysis system based on deep learning, which is applied to an electronic device, a method flowchart of the method is as shown in S401: acquiring a real-time monitoring video stream to be analyzed, and extracting a target video frame sequence in the real-time monitoring video stream.
[0074] S402: processing the target video frame sequence by using a pre-trained scene parsing model, to obtain a dynamic object mask and a scene semantic graph corresponding to the real-time monitoring video stream.
[0075] S403: acquiring thermal infrared features and dynamic environment variable information of a monitoring area corresponding to the real-time monitoring video stream. S404: inputting the dynamic object mask, the scene semantic graph, the thermal infrared features, and the dynamic environment variable information into a pre-trained semantic segmentation model, to obtain a semantic segmentation result output by the semantic segmentation model.
[0076] It should be noted that each of the embodiments in the specification adopts a progressive manner for description, and each embodiment focuses on the difference from other embodiments, and the same and similar parts of each embodiment can be referred to each other.
[0077] Finally, it should be noted that, in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between the entities or operations.
[0078] For the convenience of description, the above system is described as various units in function. Of course, the functions of each unit can be realized in the same or multiple software and / or hardware in the implementation of the present application.
[0079] Those skilled in the art can clearly understand the application by the description of the above embodiments that the application can be implemented by means of software and the necessary universal hardware platform. Based on such understanding, the technical solutions of the application can be embodied in the form of a software product in essence or in the part of the prior art that makes a contribution. The computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the method described in each embodiment or some part of the embodiments of the application.
[0080] The above describes in detail the real-time video monitoring intelligent analysis system based on deep learning provided by the application, and the principle and implementation of the application are described by applying specific examples. The above embodiment is only used to help understand the method and core idea of the application. Meanwhile, for those skilled in the art, the specific implementation and application range will be changed according to the idea of the application. In summary, the content of the specification should not be understood as a limitation of the application.
Claims
1. A real-time video surveillance intelligent analysis system based on deep learning, characterized in that, include: The video acquisition unit is used to acquire the real-time monitoring video stream to be analyzed and extract the target video frame sequence from the real-time monitoring video stream. The scene parsing unit is used to process the target video frame sequence using a pre-trained scene parsing model to obtain the dynamic object mask and scene semantic map corresponding to the real-time monitoring video stream. The acquisition unit is used to acquire the thermal infrared characteristics and dynamic environmental variable information of the monitoring area corresponding to the real-time monitoring video stream. The semantic segmentation unit is used to input the dynamic object mask, the scene semantic map, the thermal infrared features, and the dynamic environmental variable information into a pre-trained semantic segmentation model to obtain the semantic segmentation result output by the semantic segmentation model.
2. The system according to claim 1, characterized in that, The scene parsing unit includes: The scene parsing subunit is used to input the target video frame sequence into the scene parsing model, so that the backbone network in the scene parsing model performs convolution processing on the target video frame sequence to obtain convolution features; the multi-scale feature extraction network in the scene parsing model outputs a basic feature map based on the convolution features; the temporal difference convolutional network in the scene parsing model outputs a dynamic object mask and dynamic context features based on the convolution features and the basic feature map; and the semantic enhancement network in the scene parsing model obtains a scene semantic map based on the basic feature map, the dynamic object mask, and the dynamic context features.
3. The system according to claim 1, characterized in that, Also includes: First model training unit; The first model training unit is used to acquire the first training sample set and the initial scene parsing model; The first training sample set includes multiple first training samples and sample labels for each first training sample; The first training sample includes a sequence of historical target surveillance video frames; the initial scene analysis model is trained using the first training sample set.
4. The system according to claim 1, characterized in that, The semantic segmentation unit includes: The semantic segmentation subunit is used to input the dynamic object mask, the scene semantic map, the thermal infrared features, and the dynamic environmental variable information into a pre-trained semantic segmentation model. The feature alignment module in the semantic segmentation model performs spatial registration and cross-modal association on the thermal infrared features and the scene semantic map to obtain fused features. The spatiotemporal convolutional network in the semantic segmentation model outputs spatiotemporal enhanced features based on the fused features. The dynamic gating network in the semantic segmentation model outputs the semantic segmentation result based on the spatiotemporal enhanced features, the dynamic environmental variable information, and the dynamic object mask.
5. The system according to claim 1, characterized in that, Also includes: Second model training unit; The second model training unit is used to acquire a second training sample set and an initial semantic segmentation model; the second training sample set includes multiple second training samples and sample labels for each second training sample; the second training samples include historical dynamic object masks, historical scene semantic maps, historical thermal infrared features, and historical dynamic environmental variable information; the initial semantic segmentation model is trained using the second training sample set.
6. The system according to claim 1, characterized in that, The acquisition unit includes: A receiving subunit is used to receive thermal imaging data streams of the monitoring area corresponding to the real-time monitoring video stream; The feature extraction subunit is used to extract features from the thermal imaging data stream through a convolutional network to obtain thermal infrared features.
7. The system according to claim 1, characterized in that, The acquisition unit includes: The acquisition subunit is used to acquire the illumination data of the monitoring area corresponding to the real-time monitoring video stream, and calculate the spatial gradient matrix based on the Sobel operator. The identification subunit is used to identify the motion state data of the target tracking object in the target video frame sequence, and calculate the interaction energy matrix based on the motion state data and the dynamic object mask. The first execution subunit is used to receive the audio stream collected by the microphone array in the monitored area and obtain the sound field fluctuation coefficient based on the audio stream; The second execution subunit is used to compose the dynamic environmental variable information based on the spatial gradient matrix, the interaction energy matrix, and the sound field fluctuation coefficient.
8. The system according to claim 1, characterized in that, Also includes: Transmission unit; The transmission unit is used to encrypt the semantic segmentation result and send the encrypted semantic segmentation result to the monitoring center server.
9. The system according to claim 1, characterized in that, Also includes: Abnormal alarm unit; The anomaly alarm unit is used to output alarm information when the anomaly confidence map in the semantic segmentation result exceeds a preset anomaly threshold.
10. The system according to claim 1, characterized in that, It also includes a multimodal event generation unit; The multimodal event generation unit is used to match the semantic segmentation result with a preset knowledge graph to generate a structured event description.
Citation Information
Cited By
Monitoring video analysis method and system
CN121505524A
Expressway scene video segmentation method and device and electronic equipment
CN121600450A