Emergency monitoring system based on video analysis

By using DCNN network and graph convolutional neural network in the video analysis system, combined with the Retinex algorithm for data enhancement and lighting compensation, the existing system's lack of real-time and analysis accuracy in complex monitoring scenarios is solved, and more efficient abnormal alarms and more interpretable event information generation is achieved.

CN120107884APending Publication Date: 2025-06-06HUANENG CHONGQING LIANGJIANG GAS TURBINE POWER GENERATION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510160587.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

Smart Images

  • Figure CN120107884A_ABST
    Figure CN120107884A_ABST
Patent Text Reader

Abstract

The invention discloses an emergency monitoring system based on video analysis, and the system comprises a data collection and preprocessing module which is used for collecting and preprocessing real-time video stream data, and obtaining initial video data; the feature extraction module is used for extracting feature vectors from the initial video data; and the condition pre-recognition module is used for fusing the feature vectors to obtain fused feature vectors, and performing classification prediction on the fused feature vectors by using a graph convolutional neural network and a random forest classifier to obtain an emergency monitoring result based on video analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information technology, and in particular relates to an emergency monitoring system based on video analysis. Background Art

[0002] In the emergency monitoring system, real-time video analysis is a key link. The video analysis module needs to process and analyze massive amounts of video data in real time to accurately identify various abnormal events and emergencies. This requires the video analysis algorithm to have extremely high accuracy and real-time performance. However, in actual complex monitoring scenarios, there are often a large number of interference factors in the video screen, such as changes in lighting, weather, personnel occlusion, complex background, etc., which makes it difficult to accurately identify abnormal events. At the same time, the characteristics of different abnormal events vary greatly, and a single analysis algorithm is difficult to fully cover them.

[0003] In addition, real-time performance is also a major challenge. Video surveillance often requires the detection of abnormal situations as soon as possible, but the processing and analysis of massive video data is very time-consuming. How to maximize the processing speed while taking into account the accuracy of analysis and achieve real-time abnormal alarm is a technical difficulty.

[0004] Finally, the interpretability of video analysis results is also critical. When an abnormal event is identified, the system needs to provide detailed information such as the time, location, people involved, and type of event so that the commander can respond quickly. However, video analysis algorithms are usually a "black box" that cannot generate this information, making it difficult to use the analysis results directly. How to enhance the interpretability of the algorithm so that its output meets the needs of actual applications is also a problem to be solved. Summary of the invention

[0005] To solve the above problems, the present invention proposes an emergency monitoring system based on video analysis, the system comprising:

[0006] The data acquisition and preprocessing module is used to acquire real-time video stream data and perform preprocessing to obtain initial video data;

[0007] A feature extraction module, which uses a DCNN network model to extract feature vectors from the initial video data;

[0008] The situation pre-identification module is used to fuse the feature vectors to obtain a fused feature vector, and use a graph convolutional neural network and a random forest classifier to classify and predict the fused feature vector to obtain an emergency monitoring result based on video analysis.

[0009] Optionally, the preprocessing process specifically includes:

[0010] The data enhancement unit is used to perform data enhancement on the video stream data using an adaptive Retinex algorithm to obtain a data enhanced image;

[0011] The data compensation unit is used to perform adaptive illumination compensation on the data enhanced image to obtain dynamic compensation data;

[0012] The foreground separation unit is used to screen a foreground video frame sequence with balanced illumination based on the motion compensation data to obtain initial video data.

[0013] Optionally, the operation process of the data enhancement unit specifically includes:

[0014] Decomposing the video stream data using adaptive Retinex to obtain a reflection map and an illumination map;

[0015] The reflection map and the illumination map are input into a U-Net network for feature extraction, and the extracted features are fused to obtain a data enhanced image.

[0016] Optionally, the reflection map and illumination map formulas are specifically:

[0017] R(x,y)=log[I(X,Y)]-log[G(x,y)*S(x,y)]

[0018] L(x,y)=log[G(x,y)]

[0019] Wherein, R(x,y) is the reflection map, L(x,y) is the illumination map, I(x,y) is the brightness value of the video stream data at the position (x,y), G(x,y) is the illumination intensity of the image at (x,y), S(x,y) is the local contrast of the video stream data at the position (x,y), and * is the convolution operation.

[0020] Optionally, the workflow of the feature extraction module includes:

[0021] Inputting the initial video data into a DCNN network model for feature extraction;

[0022] The DCNN network model includes an input layer, a hidden layer and an output layer, wherein the hidden layer includes a convolution layer with a convolution kernel size of 40, a pooling layer, a convolution layer with a convolution kernel size of 30, a fully connected layer and a perception layer.

[0023] Optionally, the situation pre-identification module includes a feature fusion submodule, a classification submodule and a prediction submodule;

[0024] The feature fusion submodule is used to fuse the feature vectors of all initial video data to obtain fused features;

[0025] The feature optimization submodule is used to perform feature optimization on the content of the fusion feature using a graph convolutional network;

[0026] The prediction submodule is used to perform feature classification on the result after feature optimization using a random forest classifier to obtain a classification result.

[0027] Optionally, the workflow of the feature fusion submodule specifically includes:

[0028] Use the weighted attention mechanism to calculate the weight of each feature vector;

[0029] Use the softmax function to normalize the weighted feature vector to obtain the normalized weight;

[0030] The normalized weight is weighted summed with the feature vectors of all the initial video data to obtain a fusion feature.

[0031] Optionally, the workflow of the classification submodule specifically includes:

[0032] Based on the node feature matrix, normalized adjacency matrix, adjacency matrix and degree matrix of the graph in the fusion features, feature optimization is completed using a graph convolutional neural network.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention uses the Retinex algorithm to process the video stream. In the monitoring scene, when strong light shines on the surface of the object, the reflection map can better retain the detailed texture information of the surface of the object, and the illumination map reflects the influence of the illumination on the brightness. This decomposition method is helpful for the subsequent targeted processing of different components in the video data. In addition, the use of U-Net can enhance the key information such as the edge and texture of the object in the video data, so that the video image can maintain a good quality under different illumination conditions, and provide a clearer and more accurate data basis for subsequent feature extraction and situation pre-identification. The data compensation unit performs adaptive illumination compensation on the data enhanced image. In the actual video monitoring scene, the illumination conditions often change, such as the switch of indoor lights, the intensity change of outdoor daylight, etc. Adaptive illumination compensation can automatically adjust the brightness and contrast of the image according to the illumination characteristics of the video image, so that the image can maintain a relatively stable visual effect under different illumination environments. For example, in the outdoor monitoring scene where cloudy and sunny days alternate, the video image after illumination compensation can better maintain the recognizability of the object and reduce the misjudgment caused by illumination changes.

[0035] Finally, the classification submodule uses a graph convolutional neural network to optimize the fused features. The graph convolutional neural network can fully consider the topological relationship between features. In the pre-identification of surveillance video situations, there may be correlations between different features, such as the shape features and motion trajectory features of a vehicle may affect each other. Based on the node feature matrix and the standardized adjacency matrix of the graph in the fused features, the graph convolutional neural network can explore the intrinsic connections between these features and optimize the features. For example, it can strengthen features that have important connections in the topological structure and weaken the influence of noise features, so that the optimized features can better reflect the essence of the actual situation, thereby improving the accuracy of classification predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings constituting a part of the present invention are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0037] Figure 1 The present invention is a system structure diagram of an emergency monitoring system based on video analysis according to an embodiment of the present invention. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0039] First, some of the technical terms used in the present invention are explained:

[0040] Convolutional Neural Network (CNN) is a deep learning algorithm that is mainly used to process data with a grid structure, such as images, videos, etc. The following is a detailed introduction to convolutional neural networks:

[0041] Convolutional layer, the convolutional layer is the core part of CNN. It extracts the features of the input data through the convolution kernel (also called filter). The convolution kernel is a small weight matrix that slides over the input data and performs a weighted sum operation on each position.

[0042] Pooling layer, the pooling layer usually follows the convolution layer. Its function is to reduce the spatial dimension of the data while retaining important information. Common pooling operations include maximum pooling and average pooling.

[0043] Fully connected layer, fully connected layer is located at the end of CNN. In the fully connected layer, each neuron is connected to all neurons in the previous layer. Its function is to integrate the features extracted by the previous convolutional layer and pooling layer, and output the final classification result or the output required for other tasks.

[0044] Random forest is an ensemble learning algorithm, mainly used for classification and regression tasks, with the following main functions and features: Random forest is a model composed of multiple decision trees. When constructing a random forest, first generate multiple sub-datasets from the original training data set through bootstrap sampling. For example, the original data set has 1000 samples, and 1000 samples are extracted each time (with possible duplication) to form a sub-dataset. A total of 100 samples are extracted, and 100 sub-datasets are obtained. For each sub-dataset, a decision tree is constructed. In the process of constructing a decision tree, each time a node is split, instead of considering all features, a part of the features (such as the square root of the total number of features) is randomly selected from all features for splitting. In this way, each decision tree uses different data and feature subsets when constructing, which increases the differences between decision trees. When a new sample is classified, each decision tree in the random forest will classify the sample. For example, for a binary classification problem, some decision trees may classify the sample as category A, and some decision trees may classify the sample as category B. Then the random forest determines the final classification result based on the principle of majority voting. If most decision trees classify the sample as category A, then the random forest classifies the sample as category A.

[0045] Example

[0046] like Figure 1 As shown, in this embodiment, a system for monitoring emergencies based on video analysis is provided, and the system includes:

[0047] The data acquisition and preprocessing module is used to acquire real-time video stream data and perform preprocessing to obtain initial video data. The preprocessing process specifically includes:

[0048] The data enhancement unit is used to perform data enhancement on the video stream data using an adaptive Retinex algorithm to obtain a data enhanced image;

[0049] The operation process of the data enhancement unit specifically includes:

[0050] Decomposing the video stream data using adaptive Retinex to obtain a reflection map and an illumination map;

[0051] The specific formulas for reflection map and illumination map are:

[0052] R(x,y)=log[I(X,Y)]-log[G(x,y)*S(x,y)]

[0053] L(x,y)=log[G(x,y)]

[0054] Wherein, R(x,y) is the reflection map, L(x,y) is the illumination map, I(x,y) is the brightness value of the video stream data at the position (x,y), G(x,y) is the illumination intensity of the image at (x,y), S(x,y) is the local contrast of the video stream data at the position (x,y), and * is the convolution operation.

[0055] The reflection map and the illumination map are input into the U-Net network for feature extraction, and the extracted features are fused to obtain a data enhanced image.

[0056] The data compensation unit is used to perform adaptive illumination compensation on the data enhanced image to obtain dynamic compensation data; the method of adaptively adjusting the weight of the multi-scale Retinex operator, combined with the image detail information, can effectively solve the contradiction between color distortion and contrast enhancement in the downhole image, which can not only enhance the detail information of the downhole image, but also avoid color distortion. The algorithm has been verified through multiple groups of experiments, showing excellent performance in removing the dark background of low-light images, eliminating non-uniform illumination, and enhancing image details. In order to balance local contrast and color fidelity in the image enhancement process and improve the image enhancement effect to a certain extent, this study synthesizes the processing results of Retinex algorithms of different scales, selects different weight values ​​according to the processing effect, and then uses the weighted average method to fuse the processing results of Retinex algorithms of different scales. This method can take into account the advantages of multiple scales, thereby achieving better image enhancement effects.

[0057]

[0058] Among them, R k (x, y) is the enhanced image in the kth band. In order to further improve the low-light image enhancement effect, a multi-scale attention mechanism is introduced to adapt to the different scales of information in the image. In the image enhancement process, the attention mechanism can help the algorithm focus on and strengthen the key information in the image, while reducing the attention to irrelevant information. Specifically, the attention mechanism is used to learn the correlation between image features and adaptively adjust the weights to highlight the focus on important features. Before the Unet network, this study introduced an attention mechanism module, whose input is a feature map and output is a weighted feature map.

[0059] The foreground separation unit is used to screen a foreground video frame sequence with balanced illumination based on the motion compensation data to obtain initial video data.

[0060] The feature extraction module uses the DCNN network model to extract feature vectors from the initial video data. The workflow of the feature extraction module includes:

[0061] Input the initial video data into the DCNN network model for feature extraction;

[0062] The DCNN network model includes an input layer, a hidden layer and an output layer. The hidden layer includes a convolution layer with a convolution kernel size of 40, a pooling layer, a convolution layer with a convolution kernel size of 30, a fully connected layer and a perception layer.

[0063] The situation pre-identification module is used to fuse the feature vectors to obtain a fused feature vector, and use a graph convolutional neural network and a random forest classifier to classify and predict the fused feature vector to obtain an emergency monitoring result based on video analysis.

[0064] The situation pre-identification module includes a feature fusion submodule, a classification submodule, and a prediction submodule;

[0065] The feature fusion submodule is used to fuse the feature vectors of all the initial video data to obtain fused features, including:

[0066] Use the weighted attention mechanism to calculate the weight of each feature vector;

[0067] Use the softmax function to normalize the weighted feature vector to obtain the normalized weight;

[0068] The normalized weight is weighted summed with the feature vectors of all the initial video data to obtain a fusion feature.

[0069] The feature optimization submodule is used to optimize the content of the fused features using a graph convolutional network;

[0070] The prediction submodule is used to classify the results after feature optimization using a random forest classifier to obtain classification results.

[0071] The workflow of the classification submodule specifically includes:

[0072] Based on the node feature matrix, normalized adjacency matrix, adjacency matrix and degree matrix of the graph in the fusion features, feature optimization is completed using a graph convolutional neural network.

[0073] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. An emergency monitoring system based on video analysis, characterized in that: The system specifically comprises: The data acquisition and preprocessing module is used to acquire real-time video stream data and perform preprocessing to obtain initial video data; A feature extraction module, used for extracting feature vectors from the initial video data using a DCNN network model; The situation pre-identification module is used to fuse the feature vectors to obtain a fused feature vector, and use a graph convolutional neural network and a random forest classifier to classify and predict the fused feature vector to obtain an emergency monitoring result based on video analysis.

2. The emergency situation monitoring system based on video analysis according to claim 1, characterized in that: The pretreatment process specifically includes: The data enhancement unit is used to perform data enhancement on the video stream data using an adaptive Retinex algorithm to obtain a data enhanced image; The data compensation unit is used to perform adaptive illumination compensation on the data enhanced image to obtain dynamic compensation data; The foreground separation unit is used to screen a foreground video frame sequence with balanced illumination based on the motion compensation data to obtain initial video data.

3. The emergency situation monitoring system based on video analysis according to claim 2, characterized in that: The operation process of the data enhancement unit specifically includes: Decomposing the video stream data using adaptive Retinex to obtain a reflection map and an illumination map; The reflection map and the illumination map are input into a U-Net network for feature extraction, and the extracted features are fused to obtain a data enhanced image.

4. The emergency situation monitoring system based on video analysis according to claim 3, characterized in that: The reflection map and illumination map formulas are specifically: R(x,y)=log[I(X,Y)]-log[G(x,y)*S(x,y)] L(x,y)=log[G(x,y)] Wherein, R(x,y) is the reflection map, L(x,y) is the illumination map, I(x,y) is the brightness value of the video stream data at the position (x,y), G(x,y) is the illumination intensity of the image at (x,y), S(x,y) is the local contrast of the video stream data at the position (x,y), and * is the convolution operation.

5. The emergency situation monitoring system based on video analysis according to claim 1, characterized in that: The workflow of the feature extraction module includes: Inputting the initial video data into a DCNN network model for feature extraction; The DCNN network model includes an input layer, a hidden layer and an output layer, wherein the hidden layer includes a convolution layer with a convolution kernel size of 40, a pooling layer, a convolution layer with a convolution kernel size of 30, a fully connected layer and a perception layer.

6. The emergency situation monitoring system based on video analysis according to claim 1, characterized in that: The situation pre-identification module includes a feature fusion submodule, a classification submodule and a prediction submodule; The feature fusion submodule is used to fuse the feature vectors of all initial video data to obtain fused features; The feature optimization submodule is used to perform feature optimization on the content of the fusion feature using a graph convolutional network; The prediction submodule is used to perform feature classification on the result after feature optimization using a random forest classifier to obtain a classification result.

7. The emergency situation monitoring system based on video analysis according to claim 6, characterized in that: The workflow of the feature fusion submodule specifically includes: Use the weighted attention mechanism to calculate the weight of each feature vector; Use the softmax function to normalize the weighted feature vector to obtain the normalized weight; The normalized weight is weighted summed with the feature vectors of all the initial video data to obtain a fusion feature.

8. The emergency situation monitoring system based on video analysis according to claim 6, characterized in that: The workflow of the classification submodule specifically includes: Based on the node feature matrix, normalized adjacency matrix, adjacency matrix and degree matrix of the graph in the fusion features, feature optimization is completed using a graph convolutional neural network.