An intelligent processing system for monitoring video information based on cloud computing
Through the intelligent monitoring video information processing system based on cloud computing, image preprocessing and abnormal detection is used using deep learning and machine learning technology, the problem of insufficient accuracy in target object recognition under light changes and occlusion is solved, and efficient and accurate security monitoring is achieved.
Patent Information
- Application Number
- CN202411871201.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing surveillance video information is susceptible to factors such as light changes and occlusion when obtaining, resulting in insufficient accuracy in identifying target objects and false alarms.
The intelligent processing system for monitoring video information based on cloud computing is adopted, including video data acquisition, data transmission, cloud storage, data stream processing, target analysis and abnormal detection modules. Deep learning and machine learning technology are used to perform image preprocessing, background difference, dynamic target separation, abnormal behavior detection, etc., to improve the accuracy of target recognition and reduce false positives and missed reports.
It significantly improves the recognition accuracy of the target object of the monitoring video, reduces false alarms and missed reports, and achieves rapid response to abnormal behavior and efficient security monitoring.
Smart Images

Figure CN119672613B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of monitoring video information processing, and particularly relates to an intelligent processing system for monitoring video information based on cloud computing. Background Art
[0002] Cloud computing, as a model for providing computing resources and services through a network, has main characteristics such as virtualization, elastic expansion, pay-per-use, and automated management. With the rapid development of cloud computing technology, its powerful computing power and storage capacity provide strong support for the intelligent processing of monitoring video information. With the acceleration of the urbanization process, the requirements for video surveillance systems are also getting higher and higher. Along with the continuous evolution of security technologies, video surveillance technology has developed from analog and digital to networked and intelligent directions. The introduction of cloud computing technology provides a new solution for the intelligent processing of video surveillance systems.
[0003] In the prior art, when acquiring monitoring video information, it is easily affected by various factors such as light changes and occlusion, which leads to insufficient accuracy during analysis and processing, resulting in false alarms and missed alarms when the monitoring video identifies target objects. Therefore, how to weaken the influence of environmental factors on monitoring video information and improve the accuracy of identifying target objects in monitoring videos is the problem we need to solve. For this reason, an intelligent processing system for monitoring video information based on cloud computing is proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent processing system for monitoring video information based on cloud computing to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is:
[0006] An intelligent processing system for monitoring video information based on cloud computing, including a monitoring information processing platform, which is communicatively connected to a video data acquisition module, a data transmission module, a cloud data storage module, a data stream processing module, a target analysis module, an anomaly detection module, and an intelligent alarm management module;
[0007] The video data acquisition module is used to acquire monitoring video data from multiple cameras and sensors. The cameras are distributed in different monitoring areas to capture video signals in real time, ensuring that the system can receive high-quality original video streams;
[0008] The data transmission module is used to transmit the acquired video data to the cloud server through a wireless network, and it is necessary to ensure the stable and fast transmission of data for subsequent processing, ensuring the timeliness and integrity of the video data, and providing a reliable data source for subsequent intelligent processing;
[0009] The cloud data storage module is used to utilize the distributed storage capacity of the cloud computing platform to store video data transmitted to the cloud server on a large scale and efficiently, solve the capacity limitation and cost problems faced by traditional storage methods, and perform classification, indexing, and retrieval management operations on the stored video data to facilitate users to quickly find the required videos, provide almost unlimited storage space, reduce storage costs, and improve data access efficiency;
[0010] The data stream processing module is used to perform preprocessing operations on video data to improve image quality and provide clearer images for subsequent analysis. The preprocessing operations include noise reduction, deblurring, contrast enhancement, etc.;
[0011] The target analysis module is used to analyze video data using deep learning technology, identify and remove static backgrounds in the video data, separate dynamic targets, reduce false alarms caused by background changes, and perform adaptive processing on environmental factors such as light changes and occlusions, significantly improving the accuracy of target object recognition and reducing the occurrence of missed alarms and false alarms;
[0012] The anomaly detection module is used to perform real-time tracking of target objects in the video, detect abnormal behaviors or events in the video through in-depth analysis of video content, discover potential security hazards and abnormal situations, quickly respond to abnormal situations, and reduce missed alarms;
[0013] The intelligent alarm management module is used to generate alarm information according to the anomaly detection results and preset rules, and manage the alarm process to improve the accuracy and response speed of alarms.
[0014] A further improvement of the technical solution of the present invention lies in: in the video data acquisition module, the acquisition process of monitoring video data is as follows:
[0015] Deploy cameras and sensors in different monitoring areas to ensure coverage of all areas that need to be monitored, and use the cameras to capture video signals in the monitoring areas in real time, and cooperate with the sensors to capture sensor data in the monitoring areas. Among them, the sensors include infrared sensors and motion sensors to provide more comprehensive environmental information;
[0016] Perform preliminary processing on the captured video signals, perform automatic exposure adjustment and white balance adjustment to improve image quality, and then convert the preliminarily processed video signals into digital signals, perform encoding and compression, and effectively reduce the file size while ensuring video quality for easy storage and transmission;
[0017] Through the wireless network, encapsulate the encoded video data and sensor data into network data packets and transmit them to the data transmission module.
[0018] A further improvement of the technical solution of the present invention lies in: In the cloud data storage module, the process of storing video data is as follows:
[0019] The cloud server receives video data and sensor data from the data transmission module, and preliminarily processes the received video data and sensor data, including data format verification and integrity check, to ensure the accuracy and integrity of the data;
[0020] Perform data sharding on the video data and sensor data, divide the video data and sensor data into multiple data blocks, and use redundant storage technology to copy each data block to multiple storage nodes for distributed storage on multiple storage nodes to prevent data loss caused by single-point failures and help improve data reliability and access speed;
[0021] Analyze the video data processed by data sharding, classify it according to the content, source, and time attributes of the video data, and build an index for the classified video data to establish index information, including keyword index, time index, and metadata index, for quick retrieval;
[0022] Set up a user interface to enable users to retrieve in multiple ways such as time range, camera location, and event type by using the index and metadata, quickly retrieve the required video segments, and provide preview and extraction functions for the video segments, enabling users to select and extract the required video segments as needed.
[0023] A further improvement of the technical solution of the present invention lies in: In the data stream processing module, the process of preprocessing video data is as follows:
[0024] The data stream processing module receives the video data stream and sensor data from the cloud server, and decodes the encoded and compressed video data to restore it to the original video format;
[0025] Extract individual video frames from the video stream and apply a smoothing filter to reduce noise and artifacts in the image;
[0026] If there are problems such as blurring or motion blurring in the video data, then use a blind deblurring algorithm to estimate the blur kernel through an iterative optimization algorithm and restore a clear image;
[0027] Enhance the contrast of the image by adjusting the contrast of the video frame to make the details in the image more prominent and obvious, optimize the visual effect of the image, and automatically adjust the brightness according to the video content and ambient light conditions;
[0028] Re-encapsulate and combine the preprocessed video frames into a video stream, and transmit the preprocessed video data to the target analysis module and the anomaly detection module for subsequent analysis and processing.
[0029] A further improvement of the technical solution of the present invention lies in that in the target analysis module, the process of separating dynamic targets is as follows:
[0030] Receive the preprocessed video data stream from the data stream processing module, and decode the video data stream to obtain continuous video image frames;
[0031] Use a convolutional neural network to learn the background frames in the video data stream, establish a static background model, and identify the static background;
[0032] For the subsequent continuous video image frames of the video data stream, compare them with the static background model, separate the dynamic targets by background difference method, and use a target tracking algorithm to track the dynamic targets, and continuously identify the same target in the video data stream;
[0033] Adjust the detection result of the target object according to the illumination condition of the image through brightness normalization, estimate the illumination component and reflection component of the scene, perform adaptive processing of illumination, improve the illumination condition of the image, and use a deep learning model to analyze the partial occlusion situation of the target, identify the occlusion area in the video image frame, and restore the occluded target information. The expression for estimating the illumination component and reflection component of the scene and performing adaptive processing of illumination is:
[0034] ;
[0035] Where represents the adjusted pixel value, represents the pixel value of the original image, represents the estimated reflection component;
[0036] Use a deep learning model to identify the separated targets, determine the category and attributes of the targets, and extract the feature vectors of the targets, and output the analysis and recognition results to the anomaly detection module to provide a basis for further analysis and decision-making. Among them, the feature vectors of the targets include the motion features and appearance features of the targets.
[0037] A further improvement of the technical solution of the present invention lies in that the process of establishing a static background model and identifying the static background is as follows:
[0038] Select an initial background frame from the video data stream, and preprocess the selected background frame, including scaling, normalization, data augmentation, etc., to improve the generalization ability of the model;
[0039] Construct a CNN (Convolutional Neural Network) architecture for background learning, including multiple convolutional layers, activation layers, pooling layers, and fully connected layers. Among them, the convolutional layer extracts image features through filters, and there are multiple filters, each filter responsible for extracting specific features in the image. The activation layer introduces non-linearity to enable the model to learn complex features. The pooling layer is used to reduce the spatial dimension of features and increase invariance to position changes. The fully connected layer is used to map the features extracted by the convolutional layer and the pooling layer to the final output;
[0040] Define a loss function to measure the difference between the model prediction and the actual label, and input the preprocessed background frame into the CNN. Calculate the output through forward propagation, calculate the loss, and update the network weights through backpropagation;
[0041] Use the trained background model to compare with the new frame and detect dynamic targets through background subtraction;
[0042] The expression of the convolutional layer is:
[0043] ;
[0044] Among them, represents the output of the th layer at position , represents the input of the th layer at position , represents the convolutional kernel weight of the th layer, and represent the offsets of the convolutional kernel in the horizontal and vertical directions, represents the bias term of the th layer, represents the activation function, including any one of ReLU , Sigmoid , Tanh ;
[0045] The pooling layer selects max pooling, and its expression is:
[0046] ;
[0047] Among them, represents the pooling output of the th layer at position , represents the input of the th layer at position , and Indicates the position offset within the pooling window.
[0048] A further improvement of the technical solution of the present invention lies in that: the process of continuously identifying the same target in the video data stream is as follows:
[0049] Combined with training to obtain a static background model, for each new video image frame, update the background model, and calculate the difference between the current frame and the background model to obtain a difference image;
[0050] Set a dynamic analysis threshold, analyze the correlation between the difference image and the dynamic analysis threshold, pixels higher than the dynamic analysis threshold are considered dynamic targets, and generate a foreground mask to separate the dynamic targets;
[0051] Perform connected component analysis on the foreground mask, identify independent objects in the foreground area, detect the target area in the foreground mask, extract the features of the target, and use the feature matching algorithm of Kalman filtering to match the target in consecutive frames;
[0052] Track the Kalman filter to update the position and features of the target, so as to continuously identify the same target in the video sequence;
[0053] The update formula of the background model is:
[0054] ;
[0055] Among them, represents the pixel value of the background model, represents the pixel value of the current frame, represents the update factor ;
[0056] The calculation expression of the difference image is:
[0057] ;
[0058] Among them, represents the difference result, represents the pixel value of the current frame, represents the pixel value of the background model;
[0059] The judgment expression for generating the foreground mask is:
[0060] ;
[0061] Among them, represents the pixel value of the foreground mask, represents the dynamic analysis threshold.
[0062] A further improvement of the technical solution of the present invention lies in that: in the abnormal detection module, the detection process of abnormal behaviors or events in the video is as follows:
[0063] Using the detection results output by the target analysis module, determine the target objects in the video, and extract the feature vectors of the targets, including the motion features and appearance features of the targets. By comparing the distances of the feature vectors, match the target features in consecutive frames. Among them, the appearance features include color histograms, texture features, and shape descriptors, and the motion features include the speed, acceleration, and trajectory of the targets;
[0064] Use machine learning to model normal behaviors, construct a behavior model, identify abnormal behaviors by comparing the real-time video content with the predefined behavior model, and calculate the deviation between the real-time behavior and the normal model to generate an anomaly score;
[0065] Set a threshold for the anomaly score according to historical data. When the anomaly score exceeds the threshold, it is determined as an abnormal behavior. And in multi-target tracking, remove the overlapping detection results and retain the best prediction results;
[0066] Once an abnormal behavior is detected, generate an alarm message, send the alarm message to relevant personnel, and provide detailed information about the abnormal behavior, including the time, location, and involved targets, for quick response and handling.
[0067] A further improvement of the technical solution of the present invention lies in that: the calculation process of the anomaly score is as follows:
[0068] Collect video data of normal behaviors, annotate them, and extract the appearance features and motion features of the targets;
[0069] Select an SVM (Support Vector Machine) machine learning model and use the annotated normal behavior video data to train the model to distinguish normal behaviors from abnormal behaviors;
[0070] Extract the features of the targets in the real-time video to obtain the feature vectors of the real-time behavior, input the feature vectors of the real-time behavior into the trained model for analysis, calculate the deviation between the real-time behavior features and the normal model, and calculate the anomaly score according to the deviation.
[0071] A further improvement of the technical solution of the present invention lies in that: the expression for training the behavior model is:
[0072] ;
[0073] Among them, represents the decision function, represents the test sample, represents the kernel function, which calculates the similarity between the sample and the test sample and, represents the The feature vector of a training sample, denotes the label of the th training sample, indicating normal behavior or abnormal behavior, denotes the coefficient of the th support vector,
[0074] The recognition of abnormal behavior is achieved by calculating the deviation between the real-time behavior features and the normal model, and cooperating with the cosine similarity judgment. Its expression is:
[0075] ;
[0076] wherein, denotes the cosine similarity, measuring the similarity degree between two vectors, and the value range is from -1 to 1, denotes the target feature vector extracted from the real-time video frame, denotes the feature vector of the normal behavior model, and respectively denote the and th eigenvalue in the vectors and denotes the dimension of the feature vector, that is, the total number of eigenvalues in the vector, used to describe the total feature quantity of the target feature;
[0077] The expression for calculating the abnormal score is:
[0078] ;
[0079] wherein, denotes the abnormal score, denotes the cosine similarity.
[0080] Due to the adoption of the above technical solution, the technical progress achieved by the present invention compared with the prior art is:
[0081] The present invention provides an intelligent processing system for monitoring video information based on cloud computing. By training a deep neural network model to learn the characteristics of normal and abnormal behaviors, it can achieve accurate recognition and tracking of target objects in the video, detect dynamic targets in the video, extract their motion and appearance features, identify abnormal behaviors through behavior analysis, calculate an abnormal score, and generate alarm information in a timely manner. This not only improves the accuracy of the monitoring system but also reduces false alarms and missed detections, making security monitoring more efficient and reliable. The present invention provides an intelligent processing system for monitoring video information based on cloud computing. By analyzing video data in real time, it can quickly and accurately detect abnormal behaviors in the video. Through the cloud computing platform, it can parallel-process the data streams of multiple cameras, effectively shortening the response time and improving the detection efficiency. At the same time, by collecting and analyzing a large amount of historical video data, it uses machine learning algorithms to train a precise behavior pattern recognition model, further learning and understanding the characteristics and laws of normal behaviors, so as to more accurately identify abnormal behaviors and predict future behaviors of the recognized behavior patterns, early warning potential risk points, and providing users with a more comprehensive and in-depth security analysis report, which helps to better understand the security status of the monitored area. The present invention provides an intelligent processing system for monitoring video information based on cloud computing. Based on the cloud computing platform, by centrally storing and processing video data, it simplifies the deployment and management of the system. Through the management interface in the cloud, it can monitor the status of all monitoring points in real time, quickly locate and solve problems. At the same time, the resources of the cloud computing platform can be allocated on demand, avoiding resource waste and ensuring that each monitoring point can obtain sufficient computing power, thereby improving the overall monitoring efficiency and effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0083] Figure 1 It is a block diagram of the components of the present invention;
[0084] Figure 2 It is a flowchart for separating dynamic targets of the present invention;
[0085] Figure 3 It is a flowchart for continuously recognizing the same target in the video data stream of the present invention;
[0086] Figure 4 It is a flowchart for detecting abnormal behaviors or events in the video of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0087] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] An embodiment is as Figure 1 shown. The present invention provides an intelligent processing system for monitoring video information based on cloud computing, including a monitoring information processing platform, characterized in that: the monitoring information processing platform is communicatively connected to a video data acquisition module, a data transmission module, a cloud data storage module, a data stream processing module, a target analysis module, an anomaly detection module, and an intelligent alarm management module.
[0089] The video data acquisition module is used to acquire monitoring video data from multiple cameras and sensors. The cameras are distributed in different monitoring areas to capture video signals in real time, ensuring that the system can receive high-quality original video streams.
[0090] The data transmission module is used to transmit the acquired video data to the cloud server through a wireless network, and it is necessary to ensure stable and fast data transmission for subsequent processing, ensuring the timeliness and integrity of the video data, and providing a reliable data source for subsequent intelligent processing.
[0091] The cloud data storage module is used to utilize the distributed storage ability of the cloud computing platform to store the video data transmitted to the cloud server on a large scale and efficiently, solve the capacity limitation and cost problems faced by traditional storage methods, and perform classification, indexing, and retrieval management operations on the stored video data to facilitate users to quickly find the required videos, provide almost unlimited storage space, reduce storage costs, and improve data access efficiency.
[0092] The data stream processing module is used to perform preprocessing operations on the video data to improve the image quality and provide a clearer image for subsequent analysis. The preprocessing operations include noise reduction, deblurring, contrast enhancement, etc.
[0093] The target analysis module is used to analyze the video data using deep learning technology, identify and remove the static background in the video data, separate the dynamic targets, reduce false alarms caused by background changes, and perform adaptive processing on environmental factors such as light changes and occlusion, significantly improving the accuracy of target object recognition and reducing the occurrence of missed alarms and false alarms.
[0094] Anomaly detection module, used for real-time tracking of target objects in videos. Through in-depth analysis of video content, it detects abnormal behaviors or events in the videos, discovers potential security hazards and abnormal situations, quickly responds to abnormal situations, and reduces false negatives.
[0095] Intelligent alarm management module, used for generating alarm information according to the anomaly detection results and preset rules, and managing the alarm process, so as to improve the accuracy and response speed of alarms.
[0096] The intelligent processing system for surveillance video information based on cloud computing of the present invention realizes precise identification and tracking of target objects in videos by training a deep neural network model to learn the characteristics of normal and abnormal behaviors, detects dynamic targets in the videos, extracts their motion and appearance characteristics, identifies abnormal behaviors through behavior analysis, calculates anomaly scores, and generates alarm information in a timely manner. It not only improves the accuracy of the surveillance system, but also reduces the situations of false alarms and false negatives, making security surveillance more efficient and reliable.
[0097] In some embodiments, in the video data acquisition module, the process of acquiring surveillance video data is as follows:
[0098] Deploy cameras and sensors in different surveillance areas to ensure coverage of all areas that need to be monitored, and use the cameras to capture video signals in the surveillance areas in real time, and cooperate with the sensors to capture sensor data in the surveillance areas. Among them, the sensors include infrared sensors and motion sensors to provide more comprehensive environmental information; perform preliminary processing on the captured video signals, perform automatic exposure adjustment and white balance adjustment to improve the image quality, and then convert the preliminarily processed video signals into digital signals, perform encoding and compression, while ensuring the video quality, effectively reducing the file size for easy storage and transmission; through the wireless network, encapsulate the encoded video data and sensor data into network data packets and transmit them to the data transmission module.
[0099] In some embodiments, in the cloud data storage module, the process of storing video data is as follows:
[0100] The cloud server receives video data and sensor data from the data transmission module, and performs preliminary processing on the received video data and sensor data. Among them, the preliminary processing includes data format verification and integrity check to ensure the accuracy and integrity of the data; performs data sharding processing on the video data and sensor data, divides the video data and sensor data into multiple data blocks, and uses redundant storage technology to copy each data block to multiple storage nodes for distributed storage on multiple storage nodes to prevent data loss caused by single-point failures, which helps improve the reliability and access speed of the data; analyzes the sharded video data, classifies it according to the content, source, and time attributes of the video data, and constructs an index for the classified video data to establish index information. The index information includes keyword index, time index, and metadata index for quick retrieval; sets up a user interface, enabling users to retrieve in multiple ways such as time range, camera location, and event type by using the index and metadata, quickly retrieve the required video segments, and provide preview and extraction functions for the video segments, allowing users to select and extract the required video segments as needed.
[0101] In some embodiments, in the data stream processing module, the process of preprocessing video data is as follows:
[0102] The data stream processing module receives the video data stream and sensor data from the cloud server, and decodes the encoded and compressed video data to restore it to the original video format; extracts individual video frames from the video stream and applies a smoothing filter to reduce noise and artifacts in the image, improving the clarity and visual quality of the image; if there are problems such as blurring or motion blurring in the video data, a blind deblurring algorithm is used to estimate the blur kernel through an iterative optimization algorithm and restore the clear image; enhances the contrast of the image by adjusting the contrast of the video frame, making the details in the image more prominent, optimizing the visual effect of the image for subsequent analysis, and automatically adjusting the brightness according to the video content and ambient light conditions to ensure the visual effect of the image; repackages and combines the preprocessed video frames into a video stream, and transmits the preprocessed video data to the target analysis module and the anomaly detection module for subsequent analysis and processing.
[0103] In some embodiments, as Figure 2 shown, in the target analysis module, the process of separating dynamic targets is as follows:
[0104] Receive the pre - processed video data stream from the data stream processing module, decode the video data stream to obtain consecutive video image frames; use a convolutional neural network to learn the background frames in the video data stream, establish a static background model, and identify the static background; for the consecutive video image frames of the subsequent video data stream, compare them with the static background model, separate the dynamic objects through background difference method, and use an object tracking algorithm to track the dynamic objects, continuously identify the same object in the video data stream; adjust the detection result of the target object according to the illumination condition of the image through brightness normalization, estimate the illumination component and reflection component of the scene, perform adaptive processing of illumination, improve the illumination condition of the image, and use a deep learning model to analyze the partial occlusion situation of the target, identify the occlusion area in the video image frame, and restore the occluded target information; use a deep learning model to identify the separated target, determine the category and attributes of the target, extract the feature vector of the target, and output the analysis and recognition result to the anomaly detection module to provide a basis for further analysis and decision - making, where the feature vector of the target includes the motion feature and appearance feature of the target.
[0105] Further, in some embodiments, the expressions for estimating the illumination component and reflection component of the scene and performing adaptive processing of illumination are:
[0106] ;
[0107] where, represents the adjusted pixel value, represents the pixel value of the original image, represents the estimated reflection component.
[0108] Even further, the process of establishing a static background model and identifying the static background is as follows:
[0109] Select the initial background frame from the video data stream, and pre - process the selected background frame, including scaling, normalization, data augmentation, etc., to improve the generalization ability of the model;
[0110] Construct a CNN architecture for background learning, including multiple convolutional layers, activation layers, pooling layers and fully - connected layers. Among them, the convolutional layer extracts image features through filters, and there are multiple filters, each filter is responsible for extracting specific features in the image, the activation layer introduces non - linearity to enable the model to learn complex features, the pooling layer is used to reduce the spatial dimension of the features and increase the invariance to position changes, and the fully - connected layer is used to map the features extracted by the convolutional layer and pooling layer to the final output;
[0111] Define a loss function to measure the difference between the model prediction and the actual label, input the pre - processed background frame into the CNN, calculate the output through forward propagation, calculate the loss, and update the network weights through backpropagation;
[0112] Compare the trained background model with the new frame and detect dynamic objects through background subtraction method;
[0113] The expression of the convolutional layer is:
[0114] ;
[0115] Among them, represents the output of the th layer at position , represents the input of the th layer at position , represents the convolutional kernel weight of the th layer, and represent the offsets of the convolutional kernel in the horizontal and vertical directions, represents the bias term of the th layer, represents the activation function, including ReLU , Sigmoid , Tanh Any one of them. It should be noted that the convolutional layer extracts useful information by learning the features of the input data and maps these features to a higher-level representation;
[0116] The pooling layer selects max pooling, and its expression is:
[0117] ;
[0118] Among them, represents the pooling output of the th layer at position , represents the input of the th layer at position , and represent the position offsets within the pooling window. It should be noted that the pooling layer increases the invariance to position changes by reducing the spatial dimension of the features, while reducing the computational amount and preventing overfitting.
[0119] In addition, as Figure 3 shown, the process of continuously identifying the same object in the video data stream is:
[0120] Combine the trained static background model, update the background model for each new video image, and calculate the difference between the current frame and the background model to obtain the difference image;
[0121] Set the dynamic analysis threshold, analyze the correlation between the difference image and the dynamic analysis threshold. Pixels above the dynamic analysis threshold are considered dynamic targets, and a foreground mask is generated to isolate the dynamic targets;
[0122] Perform connected component analysis on the foreground mask to identify independent objects in the foreground region, detect the target region in the foreground mask, extract the features of the target, and use the feature matching algorithm of Kalman filter to match the target in consecutive frames;
[0123] Track the Kalman filter to update the position and features of the target, enabling continuous identification of the same target in the video sequence.
[0124] The update formula for the background model is:
[0125] ;
[0126] where, represents the pixel value of the background model, represents the pixel value of the current frame, represents the update factor ;
[0127] The calculation expression for the difference image is:
[0128] ;
[0129] where, represents the difference result, represents the pixel value of the current frame, represents the pixel value of the background model;
[0130] The judgment expression for generating the foreground mask is:
[0131] ;
[0132] where, represents the pixel value of the foreground mask, represents the dynamic analysis threshold.
[0133] In some embodiments, as Figure 4 shown, in the anomaly detection module, the detection process of abnormal behaviors or events in the video is:
[0134] Using the detection results output by the target analysis module, determine the target objects in the video, and extract the feature vectors of the targets, including the motion features and appearance features of the targets. By comparing the distances of the feature vectors, match the target features in consecutive frames. Among them, the appearance features include color histograms, texture features, and shape descriptors, and the motion features include the speed, acceleration, and trajectory of the targets; Use machine learning to model normal behaviors, construct a behavior model, identify abnormal behaviors by comparing the real-time video content with the predefined behavior model, calculate the deviation between the real-time behavior and the normal model, and generate an anomaly score; Set a threshold for the anomaly score according to historical data. When the anomaly score exceeds the threshold, it is determined as an abnormal behavior. And in multi-object tracking, remove the overlapping detection results and retain the best prediction results. Once an abnormal behavior is detected, generate an alarm message and send the alarm message to relevant personnel, providing detailed information about the abnormal behavior, including the time, location, and involved targets, for quick response and handling.
[0135] Furthermore, the calculation process of the anomaly score is as follows:
[0136] Collect video data of normal behaviors, annotate them, and extract the appearance features and motion features of the targets;
[0137] Select the SVM machine learning model and use the annotated normal behavior video data to train the model to distinguish normal behaviors from abnormal behaviors;
[0138] Extract the features of the targets in the real-time video to obtain the feature vectors of the real-time behavior. Input the feature vectors of the real-time behavior into the trained model for analysis, calculate the deviation between the real-time behavior features and the normal model, and calculate the anomaly score according to the deviation;
[0139] Even further, the expression for behavior model training is:
[0140] ;
[0141] Among them, represents the decision function, represents the test sample, represents the kernel function, calculating the similarity between the sample and the test sample ; represents the th feature vector of the th training sample, represents the label of the th training sample, referring to normal behavior or abnormal behavior, represents the coefficient of the th support vector,
[0142] The recognition of abnormal behavior is achieved by calculating the deviation between real-time behavior features and the normal model, in conjunction with the cosine similarity judgment. Its expression is:
[0143] ;
[0144] wherein, represents the cosine similarity, which measures the similarity degree between two vectors, and its value range is from -1 to 1. represents the target feature vector extracted from the real-time video frame. represents the feature vector of the normal behavior model. and respectively represent the and th eigenvalue in the vectors . It should be noted that represents the dimension of the feature vector, that is, the total number of eigenvalues in the vector, which is used to describe the total feature quantity of the target feature. The closer the value of is to 1, the more similar the two vectors are, that is, the closer the target behavior is to the normal behavior model;
[0145] The expression for calculating the abnormal score is:
[0146] ;
[0147] wherein, represents the abnormal score, represents the cosine similarity.
[0148] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
Claims
1. An intelligent processing system for monitoring video information based on cloud computing, including a monitoring information processing platform, characterized in that: The monitoring information processing platform is communicatively connected with a video data acquisition module, a data transmission module, a cloud data storage module, a data stream processing module, a target analysis module, an anomaly detection module, and an intelligent alarm management module; The video data acquisition module is used to acquire monitoring video data from multiple cameras and sensors. The cameras are distributed in different monitoring areas to capture video signals in real time; The data transmission module is used to transmit the acquired video data to the cloud server through a wireless network; The cloud data storage module is used to store the video data transmitted to the cloud server by utilizing the distributed storage capacity of the cloud computing platform, and perform classification, indexing, and retrieval management operations on the stored video data; The data stream processing module is used to perform preprocessing operations on the video data to improve the image quality; The target analysis module is used to analyze the video data by using deep learning technology, identify and remove the static background in the video data, separate the dynamic targets, and perform adaptive processing on environmental factors such as light changes and occlusions. In the target analysis module, the process of separating dynamic targets is as follows: Receive the preprocessed video data stream from the data stream processing module, and decode the video data stream to obtain continuous video image frames; Use a convolutional neural network to learn the background frames in the video data stream, establish a static background model, and identify the static background; For the subsequent continuous video image frames of the video data stream, compare them with the static background model, separate the dynamic targets by the background difference method, and use a target tracking algorithm to track the dynamic targets, and continuously identify the same target in the video data stream; Adjust the detection result of the target object according to the illumination condition of the image through brightness normalization, estimate the illumination component and reflection component of the scene, perform adaptive processing of illumination, improve the illumination condition of the image, and use a deep learning model to analyze the partial occlusion situation of the target, identify the occlusion area in the video image frame, and restore the occluded target information; Use a deep learning model to identify the separated targets, determine the category and attributes of the targets, and extract the feature vectors of the targets, and output the analysis and identification results to the anomaly detection module. Among them, the feature vectors of the targets include the motion features and appearance features of the targets; The anomaly detection module is used to perform real-time tracking on the target objects in the video, detect abnormal behaviors or events in the video through the analysis of the video content, and discover abnormal situations; The intelligent alarm management module is used to generate alarm information according to the anomaly detection results and preset rules, and manage the alarm process.
2. The intelligent processing system for monitoring video information based on cloud computing according to claim 1, wherein: In the video data acquisition module, the acquisition process of the monitoring video data is as follows: Deploy cameras and sensors in different monitoring areas, and use the cameras to capture video signals in the monitoring areas in real time, and cooperate with the sensors to capture sensor data in the monitoring areas. Among them, the sensors include infrared sensors and motion sensors; Perform preliminary processing on the captured video signals, perform automatic exposure adjustment and white balance adjustment, and then convert the preliminarily processed video signals into digital signals and perform encoding and compression; Through the wireless network, the encoded video data and sensor data are encapsulated into network data packets and transmitted to the data transmission module.
3. The intelligent processing system for monitoring video information based on cloud computing according to claim 1, characterized in that: In the cloud data storage module, the process of storing video data is as follows: The cloud server receives the video data and sensor data from the data transmission module, and performs preliminary processing on the received video data and sensor data; Perform data sharding processing on the video data and sensor data, divide the video data and sensor data into multiple data blocks, and use redundant storage technology to copy each data block to multiple storage nodes for distributed storage on multiple storage nodes; Analyze the video data after data sharding processing, classify it according to the content, source, and time attributes of the video data, and build an index for the classified video data to establish index information. The index information includes keyword index, time index, and metadata index; Set up a user interface that enables users to retrieve in multiple ways such as time range, camera location, and event type by using the index and metadata, quickly retrieve the required video clips, and provide preview and extraction functions for the video clips, enabling users to select and extract the required video clips according to their needs.
4. The intelligent processing system for monitoring video information based on cloud computing according to claim 1, characterized in that: In the data stream processing module, the process of preprocessing video data is as follows: The data stream processing module receives the video data stream and sensor data from the cloud server, and decodes the encoded and compressed video data to restore it to the original video format; Extract individual video frames from the video stream and apply a smoothing filter to reduce noise and artifacts in the image; If there is a blur problem in the video data, use a blind deblurring algorithm to estimate the blur kernel through an iterative optimization algorithm and restore a clear image; Enhance the contrast of the image by adjusting the contrast of the video frame, optimize the visual effect of the image, and automatically adjust the brightness according to the video content and ambient light conditions; Re-encapsulate and combine the preprocessed video frames into a video stream, and transmit the preprocessed video data to the target analysis module and anomaly detection module for subsequent analysis and processing.
5. The intelligent processing system for monitoring video information based on cloud computing according to claim 1, characterized in that: The process of establishing a static background model and identifying the static background is as follows: Select an initial background frame from the video data stream and perform preprocessing on the selected background frame; Construct a CNN architecture for background learning, including multiple convolutional layers, activation layers, pooling layers, and fully connected layers. Among them, the convolutional layer extracts image features through filters, and there are multiple filters, and each filter is responsible for extracting specific features in the image; Define a loss function to measure the difference between the model prediction and the actual label, input the preprocessed background frame into the CNN, calculate the output through forward propagation, calculate the loss, and update the network weights through backpropagation; Use the trained background model to compare with new frames and detect dynamic targets through background difference method; The expression of the convolutional layer is: Among them, represents the output of the l-th layer at position (i, j), represents the input of the (l - 1)-th layer at position (i + m, j + n), represents the convolutional kernel weight of the l-th layer, m and n represent the offsets of the convolutional kernel in the horizontal and vertical directions, b l represents the bias term of the l-th layer, and f represents the activation function; The pooling layer selects max pooling, and its expression is: Among them, represents the pooling output of the l-th layer at position (i, j), represents the input of the (l - 1)-th layer at position (i + k, j + l), where k and l represent the position offsets within the pooling window.
6. The intelligent processing system for monitoring video information based on cloud computing according to claim 1, characterized in that: The process of continuously identifying the same target in the video data stream is as follows: Combined with the trained static background model, for each new video image frame, update the background model, calculate the difference between the current frame and the background model to obtain a difference image; Set the dynamic analysis threshold, analyze the correlation between the difference image and the dynamic analysis threshold. Pixels above the dynamic analysis threshold are considered dynamic targets, and a foreground mask is generated to isolate the dynamic targets; Perform connected component analysis on the foreground mask to identify independent objects in the foreground region, detect the target region in the foreground mask, extract the features of the target, and use the feature matching algorithm of Kalman filter to match the target in consecutive frames; Track the Kalman filter to update the position and features of the target, so as to continuously identify the same target in the video sequence; The update formula of the background model is: B ij ← αB ij +(1 - α)I ij ; Among them, B ij represents the pixel value of the background model, I ij represents the pixel value of the current frame, and α represents the update factor where 0 < α < 1; The calculation expression of the difference image is: D ij = |I ij - B ij |; Among them, D ij represents the differential result, I ij represents the pixel value of the current frame, B ij represents the pixel value of the background model; The judgment expression for generating the foreground mask is: where M ij represents the pixel value of the foreground mask, and T represents the dynamic analysis threshold.
7. The intelligent processing system for monitoring video information based on cloud computing according to claim 1, wherein: In the anomaly detection module, the detection process of abnormal behaviors or events in the video is: Utilize the detection results output by the target analysis module to determine the target objects in the video, and extract the feature vectors of the targets, including the motion features and appearance features of the targets. By comparing the distances of the feature vectors, match the target features in consecutive frames, where the appearance features include color histograms, texture features, and shape descriptors, and the motion features include the speed, acceleration, and trajectory of the targets; Use machine learning to model normal behaviors, construct a behavior model, identify abnormal behaviors by comparing the real-time video content with the predefined behavior model, calculate the deviation between the real-time behavior and the normal model, and generate an anomaly score; Set the threshold of the anomaly score according to historical data. When the anomaly score exceeds the threshold, it is determined as an abnormal behavior. And in multi-object tracking, remove the overlapping detection results and retain the best prediction result; Once an abnormal behavior is detected, generate an alarm message, send the alarm message to relevant personnel, and provide detailed information about the abnormal behavior, including the occurrence time, location, and involved targets.
8. The intelligent processing system for monitoring video information based on cloud computing according to claim 7, wherein: The calculation process of the anomaly score is: Collect video data of normal behaviors, annotate them, and extract the appearance features and motion features of the targets; Select the SVM machine learning model and use the annotated video data of normal behaviors to train the model to distinguish normal behaviors from abnormal behaviors; Extract the features of the targets in the real-time video to obtain the feature vectors of the real-time behavior. Input the feature vectors of the real-time behavior into the trained model for analysis, calculate the deviation between the real-time behavior features and the normal model, and calculate the anomaly score according to the deviation.
9. The intelligent processing system for monitoring video information based on cloud computing according to claim 8, characterized in that: The expression for training the behavior model is: Among them, f(x) represents the decision function, x represents the test sample, K(x i , x) represents the kernel function, which calculates the similarity between the sample x i and the test sample x, x i represents the feature vector of the i-th training sample, y i represents the label of the i-th training sample, indicating normal behavior or abnormal behavior, α i represents the coefficient of the i-th support vector, and b represents the bias term, which is a constant term in the SVM model; The identification of abnormal behaviors is achieved by calculating the deviation between the real-time behavior features and the normal model and cooperating with the cosine similarity judgment. Its expression is: Among them, S represents the cosine similarity, which measures the similarity degree of two vectors and ranges from -1 to 1. s represents the target feature vector extracted from the real-time video frame, and d represents the feature vector of the normal behavior model. s t and d t respectively represent the t-th eigenvalue in the vectors s and d. z represents the dimension of the feature vector, that is, the total number of eigenvalues in the vector, which is used to describe the total feature quantity of the target feature; The expression for calculating the anomaly score is: Score = 1 - S; Where Score represents the anomaly score and S represents the cosine similarity.
Citation Information
Patent Citations
Smart city monitoring system and method based on AI
CN117312801A
Indoor surveillance system and indoor surveillance method
US20140055610A1