Video monitoring system based on AI
By introducing a long and short-term memory network timing model and decision-making response module into the video surveillance system, the problem of existing systems not being able to identify and classify abnormal behaviors of videos and setting response priorities is solved, and efficient decision-making response and early warning push is achieved.
Patent Information
- Application Number
- CN202510290168.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI-based video surveillance systems cannot build behavior abnormality recognition models based on data characteristics in video images, classify abnormal behaviors in videos, and cannot set response priority for abnormal behaviors, resulting in low efficiency in decision response and early warning push.
An AI-based video surveillance system is designed, including front-end acquisition components, back-end processing module, decision-making response module, interactive supervision platform and data management module. The system constructs a timing model through long and short-term memory networks, identifies and classifies abnormal behaviors, and sets response priority according to the urgency level to achieve real-time early warning push.
It realizes the rapid identification and classification of video behavior, improves decision response speed and early warning push efficiency, and ensures the security and accuracy of the system.
Smart Images

Figure CN120220058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security monitoring, and specifically to an AI-based video monitoring system. Background Art
[0002] AI video monitoring systems are widely used in multiple industries such as retail, security, healthcare, education, and transportation. By combining artificial intelligence and video monitoring technologies, they can identify abnormal behaviors in real time, enhance security, optimize operational efficiency, and help enterprises make accurate decisions, not only improving the operational efficiency of enterprises but also enhancing the customer experience and security.
[0003] After retrieval, a patent for an invention with the Chinese patent publication number CN118155030A discloses an AI-based video monitoring system and method. This video monitoring system can clearly determine abnormal features by constructing a box plot, thereby promptly discovering abnormalities in video monitoring, effectively preventing the occurrence of emergencies, preventing losses, and through constructing a growth channel, it can perform self-learning based on the number of occurrences of abnormal features. By constructing an isolation tree and leaf nodes, corresponding adjustments can be made to the monitoring devices according to abnormal features, improving the security of the monitoring devices and reducing the misjudgment probability;
[0004] However, this AI-based video monitoring system cannot construct a behavior anomaly recognition model based on the data features in the video image to classify abnormal behaviors in the video, and at the same time, it cannot set response priorities for abnormal behaviors. Therefore, there is still room for further improvement in terms of improving the system's decision-making response and early warning push efficiency. Thus, an AI-based video monitoring system is proposed. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the deficiencies of the prior art, the present invention provides an AI-based video monitoring system, which has the advantages of automatically identifying and analyzing abnormal behavior patterns and setting response priorities to improve the decision-making response and early warning push efficiency, and solves the problem that the AI-based video monitoring system and method in the above background art cannot set response priorities based on a behavior anomaly recognition model and needs to further improve the decision-making response and early warning push efficiency.
[0007] (2) Technical Solutions
[0008] To achieve the above purpose of automatically identifying and analyzing abnormal behavior patterns and setting response priorities to improve the decision-making response and early warning push efficiency, the present invention provides the following technical solution: An AI-based video monitoring system includes a front-end acquisition component for real-time monitoring and video data acquisition, and through preliminary data processing to obtain a high-resolution video stream;
[0009] The back-end processing module is used to receive the transmitted data from the front-end and analyze it in real time, and detect and identify abnormal behaviors in a timely manner;
[0010] The decision-making response module is used to classify events for different types of abnormal behaviors and set response priorities according to the urgency;
[0011] The interactive supervision platform is used to automatically execute response actions according to the event type and priority, and interact with the user side to push the early warning mechanism in real time;
[0012] The data control module is used to store and supervise the data after encryption processing to prevent unauthorized access and data leakage.
[0013] Preferably, the front-end acquisition component includes a video data acquisition layer and a data preprocessing layer. The video data acquisition layer is composed of a high-definition camera and an edge computing node device. The edge computing node device receives video stream data from the data transmission optical fiber according to the streaming media protocol. The data preprocessing layer includes data cleaning and image enhancement processing. The specific steps for cleaning the video stream data are as follows:
[0014] 1) Missing value processing: a. Directly delete the sample data with a large amount of missing values; b. Fill the sample data with a small amount of missing values according to the mean interpolation method;
[0015] 2) Outlier processing: a. Directly delete the outliers with a small number; b. Process the outliers with a large amount of data according to the 3σ principle;
[0016] 3) Duplicate value processing: a. Directly delete the duplicate data with a small number; b. Merge and mark the duplicate data containing valid information.
[0017] Preferably, the image data after data cleaning is enhanced based on the Retinex algorithm. The specific steps are as follows:
[0018] 1) The Retinex algorithm formula is I(x,y) = R(x,y)L(x,y), where I(x,y) is the image seen by the human eye, R(x,y) is the reflection component of the object, L(x,y) is the ambient light illumination component, and (x,y) represents the position corresponding to the two-dimensional image;
[0019] 2) Estimate the value of L(x,y) and obtain R(x,y) by performing convolution operations on I(x,y) through Gaussian blur. It is expressed by the formula:
[0020] log(R) = log(I) - log(L),
[0021] L = G * I;
[0022] Among them, G is the filter for Gaussian blur, and * represents the convolution operation;
[0023] 3) The filter calculation formula is where σ is the Gaussian surrounding space constant. For a two-dimensional image r 2 equals the corresponding position, which can be expressed as x 2 +y 2 , and generally, it is considered that the illumination component L(x, y) is the result of the original image after Gaussian filtering;
[0024] 4) The weight calculation formula is:
[0025] log(R) = log(R) + Weight(i) * (log(I i ) - log(L i ))
[0026] where Weight(i) represents the weight corresponding to each scale, and it is required that the sum of the weights of each scale must be 1. The classic value is equal weights;
[0027]
[0028] In the above formula, I i is the original input image, G n is the filtering function, N is the number of scales, w n is the weight of each scale, generally expressed as 1 / N, and R MSR represents the output after image enhancement in the logarithmic domain.
[0029] Preferably, the backend processing module includes a data optimization layer and an anomaly detection layer. The image data is optimized in the data optimization layer, specifically including:
[0030] a. Dynamic range optimization: Based on automatic exposure adjustment and gamma correction technologies to ensure clear images can be obtained under different lighting conditions;
[0031] b. Image stabilization and alignment: Improve the video stability through image registration technology;
[0032] The anomaly detection layer includes a target detection unit, a behavior recognition unit, and an anomaly detection unit. Behavioral anomaly detection and behavior type analysis are performed on the image data in the anomaly detection layer. The target detection unit detects targets based on the YOLO algorithm to identify people and objects in video images. The specific steps include:
[0033] 1) Image segmentation: Divide the input image into a fixed-size N×N grid, and each grid is responsible for detecting the targets within the delimited area to adapt to the network input and capture the subtle features in the image;
[0034] 2) Feature extraction: Extract image features through a convolutional neural network, where the convolutional neural network contains 24 convolutional layers and 2 fully connected layers, and classify the objects within each bounding box through the fully connected layer;
[0035] 3) Bounding box prediction: Predict multiple bounding boxes, each represented by the center coordinates (x, y), width w, and height h, and contain a confidence score to indicate whether there is an object within the bounding box and its accuracy;
[0036] 4) Class probability prediction: Each grid also predicts the probability of each class, that is, the probability that the target belongs to each class given that the bounding box contains the target;
[0037] 5) Non-maximum suppression: Use non-maximum suppression to process multiple overlapping bounding boxes to ensure that each target retains one best bounding box.
[0038] Preferably, the behavior recognition unit includes:
[0039] 1) Label behavior information: Identify human body information based on the YOLO algorithm, and accurately label the behaviors in the video, including behavior categories, start time, and end time information;
[0040] 2) Extract spatio-temporal features: Use a convolutional neural network to extract spatio-temporal features, and extract key features from the video images to reflect the behavior patterns, including features such as shape and texture in space, as well as motion information and change trends in time;
[0041] 3) Build a time series model: Build a time series model based on the long short-term memory network, and input the extracted spatio-temporal features into the model to capture the time-dependent relationships and sequence information between video frames;
[0042] 4) Model training: Train the time series model according to the labeled behavior data, and optimize the parameters of the model through the backpropagation algorithm to make the model accurately predict the occurrence probability of behavior categories or behavior patterns in the video;
[0043] 5) Behavior recognition and classification: Input the video to be recognized into the trained time series model, and the model will output a prediction result, indicating the probability distribution of the behavior categories or behavior patterns that may be included in the video. At the same time, according to the prediction result of the model, select the behavior category with the highest probability as the final recognition result to determine the continuous behavior pattern in the video;
[0044] The anomaly detection unit establishes a normal behavior model based on the training set, classifies and marks behaviors that deviate from the normal pattern. The abnormal behavior patterns include physical conflicts, abnormal movements, and abnormal behaviors in emotional fluctuations.
[0045] Preferably, the anomaly detection unit and the behavior recognition unit classify events for different types of abnormal behaviors, which include violent behaviors, personnel gathering, and abnormal movement, and set the priority of event response according to the urgency. The interactive supervision platform, based on the real-time push and alarm mechanism, automatically sends an alarm notification to the user terminal according to the identified abnormal behavior and response priority. The alarm mechanism includes remote alarm, sound prompt, and early warning push to a third-party platform.
[0046] Preferably, the data control module includes a local database and a private cloud database, specifically including:
[0047] 1) Privacy supervision: Set the data privacy supervision level. For data with a higher supervision level, use the synchronous mode of local storage and private cloud deployment to ensure data security and controllability;
[0048] 2) Video encryption and data protection: Use the AES encryption algorithm to encrypt the stored video data to ensure data privacy and security;
[0049] 3) Redundant backup and disaster recovery solution: Support local or distributed storage solutions to ensure the high availability and disaster tolerance of video data;
[0050] 4) Compliance requirements: Ensure that the storage and use of monitoring data comply with local laws, regulations, and data protection regulations.
[0051] (III) Beneficial effects
[0052] Compared with the prior art, the present invention provides an AI-based video monitoring system with the following
[0053] beneficial effects:
[0054] 1. For the AI-based video monitoring system, by constructing a time series model based on the long short-term memory network to capture the temporal dependence and sequence information between video frames, annotating video behavior data and using the annotated data to train the model, identifying and classifying video behaviors, and establishing a normal behavior model to classify and mark abnormal behavior patterns, rapid identification and detection of abnormal behaviors can be achieved.
[0055] 2. For the AI-based video monitoring system, by classifying events for different types of abnormal behaviors, setting the priority of event response according to the urgency, and the interactive supervision platform automatically sending an alarm notification based on the real-time push and alarm mechanism and according to the identified abnormal behavior and response priority, hierarchical and rapid early warning can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of the system module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments and drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Please refer to Figure 1 , a video surveillance system based on AI, including a front-end acquisition component for real-time monitoring and video data acquisition, and high-resolution video streams through preliminary data processing;
[0059] A back-end processing module for receiving transmission data from the front-end and performing real-time analysis to detect and identify abnormal behaviors in a timely manner;
[0060] A decision response module for classifying events for different types of abnormal behaviors and setting response priorities according to the urgency;
[0061] An interactive supervision platform for automatically executing response actions according to the event type and priority, and interacting with the user side to push the early warning mechanism in real time;
[0062] A data control module for storing and supervising the data after encryption processing to prevent unauthorized access and data leakage.
[0063] Furthermore, the front-end acquisition component includes a video data acquisition layer and a data preprocessing layer. The video data acquisition layer is composed of high-definition cameras and edge computing node devices. The edge computing node devices receive video stream data from the data transmission optical fiber according to the streaming media protocol. The data preprocessing layer includes data cleaning and image enhancement processing. The specific steps for cleaning the video stream data include:
[0064] 1) Missing value processing: a. Directly deleting the sample data with a large amount of missing values; b. Filling the sample data with a small amount of missing values according to the mean interpolation method;
[0065] 2) Outlier processing: a. Directly deleting the outliers with a small number; b. Processing the outliers with a large amount of data according to the 3σ principle;
[0066] 3) Duplicate value processing: a. Directly deleting the duplicate data with a small number; b. Merging and marking the duplicate data containing valid information.
[0067] Specifically, mean imputation is to use the sample data mean or mode as its replacement value to interpolate the data, so as to fill the missing values of the data;
[0068] The 3σ principle method for dealing with outliers is to first calculate the mean μ, and then calculate the standard deviation of the data, expressed as According to the formula Perform data transformation processing. At this time, the mean of the transformed data is 0 and the standard deviation is 1. Construct a normal distribution graph and set the intervals (μ - σ, μ + σ), (μ - 2σ, μ + 2σ), and (μ - 3σ, μ + 3σ). At this time, the outliers outside the interval (μ - 3σ, μ + 3σ) are regarded as missing value data and processed according to the missing value method.
[0069] Furthermore, the image data after data cleaning is enhanced based on the Retinex algorithm. The specific steps include:
[0070] 1) The Retinex algorithm formula is I(x, y) = R(x, y)L(x, y), where I(x, y) is the image seen by the human eye, R(x, y) is the reflection component of the object, L(x, y) is the ambient light illumination component, and (x, y) represents the position corresponding to the two-dimensional image;
[0071] 2) Calculate R(x, y) by estimating the value of L(x, y) through Gaussian blur and performing convolution operation on I(x, y), which is expressed by the formula:
[0072] log(R) = log(I) - log(L),
[0073] L = G * I;
[0074] where G is the filter for Gaussian blur, and * represents convolution operation;
[0075] 3) The filter calculation formula is where σ is the Gaussian surrounding space constant. For a two-dimensional image r 2 equals the corresponding position, which can be expressed as x 2 + y 2 , and it is generally considered that the illumination component L(x, y) is the result of the original image after Gaussian filtering;
[0076] 4) The weight calculation formula is:
[0077] log(R) = log(R) + Weight(i) * (log(I i ) - log(L i ))
[0078] where Weight(i) represents the weight corresponding to each scale, and it is required that the sum of the weights of each scale must be 1. The classic value is equal weight;
[0079]
[0080] In the above formula, Ii is the original input image, G n is the filtering function, N is the number of scales, w n is the weight for each scale, generally expressed as 1 / N, R MSR represents the output after image enhancement in the logarithmic domain.
[0081] Specifically, first perform a convolution operation on I(x, y) and obtain R(x, y) through Gaussian blur, then filter the image according to the obtained filter, calculate each weight respectively, and output the enhanced image.
[0082] Furthermore, the backend processing module includes a data optimization layer and an anomaly detection layer. The image data is optimized in the data optimization layer, specifically including:
[0083] a. Dynamic range optimization: Based on automatic exposure adjustment and gamma correction techniques to ensure clear images can be obtained under different lighting conditions;
[0084] b. Image stabilization and alignment: Improve video stability through image registration techniques;
[0085] The anomaly detection layer includes an object detection unit, a behavior recognition unit, and an anomaly detection unit. Behavioral anomaly detection and behavior type analysis are performed on the image data in the anomaly detection layer. The object detection unit detects objects based on the YOLO algorithm to identify people and objects in video images. The specific steps include:
[0086] 1) Image segmentation: Divide the input image into a fixed-size N×N grid, and each grid is responsible for detecting objects within the delimited area to adapt to the network input and capture subtle features in the image;
[0087] 2) Feature extraction: Extract image features through a convolutional neural network, where the convolutional neural network contains 24 convolutional layers and 2 fully connected layers, and classify the objects within each bounding box through the fully connected layer;
[0088] 3) Bounding box prediction: Predict multiple bounding boxes, each bounding box is represented by the center coordinates (x, y), width w, and height h, and contains a confidence score, which is used to indicate whether there is an object within the bounding box and its accuracy;
[0089] 4) Class probability prediction: Each grid also predicts the probability of each class, that is, the probability that the object belongs to each class given that the bounding box contains the target;
[0090] 5) Non-maximum suppression: Use non-maximum suppression to process multiple overlapping bounding boxes to ensure that each target retains one best bounding box.
[0091] Specifically, through the data optimization layer, based on exposure adjustment and gamma correction techniques, the clarity of the image is improved. Then, through the registration technique, the image is registered and processed to improve the image stability. Based on the YOLO algorithm, segmentation and feature extraction processing are performed on the image. Then, probability prediction is carried out for each target category in each bounding box to detect the target type of the image, so as to identify people and objects in the video image.
[0092] Furthermore, the behavior recognition unit includes:
[0093] 1) Label behavior information: Based on the YOLO algorithm, human body information is recognized, and the behaviors in the video are accurately labeled, including behavior categories, start time, and end time information;
[0094] 2) Extract spatio-temporal features: Use a convolutional neural network to extract spatio-temporal features, and extract key features from the video image to reflect the behavior pattern, including features such as shape and texture in space, as well as motion information and change trends in time;
[0095] 3) Build a time series model: Based on the long short-term memory network, build a time series model, and input the extracted spatio-temporal features into the model to capture the time-dependent relationship and sequence information between video frames;
[0096] 4) Model training: Train the time series model according to the labeled behavior data, and optimize the parameters of the model through the backpropagation algorithm to make the model accurately predict the occurrence probability of behavior categories or behavior patterns in the video;
[0097] 5) Behavior recognition and classification: Input the video to be recognized into the trained time series model. The model will output a prediction result, indicating the probability distribution of behavior categories or behavior patterns that may be included in the video. At the same time, according to the prediction result of the model, select the behavior category with the highest probability as the final recognition result, so as to determine the continuous behavior pattern in the video;
[0098] The anomaly detection unit establishes a normal behavior model based on the training set, classifies and marks behaviors that deviate from the normal pattern. The abnormal behavior patterns include physical conflicts, abnormal movements, and abnormal behaviors in emotional fluctuations.
[0099] Specifically, after recognizing the human body information in the image based on the YOLO algorithm, extract the key features in the human body information that can reflect the behavior pattern, use the spatio-temporal features to build a time series model, and use the labeled behavior data to train the model, so that the model can accurately predict the behavior categories and behavior pattern probabilities in the behavior data.
[0100] Furthermore, based on the anomaly detection unit and the behavior recognition unit, event classification is performed on different types of abnormal behaviors. The abnormal behaviors include violent behaviors, personnel gathering, and abnormal movement. The priority of event response is set according to the urgency level. The interactive supervision platform, based on the real-time push and alarm mechanism, automatically sends alarm notifications to the user terminal according to the identified abnormal behaviors and response priorities. The alarm mechanism includes remote alarm, sound prompt, and early warning push from third-party platforms.
[0101] Specifically, event classification is carried out through different types of abnormal behaviors. According to whether there is abnormal movement of personnel, for abnormal movement of personnel, whether there is personnel gathering, and for personnel gathering, whether there is violent behavior is judged, so as to classify and identify different behavior types;
[0102] The priority of event response is determined according to the priority level of each behavior type. The interactive supervision platform realizes the real-time push and alarm mechanism according to the priority of event response, and automatically sends alarm notifications to the user terminal. The alarm mechanism includes remote alarm, in-range sound prompt, and push of early warning information by means such as WeChat, SMS, and email.
[0103] Furthermore, the data control module includes a local database and a private cloud database, specifically including;
[0104] 1) Privacy supervision: Set the data privacy supervision level. For data with a higher supervision level, adopt the synchronous mode of local storage and private cloud deployment to ensure data security and controllability;
[0105] 2) Video encryption and data protection: Use the AES encryption algorithm to encrypt the stored video data to ensure data privacy and security;
[0106] 3) Redundant backup and disaster recovery solution: Support local or distributed storage solutions to ensure the high availability and disaster recovery ability of video data;
[0107] 4) Compliance requirements: Ensure that the storage and use of monitoring data comply with local laws, regulations, and data protection regulations.
[0108] Specifically, the video stream data is encrypted through the AES encryption algorithm, and both local and distributed storage solutions are supported to ensure data security and controllability.
[0109] To sum up, this AI-based video surveillance system constructs a time series model based on the long short-term memory network to capture the temporal dependence relationship and sequence information between video frames, annotates the video behavior data and uses the annotated data to train the model, identifies and classifies the video behaviors, and classifies and marks the abnormal behavior patterns by establishing a normal behavior model, so as to realize the rapid identification and detection of abnormal behaviors;
[0110] By classifying events for different types of abnormal behaviors and setting the priorities of event responses according to the urgency level, the interactive supervision platform automatically sends alarm notifications based on the real-time push and alarm mechanism and according to the identified abnormal behaviors and response priorities, so as to achieve hierarchical and rapid early warning.
[0111] All relevant modules involved in this system are hardware system modules or functional modules that combine computer software programs or protocols in the prior art with hardware. The computer software programs or protocols themselves involved in this functional module are all well-known technologies to those skilled in the art, and they are not the improvements of this system; the improvement of this system lies in the interaction relationship or connection relationship between each module, that is, the overall structure of the system is improved to solve the corresponding technical problems to be solved by this system.
[0112] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-based video surveillance system, characterized in that: It includes front-end acquisition components for real-time monitoring and video data acquisition, and high-resolution video streams through preliminary data processing; The back-end processing module is used to receive the transmission data from the front-end and analyze it in real time to detect and identify abnormal behaviors in a timely manner; The decision response module is used to classify events according to different types of abnormal behaviors and set response priorities based on the urgency; Interactive supervision platform, used to automatically execute response actions according to event type and priority, and interact with user-side information to push early warning mechanisms in real time; The data control module is used to encrypt and store data for supervision to prevent unauthorized access and data leakage.
2. The AI-based video surveillance system according to claim 1, characterized in that: The front-end acquisition component includes a video data acquisition layer and a data preprocessing layer. The video data acquisition layer is composed of a high-definition camera and an edge computing node device. The edge computing node device receives video stream data from the data transmission optical fiber according to the streaming media protocol. The data preprocessing layer includes data cleaning and image enhancement processing. The specific steps of cleaning the video stream data include: 1) Missing value processing: a. directly delete the sample data with large missing values; b. fill in the sample data with small missing values using the mean interpolation method; 2) Outlier processing: a. Delete outliers with a small number directly; b. Process outliers with large data according to the 3σ principle; 3) Duplicate value processing: a. Directly delete the duplicate data with a small amount; b. Merge and mark the duplicate data containing valid information.
3. The AI-based video surveillance system according to claim 2, characterized in that: The cleaned image data is enhanced based on the Retinex algorithm. The specific steps include: 1) The formula of the Retinex algorithm is I(x,y)=R(x,y)L(x,y), where I(x,y) is the image seen by the human eye, R(x,y) is the reflection component of the object, L(x,y) is the ambient light illumination component, and (x,y) represents the position corresponding to the two-dimensional image; 2) According to the estimated value of L(x,y), Gaussian blur and convolution operation are performed on I(x,y) to obtain R(x,y), which can be expressed as: log(R)=log(I)-log(L), L = G * I; Where G is the Gaussian blur filter, and * represents the convolution operation; 3) The filter calculation formula is: Where σ is the space constant around the Gaussian, for the two-dimensional image r 2 is equal to the corresponding position, which can be expressed as x 2 +y 2 , it is generally believed that the illumination component L(x,y) is the result of Gaussian filtering of the original image; 4) The weight calculation formula is: log(R)=log(R)+Weight(i)*(log(I i )-log(L i )), Among them, Weight(i) represents the weight corresponding to each scale, requiring that the sum of the weights of each scale must be 1, and the classic value is equal weight; In the above formula, I i is the original input image, G n is the filter function, N is the number of scales, w n is the weight of each scale, generally expressed as 1 / N, R MSR represents the output after image enhancement in the logarithmic domain.
4. The AI-based video surveillance system according to claim 1, characterized in that: The backend processing module includes a data optimization layer and an anomaly detection layer. The image data is optimized and processed in the data optimization layer, specifically including: a. Dynamic range optimization: Based on automatic exposure adjustment and gamma correction technology to ensure clear images under different lighting conditions; b. Image stabilization and alignment: Improve video stability through image registration technology; The anomaly detection layer includes a target detection unit, a behavior recognition unit, and an anomaly detection unit. The anomaly detection layer performs behavior anomaly detection and behavior type analysis on the image data. The target detection unit detects targets based on the YOLO algorithm to identify people and objects in the video image. The specific steps include: 1) Image segmentation: The input image is divided into a fixed-size N×N grid, each grid is responsible for detecting the object within the defined area to adapt to the network input and capture the subtle features in the image; 2) Feature extraction: Image features are extracted through a convolutional neural network, where the convolutional neural network contains 24 convolutional layers and 2 fully connected layers. The fully connected layers are used to classify objects within each bounding box. 3) Bounding box prediction: predict multiple bounding boxes, each of which is represented by the center coordinates (x, y), width w and height h, and contains a confidence score to indicate whether the bounding box contains an object and its accuracy; 4) Category probability prediction: Each grid also predicts the probability of each category, that is, the probability that the object belongs to each category given that the bounding box contains the object; 5) Non-maximum suppression: Use non-maximum suppression to process multiple overlapping bounding boxes to ensure that each object retains an optimal bounding box.
5. The AI-based video surveillance system according to claim 4, characterized in that: The behavior recognition unit comprises: 1) Labeling behavior information: Based on the YOLO algorithm, human body information is identified and the behaviors in the video are accurately labeled, including behavior category, start time, and end time information; 2) Extracting spatiotemporal features: Using convolutional neural networks to extract spatiotemporal features, the key features that reflect the behavior patterns are extracted from video images, including spatial shape, texture and other features, as well as temporal motion information and change trends; 3) Constructing a temporal model: A temporal model is constructed based on a long short-term memory network, and the extracted spatiotemporal features are input into the model to capture the temporal dependency and sequence information between video frames; 4) Model training: Train the time series model based on the labeled behavior data, and optimize the model parameters through the back propagation algorithm so that the model can accurately predict the occurrence probability of the behavior category or behavior pattern in the video; 5) Behavior recognition and classification: The video to be recognized is input into the trained time series model. The model will output a prediction result, which indicates the probability distribution of the behavior categories or behavior patterns that may be contained in the video. At the same time, based on the prediction result of the model, the behavior category with the highest probability is selected as the final recognition result, thereby determining the continuous behavior pattern in the video. The abnormality detection unit establishes a normal behavior model based on the training set, and classifies and marks behaviors that deviate from the normal pattern. The abnormal behavior pattern includes physical conflict, abnormal movement, and abnormal behavior in emotional fluctuations.
6. The AI-based video surveillance system according to claim 5, characterized in that: Based on the anomaly detection unit and the behavior recognition unit, different types of abnormal behaviors are classified into event categories, including violent behavior, crowd gathering and abnormal movement, and the priority of event response is set according to the degree of urgency. The interactive supervision platform is based on the real-time push and alarm mechanism. According to the identified abnormal behavior and response priority, it automatically sends an alarm notification to the user terminal. The alarm mechanism includes remote alarm, sound prompt and third-party platform push warning.
7. The AI-based video surveillance system according to claim 1, characterized in that: The data control module includes a localized database and a private cloud database, specifically including: 1) Privacy supervision: Set the data privacy supervision level, and use local storage and private cloud deployment synchronization mode for data with a higher supervision level to ensure data security and controllability; 2) Video encryption and data protection: Use the AES encryption algorithm to encrypt the stored video data to ensure data privacy and security; 3) Redundant backup and disaster recovery solutions: Support local or distributed storage solutions to ensure high availability and disaster recovery capabilities of video data; 4) Compliance requirements: Ensure that the storage and use of monitoring data complies with local laws, regulations and data protection regulations.
Citation Information
Patent Citations
Video monitoring system and method based on AI
CN118155030A
Cited By
Microgrid digital twin modeling method
CN120764117A
Family safety AI monitoring method and system based on multi-modal algorithm fusion
CN120976853A