AI video analysis monitoring early warning platform and early warning method thereof
By designing an AI video analysis monitoring and early warning platform, using spatiotemporal feature compression, multimodal information fusion and comparison learning algorithms, the problem of inconsistent detection performance under different lighting conditions is solved, efficient and accurate abnormal behavior detection and early warning is achieved, and false alarm rate and computing resource requirements are reduced.
Patent Information
- Application Number
- CN202510645644.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has inconsistent detection performance under different lighting conditions, high computing resources consumption, high false alarm rate, and difficult to meet the needs of 24-hour uninterrupted monitoring.
An AI video analysis monitoring and early warning platform was designed, including video acquisition module, feature processing module, multimodal fusion module, behavior mapping module and abnormal detection module. Through spatiotemporal feature compression, illumination constant feature extraction, multimodal information fusion, comparative learning algorithms and adaptive threshold determination, unified anomaly detection under different lighting conditions is achieved.
It realizes integrated day-night detection capabilities, reduces computing resource requirements, reduces false alarm rates, simplifies the system architecture, and reduces data requirements and maintenance costs.
Smart Images

Figure CN120182897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, computer vision, and video surveillance, and more specifically, it relates to an AI video analysis and monitoring warning platform and its warning method. Background Art
[0002] With the continuous improvement of public safety requirements, video surveillance systems have been widely deployed in important places such as urban public areas, transportation hubs, and commercial venues. Traditional video surveillance systems mainly rely on manual viewing of surveillance footage to identify abnormal situations, which have problems such as low efficiency, many missed reports, and high labor costs. In recent years, video analysis technologies based on artificial intelligence have gradually been applied to automatic abnormal behavior detection, but still face many challenges in practical applications.
[0003] Existing technologies usually use deep learning models to analyze video content, but consume a large amount of computing resources when processing high-resolution and long-duration videos, making it difficult to meet the requirements of real-time monitoring. In addition, existing methods have difficulties in distinguishing normal behavior changes from real abnormal behaviors, resulting in a high false alarm rate and increasing the workload of manual verification.
[0004] Existing video anomaly detection technologies show significant differences in performance under different lighting conditions. In night or low-light environments, the quality of video images decreases, leading to a significant reduction in detection performance. Many systems have to train different detection models for day and night respectively, increasing the system complexity and maintenance costs. Existing technologies are difficult to extract light-invariant features and lack an effective day-night unified mapping mechanism, resulting in poor all-weather monitoring effects.
[0005] Currently, the industry lacks a technical solution that can efficiently and accurately detect abnormal behaviors under various lighting conditions, making it difficult to meet the requirements of 24-hour uninterrupted monitoring. There is a need to develop new algorithms and systems to achieve day-night integrated video abnormal behavior detection and warning functions. Summary of the Invention
[0006] The present invention provides an AI video analysis and monitoring warning platform and its warning method to solve the technical problems in related technologies such as inconsistent detection performance under different lighting conditions, large consumption of computing resources, and high false alarm rate.
[0007] The present invention provides an AI video analysis and monitoring warning platform, including:
[0008] A video acquisition module for receiving an input video stream;
[0009] A feature processing module for performing spatio-temporal feature compression and extraction on the video stream to obtain a low-dimensional spatio-temporal feature vector and light-invariant features;
[0010] A multi-modal fusion module, which is used to perform multi-modal information fusion and enhancement on the video stream according to the current lighting conditions to obtain fused features;
[0011] A behavior mapping module, which is used to map the fused features to a unified feature space through a contrastive learning algorithm, establish a unified mapping relationship of the same behavior under different lighting conditions, and construct a day-night behavior mapping model;
[0012] An anomaly detection module, which is used to establish a normal behavior pattern library based on historical data, calculate the deviation degree from the normal pattern for the mapped feature vectors to form an anomaly metric score, dynamically adjust the anomaly determination threshold according to the scene complexity and lighting conditions, compare the anomaly metric score with the threshold, and generate an anomaly behavior warning message.
[0013] In a preferred embodiment, the spatio-temporal feature compression and extraction in the feature processing module includes:
[0014] Segment the input video stream into a video frame sequence of a fixed length;
[0015] Apply a convolutional neural network to each frame image in the video frame sequence to extract spatial features;
[0016] Connect the extracted spatial features along the time dimension to form a three-dimensional feature tensor;
[0017] Apply spatio-temporal convolution operations to compress the feature tensor to obtain a low-dimensional spatio-temporal feature vector.
[0018] In a preferred embodiment, the extraction of illumination-invariant features in the feature processing module includes:
[0019] Perform a frequency-domain transform on the input video sequence to obtain a spectral representation;
[0020] Calculate the gradient information of each frame image;
[0021] Combine the frequency-domain information and the gradient information through a fusion transform to generate an illumination-invariant feature representation, and the fusion transform is implemented by a weighted combination method.
[0022] In a preferred embodiment, the information fusion and enhancement in the multi-modal fusion module includes:
[0023] Calculate the average brightness value of the video frames;
[0024] Judge whether the current lighting condition is a low-light environment according to a preset threshold;
[0025] When it is determined to be a low-light environment, automatically call a thermal imaging or infrared sensor to collect data of the same scene;
[0026] Extract features of different modalities and apply a multimodal fusion function to combine the features of different modalities to generate fused features.
[0027] In a preferred embodiment, the fusion function is implemented in an adaptive weight manner. By assigning dynamically adjusted weight coefficients to the visual enhancement features and the thermal imaging or infrared features, the contribution ratio of each feature is automatically adjusted according to the current illumination condition to form the final fused features.
[0028] In a preferred embodiment, the day-night behavior mapping and modeling in the behavior mapping module include:
[0029] Screen video clips of the same behavior under different illumination conditions in the same scene from the training dataset to form a set of behavior sample pairs;
[0030] Define a feature extraction network to map the input samples to the feature space;
[0031] Define a similarity function to calculate the similarity between feature vectors;
[0032] Construct and minimize the contrast loss function;
[0033] Construct a unified day-night behavior mapping space.
[0034] In a preferred embodiment, the non-parametric anomaly metric modeling in the anomaly detection module includes:
[0035] Collect video clips of normal behaviors and extract their unified feature representations;
[0036] Apply the kernel density estimation method to construct the normal behavior feature distribution;
[0037] For the newly input video clips, calculate the probability density in the normal behavior distribution and convert it into an anomaly metric score;
[0038] Perform temporal smoothing on the anomaly scores of the continuous video frame sequence and optimize the detection results through temporal consistency verification.
[0039] In a preferred embodiment, the adaptive threshold determination in the anomaly detection module includes:
[0040] Calculate the number of targets, the dynamic change frequency, and environmental factors in the scene, and comprehensively form a scene complexity index;
[0041] Set a base threshold and calculate the adjustment coefficient according to the scene complexity and illumination condition;
[0042] Calculate the dynamic threshold and compare the smoothed anomaly score with the dynamic threshold to determine whether to trigger an anomaly warning;
[0043] For abnormal events that trigger early warnings, generate early warning information including the type, location, time, and confidence level of the abnormality.
[0044] In a preferred embodiment, the calculation of the dynamic threshold further includes a time adaptation mechanism. By comparing the mean of historical anomaly scores in the current period with the mean of historical anomaly scores throughout the day, and combining with a time adjustment coefficient, the basic threshold is dynamically adjusted to enable the system to adapt to changes in normal behavior patterns in different periods.
[0045] In a preferred embodiment, an AI video analysis monitoring and early warning method is applied to an AI video analysis monitoring and early warning platform, including the following steps:
[0046] Receive the input video stream, and perform spatio-temporal feature compression and extraction on the video stream to obtain a low-dimensional spatio-temporal feature vector and illumination-invariant features;
[0047] According to the current illumination conditions, perform multi-modal information fusion and enhancement on the video stream to obtain fused features;
[0048] Through a contrastive learning algorithm, map the fused features to a unified feature space, establish a unified mapping relationship of the same behavior under different illumination conditions, and construct a day-night behavior mapping model;
[0049] Based on historical data, establish a normal behavior pattern library, calculate the deviation degree from the normal pattern for the mapped feature vector, and form an anomaly metric score;
[0050] Dynamically adjust the anomaly determination threshold according to the scene complexity and illumination conditions, compare the anomaly metric score with the threshold, and generate early warning information for abnormal behaviors.
[0051] The beneficial effects of the present invention are as follows:
[0052] Day-night integrated detection ability: Through illumination-invariant feature extraction and day-night behavior mapping model, unified anomaly detection under different illumination conditions is achieved, solving the problem of a significant decline in the performance of traditional systems at night.
[0053] Efficient calculation and resource saving: By using spatio-temporal feature compression and non-parametric modeling methods, the computational resource requirements are reduced. Compared with traditional deep learning methods, the consumption of computational resources is reduced, and real-time processing can be achieved on ordinary computing hardware.
[0054] Reduction of false alarm rate: Through multi-modal information fusion, temporal consistency verification, and adaptive threshold determination, the system can effectively distinguish between normal behavior changes and true abnormal behaviors, reducing the false alarm rate and significantly reducing the workload of security personnel in handling false alarms.
[0055] System architecture simplification: Replace multiple detection models for different lighting conditions with a single day-night integrated model, reducing the complexity of model maintenance and update and lowering the system management cost.
[0056] Reduced data requirements: Through non-parametric modeling and contrastive learning methods, the system can work effectively with a small amount of labeled data, reducing the data collection and annotation costs and enabling the system to be quickly deployed to new scenarios. Description of the Drawings
[0057] Figure 1 It is a module diagram of an AI video analysis monitoring and early warning platform of the present invention. Detailed Implementation Modes
[0058] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0059] At least one embodiment of the present invention discloses an AI video analysis monitoring and early warning platform, as Figure 1 shown, including:
[0060] A video acquisition module for receiving an input video stream;
[0061] Specifically, it includes the following steps:
[0062] Step 1.1, video frame sequence conversion;
[0063] Segment the input video stream into video frame sequences of a fixed length, and each sequence contains frames (for example, , representing a 1-second video clip when the video frame rate is 30fps). For each video frame sequence , where represents a complete video frame sequence and serves as the basic unit for feature extraction; , , respectively represent the , , frame images in the sequence; represents the number of frames contained in each video sequence and is a parameter of the time window size.
[0064] Apply a low-dimensional spatio-temporal feature compression algorithm to process this sequence:
[0065] For each frame of the image Apply a convolutional neural network to extract spatial features and obtain a feature matrix ;
[0066] Connect the feature matrices of several frames along the time dimension to form a three-dimensional tensor ;
[0067] Apply spatio-temporal convolution operations to compress the tensor and obtain a low-dimensional spatio-temporal feature vector .
[0068] This process can be expressed as:
[0069] ;
[0070] where is the finally obtained low-dimensional spatio-temporal feature vector, representing the spatio-temporal features of the compressed video; represents the convolutional neural network feature extraction operation, which is used to extract spatial features from each frame of the image; represents the spatio-temporal convolution compression operation, which is used to compress the features of multiple frames into a low-dimensional representation; , , respectively represent the feature representations obtained after applying the convolutional neural network to the , , th frame of the image; n represents the number of frame images in the video sequence.
[0071] Step 1.2, Illumination-invariant feature extraction;
[0072] Aiming at the problem that the video features vary greatly under different illumination conditions, this step simultaneously performs illumination-invariant feature extraction:
[0073] Perform a frequency-domain transformation on the input video sequence to convert the image information from the spatial domain to the frequency domain and obtain a spectral representation ;
[0074] Calculate the gradient information of each frame of the image , including the horizontal and vertical gradients, which are insensitive to illumination changes;
[0075] Combine the frequency-domain information and the gradient information through a fusion transformation to generate an illumination-invariant feature representation.
[0076] This process can be expressed as:
[0077] ;
[0078] Among them, represents the illumination-invariant feature of the image, which is a feature representation extracted from the th frame image and is not affected by illumination changes; frame image extraction of the illumination change-free feature representation; is the image gradient, which represents the brightness change rate of the image at each position and is insensitive to illumination changes; is the result of the frequency-domain transformation, which represents the form of the image after being transformed from the spatial domain to the frequency domain and can capture the frequency characteristics of the image; is the fusion transformation function, which is used to perform weighted combination of the gradient information and the frequency-domain information to generate the final illumination-invariant feature.
[0079] Fusion transformation function is implemented by using a weighted combination method:
[0080] ;
[0081] Among them, represents the fusion transformation function, which is used to fuse the gradient information and the frequency-domain information; represents the image gradient information, which contains features such as edges and textures that are insensitive to illumination changes; represents the image frequency-domain transformation result; represents the normalization operation, which is used to map features of different scales to the same numerical range; represents the low-frequency part information after the frequency-domain transformation. The low-frequency part contains the main structure information of the image and is relatively robust to noise interference; , respectively represent the weight coefficients of the gradient information and the frequency-domain information, which are used to control the relative importance of the gradient information and the frequency-domain information in the fusion process and satisfy .
[0082] The frequency-domain transformation FT is implemented by using the two-dimensional discrete Fourier transform (Two-Dimensional Discrete Fourier Transform, 2D-DFT) to transform the image from the spatial domain to the frequency domain:
[0083] ;
[0084] Among them, represents the result of performing the two-dimensional discrete Fourier transform on the image ; is the pixel value of the image at the position ; and are respectively the width and height of the image, representing the spatial size of the image; is the coordinate in the frequency domain, ranges from to , ranges from to ; is the complex exponential function, represents the imaginary unit, which is used to transform spatial domain information into the frequency domain; represents the double summation operation over all pixels of the image.
[0085] Image gradient includes the horizontal gradient and the vertical gradient , which are calculated by the Sobel operator:
[0086] ;
[0087] ;
[0088] where, represents the convolution operation, which is a basic operation for feature extraction in image processing; and are the horizontal and vertical templates of the Sobel operator respectively, which are used to detect horizontal and vertical edges in the image; represents the gradient component of the image in the horizontal direction, which reflects the rate of change of image brightness in the horizontal direction; represents the gradient component of the image in the vertical direction, which reflects the rate of change of image brightness in the vertical direction; represents the th frame image of the input.
[0089] The final gradient magnitude is calculated as:
[0090] ;
[0091] where, represents the gradient magnitude of the image , which is a measure of the intensity of brightness change at each position in the image; represents the gradient component in the horizontal direction, which reflects the rate of change of image brightness in the horizontal direction; represents the gradient component in the vertical direction, which reflects the rate of change of image brightness in the vertical direction; represents the Euclidean norm of the gradient, which is obtained by calculating the square root of the sum of the squares of the horizontal and vertical gradient components. This calculation method can comprehensively consider the brightness changes in all directions.
[0092] By combining frequency domain information and gradient information, the system can extract features that are insensitive to illumination changes, providing a stable feature representation for day-night integrated abnormal behavior detection.
[0093] A feature processing module for performing spatio-temporal feature compression and extraction on the video stream to obtain a low-dimensional spatio-temporal feature vector and illumination-invariant features;
[0094] Specifically, it includes the following steps:
[0095] Step 2.1, Illumination condition detection and sensor invocation;
[0096] The system first evaluates the illumination condition of the input video to determine whether to invoke auxiliary sensor data:
[0097] Calculate the average brightness value of the video frame and the brightness variance ;
[0098] According to a preset threshold Judge the current illumination condition:
[0099] If , it is considered a low-illumination environment;
[0100] If , it is considered a normal-illumination environment.
[0101] For a low-illumination environment, the system automatically invokes an auxiliary sensor (such as a thermal imager or an infrared sensor) to collect data of the same scene.
[0102] Step 2.2, Multi-modal feature extraction and fusion;
[0103] For video data in a low-illumination environment, the system obtains additional information from the auxiliary sensor:
[0104] Obtain data from a thermal imager or an infrared sensor , extract thermal features ;
[0105] Perform enhancement processing on the original video image to obtain an enhanced image and its features ;
[0106] Apply a multi-modal fusion function to combine features of different modalities and generate fused features .
[0107] The fusion process can be expressed as:
[0108] ;
[0109] where It represents the feature vector obtained after multimodal fusion, which is the basic input for subsequent processing of the system; It is a visual feature processed by image enhancement technology, which contains the visual information extracted from the original video after contrast enhancement, denoising and other processing; It is a thermal signature obtained from thermal imaging or infrared sensors, which can provide stable target outline and motion information in low-light environments; It is a fusion function used to effectively combine the features of different modalities. It can be implemented by weighted summation, attention mechanism or neural network.
[0110] Fusion Function Adopt adaptive weight method to achieve:
[0111] ;
[0112] in, represents the fusion function, which is used to perform weighted combination of features of different modalities; Represents the feature vector extracted from the enhanced visual image, containing information of the visual modality; Represents a feature vector obtained from a thermal imaging or infrared sensor, containing data of the thermal information modality; Represents the weight coefficient of the visual feature, which determines the proportion of visual information in the fusion result; The weight coefficient of the thermal imaging feature determines the proportion of thermal information in the fusion result.
[0113] These adaptive weights and It will be dynamically adjusted according to the lighting conditions of the current scene:
[0114] ;
[0115] in, Represents the weight coefficient of visual features; Represents the weight coefficient of thermal imaging features; It is an S-type activation function that maps the input to between 0 and 1, ensuring that the weight value is within a reasonable range; It is the parameter that controls the slope of the sigmoid function and determines the sensitivity of the weight to changes in illumination. Values will make the weight change more drastically; is the bias parameter of the sigmoid function, which is used to adjust the center point of weight change. Values of will cause the visual features to retain a higher weight in lower lighting conditions; It is the average brightness value of the current video frame, which reflects the overall illumination level of the scene and is the key input for adaptive weight adjustment.
[0116] Parameter and is adjusted according to the actual application scenario, so that when the lighting conditions are extremely poor it is close to 1 (mainly relying on thermal imaging data), and when the lighting conditions improve it gradually increases (more relying on visual data).
[0117] For example, in the monitoring scenario of the entrance and exit of a subway station, when the night lights are flickering, relying solely on visual data may lead to unstable detection results. At this time, through multimodal fusion, combining the thermal imaging sensor data installed in the same area, the system can stably capture the pedestrian contour and action features, and maintain the detection performance even under complex lighting conditions.
[0118] Step 2.3, Feature quality evaluation and optimization;
[0119] To ensure the quality of the fused features, the system performs feature quality evaluation and optimization:
[0120] Calculate the signal-to-noise ratio SNR and the feature discriminability index of the fused features ;
[0121] If the index is lower than the preset threshold, adjust the fusion parameters and re-perform the fusion;
[0122] Generate the final multimodal fusion features .
[0123] The final fused features can be expressed as:
[0124] ;
[0125] where represents the final multimodal fusion features, is the input video, is the enhanced visual features, is the multimodal supplementary features, is the optimized fusion function.
[0126] Therefore, through this step, the system can obtain high-quality feature representations even in low-light environments, providing reliable inputs for subsequent abnormal behavior detection.
[0127] The multimodal fusion module is used to perform multimodal information fusion and enhancement on the video stream according to the current lighting conditions to obtain fused features;
[0128] Specifically, it includes the following steps:
[0129] Step 3.1, Construction of behavior sample pairs;
[0130] To train the day-night integrated model, it is first necessary to construct sample pairs of the same behavior under different lighting conditions:
[0131] Screen video clips of the same behavior under different lighting conditions (daytime and nighttime) from the training dataset in the same scene;
[0132] For each behavior category , collect multiple positive sample pairs , where is the daytime sample, is the corresponding nighttime sample;
[0133] Form a set of behavior sample pairs , where represents the set of all behavior sample pairs; , , respectively represent the , , th behavior sample pair, containing a pair of the same behavior under different lighting conditions; represents the total number of sample pairs, which determines the scale of the training data.
[0134] Step 3.2, Contrastive learning model training;
[0135] Based on the constructed set of sample pairs, apply the contrastive learning algorithm to train the day-night behavior mapping model:
[0136] Define a feature extraction network , which maps the input samples to the feature space;
[0137] For each sample pair , calculate their representations in the feature space and ;
[0138] Define a similarity function , and calculate the cosine similarity between feature vectors:
[0139] ;
[0140] Among them, represents the cosine similarity between the feature vectors and , which is used to measure the similarity degree of the directions of two vectors; represents the dot product (inner product) of the vectors and ; represents the Euclidean norm (L2 norm) of the vector , that is, the vector length; represents the vector Euclidean norm; represents the normalized dot product, and the result ranges from [-1, 1]. The closer the value is to 1, the more similar the directions of the two vectors are. The closer the value is to -1, the more opposite the directions are. A value of 0 indicates orthogonality (no correlation).
[0141] Construct a contrastive loss function , minimize the distance between positive sample pairs and maximize the distance from other samples:
[0142] ;
[0143] where, represents the contrastive loss function, which is used to measure the distinguishability between positive sample pairs and other samples; represents the feature representation of sample under standard lighting conditions (daytime); represents the corresponding feature representation of the same behavior sample under low-light conditions (nighttime); represents the feature representation of other samples in the training set; represents the cosine similarity function between feature vectors; is the temperature parameter, which controls the smoothness of the feature distribution. A smaller value (such as 0.07) makes the model more sensitive to similarity differences and enhances the discriminative ability of features; represents the natural exponential function; represents the natural logarithm function. The numerator part calculates the similarity score of the target sample pair, and the denominator part calculates the sum of the similarity scores of the target sample and all other samples. The overall expression measures the similarity degree of the positive sample pair relative to other sample pairs.
[0144] Optimize the network parameters through gradient descent so that the feature representations of the same behavior under different lighting conditions are as close as possible.
[0145] Feature extraction network consists of a Convolutional Neural Network (CNN) and a Temporal Convolutional Network (TCN), and the structure is as follows:
[0146] Input layer: Receive a sequence of video clips;
[0147] Spatial feature extraction layer: Consists of multiple convolutional layers to extract the spatial features of each frame;
[0148] Temporal feature extraction layer: Consists of multiple temporal convolutional layers to capture the temporal dependencies across frames;
[0149] Feature Fusion Layer: Combines spatial and temporal features;
[0150] Mapping Layer: Maps the fused features to a unified feature space;
[0151] The network optimizes the parameters through the backpropagation algorithm , minimizing the contrast loss function :
[0152] ;
[0153] Among them, represents the set of optimized model parameters, that is, the network weights and bias values finally obtained through training; represents finding the parameters that minimize the objective function ; represents the contrast loss function, which is a function of the parameter , measuring the performance of the model under the current parameters. The smaller the value, the closer the feature representations of the same behavior under different lighting conditions; represents the set of trainable parameters of the model, including the weight matrices and bias vectors in the convolutional layer, temporal convolutional layer, and mapping layer.
[0154] Among them, the optimization adopts Stochastic Gradient Descent (SGD) with momentum:
[0155] ;
[0156] Among them, represents the model parameters after the th iteration; represents the model parameters of the th iteration; represents the model parameters of the th iteration; is the learning rate, controlling the step size of each parameter update; is the momentum coefficient, used to accelerate convergence and help escape local minima; is the loss function with respect to the parameter , representing the change direction and rate of the loss function at the current parameter position; represents the change amount of the previous parameter update, which forms a momentum term after being multiplied by the momentum coefficient, helping to maintain the inertia of parameter updates.
[0157] In practical applications, such as in shopping mall surveillance scenarios, the model can map the same behaviors (such as customers wandering, gathering, etc.) during the day and at night to similar feature representations. After training, even when the lighting conditions change at night, the system can still accurately identify the same behavior patterns as during the day, reducing false alarms caused by lighting changes.
[0158] Step 3.3, Unified mapping of behavioral features;
[0159] After training, the system constructs a unified day-night behavior mapping space:
[0160] For any input video clip , regardless of its lighting conditions, apply the trained feature extraction network to generate a unified feature representation ;
[0161] Build a behavioral prototype library , where represents the set containing the prototype features of all behavior categories; , , respectively represent the prototype features of the , , th category of behavior; represents the total number of behavior categories defined in the system;
[0162] Each prototype feature is calculated from the feature mean of all samples in this category, that is:
[0163] ;
[0164] where, represents the th prototype feature, represents the total number of samples of the th category of behavior, represents the th sample of the th category of behavior, represents the feature representation obtained by this sample through the feature extraction network;
[0165] For the feature representation of the newly input sample, calculate its similarity with each prototype, which can be used for subsequent behavior recognition and anomaly detection.
[0166] In addition, through this step, the system realizes the unified representation of behavioral features under different lighting conditions, enabling subsequent anomaly detection to be carried out in a unified feature space without maintaining multiple sets of models for different lighting conditions.
[0167] A behavior mapping module, which is used to map the fused features to a unified feature space through a contrastive learning algorithm, establish a unified mapping relationship of the same behavior under different lighting conditions, and construct a day-night behavior mapping model;
[0168] Specifically, it includes the following steps:
[0169] Step 4.1, construction of a normal behavior pattern library;
[0170] The system first establishes a normal behavior pattern library based on historical data:
[0171] Collect a large number of video clips of normal behaviors and extract their unified feature representations through the aforementioned steps ;
[0172] Apply a density estimation method to construct the normal behavior feature distribution ;
[0173] Adopt the Kernel Density Estimation (KDE) method to avoid making parametric assumptions about the data distribution:
[0174] ;
[0175] Among them, represents the feature in the probability density value of the normal behavior distribution; represents the feature representation vector of the sample to be evaluated; represents the total number of normal samples; represents the bandwidth parameter, which controls the smoothness of the kernel function. A larger value will make the probability density estimation smoother, and a smaller value will make the estimation closer to the sample data; represents the summation over all normal samples; is the kernel function (usually a Gaussian kernel), which is used to measure the similarity between two points in the feature space; is the th feature representation vector of the normal sample; represents normalizing the feature difference between the sample to be evaluated and the normal sample and then inputting it into the kernel function to calculate the similarity contribution.
[0176] In specific implementation, the Gaussian kernel function is expressed as:
[0177] ;
[0178] Among them, represents the Gaussian kernel function, which is used to measure the similarity between points in the feature space; Denotes the standardized distance, representing the normalized distance between the current sample feature and the reference sample feature; is the normalization constant to ensure that the integral of the Gaussian kernel function is 1; is the exponential function part, which decays rapidly as the distance increases, reflecting the characteristic that the similarity is lower for greater distances.
[0179] Bandwidth parameter is adaptively determined by the Silverman rule:
[0180] ;
[0181] where, is the bandwidth parameter that controls the smoothness of the kernel density estimation. A larger value makes the estimation smoother but may lose details, while a smaller value preserves more details but may introduce noise; is the sample standard deviation, a statistic representing the degree of data dispersion; is the sample interquartile range, i.e., the difference between the third quartile and the first quartile, a robust statistic for measuring the data dispersion and insensitive to outliers; is the sample size, representing the total number of normal behavior samples used to build the model; Choose the smaller value between the standard deviation and the normalized interquartile range to improve the robustness of the estimation; is a power function of the sample size. As the sample size increases, the bandwidth will decrease appropriately; is the empirical adjustment coefficient used to fine-tune the final bandwidth value.
[0182] To improve the computational efficiency, the system uses a BallTree data structure to store the normal sample features and accelerate the nearest neighbor search process. This structure divides the feature space into nested hyperspheres to enable fast queries for high-dimensional features.
[0183] In some application scenarios, such as airport security area monitoring, normal behavior patterns may have obvious time period characteristics. For example, the passenger flow patterns during peak and off-peak hours are quite different. In response to this situation, the system can build a time period-related pattern library:
[0184] ;
[0185] where, represents the conditional probability density value of the feature in the normal behavior distribution under the condition of time period ; represents the feature representation vector of the sample to be evaluated; Represents a specific time period (such as the morning rush hour, off-peak period, etc.); Represents the time period The total number of samples within; Represents the bandwidth parameter, which controls the smoothness of the kernel function; Represents the time period The sample set within; Represents the sum of all samples within the time period ; Represents the kernel function, usually the Gaussian kernel function is selected; Represents the time period The th feature representation vector of the normal sample within; Represents inputting the standardized feature difference between the sample to be evaluated and the normal sample into the kernel function. This enables the system to select an appropriate normal behavior model according to the current time period, further improving the detection accuracy.
[0186] For applications with a large monitoring area and complex scenarios, the system can adopt a partitioned modeling strategy, dividing the monitoring area into multiple sub-regions, and constructing an independent normal behavior model for each sub-region. This method can capture the behavior characteristics of different regions more precisely and reduce the interference caused by spatial location differences.
[0187] Step 4.2, Abnormality metric calculation;
[0188] For the newly input video clip, the system calculates its abnormality metric score:
[0189] Extract the unified feature representation of the input video clip ;
[0190] Calculate the probability density of this feature in the normal behavior distribution ;
[0191] Convert it to an abnormality metric score :
[0192] ;
[0193] Among them, Represents the abnormality score of the newly input sample , and the higher the value, the more likely the sample is an abnormal behavior; Represents the probability density value of the new sample feature in the normal behavior distribution. The lower the probability density value, the less likely the sample feature appears in the normal behavior distribution; Is the negative logarithm transformation, which converts the probability density value into a non-negative abnormality score, and at the same time amplifies the difference of low-probability events, making the lower the probability, the higher the converted abnormality score.
[0194] The lower the probability density value, the higher the corresponding anomaly score.
[0195] Step 4.3, Temporal consistency verification;
[0196] To reduce false alarms, the system optimizes the anomaly detection results through temporal consistency verification:
[0197] For a continuous video frame sequence, record the time series of its anomaly scores ;
[0198] Apply a temporal smoothing filter to reduce random fluctuations:
[0199] ;
[0200] where, represents the smoothed anomaly score at time point ; represents the original anomaly score at time point , that is, the anomaly score by looking back time units; represents the weighted sum of all points within the considered time window; is the smoothing weight coefficient, used to control the contribution degree of anomaly scores at different time points to the current smoothing result. Usually, the anomaly scores in the recent period have higher weights, and the anomaly scores in the earlier period have lower weights, satisfying , ensuring weight normalization; represents the time offset, which defines the range of the time window.
[0201] Detect the mutation pattern of the anomaly score to distinguish real anomalies from random fluctuations:
[0202] If the anomaly score suddenly rises and lasts for a period of time, it is determined as a real anomaly;
[0203] If the anomaly score is only briefly fluctuating, it may be a false alarm and is filtered out.
[0204] In addition, in the specific implementation, the exponentially weighted moving average method is used to calculate the smoothed anomaly score:
[0205] ;
[0206] where, represents the smoothed anomaly score at the current time , which is the anomaly value after temporal smoothing; represents the original anomaly score at the current time , which is directly calculated by the system; represents the previous time The smoothed anomaly score is used for weighted averaging with the current score; is the smoothing factor that controls the weight ratio between the current observation and the historical smoothed value. A smaller value makes the smoothing result more dependent on historical values, producing a smoother time series. A larger value makes the result respond more quickly to new changes; is the weight coefficient of the historical smoothed value, ensuring that the current smoothed value comprehensively considers historical information.
[0207] It can be seen that through non-parametric modeling and time series consistency verification, the system can efficiently distinguish changes in normal behavior from true abnormal behavior, significantly reducing the false alarm rate while maintaining a high detection rate for real anomalies.
[0208] The anomaly detection module is used to establish a normal behavior pattern library based on historical data, calculate the deviation from the normal pattern for the mapped feature vectors to form an anomaly metric score, dynamically adjust the anomaly determination threshold according to the scene complexity and lighting conditions, compare the anomaly metric score with the threshold, and generate an early warning message for abnormal behavior.
[0209] Specifically, it includes the following steps:
[0210] Step 5.1, Scene complexity assessment;
[0211] The system first assesses the complexity of the monitored scene:
[0212] Calculate the number of targets in the scene , the dynamic change frequency and environmental factors (such as light changes, weather conditions, etc.);
[0213] Calculate the scene complexity index by integrating these factors :
[0214] ;
[0215] Among them, represents the scene complexity index, which is used to quantify the complexity of the monitored scene; represents the number of target objects in the scene. The larger the number, the more complex the scene; represents the frequency of dynamic changes in the scene. The more frequent the changes, the more complex the scene; represents environmental factors, including external environmental impacts such as light changes and weather conditions; , , respectively represent the weight coefficients of the number of targets, dynamic change frequency, and environmental factors, which are adjusted according to the actual application scenario;
[0216] The higher the complexity of the scenario, the greater the variability of normal behavior, and the anomaly determination threshold needs to be adjusted accordingly.
[0217] Step 5.2, threshold dynamic adjustment;
[0218] Based on scene complexity and lighting conditions, the system dynamically adjusts the anomaly determination threshold:
[0219] Setting Baseline Thresholds , usually determined by optimizing performance on a validation set;
[0220] According to the complexity of the scene and lighting conditions Calculate the adjustment factor :
[0221] ;
[0222] in, is the adjustment function, which can adopt linear or nonlinear mapping; is the threshold adjustment factor, used to scale the base threshold; It is a scene complexity indicator, reflecting the comprehensive evaluation value of the number of targets, dynamic change frequency and environmental factors in the monitoring scene; It is the average light intensity of the scene, which is used to measure the lighting conditions of the current monitoring environment.
[0223] Calculating dynamic thresholds :
[0224] ;
[0225] in, is the dynamic threshold used in the end, It is the basic threshold preset by the system. It is an adjustment factor calculated based on scene complexity and lighting conditions, and is used to scale the basic threshold to meet the needs of different monitoring environments.
[0226] Adjustment function The specific implementation is:
[0227] ;
[0228] in, is the threshold adjustment function, which is used to calculate the adjustment coefficient according to the scene complexity and lighting conditions; It is the complexity index of the current scene. The larger the value, the more complex the scene. is the average light intensity of the current scene. The smaller the value, the worse the lighting conditions. It is the standard scene complexity benchmark value, which serves as a reference standard for scene complexity; is the standard light intensity threshold, serving as a reference standard for lighting conditions; is the scene complexity adjustment parameter, controlling the degree of influence of scene complexity on the threshold; is the lighting condition adjustment parameter, controlling the degree of influence of lighting conditions on the threshold.
[0229] When the monitored scene complexity is higher than the standard value ( ), or the lighting condition is lower than the standard value ( ), the value will increase, resulting in an increase in the abnormal determination threshold, thereby reducing system false alarms.
[0230] The system also implements a time - adaptive mechanism for the threshold, which is fine - tuned according to the historical statistical characteristics of different time periods:
[0231] ;
[0232] Among them, represents the final dynamic threshold at time , which is the threshold value after time - adaptive adjustment; represents the initial dynamic threshold calculated based on scene complexity and lighting conditions; is the mean value of historical abnormal scores in the current time period, used to characterize the average level of abnormal scores within a specific time period; is the mean value of historical abnormal scores throughout the day, representing the overall average level of abnormal scores within 24 hours of the whole day; is the time adjustment coefficient, controlling the sensitivity of the time factor to threshold adjustment. A larger value will make the threshold more sensitive to time - period differences; represents the deviation rate of the mean value of abnormal scores in the current time period relative to the mean value throughout the day. A positive value indicates that the abnormal scores in the current time period are generally higher than the average level throughout the day, and a negative value indicates lower than the average level. This time - adaptive mechanism enables the system to automatically adjust the abnormal determination criteria according to the behavioral pattern characteristics of different time periods, further reducing the false alarm rate.
[0233] In practical applications, such as hospital ward monitoring, this adaptive threshold mechanism can effectively handle detection requirements in different scenarios. During visiting hours, due to more personnel flow, the system will automatically increase the abnormal determination threshold to avoid normal visiting activities being misjudged as abnormal; while during the quiet night period, the system will automatically lower the threshold to be sensitive to minor abnormal behaviors and ensure patient safety.
[0234] The system can also implement a feedback - based threshold optimization mechanism. When the operator confirms or negates the system warning, this feedback information is recorded and used to continuously optimize the threshold parameters:
[0235] If the warning is confirmed as a real anomaly, slightly reduce the threshold in similar scenarios;
[0236] If the warning is marked as a false alarm, appropriately increase the threshold in similar scenarios.
[0237] Through this closed-loop optimization, the system can continuously improve the accuracy of warnings over time and reduce the need for manual intervention.
[0238] Step 5.3, Abnormal warning generation;
[0239] When abnormal behavior is detected, the system generates detailed warning information:
[0240] Compare the smoothed anomaly score with the dynamic threshold :
[0241] If , an abnormal warning is triggered;
[0242] Otherwise, it is regarded as normal behavior.
[0243] For abnormal events that trigger warnings, the system generates warning information including:
[0244] Anomaly type: determined according to the matching degree between the characteristics of abnormal behavior and predefined patterns;
[0245] Anomaly location: determine the exact physical location where the anomaly occurs through spatio-temporal coordinate mapping;
[0246] Anomaly time: record the exact time point and duration when the anomaly occurs;
[0247] Confidence level: calculate the credibility of the warning based on the anomaly score, expressed as:
[0248] ;
[0249] Among them, is the confidence level of the abnormal warning, indicating the degree of certainty of the system's judgment on this anomaly; is the anomaly score after smoothing, reflecting the degree of abnormality of the current behavior; is the dynamically adjusted anomaly determination threshold, representing the current standard line for the system to determine anomalies. The higher the confidence level, the more obvious the detected abnormal behavior exceeds the normal range, and the higher the degree of certainty of the system for this warning.
[0250] Send the warning information to the relevant processing terminals or storage systems according to the priority.
[0251] Therefore, through adaptive threshold determination, the system can dynamically adjust the abnormal determination criteria according to different scenarios and environmental conditions, effectively reducing the false alarm rate while maintaining sensitivity to real anomalies and providing accurate and timely warning information.
[0252] Application example of this embodiment:
[0253] Application in urban traffic monitoring scenario:
[0254] Taking the urban traffic monitoring scenario as an example, the application process of this embodiment is as follows:
[0255] The system is deployed at an intersection of a traffic hub in the city, equipped with 4K high-definition cameras and auxiliary thermal imaging sensors. The monitoring area covers the sidewalk, motor vehicle lane, and non-motor vehicle lane, and the goal is to detect traffic abnormal behaviors.
[0256] Example of the implementation process:
[0257] Space-time feature compression and extraction: The system processes 30 frames of 4K video per second. First, the video is segmented into 1-second segments (30 frames), and the ResNet-18 convolutional network is applied to each frame to extract features, obtaining 2048-dimensional feature vectors. These feature vectors are compressed through space-time convolution to obtain a 256-dimensional low-dimensional representation.
[0258] At the same time, the system extracts illumination-invariant features, especially for scenes with low light at night and drastic light changes (such as headlight illumination, shadow changes, etc.).
[0259] Multi-modal information fusion and enhancement: In a low-light environment at night (average brightness value below 50), the system automatically activates the thermal imaging sensor. When pedestrians or vehicles are detected in the night environment, the visual features and thermal imaging features are fused according to an adaptive weight ratio of 70%:30% to improve the feature quality.
[0260] For example, when a vehicle driving with its lights off enters the monitoring area at night, although it is almost invisible in a normal camera, the system can still capture its motion features through the supplement of thermal imaging features.
[0261] Day-night behavior mapping and modeling: The system trains a contrastive learning model based on 500 pairs of samples of the same behavior during the day and at night (such as the behavior of pedestrians crossing the road at the same location). The model uses a temperature parameter , and is trained for 100 rounds through the SGD optimizer. The initial learning rate is set to 0.01 and the cosine annealing strategy is adopted.
[0262] After training, whether in the bright environment during the day or the dim condition at night, the system can map the same behavior to similar positions in a unified feature space.
[0263] Non-parametric Anomaly Metric Modeling: The system uses the normal traffic behavior data collected within two weeks (about 100 hours) to build a library of normal behavior patterns. A Gaussian kernel (bandwidth ) is used to perform kernel density estimation and distinguish different time period patterns (such as morning and evening rush hours and off-peak periods). When the probability density of the newly input behavior features in the normal distribution is lower than the threshold, anomaly detection is triggered. For example, behaviors such as pedestrians staying in the middle of the lane, vehicles suddenly changing lanes or driving in reverse will generate characteristic representations that significantly deviate from the normal pattern.
[0264] Adaptive Threshold Determination: The system dynamically adjusts the anomaly determination threshold. The basic threshold is set to 0.85 and is adjusted according to the current scene complexity (such as traffic flow, pedestrian density) and lighting conditions.
[0265] In complex scenarios (such as rush hours), the threshold is increased by 10% - 15%;
[0266] In low-light environments, the threshold is increased by 5% - 10% to reduce false alarms.
[0267] At the same time, the system further fine-tunes the threshold according to the characteristics of different time periods. For example, at the moment of traffic signal switching, the threshold is temporarily increased to allow for normal traffic pattern changes.
[0268] Verification of Technical Effects:
[0269] Through the analysis of the operation data of this traffic hub monitoring system for one month, the following two key technical effects are verified:
[0270] All-weather Detection Ability: The comparison of the anomaly detection accuracy rates of the system under different lighting conditions is shown in Table 1:
[0271] Table 1: Comparison of anomaly detection accuracy rates of the system under different lighting conditions;
[0272]
[0273] The results show that this system maintains a high detection accuracy rate under various lighting conditions. Especially in low-light environments, there is a significant improvement compared with traditional methods, achieving a true all-weather monitoring ability.
[0274] System Resource Efficiency: The comparison of resource utilization efficiency is shown in Table 2:
[0275] Table 2: Comparison of resource utilization efficiency;
[0276]
[0277] The results show that this system significantly reduces the computing resource requirements, enabling a single server to process more video streams simultaneously, and greatly reducing the deployment and operation and maintenance costs.
[0278] The embodiments of the present invention have been described above. However, these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of these embodiments, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of these embodiments.
Claims
1. An AI video analysis monitoring and early warning platform, characterized in that: include: A video acquisition module, used for receiving an input video stream; The feature processing module is used to perform spatiotemporal feature compression and extraction on the video stream to obtain low-dimensional spatiotemporal feature vectors and illumination invariant features; The multimodal fusion module is used to perform multimodal information fusion and enhancement on the video stream according to the current lighting conditions to obtain fusion features; The behavior mapping module is used to map the fused features to a unified feature space through a contrast learning algorithm, establish a unified mapping relationship for the same behavior under different lighting conditions, and construct a day-night behavior mapping model; The anomaly detection module is used to establish a normal behavior pattern library based on historical data, calculate the deviation of the mapped feature vector from the normal pattern, form an anomaly measurement score, dynamically adjust the anomaly judgment threshold according to the scene complexity and lighting conditions, compare the anomaly measurement score with the threshold, and generate abnormal behavior warning information.
2. The AI video analysis monitoring and early warning platform according to claim 1 is characterized in that: The temporal and spatial feature compression and extraction in the feature processing module include: Split the input video stream into a fixed-length video frame sequence; Apply convolutional neural network to extract spatial features for each frame image in the video frame sequence; Connect the extracted spatial features along the time dimension to form a three-dimensional feature tensor; Apply the spatiotemporal convolution operation to compress the feature tensor and obtain a low-dimensional spatiotemporal feature vector.
3. The AI video analysis monitoring and early warning platform according to claim 1 is characterized in that: The illumination invariant feature extraction in the feature processing module includes: Perform frequency domain transformation on the input video sequence to obtain a spectrum representation; Calculate the gradient information of each frame image; The frequency domain information and gradient information are combined through fusion transformation to generate illumination invariant feature representation. The fusion transformation is implemented by weighted combination.
4. The AI video analysis monitoring and early warning platform according to claim 1 is characterized in that: The information fusion and enhancement in the multimodal fusion module includes: Calculate the average brightness value of the video frame; Determine whether the current lighting condition is a low-light environment according to a preset threshold; When it is determined to be a low-light environment, thermal imaging or infrared sensors are automatically used to collect data of the same scene; Extract features of different modalities, and apply multimodal fusion function to combine the features of different modalities to generate fused features.
5. The AI video analysis monitoring and early warning platform according to claim 4 is characterized in that: The fusion function is implemented in an adaptive weighting manner, by assigning dynamically adjusted weight coefficients to visual enhancement features and thermal imaging or infrared features, and automatically adjusting the contribution ratio of each feature according to current lighting conditions to form a final fusion feature.
6. The AI video analysis monitoring and early warning platform according to claim 1 is characterized in that: The circadian behavior mapping and modeling in the behavior mapping module include: Filter video clips of the same behavior under different lighting conditions in the same scene from the training data set to form a set of behavior sample pairs; Define a feature extraction network to map input samples to feature space; Define a similarity function and calculate the similarity between feature vectors; Construct and minimize the contrastive loss function; Constructing a unified diurnal behavior mapping space.
7. The AI video analysis monitoring and early warning platform according to claim 1 is characterized in that: The non-parametric anomaly measurement modeling in the anomaly detection module includes: Collect video clips of normal behavior and extract their unified feature representation; The kernel density estimation method is used to construct the distribution of normal behavior characteristics; For a new input video clip, calculate the probability density in the normal behavior distribution and convert it into an abnormality metric score; Temporal smoothing of anomaly scores of continuous video frame sequences is performed, and detection results are optimized through temporal consistency verification.
8. The AI video analysis monitoring and early warning platform according to claim 1 is characterized in that: The adaptive threshold determination in the anomaly detection module includes: Calculate the number of targets, dynamic change frequency and environmental factors in the scene to comprehensively form a scene complexity index; Set a basic threshold and calculate an adjustment factor based on scene complexity and lighting conditions; Calculate the dynamic threshold and compare the smoothed anomaly score with the dynamic threshold to determine whether an anomaly warning is triggered; For abnormal events that trigger early warnings, early warning information is generated including the abnormal type, location, time and confidence level.
9. The AI video analysis monitoring and early warning platform according to claim 8, characterized in that: The calculation of the dynamic threshold also includes a time adaptive mechanism, which dynamically adjusts the basic threshold by comparing the difference between the average historical anomaly score of the current period and the average historical anomaly score of the whole day, combined with the time adjustment coefficient, so that the system can adapt to the changes in normal behavior patterns in different time periods.
10. An AI video analysis monitoring and early warning method, applied to an AI video analysis monitoring and early warning platform according to any one of claims 1 to 9, characterized in that: The following steps are involved: Receive an input video stream, and perform spatiotemporal feature compression and extraction on the video stream to obtain a low-dimensional spatiotemporal feature vector and illumination invariant features; According to the current lighting conditions, multimodal information fusion and enhancement are performed on the video stream to obtain fusion features; Through contrast learning algorithm, the fused features are mapped to a unified feature space, a unified mapping relationship of the same behavior under different lighting conditions is established, and a day and night behavior mapping model is constructed; A normal behavior pattern library is established based on historical data, and the deviation of the mapped feature vector from the normal pattern is calculated to form an abnormal measurement score; The anomaly determination threshold is dynamically adjusted according to scene complexity and lighting conditions, the anomaly metric score is compared with the threshold, and abnormal behavior warning information is generated.
Citation Information
Cited By
Expressway scene video segmentation method and device and electronic equipment
CN121600450A
A highway scene video segmentation method, device and electronic equipment
CN121600450B